Working Papers

Accounting for Industry Heterogeneity in Predicting Corporate Default Probability: A Semiparametric Learning Approach

with Shaobo Li, Ben Sherwood, and Xiaorui Zhu

Revise and Resubmit, European Journal of Operational Research

Abstract: Predicting corporate default is essential for risk management, credit allocation, and investment decision making, yet most bankruptcy models overlook heterogeneity in how financial distress manifests across industries. We propose the single-index hazard model with multiple links (SIML), a semiparametric learning method that combines a common financial risk index with industry-specific flexible link functions. This structure captures nonlinear industry effects and their interactions with firm-level financial characteristics while retaining an interpretable low-dimensional index. We also develop a two-stage simulation-based testing procedure for industry heterogeneity and nonlinearity. Using a comprehensive database of 18,651 U.S. publicly traded firms from 1980 to 2023, we find strong evidence of distinct nonlinear risk patterns across industries. SIML outperforms a broad set of benchmark models in both in-sample and out-of-sample evaluations. A simulated corporate loan market and an asset-pricing analysis of the distress risk anomaly further show the economic value of accounting for industry-specific default risk.

Cost Restricted Feature Selection with Information Criterion

with Ben Sherwood and Shaobo Li

Accepted for presentation at the 2026 INFORMS Workshop on Data Science
Preparing for submission to Management Science

Abstract: We study feature selection for data acquisition when candidate variables are costly and must be selected under an acquisition budget. Existing cost-restricted methods typically focus on predictive loss subject to the budget, but they do not explicitly account for the increase in model complexity from acquiring additional variables. We propose a cost-restricted information criterion approach, CRIC, that preserves the acquisition budget as a feasibility constraint while adding an AIC- or BIC-type complexity charge for each selected feature. This allows the method to reject feasible variables whose incremental contribution is too small to justify their inclusion, rather than automatically exhausting the available budget. We develop computational procedures for Gaussian and binary-response settings and evaluate the approach through extensive simulations and empirical applications. The results show that CRIC can substantially reduce unnecessary feature acquisition and select smaller models while preserving competitive out-of-sample predictive performance, with particularly clear gains when the acquisition budget is relatively loose.

Robust Outcome Optimization via Multi-Model Inverse Classification

with Michael Lash, Qihang Lin, and Nick Street

Target journal: Journal of Machine Learning Research

Abstract: This paper develops a multi-model inverse classification framework for robust outcome optimization. Existing inverse classification methods generate intervention recommendations with respect to a single predictive model, even though practitioners often face a collection of plausible models that can disagree about the effect of an intervention. We instead optimize recommendations over the distribution of predictions generated by multiple models and allow the decision maker to choose the desired level of risk. The framework accounts for intervention costs, feasibility constraints, and indirect changes in non-actionable features, while the optimization formulation targets recommendations that remain effective across model variation rather than being tailored to a single fitted model. In an application to cardiovascular risk reduction, the method produces practical recommendations that remain effective across alternative predictive specifications, illustrating how multi-model inverse classification can make prescriptive analytics more robust to variation in the underlying prediction model.

Objective-Aligned Sparse Learning for Neyman-Pearson Classification

Manuscript in preparation

Abstract: This paper develops an objective-aligned framework for asymmetric classification in which score estimation and variable selection are learned specifically to minimize false negatives subject to a prespecified false positive rate constraint. Standard Neyman-Pearson procedures typically estimate a scoring rule under a symmetric objective and then use threshold calibration to impose asymmetric error control. Our method instead incorporates the deployment objective directly into regularized score learning, while retaining a separate calibration step for final false positive rate control. This allows the selected variables and score direction to reflect the operating point that matters for the decision problem. Simulation studies consistently show lower false negative rates than standard and penalized Neyman-Pearson classifiers, especially when the variables that matter under the asymmetric objective differ from those selected under a symmetric loss. In a bankruptcy application using 18,651 U.S. publicly traded firms from 1980 to 2023, the method reduces the false negative rate from 48.9% to 35.1% at a 5% false positive rate target, a 28% reduction in missed bankruptcies, and raises recall from 51.1% to 64.9%.

Dynamic Instrumented Principal Components for Time-Varying Risk Exposures

Manuscript in preparation

Abstract: This paper develops a dynamic extension of instrumented principal components analysis (IPCA) that allows the mapping from firm characteristics to latent factor loadings to evolve over time. Standard IPCA permits characteristics to vary across firms and time but keeps the characteristic-to-loading map fixed. Our model allows this relationship itself to change, capturing persistent shifts in how firm characteristics translate into systematic risk exposures while preserving the characteristic-based factor structure and interpretability of IPCA. Empirically, the dynamic specification substantially improves in-sample fit across latent factor dimensions. With appropriate control of dynamic variation, it also produces sizable out-of-sample gains in return and managed-portfolio prediction relative to static IPCA. The results indicate that allowing the characteristic-risk relationship to evolve improves the description of realized systematic variation and the prediction of expected returns.

Publications

Relationship between piperacillin concentrations, clinical factors and piperacillin/tazobactam-associated acute kidney injury

Tang Girdwood, S., Hasson, D., Caldwell, J. T., Slagle, C., Dong, S., Fei, L., Tang, P., Vinks, A. A., Kaplan, J., & Goldstein, S. L.

Journal of Antimicrobial Chemotherapy, 78(2), 478–487, 2023.

DOI

Diagnostic accuracy of serum matrix metalloproteinase-7 as a biomarker of biliary atresia in a large North American cohort

Pandurangi, S., Mourya, R., Nalluri, S., Fei, L., Dong, S., Harpavat, S., Guthery, S. L., Molleston, J. P., Rosenthal, P., Sokol, R. J., Wang, K. S., Ng, V., Alonso, E. M., Hsu, E. K., Karpen, S. J., Loomes, K. M., Magee, J. C., Shneider, B. L., Horslen, S. P., Teckman, J. H., & Bezerra, J. A.

Hepatology, 2024.

DOI

Maternal metabolic dysfunction-associated steatotic liver disease during pregnancy and its impact on offspring steatosis

Klepper, C. M., Trout, A. T., Dong, S., Fei, L., Stankiewicz, T. E., Wayland, J., Moreno-Fernandez, M. E., Divanovic, S., Xanthakos, S. A., DeFranco, E., Woo, J. G., & Mouzaki, M.

Clinical Nutrition ESPEN, 70, 327–335, 2025.

DOI