Working Papers
Accounting for Industry Heterogeneity in Predicting Corporate
Default Probability: A Semiparametric Learning Approach
with Shaobo Li, Ben Sherwood, and Xiaorui Zhu
Revise and Resubmit,
European Journal of Operational Research
Abstract:
Predicting corporate default is essential for risk management,
credit allocation, and investment decision making, yet most
bankruptcy models overlook heterogeneity in how financial distress
manifests across industries. We propose the single-index hazard
model with multiple links (SIML), a semiparametric learning method
that combines a common financial risk index with industry-specific
flexible link functions. This structure captures nonlinear industry
effects and their interactions with firm-level financial
characteristics while retaining an interpretable low-dimensional
index. We also develop a two-stage simulation-based testing
procedure for industry heterogeneity and nonlinearity. Using a
comprehensive database of 18,651 U.S. publicly traded firms from
1980 to 2023, we find strong evidence of distinct nonlinear risk
patterns across industries. SIML outperforms a broad set of
benchmark models in both in-sample and out-of-sample evaluations.
A simulated corporate loan market and an asset-pricing analysis
of the distress risk anomaly further show the economic value of
accounting for industry-specific default risk.
Cost Restricted Feature Selection with Information Criterion
with Ben Sherwood and Shaobo Li
Accepted for presentation at the 2026 INFORMS Workshop on Data Science
Preparing for submission to Management Science
Abstract:
We study feature selection for data acquisition when candidate
variables are costly and must be selected under an acquisition
budget. Existing cost-restricted methods typically focus on
predictive loss subject to the budget, but they do not explicitly
account for the increase in model complexity from acquiring
additional variables. We propose a cost-restricted information
criterion approach, CRIC, that preserves the acquisition budget as
a feasibility constraint while adding an AIC- or BIC-type complexity
charge for each selected feature. This allows the method to reject
feasible variables whose incremental contribution is too small to
justify their inclusion, rather than automatically exhausting the
available budget. We develop computational procedures for Gaussian
and binary-response settings and evaluate the approach through
extensive simulations and empirical applications. The results show
that CRIC can substantially reduce unnecessary feature acquisition
and select smaller models while preserving competitive out-of-sample
predictive performance, with particularly clear gains when the
acquisition budget is relatively loose.
Robust Outcome Optimization via Multi-Model Inverse Classification
with Michael Lash, Qihang Lin, and Nick Street
Target journal: Journal of Machine Learning Research
Abstract:
This paper develops a multi-model inverse classification framework
for robust outcome optimization. Existing inverse classification
methods generate intervention recommendations with respect to a
single predictive model, even though practitioners often face a
collection of plausible models that can disagree about the effect
of an intervention. We instead optimize recommendations over the
distribution of predictions generated by multiple models and allow
the decision maker to choose the desired level of risk. The
framework accounts for intervention costs, feasibility constraints,
and indirect changes in non-actionable features, while the
optimization formulation targets recommendations that remain
effective across model variation rather than being tailored to a
single fitted model. In an application to cardiovascular risk
reduction, the method produces practical recommendations that
remain effective across alternative predictive specifications,
illustrating how multi-model inverse classification can make
prescriptive analytics more robust to variation in the underlying
prediction model.
Objective-Aligned Sparse Learning for Neyman-Pearson Classification
Manuscript in preparation
Abstract:
This paper develops an objective-aligned framework for asymmetric
classification in which score estimation and variable selection are
learned specifically to minimize false negatives subject to a
prespecified false positive rate constraint. Standard Neyman-Pearson
procedures typically estimate a scoring rule under a symmetric
objective and then use threshold calibration to impose asymmetric
error control. Our method instead incorporates the deployment
objective directly into regularized score learning, while retaining
a separate calibration step for final false positive rate control.
This allows the selected variables and score direction to reflect
the operating point that matters for the decision problem.
Simulation studies consistently show lower false negative rates
than standard and penalized Neyman-Pearson classifiers, especially
when the variables that matter under the asymmetric objective differ
from those selected under a symmetric loss. In a bankruptcy
application using 18,651 U.S. publicly traded firms from 1980 to
2023, the method reduces the false negative rate from 48.9% to 35.1%
at a 5% false positive rate target, a 28% reduction in missed
bankruptcies, and raises recall from 51.1% to 64.9%.
Dynamic Instrumented Principal Components for Time-Varying Risk Exposures
Manuscript in preparation
Abstract:
This paper develops a dynamic extension of instrumented principal
components analysis (IPCA) that allows the mapping from firm
characteristics to latent factor loadings to evolve over time.
Standard IPCA permits characteristics to vary across firms and time
but keeps the characteristic-to-loading map fixed. Our model allows
this relationship itself to change, capturing persistent shifts in
how firm characteristics translate into systematic risk exposures
while preserving the characteristic-based factor structure and
interpretability of IPCA. Empirically, the dynamic specification
substantially improves in-sample fit across latent factor dimensions.
With appropriate control of dynamic variation, it also produces
sizable out-of-sample gains in return and managed-portfolio
prediction relative to static IPCA. The results indicate that
allowing the characteristic-risk relationship to evolve improves
the description of realized systematic variation and the prediction
of expected returns.
Publications
Relationship between piperacillin concentrations, clinical factors
and piperacillin/tazobactam-associated acute kidney injury
Tang Girdwood, S., Hasson, D., Caldwell, J. T., Slagle, C.,
Dong, S., Fei, L., Tang, P., Vinks, A. A.,
Kaplan, J., & Goldstein, S. L.
Journal of Antimicrobial Chemotherapy,
78(2), 478–487, 2023.
DOI
Diagnostic accuracy of serum matrix metalloproteinase-7 as a
biomarker of biliary atresia in a large North American cohort
Pandurangi, S., Mourya, R., Nalluri, S., Fei, L.,
Dong, S., Harpavat, S., Guthery, S. L.,
Molleston, J. P., Rosenthal, P., Sokol, R. J., Wang, K. S.,
Ng, V., Alonso, E. M., Hsu, E. K., Karpen, S. J.,
Loomes, K. M., Magee, J. C., Shneider, B. L., Horslen, S. P.,
Teckman, J. H., & Bezerra, J. A.
Hepatology, 2024.
DOI
Maternal metabolic dysfunction-associated steatotic liver disease
during pregnancy and its impact on offspring steatosis
Klepper, C. M., Trout, A. T., Dong, S.,
Fei, L., Stankiewicz, T. E., Wayland, J.,
Moreno-Fernandez, M. E., Divanovic, S., Xanthakos, S. A.,
DeFranco, E., Woo, J. G., & Mouzaki, M.
Clinical Nutrition ESPEN,
70, 327–335, 2025.
DOI