Academic Journal

Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods.
Συγγραφείς: Vey JA; Institute of Medical Biometry, University of Heidelberg, Im Neuenheimer Feld 130.3, Heidelberg, 69120, Germany. vey@imbi.uni-heidelberg.de., Heinze G; Institute of Clinical Biometrics, Center for Medical Data Science, Medical University of Vienna, Spitalgasse 23, Vienna, 1090, Austria., Kieser M; Institute of Medical Biometry, University of Heidelberg, Im Neuenheimer Feld 130.3, Heidelberg, 69120, Germany.
Πηγή: BMC medical research methodology [BMC Med Res Methodol] 2026 Jul 07; Vol. 26 (1). Date of Electronic Publication: 2026 Jul 07.
Τύπος έκδοσης: Journal Article; Comparative Study
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: BioMed Central Country of Publication: England NLM ID: 100968545 Publication Model: Electronic Cited Medium: Internet ISSN: 1471-2288 (Electronic) Linking ISSN: 14712288 NLM ISO Abbreviation: BMC Med Res Methodol Subsets: MEDLINE
Imprint Name(s): Original Publication: London : BioMed Central, [2001-
Ιατρικοί όροι (MeSH): Machine Learning* , Computer Simulation*, Humans ; Predictive Learning Models ; Prediction Algorithms ; Boosting Machine Learning Algorithms ; Linear Models ; Random Forest ; Regression Analysis ; Algorithms ; Soft Computing
Περίληψη: Purpose: A wide range of methods exist for developing a clinical prediction model (CPM) and for performing variable selection. Our purpose was to develop a fair simulation study design and to investigate the properties, strengths, and weaknesses of different methods to predict a continuous outcome in low-dimensional data situations.
Methods: In this simulation study, we conducted a neutral comparison of traditional (linear regression with stepwise selection) and machine learning (regularized regression with elastic net, gradient boosting, random forest) variable selection strategies to derive a CPM. The generated datasets included a total of 15 variables, with 8 of those being predictor variables. Four data- and outcome-generating mechanisms with increasing complexity produced data structures typical for biomedicine covering linear associations and gradually introducing non-linear and non-additive elements into the data structure.
Results: All methods generally performed better with increasing sample size and less noise in the data. Gradient boosting with regression models and with trees as base learners, and the elastic net regularized regression included nearly all variables (i.e., both the predictor and non-predictor variables), especially with increasing sample size. The linear regression model with stepwise selection (LMSS) showed the best trade-off between correctly including the predictors and excluding the non-predictor variables in most of the scenarios, even when the functional form of continuous predictors deviated from linearity. In more complex data, variable selection using the Boruta or Hapfelmeier approach for random forest performed similar to LMSS.
Conclusion: The sample size must be sufficiently large to enable the methods to reliably identify the predictor variables and to ensure that the developed CPMs are accurate and well-calibrated. LMSS revealed good properties and the random forest with the Boruta or Hapfelmeier approach are suitable alternatives if complex associations between predictors and outcomes are assumed.
(© 2026. The Author(s).)
Competing Interests: Declarations. Ethics approval and consent to participate: Not applicable. Consent for publication: Not applicable. Competing interests: The authors declare no competing interests.
References: Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD). Circulation. 2015;131(2):211–9. https://doi.org/10.1161/CIRCULATIONAHA.114.014508 . (PMID: 10.1161/CIRCULATIONAHA.114.014508255615164297220)
Hippisley-Cox J, Coupland C, Brindle P. Development and validation of QRISK3 risk prediction algorithms to estimate future risk of cardiovascular disease: prospective cohort study. BMJ. 2017;357. https://doi.org/10.1136/bmj.j2099 .
Subramanian J, Simon R. Overfitting in prediction models – Is it a problem only in high dimensions? Contemp Clin Trials. 2013;36(2):636–41. https://doi.org/10.1016/j.cct.2013.06.011 . (PMID: 10.1016/j.cct.2013.06.01123811117)
Harrell FE Jr, Lee KL, Mark DB. Multivariable prognostic models: Issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors. Stat Med. 1996;15(4):361–87. https://doi.org/10.1002/(SICI)1097-0258(19960229)15:4%253C361::AID-SIM168%3E3.0.CO;2-4 .
Riley RD, Ensor J, Snell KIE, Harrell FE, Martin GP, Reitsma JB, et al. Calculating the sample size required for developing a clinical prediction model. BMJ. 2020;368:m441. https://doi.org/10.1136/bmj.m441 . (PMID: 10.1136/bmj.m44132188600)
Tibshirani R. Regression shrinkage and selection via the lasso. J Roy Stat Soc: Ser B (Methodol). 1996;58(1):267–88. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x . (PMID: 10.1111/j.2517-6161.1996.tb02080.x)
Zou H, Hastie T. Regularization and variable selection via the elastic net. J R Stat Soc Ser B (Stat Methodol). 2005;67(2):301–20. (PMID: 10.1111/j.1467-9868.2005.00503.x)
Breiman L. Random Forests. Mach Learn. 2001;45:5–32. https://doi.org/10.1023/A:1010933404324 . (PMID: 10.1023/A:1010933404324)
Friedman J, Hastie T, Tibshirani R. Additive logistic regression: a statistical view of boosting. Ann Statist. 2000;28(2):337–407. https://doi.org/10.1214/aos/1016218223 . With discussion and a rejoinder by the authors.
Friedman JH. Greedy function approximation: a gradient boosting machine. Ann Statist. 2001;29(5):1189–232. https://doi.org/10.1214/aos/1013203451 . (PMID: 10.1214/aos/1013203451)
Bühlmann P, Yu B. Boosting with the L2 loss. J Am Stat Assoc. 2003;98(462):324–39. https://doi.org/10.1198/016214503000125 . (PMID: 10.1198/016214503000125)
Boulesteix AL, Wilson R, Hapfelmeier A. Towards evidence-based computational statistics: Lessons from clinical research on the role and design of real-data benchmark studies. BMC Med Res Methodol. 2017;17:1–12. https://doi.org/10.1186/S12874-017-0417-2 . (PMID: 10.1186/S12874-017-0417-2)
R Core Team. R: A language and environment for statistical computing. Vienna, Austria; 2025. https://www.R-project.org/ . Accessed date 24 Aug 2025.
Heinze G, Wallisch C, Dunkler D. Variable selection - A review and recommendations for the practicing statistician. Biom J. 2018;60:431–49. https://doi.org/10.1002/bimj.201700067 . (PMID: 10.1002/bimj.201700067292925335969114)
Hoerl AE, Kennard RW. Ridge Regression: Biased Estimation for Nonorthogonal Problems. Technometrics. 1970;12(1):55–67. https://doi.org/10.1080/00401706.1970.10488634 . (PMID: 10.1080/00401706.1970.10488634)
Friedman J, Tibshirani R, Hastie T. Regularization paths for generalized linear models via coordinate descent. J Stat Softw. 2010;33(1):1–22. https://doi.org/10.18637/jss.v033.i01 . (PMID: 10.18637/jss.v033.i01208087282929880)
Box GEP, Tidwell PW. Transformation of the independent variables. Technometrics. 1962;4(4):531–50. (PMID: 10.1080/00401706.1962.10490038)
Royston P, Sauerbrei W. MFP: Multivariable model-building with fractional polynomials. In: Multivariable Model-Building. Wiley; 2008. p. 115–50. https://doi.org/10.1002/9780470770771.ch6 .
Mayr A, Hofner B, Schmid M. The importance of knowing when to stop: A sequential stopping rule for component-wise gradient boosting. Methods Inf Med. 2012;51:178–86. https://doi.org/10.3414/ME11-02-0030 . (PMID: 10.3414/ME11-02-003022344292)
Mayr A, Hofner B. Boosting for statistical modelling: A non-technical introduction. Stat Model. 2018;18:365–84. https://doi.org/10.1177/1471082X17748086 . (PMID: 10.1177/1471082X17748086)
Hapfelmeier A, Ulm K. A new variable selection approach using random forests. Comput Stat Data Anal. 2013;60:50–69. https://doi.org/10.1016/J.CSDA.2012.09.020 . (PMID: 10.1016/J.CSDA.2012.09.020)
Kursa MB, Rudnicki WR. Feature selection with the Boruta package. 2010. http://www.jstatsoft.org/ . Accessed date 24 Aug 2025.
Rothacher Y, Strobl C. Identifying informative predictor variables with random forests. J Educ Behav Stat. 2023. https://doi.org/10.3102/10769986231193327 .
O’Connell NS, Jaeger BC, Bullock GS, Speiser JL. A comparison of random forest variable selection methods for regression modeling of continuous outcomes. Brief Bioinforma. 2025;26. https://doi.org/10.1093/BIB/BBAF096.
Vey JA. Simulation study protocol. 2025. OSF, 23 Jan. 2025. https://doi.org/10.17605/OSF.IO/8DR3K.
Morris TP, White IR, Crowther MJ. Using simulation studies to evaluate statistical methods. Stat Med. 2019;38(11):2074–102. https://doi.org/10.1002/sim.8086 . (PMID: 10.1002/sim.8086306523566492164)
Binder H, Sauerbrei W, Royston P. Multivariable model-building with continuous covariates: 1. Performance measures and simulation design. 2011. https://www.fdm.uni-freiburg.de/publications-preprints/papers/pre105 . Accessed date 24 Aug 2025.
Schumacher M, Bastert G, Bojar H, Hübner K, Olschewski M, Sauerbrei W. Randomized 2 x 2 trial evaluating hormonal treatment and the duration of chemotherapy in node-positive breast cancer patients. German Breast Cancer Study Group. J Clin Oncol Offi J Am Soc Clin Oncol. 1994;12:2086–93. https://doi.org/10.1200/JCO.1994.12.10.2086 . (PMID: 10.1200/JCO.1994.12.10.2086)
Royston P, Sauerbrei W. Handling categorical and continuous predictors. In: Multivariable Model-Building. Wiley; 2008. p. 53–70. https://doi.org/10.1002/9780470770771.ch3 .
Wünsch M, Herrmann M, Noltenius E, Mohr M, Morris TP, Boulesteix AL. Rethinking the Handling of Method Failure in Comparison Studies. Stat Med. 2025;44:e70257. https://doi.org/10.1002/SIM.70257 . (PMID: 10.1002/SIM.702574106533912509789)
Calster BV, van Smeden M, Cock BD, Steyerberg EW. Regression shrinkage methods for clinical prediction models do not guarantee improved performance: Simulation study. Stat Methods Med Res. 2020;29:3166–78. https://doi.org/10.1177/0962280220921415 . (PMID: 10.1177/096228022092141532401702)
Fox J, Weisberg S. Visualizing fit and lack of fit in complex regression models with predictor effect plots and partial residuals. J Stat Softw. 2018;87:1–27. https://doi.org/10.18637/JSS.V087.I09 . (PMID: 10.18637/JSS.V087.I09)
Dehne S, Riede C, Feisst M, Klotz R, Etheredge M, Hölle T, et al. Tranexamic acid administration during liver transplantation is not associated with lower blood loss or with reduced utilization of red blood cell transfusion. Anesth Analg. 2024;139:598–608. https://doi.org/10.1213/ANE.0000000000006804 . (PMID: 10.1213/ANE.000000000000680438236761)
Meinshausen N, Bühlmann P. High-dimensional graphs and variable selection with the lasso. Ann Stat. 2006;34(3):1436–62. (PMID: 10.1214/009053606000000281)
Meinshausen N, Bühlmann P. Stability selection. J R Stat Soc Ser B Stat Methodol. 2010;72(4):417–73. https://doi.org/10.1111/j.1467-9868.2010.00740.x . (PMID: 10.1111/j.1467-9868.2010.00740.x)
Hofner B, Boccuto L, Göker M. Controlling false discoveries in high-dimensional situations: Boosting with stability selection. BMC Bioinformatics. 2015;16:1–17. https://doi.org/10.1186/S12859-015-0575-3 . (PMID: 10.1186/S12859-015-0575-3)
Ploeg TVD, Austin PC, Steyerberg EW. Modern modelling techniques are data hungry: A simulation study for predicting dichotomous endpoints. BMC Med Res Methodol. 2014;14:1–13. https://doi.org/10.1186/1471-2288-14-137 . (PMID: 10.1186/1471-2288-14-137)
Riley RD, Snell KIE, Ensor J, Burke DL, Harrell FE, Moons KGM, et al. Minimum sample size for developing a multivariable prediction model: Part I - Continuous outcomes. Stat Med. 2019;38:1262–75. https://doi.org/10.1002/SIM.7993 . (PMID: 10.1002/SIM.799330347470)
Barreñada L, Dhiman P, Timmerman D, Boulesteix AL, Calster BV. Understanding overfitting in random forest for probability estimation: A visualization and simulation study. Diagnostic and Prognostic Research 2024 8:1, vol. 8. 2024. p. 1–14. https://doi.org/10.1186/S41512-024-00177-1 .
Boulesteix AL, Binder H, Abrahamowicz M, Sauerbrei W. On the necessity and design of studies comparing statistical methods. Biom J. 2018;60:216–8. https://doi.org/10.1002/BIMJ.201700129 . (PMID: 10.1002/BIMJ.20170012929193206)
Riley RD, Snell KIE, Martin GP, Whittle R, Archer L, Sperrin M, et al. Penalization and shrinkage methods produced unreliable clinical prediction models especially when sample size was small. J Clin Epidemiol. 2021;132:88–96. https://doi.org/10.1016/j.jclinepi.2020.12.005 . (PMID: 10.1016/j.jclinepi.2020.12.00533307188)
Šinkovec H, Heinze G, Blagus R, Geroldinger A. To tune or not to tune, a case study of ridge logistic regression in small or sparse datasets. BMC Med Res Methodol. 2021;21:1–15. https://doi.org/10.1186/S12874-021-01374-Y . (PMID: 10.1186/S12874-021-01374-Y)
Riley RD, Collins GS, Riley CDR. Stability of clinical prediction models developed using statistical or machine learning methods. Biom J. 2023;2200302. https://doi.org/10.1002/BIMJ.202200302 .
Contributed Indexing: Keywords: Clinical prediction models; Machine learning; Neutral comparison study; Regression models; Variable selection
Entry Date(s): Date Created: 20260707 Date Completed: 20260708 Latest Revision: 20260726
Update Code: 20260726
PubMed Central ID: PMC13340120
DOI: 10.1186/s12874-026-02930-0
PMID: 42414897
Βάση Δεδομένων: MEDLINE