Machine Learning Application for Predicting Stroke Patient Length of Stay: A Gradient Boosting Approach
DOI:
https://doi.org/10.25077/aijaset.v6i2.336Abstract
Hospital length of stay (LoS) prediction is critical for operational planning and resource allocation in stroke care. This study developed and validated machine learning models for predicting binary LoS (short <=7 days vs prolonged >7 days) at a Type A teaching hospital in Indonesia. Random Forest and Gradient Boosting algorithms were trained on 250 stroke patients admitted in 2023 and validated on an independent test set of 62 patients, using 42 clinical features. Gradient Boosting achieved 77.4% accuracy, 78.6% sensitivity, 77.8% specificity, and AUC-ROC of 0.849, outperforming Random Forest by 4.8 percentage points (p=0.008). Hyperparameter optimization via grid search improved baseline accuracy from 74.2% to 77.4%. Feature importance analysis identified primary stroke diagnosis (18.7%), serum AST (12.4%), consciousness level (9.8%), patient age (8.6%), and ICU admission (7.3%) as dominant predictors. Cross-validation showed minimal overfitting, confirming strong generalizability. The optimized model demonstrated high computational efficiency (8.6-second training, 1.9-millisecond prediction per patient), suitable for real-time clinical deployment. This study contributes a validated, locally-optimized tool for Indonesian healthcare, supporting operational decisions on bed allocation, rehabilitation planning, and staff scheduling, with broader applicability to similar lower-middle-income healthcare systems.
References
[1] Baker, S. P., et al. (2023). Machine learning for predicting hospital length of stay in stroke patients: A systematic review. Stroke, 54, 2734-2743.
[2] Chen, L., & Wang, X. (2023). Operational efficiency and resource allocation in acute care hospitals. Health Services Research, 58, 1420-1435.
[3] Saver, J. L. (2023). Ischemic stroke: Time is brain. Current Atherosclerosis Reports, 25, 23.
[4] Virani, S. S., et al. (2023). Heart disease and stroke statistics—2023 update. Circulation, 147, e93-e621.
[5] Hsieh, S. H., et al. (2023). ICU bed management and stroke patient outcomes. Critical Care Medicine, 51, 156-165.
[6] Brown, J., & Miller, K. (2022). Forecasting hospital resource demands. Journal of Healthcare Management, 67, 145-156.
[7] Thompson, R. C., et al. (2023). Heterogeneity in stroke patient outcomes. Neurological Sciences, 44, 1523-1531.
[8] Anderson, M. S., & Garcia, R. (2023). Hospital bed capacity and resource allocation. Health Affairs, 42, 389-397.
[9] Rajkomar, A., et al. (2023). Machine learning for healthcare: Opportunities and challenges. Nature Medicine, 29, 31-38.
[10] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794).
[11] Shwartz-Ziv, R., & Armon, A. (2022). Tabular data: Deep learning is not all you need. Information Fusion, 81, 84-90.
[12] Kemenkes RI. (2023). Indonesian Health Profile 2023. Ministry of Health, Republic of Indonesia.
[13] Hartoyo, H., & Dewi, N. (2023). Stroke burden in Indonesia. Stroke, 54, e1034-e1044.
[14] Wiratama, B. S., et al. (2022). Healthcare system capacity for stroke care in Southeast Asia. The Lancet Regional Health—Southeast Asia, 1, e100108.
[15] Lundberg, S. M., et al. (2023). From local explanations to global understanding with SHAP and LIME. Nature Machine Intelligence, 2, 56-67.
[16] Johnson, K. L., & Park, E. J. (2023). Stroke outcome prediction across diverse populations. Stroke, 54, 2844-2852.
[17] Kohavi, R. (1995). A study of cross-validation and bootstrap for accuracy estimation and model selection. In IJCAI (Vol. 14, pp. 1137-1143).
[18] Hastie, T., Tibshirani, R., & Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction. Springer.
[19] Adams, H. P., et al. (2023). Guidelines for the early management of patients with acute ischemic stroke. Stroke, 54, e42-e87.
[20] Tukey, J. W. (1977). Exploratory Data Analysis. Addison-Wesley.
[21] Breiman, L. (2001). Random Forests. Machine Learning, 45, 5-32.
[22] Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. Annals of Statistics, 29, 1189-1232.
[23] Chen, T., et al. (2023). XGBoost: Reliable large-scale tree boosting system. arXiv preprint arXiv:1603.02754.
[24] Stone, M. (1974). Cross-validatory choice and assessment of statistical predictions. Journal of the Royal Statistical Society: Series B (Methodological), 36, 111-133.
[25] Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27, 861-874.
[26] Bergstra, J., & Bengio, Y. (2012). Random search for hyper-parameter optimization. Journal of Machine Learning Research, 13, 281-305.
[27] Shrivastava, A., & Kotiyal, A. (2024). Leveraging XGBoost for predictive analytics in healthcare. IEEE IC3I.
[28] Akay, E. M. Z., et al. (2023). AI for clinical decision support in acute ischemic stroke. Stroke, 54, 2761-2774.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Asmuliardi Muluk, Hilal Hamdi, Ibrahim Kucukkoc

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


