Comparative Performance of Gradient Boosting Algorithms for Household Food Resilience Classification during The COVID-19
DOI:
https://doi.org/10.24036/eksakta/vol27-iss04/667Keywords:
COVID-19, Food Resilience, Gradient Boosting, Machine Learning, SHAPAbstract
The need for data-driven analysis to understand socio-economic vulnerability, particularly in relation to food security, has intensified due to the pandemic. The growing volume of survey data demands analytical methods that can capture multidimensional relationships. Over time, food security evolves into the concept of food resilience, reflecting a household's capacity to withstand or recover from adverse conditions. This study uses machine learning techniques to categorize household food resilience, based on data from the World Bank's High Frequency Phone Survey (HFPS), covered 2,868 households in Indonesia. A comparative evaluation of gradient boosting algorithms (XGBoost, LightGBM and CatBoost) was conducted. Model performance was evaluated using accuracy, sensitivity, specificity, and AUC across ten repeated train-test splits, with statistical significance assessed using Friedman and Wilcoxon tests. The results show that CatBoost performed best and most consistently, achieving mean accuracy of 0.7592 and mean AUC of 0.8331, which is significantly higher than that of competing models. SHAP analysis further indicates that baseline vulnerability, financial concerns, income capacity, food price shocks and unmet healthcare needs are important features for identifying resilient households. These findings demonstrate that gradient boosting, particularly CatBoost, provides strong predictive power and interpretability to support data-driven decision-making in classifying food resilience.
Downloads
References
[1] KC, U., Campbell-Ross, H., Godde, C., Friedman, R., Lim-Camacho, L., & Crimp, S. (2024). A systematic review of the evolution of food system resilience assessment. Global Food Security, 40, 100744.
[2] Yenerall, J., & Jensen, K. (2022). Food security, financial resources, and mental health: Evidence during the COVID-19 pandemic. Nutrients, 14(1), 161.
[3] Wambede, N. M., Sharon, A., JoyFred, A., Mulabbi, A., & Robert, T. (2025). Adaptations to climate variability in Northern Uganda: Implications for food security. Journal of Tropical Crop Science, 12(2), 261–273.
[4] Rahardiantoro, S., & Sakamoto, W. (2021). Clustering regions based on socio-economic factors which affected the number of COVID-19 cases in Java Island. Journal of Physics: Conference Series, 1863(1), 012014.
[5] Satria, A., Anggraini, E., Widyastutik, Nuryartono, N., Amaliah, S., Helmi, A., et al. (2022). Membangun resiliensi sistem pangan Indonesia. Policy Brief Pertanian, Kelautan Dan Biosains Tropika, 4(4), 338–345.
[6] Oktarina, S. D., Nurkhoiry, R., Amalia, R., Pradikto, I., & Rahutomo, S. (2022). Dampak ketidakpastian COVID-19, iklim, dan kompleksitas lainnya pada industri kelapa sawit. Warta PPKS, 27(2), 70–77.
[7] BPS-Statistics Indonesia. (2025). Prevalence of undernourishment. Jakarta: BPS.
[8] Akbar, A., Darma, R., Fahmid, I. M., & Irawan, A. (2023). Determinants of household food security during the COVID-19 pandemic in Indonesia. Sustainability, 15(5), 4131.
[9] Hassan, R., Di Martino, M., & Daher, B. (2025). Quantifying sustainability and resilience in food systems: A systematic analysis for evaluating the convergence of current methodologies and metrics. Frontiers in Sustainable Food Systems, 9, 1479691.
[10] Villacis, A. H., Badruddoza, S., & Mishra, A. K. (2024). A machine learning-based exploration of resilience and food security. Applied Economic Perspectives and Policy, 46(4), 1479–1505.
[11] Ronalia, P., Hartono, D., & Misdawita, M. (2023). The impact of resilience on household food insecurity in Indonesia. Jurnal Ekonomi Pembangunan, 21(1), 25–38.
[12] Machefer, M., Thomas, A. C., Meroni, M., Veiga Lopez Pena, J. M., Ronco, M., Corbane, C., et al. (2025). Potential and limitations of machine learning modeling for forecasting acute food insecurity. Global Food Security, 45, 100859.
[13] Aguilera, T., & Jatmiko, Y. A. (2023). Pemodelan tingkat kerawanan pangan rumah tangga di Indonesia tahun 2021 dengan pendekatan regresi logistik ordinal. Indonesian Journal of Applied Statistics, 5(2), 90.
[14] Rahardiantoro, S., Juhanda, A. R. N., Kurnia, A., Aswi, A., Sartono, B., Handayani, D., et al. (2024). Spatio-temporal modeling to identify factors associated with stunting in Indonesia using a modified generalized Lasso. Spatial and Spatio-temporal Epidemiology, 51, 100694.
[15] Hartono, A., Dewi, L. A., Yuniarti, E., Putri, S. T. H., & Harahap, T. S. (2024). Machine learning classification for detecting heart disease with K-NN algorithm, decision tree and random forest. Eksakta: Berkala Ilmiah Bidang MIPA, 24(4), 513–522.
[16] Kyriazos, T., & Poga, M. (2024). Application of machine learning models in social sciences: Managing nonlinear relationships. Encyclopedia, 4(4), 1790–1805.
[17] Kern, C., Klausch, T., & Kreuter, F. (2019). Tree-based machine learning methods for survey research. https://doi.org/10.31235/osf.io/7w3sc .
[18] Dharmawan, H., Sartono, B., Kurnia, A., Hadi, A. F., & Ramadhani, E. (2022). A study of machine learning algorithms to measure the feature importance in class-imbalance data of food insecurity cases in Indonesia. Communications in Mathematical Biology and Neuroscience, 2022, 1-15.
[19] Khikmah, K. N., Sartono, B., Susetyo, B., & Dito, G. A. (2024). Performance comparative study of machine learning classification algorithms for food insecurity experience by households in West Java. Jurnal Online Informatika, 9(1), 128–137.
[20] Fransiska, H., Soleh, A. M., Notodiputro, K. A., & Erfiani. (2025). Evaluation of machine learning models based on household food insecurity data in Indonesia. BIO Web of Conferences, 171, 02011.
[21] Abd El-Ghani, S. S., Mansour, T. G. I., & Esleem, S. A. (2025). The most important economic and social indicators of the challenges facing food security for the most important crops in Egypt. Environmental and Sustainability Indicators, 27, 100808.
[22] Yuliani, E., Sartono, B., Wijayanto, H., Hadi, A. F., & Ramadhani, E. (2022). Study of features importance level identification of machine learning classification model in sub-populations for food insecurity. AIP Conference Proceedings, 2668(1), 050012.
[23] Borys, K., Schmitt, Y. A., Nauta, M., Seifert, C., Krämer, N., Friedrich, C. M., et al. (2023). Explainable AI in medical imaging: An overview for clinical practitioners - Saliency-based XAI approaches. European Journal of Radiology, 162, 110787.
[24] Oktora, S. I., Matualage, D., Notodiputro, K. A., & Sartono, B. (2025). Data-driven insights into underdeveloped regencies: SHAP-based explainable artificial intelligence approach. International Journal of Artificial Intelligence Research, 9(1), 1-12
[25] Saranya, A., & Subhashini, R. (2023). A systematic review of explainable artificial intelligence models and applications: Recent developments and future trends. Decision Analytics Journal, 7, 100230.
[26] Rane, N. L., Mallick, S. K., & Rane, J. (2025). Artificial intelligence and machine learning for enhancing resilience: Concepts, applications, and future directions. Deep Science Publishing.
[27] World Bank. (2023). Indonesia: High-frequency monitoring of COVID-19 impacts rounds 1-8, 2020-2023 (HIFY 2020-2023). World Bank Open Data. https://doi.org/10.48529/VC74-5V86
[28] Letta, M., Montalbano, P., Morales Opazo, C., & Petruccelli, F. (2025). Measuring and testing vulnerability to food insecurity for prediction and targeting. Economics & Human Biology, 59, 101552.
[29] Jacobsen, E., Ran, X., Liu, A., Chang, C. C. H., & Ganguli, M. (2021). Predictors of attrition in a longitudinal population-based study of aging. International Psychogeriatrics, 33(8), 767–778.
[30] Joel, L.O., Doorsamy, W. & Paul, B.S. (2025). A comparative study of imputation techniques for missing values in healthcare diagnostic datasets. Int J Data Sci Anal , 20, 6357–6373.
[31] Katya, E. (2023). Exploring feature engineering strategies for improving predictive models in data science. Research Journal of Computer Systems and Engineering, 4(2), 201–215.
[32] Safrizal, A. N., & Agustian, S. (2024). Classification of COVID-19 vaccine sentiment using K-Nearest Neighbor and Fasttext on Twitter. Eksakta: Berkala Ilmiah Bidang MIPA, 25(3), 362–371.
[33] Rahman, I. A., Siregar, P. S. G., & Suharsih, S. (2024). Analysis of the potential and volatility of large chili and cayenne chili after COVID-19 in Yogyakarta City. Journal of International Conference Proceedings, 6(6), 497–508.
[34] Cascarano, A., Mur-Petit, J., Hernández-González, J., Camacho, M., de Toro Eadie, N., Gkontra, P., et al. (2023). Machine and deep learning for longitudinal biomedical data: A review of methods and applications. Artificial Intelligence Review, 56(2), 1711–1771.
[35] Bentéjac, C., Csörgő, A., & Martínez-Muñoz, G. (2021). A comparative analysis of gradient boosting algorithms. Artificial Intelligence Review, 54(3), 1937–1967.
[36] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794.
[37] Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., et al. (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30, 3146–3154.
[38] Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., & Gulin, A. (2018). CatBoost: Unbiased boosting with categorical features. Advances in Neural Information Processing Systems, 31, 6638–6648.
[39] Syamkalla M., Khomsah S., Nur Y., (2024). Implementasi Algoritma Catboost Dan Shapley Additive Explanations (SHAP) Dalam Memprediksi Popularitas Game Indie Pada Platform Steam. Jurnal Teknologi Informasi dan Ilmu Komputer, 11(4), 777-786.
[40] Putri, N., Parmikanti, K., & Gusriani, N. (2024). Robust linear discriminant analysis with modified one-step M-estimator Qn scale for classifying financial distress in banks: Case study. Eksakta: Berkala Ilmiah Bidang MIPA, 25(2), 219–230.
[41] Ratnaningayu, N. D., Tedjo, A., & Panigoro, S. S. (2024). The implementation of machine learning algorithms for breast cancer biomarker validation in metabolomics studies. Eksakta: Berkala Ilmiah Bidang MIPA, 25(4), 468–483.
[42] Yunita, A., Pratama, M. I., Almuzakki, M. Z., Ramadhan, H., Akhir, E. A. P., Mansur, A. B. F., et al. (2025). Performance analysis of neural network architectures for time series forecasting: A comparative study of RNN, LSTM, GRU, and hybrid models. MethodsX, 15, 103462.
[43] Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14(1), 6086.
[44] Lundberg, S. M., Erion, G., Chen, H., DeGrave, A., Prutkin, J. M., Nair, B., et al. (2020). From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2(1), 56–67.
[45] Chowdhury, T. R., & Dey, S. (2024). A comparative survey of SHAP and LIME: Explaining machine learning models for transparent AI. International Journal of Innovative Research in Technology, 11(1), 827–835.
[46] Gomes Mantovani, R., Horváth, T., Rossi, A. L. D., Cerri, R., Barbon Junior, S., Vanschoren, J., et al. (2024). Better trees: An empirical study on hyperparameter tuning of classification decision tree induction algorithms. Data Mining and Knowledge Discovery, 38(3), 1364–1416.
[47] Nafis, M. A. N., Susilaningrum, D., Ulama, B., & Kusrini, D. E. (2024). Assessing the impact of household socioeconomic factors on clean and healthy living behaviors with binary logistic regression: A study in Probolinggo Regency. Eksakta: Berkala Ilmiah Bidang MIPA, 25(4), 436–447.
[48] Tang, M., Peng, Z., & Wu, H. (2021). Fault detection for pitch system of wind turbine-driven doubly fed based on IHHO-LightGBM. Applied Sciences, 11(17), 8030.
[49] Ilemobayo, J. A., Durodola, O., Alade, O., Awotunde, O. J., Olanrewaju, A. T., Falana, O., et al. (2024). Hyperparameter tuning in machine learning: A comprehensive review. Journal of Engineering Research and Reports, 26(6), 388–395.
[50] Raparthi, M., Dhabliya, D., Kumari, T., Upadhyaya, R., & Sharma, A. (2024). Implementation and performance comparison of gradient boosting algorithms for tabular data classification. In International Conference on Artificial Intelligence and Soft Computing (pp. 461–479).
[51] Hancock, J. T., & Khoshgoftaar, T. M. (2020). CatBoost for big data: An interdisciplinary review. Journal of Big Data, 7(1), 94.
[52] de Hond, A. A. H., Steyerberg, E. W., & van Calster, B. (2022). Interpreting area under the receiver operating characteristic curve. The Lancet Digital Health, 4(12), e853–e855.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Nabila Syukri, Sachnaz Desta Oktarina, Septian Rahardiantoro

This work is licensed under a Creative Commons Attribution 4.0 International License.
This is an open-access article distributed under the terms of the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (http://creativecommons.org/licenses/by/4.0/) which permits unrestricted non-commercial use, distribution and reproduction in any medium
























