Differentiated Thyroid Cancer Recurrence Classification Using Machine Learning Models and Bayesian Neural Networks with Varying Priors: A SHAP-Based Interpretation of the Best Performing Model

Authors

Keywords:

Bayesian Neural Network, Classification, Differentiated thyroid cancer recurrence, Machine Learning, SHAP, Uncertainty quantification

Abstract

Differentiated thyroid cancer (DTC) recurrence is a major public health concern, requiring classification and predictive models that are not only accurate but also interpretable and uncertainty aware. This study introduces a comprehensive framework for DTC recurrence classification using a dataset containing 383 patients and 16 clinical and pathological variables. Initially, 11 machine learning (ML) models were employed using the complete dataset, where the Support Vector Machines (SVM) model achieved the highest accuracy of 0.9481. To reduce complexity and redundancy, feature selection was carried out using the Boruta algorithm, and the same ML models were applied to the reduced dataset, where it was observed that the Logistic Regression (LR) model obtained the maximum accuracy of 0.9611. However, these ML models often lack uncertainty quantification, which is critical in clinical decision making. Therefore, to address this limitation, the Bayesian Neural Networks (BNN) with six varying prior distributions, including Normal (0,1), Normal (0,10), Laplace (0,1), Cauchy (0,1), Cauchy (0,2.5), and Horseshoe (1), were implemented on both the complete and reduced datasets. The BNN model with Normal (0,10) prior distribution exhibited maximum accuracies of 0.9740 and 0.9870 before and after feature selection, respectively. As the BNN model with N(0,10) prior distribution employed after feature selection outperformed all the other models, it was chosen as the best performing model for DTC recurrence classification. This model was further analysed using epistemic and aleatoric uncertainty, reflecting the model’s confidence in its prediction. In addition, to enhance this model’s interpretability, SHapley Additive exPlanations (SHAP) values were calculated, providing valuable insights into the contribution of key variables to the model’s output.

References

[1] S. Mitra et al., “Prospective multifunctional roles and pharmacological potential of dietary flavonoid narirutin,” Biomedicine & Pharmacotherapy, vol. 150, p. 112932, Jun. 2022, doi: 10.1016/j.biopha.2022.112932.

[2] D. Das, A. Banerjee, A. B. Jena, A. K. Duttaroy, and S. Pathak, “Essentiality, relevance, and efficacy of adjuvant/combinational therapy in the management of thyroid dysfunctions,” Biomedicine & Pharmacotherapy, vol. 146, p. 112613, Feb. 2022, doi: 10.1016/j.biopha.2022.112613.

[3] S. Bhattacharya, R. K. Mahato, S. Singh, G. K. Bhatti, S. S. Mastana, and J. S. Bhatti, “Advances and challenges in thyroid cancer: The interplay of genetic modulators, targeted therapies, and AI-driven approaches,” Life Sciences, vol. 332, p. 122110, Nov. 2023, doi: 10.1016/j.lfs.2023.122110.

[4] S. Rajoria et al., “Estrogen activity as a preventive and therapeutic target in thyroid cancer,” Biomedicine & Pharmacotherapy, vol. 66, no. 2, pp. 151–158, Mar. 2012, doi: 10.1016/j.biopha.2011.11.010.

[5] N. Avenia et al., “Thyroid cancer invading the airway: diagnosis and management,” International Journal of Surgery (London, England), vol. 28 Suppl 1, pp. S75-78, Apr. 2016, doi: 10.1016/j.ijsu.2015.12.036.

[6] A. Greco, C. Miranda, M. G. Borrello, and M. A. Pierotti, “Thyroid Cancer,” Cancer Genomics, pp. 265–280, 2014, doi: 10.1016/b978-0-12-396967-5.00016-5.

[7] S. L. Gillanders and J. P. O’Neill, “Prognostic markers in well differentiated papillary and follicular thyroid cancer (WDTC),” European Journal of Surgical Oncology, vol. 44, no. 3, pp. 286–296, Mar. 2018, doi: 10.1016/j.ejso.2017.07.013.

[8] Y. Ito and A. Miyauchi, “Prognostic factors of papillary and follicular carcinomas based on pre-, intra-, and postoperative findings,” European Thyroid Journal, Aug. 2024, doi: 10.1530/etj-24-0196.

[9] J. Sastre Marcos et al., “Differentiated thyroid carcinoma: survival and prognostic factors,” Endocrinología y Nutrición (English Edition), vol. 58, no. 4, pp. 157–162, Jan. 2011, doi: 10.1016/s2173-5093(11)70040-x.

[10] A. R. Shaha, “Recurrent differentiated thyroid cancer,” Endocrine Practice, vol. 18, no. 4, pp. 600–603, Jul. 2012. doi: 10.4158/ep12047.co

[11] D. Iqbal and A. Shahzad, “Prediction of thyroid cancer recurrence with machine learning models,” Pakistan Journal of Nuclear Medicine, p. 1, 2024, doi: 10.24911/pjnmed.175-1721068107.

[12] K. Setiawan, “Predicting recurrence in differentiated thyroid cancer: a comparative analysis of various machine learning models including ensemble methods with chi-squared feature selection,” Communications in Mathematical Biology and Neuroscience, 2024, doi: 10.28919/cmbn/8506.

[13] E. Onah, U. J. Eze, A. S. Abdulraheem, U. G. Ezigbo, K. C. Amorha, and F. Ntie-Kang, “Optimizing unsupervised feature engineering and classification pipelines for differentiated thyroid cancer recurrence prediction,” BMC Medical Informatics and Decision Making, vol. 25, no. 1, May 2025, doi: 10.1186/s12911-025-03018-3.

[14] I. Juliana, K. Meylda, and Suharjito Suharjito, “Enhancing Predictive Accuracy for Differentiated Thyroid Cancer (DTC) Recurrence Through Advanced Data Mining Techniques,” TIN Terapan Informatika Nusantara, vol. 5, no. 1, pp. 11–22, Jun. 2024, doi: 10.47065/tin.v5i1.5237.

[15] E. Clark, S. Price, T. Lucena, B. Haberlein, Abdullah Wahbeh, and Raed Seetan, “Predictive Analytics for Thyroid Cancer Recurrence: A Machine Learning Approach,” Knowledge, vol. 4, no. 4, pp. 557–570, Nov. 2024, doi: 10.3390/knowledge4040029.

[16] Irem Senyer Yapici and R. U. Arslan, “Predictive analytics for thyroid cancer recurrence: a feature selection and data balancing approach,” The European Physical Journal Special Topics, Jun. 2025, doi: 10.1140/epjs/s11734-025-01720-x.

[17] Shiva Borzooei, G. Briganti, Mitra Golparian, J. R. Lechien, and Aidin Tarokhian, “Machine learning for risk stratification of thyroid cancer patients: a 15-year cohort study,” European Archives of Oto-Rhino-Laryngology, Oct. 2023, doi: 10.1007/s00405-023-08299-w.

[18] S. Das, A. K. Chaudhuri, N. R. Choudhury, and P. Ghosh, “Identification of the Recurrence of Differentiated Thyroid Cancer by Stacking Classifier,” Research Square (Research Square), Jan. 2025, doi: 10.21203/rs.3.rs-5713674/v1.

[19] Ş. Yaşar, “Determination of Possible Biomarkers for Predicting Well-Differentiated Thyroid Cancer Recurrence by Different Ensemble Machine Learning Methods,” Middle Black Sea Journal of Health Science, vol. 10, no. 3, pp. 255–265, Aug. 2024, doi: 10.19127/mbsjohs.1498383.

[20] G. M. Idroes et al., “Prognostication of differentiated thyroid cancer recurrence: An explainable machine learning approach,” Narra X, vol. 2, no. 3, pp. e183–e183, Dec. 2024, doi: 10.52225/narrax.v2i3.183.

[21] M. A.-S. Ahmad and J. Haddad, “An Explainable AI Model for Predicting the Recurrence of Differentiated Thyroid Cancer,” arXiv (Cornell University), Oct. 2024, doi: 10.48550/arxiv.2410.10907.

[22] A. K. Arslan and C. Çolak, “Explainable Machine Learning Models for Predicting Recurrence in Differentiated Thyroid Cancer,” Medical Records, Aug. 2024, doi: 10.37990/medr.1525801.

[23] L. J. Lechuga López, S. Elsharief, D. Al Jorf, F. Darwish, C. Ma, and F. E. Shamout, “Uncertainty quantification for machine learning in healthcare: A survey,” arXiv (Cornell University), May. 2025, doi: 10.48550/arXiv.2505.02874.

[24] T. Wang et al., “From aleatoric to epistemic: Exploring uncertainty quantification techniques in artificial intelligence,” arXiv (Cornell University), Jan. 2025, doi: 10.48550/arXiv.2501.03282.

[25] “Differentiated thyroid cancer recurrence,” Kaggle, Jan. 19, 2024. [Online]. Available: https://www.kaggle.com/datasets/joebeachcapital/differentiated-thyroid-cancer-recurrence

[26] K. Rupabanta Singh and S. Dash, “Early detection of neurological diseases using machine learning and deep learning techniques: A review,” Elsevier eBooks, pp. 1–24, Jan. 2023, doi: 10.1016/b978-0-323-90277-9.00001-8.

[27] G. Cosma, D. Brown, M. Archer, M. Khan, and A. Graham Pockley, “A survey on computational intelligence approaches for predictive modeling in prostate cancer,” Expert Systems with Applications, vol. 70, pp. 1–19, Mar. 2017, doi: 10.1016/j.eswa.2016.11.006.

[28] D. Wu, Y. Ren, L. He, and J. Johnson, “The identification of risk factors associated with COVID-19 in a large inpatient cohort using machine learning approaches,” Elsevier eBooks, pp. 189–199, Jan. 2022, doi: 10.1016/b978-0-12-821318-6.00017-7.

[29] C. Zucco, “Multiple Learners Combination: Bagging,” Encyclopedia of Bioinformatics and Computational Biology, pp. 525–530, 2019, doi: 10.1016/b978-0-12-809633-8.20345-2.

[30] C. Gong, Z. Su, X. Zhang, and Y. You, “Adaptive evidential K-NN classification: Integrating neighborhood search and feature weighting,” Information sciences, vol. 648, pp. 119620–119620, Nov. 2023, doi: 10.1016/j.ins.2023.119620.

[31] I. Semanjski, “Data analytics,” Elsevier eBooks, pp. 121–170, Jan. 2023, doi: 10.1016/b978-0-12-820717-8.00008-7.

[32] M. Meloun and J. Militký, “Statistical analysis of multivariate data,” Statistical Data Analysis, pp. 151–403, 2011, doi: 10.1533/9780857097200.151.

[33] A. Malik, Yash Tejas Javeri, M. Shah, and Ramchandra Mangrulkar, “Impact analysis of COVID-19 news headlines on global economy,” Elsevier eBooks, pp. 189–206, Jan. 2022, doi: 10.1016/b978-0-12-824557-6.00001-7.

[34] W. Zhang, X. Gu, L. Hong, L. Han, and L. Wang, “Comprehensive review of machine learning in geotechnical reliability analysis: Algorithms, applications and further challenges,” vol. 136, pp. 110066–110066, Mar. 2023, doi: 10.1016/j.asoc.2023.110066.

[35] Temidayo Oluwatosin Omotehinwa, David Opeoluwa Oyewola, and Emmanuel Gbenga Dada, “A Light Gradient-Boosting Machine algorithm with Tree-Structured Parzen Estimator for breast cancer diagnosis,” pp. 100218–100218, Jun. 2023, doi: 10.1016/j.health.2023.100218.

[36] X. Ren, H. Yu, X. Chen, Y. Tang, G. Wang, and X. Du, “Application of the CatBoost Model for Stirred Reactor State Monitoring Based on Vibration Signals,” Computer modeling in engineering & sciences, vol. 140, no. 1, pp. 647–663, Jan. 2024, doi: 10.32604/cmes.2024.048782.

[37] S. Abinaya and M. K. K. Devi, “Enhancing crop productivity through autoencoder-based disease detection and context-aware remedy recommendation system,” Elsevier eBooks, pp. 239–262, Jan. 2022, doi: 10.1016/b978-0-323-90550-3.00014-x.

[38] O. Peretz, M. Koren, and O. Koren, “Naive Bayes classifier – An ensemble procedure for recall and precision enrichment,” Engineering Applications of Artificial Intelligence, vol. 136, p. 108972, Oct. 2024, doi: 10.1016/j.engappai.2024.108972.

[39] G. Manikandan, B. Pragadeesh, V. Manojkumar, A. L. Karthikeyan, R. Manikandan, and A. H. Gandomi, “Classification models combined with Boruta feature selection for heart disease prediction,” Informatics in medicine unlocked, vol. 44, pp. 101442–101442, Jan. 2024, doi: 10.1016/j.imu.2023.101442.

[40] L. V. Jospin, H. Laga, F. Boussaid, W. Buntine, and M. Bennamoun, “Hands-On Bayesian Neural Networks—A Tutorial for Deep Learning Users,” IEEE Computational Intelligence Magazine, vol. 17, no. 2, pp. 29–48, May 2022, doi: 10.1109/mci.2022.3155327.

[41] M. Magris and A. Iosifidis, “Bayesian learning for neural networks: an algorithmic survey,” Artificial Intelligence Review, Mar. 2023, doi: 10.1007/s10462-023-10443-1.

[42] M. Nishio and A. Arakawa, “Performance of Hamiltonian Monte Carlo and No-U-Turn Sampler for estimating genetic parameters and breeding values,” Genetics Selection Evolution, vol. 51, no. 1, Dec. 2019, doi: 10.1186/s12711-019-0515-1.

[43] M. D. Hoffman and A. Gelman, “The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo,” arXiv (Cornell University), Jan. 2011, doi: 10.48550/arxiv.1111.4246.

[44] A. Kendall and Y. Gal, “What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?,” arXiv (Cornell University), Jan. 2017, doi: 10.48550/arxiv.1703.04977.

[45] J. W. Lund, J. T. Groten, D. L. Karwan, and C. Babcock, “Using machine learning to improve predictions and provide insight into fluvial sediment transport,” Hydrological Processes, vol. 36, no. 8, Jul. 2022, doi: 10.1002/hyp.14648.

[46] E. Štrumbelj and I. Kononenko, “Explaining prediction models and individual predictions with feature contributions,” Knowledge and Information Systems, vol. 41, no. 3, pp. 647–665, Aug. 2013, doi: 10.1007/s10115-013-0679-x.

[47] S. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” arXiv.org, Nov. 24, 2017. [Online]. Available: https://arxiv.org/abs/1705.07874v2

[48] S. Wang, H. Peng, Q. Hu, and M. Jiang, “Analysis of runoff generation driving factors based on hydrological model and interpretable machine learning method,” Journal of Hydrology: Regional Studies, vol. 42, p. 101139, Aug. 2022, doi: 10.1016/j.ejrh.2022.101139.

[49] C. Molnar, G. Casalicchio, and B. Bischl, “Interpretable Machine Learning – A Brief History, State-of-the-Art and Challenges,” ECML PKDD 2020 Workshops, pp. 417–431, 2020, doi: 10.1007/978-3-030-65965-3_28.

[50] H. Lamane, L. Mouhir, R. Moussadek, B. Baghdad, O. Kisi, and A. El Bilali, “Interpreting machine learning models based on SHAP values in predicting suspended sediment concentration,” International Journal of Sediment Research, vol. 40, no. 1, pp. 91–107, Feb. 2025, doi: 10.1016/j.ijsrc.2024.10.002.

Additional Files

Published

16-11-2025 — Updated on 16-07-2026

How to Cite

Herath Mudiyanselage, N. S. K., Kumari, H., & Nawarathne, U. (2026). Differentiated Thyroid Cancer Recurrence Classification Using Machine Learning Models and Bayesian Neural Networks with Varying Priors: A SHAP-Based Interpretation of the Best Performing Model. International Journal of Research in Computing, 5(2), 17–42. Retrieved from https://www.ijrcom.org/index.php/ijrc/article/view/161