Beyond Average Error: Calibration, Selective Prediction, and Stress Robustness of A Three-Run Deep Learning Ensemble For Pediatric Bone Age Estimation

Authors

  • Safaa Abass Mohammed
  • Mohammed Hashim Albashir
  • Amal Yousif Ahmed
  • Suhaila M Elmudather

Keywords:

bone age; pediatric radiography; deep learning; ensemble learning; conformal prediction; uncertainty quantification; stress testing.

Abstract

Background: Bone-age models are commonly compared by average error, but reliability also depends on repeatability, interval calibration, error ranking, and image robustness.

Purpose: To evaluate the internal reliability of a reproducible three-run ensemble for pediatric bone-age estimation.

Materials and Methods: This retrospective study used 12,611 labeled radiographs from the public RSNA Pediatric Bone Age Challenge. Records were assigned to disjoint training (n=8,827), tuning (n=1,261), calibration (n=1,261), and internal-test (n=1,262) partitions, stratified by sex and age band. Three pretrained ResNet-18 regressors with a sex embedding were trained with identical settings but different random seeds; predictions were averaged. We evaluated error metrics, bootstrap confidence intervals (CIs), split-conformal intervals, disagreement-based selective prediction, and paired image-presentation stress tests.

Results: Individual test MAEs were 8.005, 7.855, and 8.364 months; ensemble MAE was 7.578 months (95% CI, 7.238–7.945), RMSE 9.911 months, and R² 0.940. The ensemble improved MAE by 0.276 months versus the best member. A fixed 90% split-conformal interval achieved 91.52% coverage with 32.49-month mean width. Disagreement weakly correlated with absolute error (Spearman ρ=0.099). Downsampling, additive noise, and horizontal flipping increased MAE by 3.690, 1.903, and 1.108 months, respectively.

Conclusion: Three-run averaging modestly improved internal accuracy and fixed conformal intervals achieved near-nominal coverage. However, disagreement was inefficient for case-level referral, and stressors exposed important fragility. External testing is required before clinical use.

 

Downloads

Published

2026-09-25

How to Cite

Mohammed, S. A., Albashir, M. H., Ahmed, A. Y., & Elmudather, S. M. (2026). Beyond Average Error: Calibration, Selective Prediction, and Stress Robustness of A Three-Run Deep Learning Ensemble For Pediatric Bone Age Estimation. Adolescência E Saúde, 846–855. Retrieved from https://adolescenciaesaude.com/index.php/aes/article/view/2388

Issue

Section

Original Articles