Artificial Intelligence for Clinical Competence Assessment Compared with Human Examiners: A Systematic Review and Exploratory Meta-Analysis of Agreement, Reliability, And Validity
Keywords:
Artificial intelligence; Clinical competence; Assessment; Reliability and agreement; Medical education.Abstract
Artificial intelligence (AI) is increasingly being used to assess clinical competence, yet its measurement performance relative to human examiners remains uncertain across professions, tasks, and assessment formats. This systematic review and exploratory meta-analysis evaluated agreement, association, reliability, validity, score differences, classification performance, and human-human benchmarking for AI-based assessment of human clinical competence. Searches identified 29 eligible reports representing 27 independent studies across medicine, dentistry, nursing, pharmacy, and paramedicine. Owing to substantial heterogeneity in measurement methods, most outcomes were synthesized narratively. Two independent studies reported Spearman correlations and were included in an exploratory random-effects meta-analysis despite differing substantially in clinical assessment context using Fisher z transformation and restricted maximum likelihood estimation. The pooled AI-human association was ρ = 0.82 (95% CI, 0.55-0.93). The two correlations were numerically similar (I² = 0%), but between-study heterogeneity could not be estimated reliably because only two studies were available. This estimate should be interpreted cautiously as a cross-context summary of association rather than a common clinical effect. Direct agreement varied markedly across studies, and favorable correlations or reliability estimates sometimes coexisted with systematic score differences or disagreement near decision thresholds. Performance appeared strongest in structured, rubric-based tasks and more variable for communication, non-technical skills, and context-dependent judgments. Human-reference quality, input equivalence, model choice, and validation design also influenced interpretability across studies. The available measurement evidence is most compatible with human-supervised formative, supplementary, second-rater, or quality-assurance applications; autonomous high-stakes use requires stronger external validation, reproducibility, and decision-level safety evidence.

