Abstract:
Objective To compare the diagnostic accuracy and differential diagnosis comprehensiveness of six general-purpose large language models (LLMs) on standardized clinical cases, and to assess their applicability for clinical reasoning education.
Methods A total of 74 standardized clinical cases published in the
BMJ "Endgames" column between March 2023 and February 2025 were included. Six general-purpose LLMs (ChatGPT-4, ChatGPT4o, DeepSeek-V3, DeepSeek-R1, Gemini2.0, and Kimi-k1.5) were prompted with identical text-only case information via their application programming interfaces (APIs). Diagnostic accuracy and differential diagnosis comprehensiveness were evaluated and compared across models. For statistical analysis, Cochran's Q test was used for overall comparisons of diagnostic accuracy, with post-hoc pairwise comparisons performed using McNemar tests; the Friedman test was used for overall comparisons of differential diagnosis comprehensiveness, with post-hoc pairwise comparisons performed using Wilcoxon signed-rank tests; all
P-values were adjusted using the Bonferroni method. Grade distribution differences were analyzed using Pearson's chi-square test.
Results Diagnostic accuracy ranged from 62.2% to 78.4% across the six models, with a statistically significant overall difference among models (
P=0.022); however, no pairwise comparison remained significant after Bonferroni correction (all
P>0.05). Stratified analysis showed that accuracy was higher for cases without images (63.3%-81.6%) than for those with images (52.0%-72.0%), with no significant differences among models in either subgroup (all
P>0.05). Regarding differential diagnosis comprehensiveness, DeepSeek-R1 achieved the highest mean coverage rate (61.2%), while ChatGPT-4o scored the lowest (50.3%), with a statistically significant overall difference among models (
χ2=16.6,
P=0.005). Post-hoc pairwise comparisons revealed that ChatGPT-4o was significantly inferior to Gemini-2.0 (adjusted
P=0.009,
r=0.436) and showed a borderline inferiority to DeepSeek-R1 (adjusted
P=0.052,
r=0.451). Regarding grade distribution, approximately 70% -80% of cases achieved only "moderate" or "limited" comprehensiveness, with no significant difference among models (
χ2=17.90,
P=0.268).
Conclusions Current general-purpose LLMs can achieve a moderate level of diagnostic accuracy on standardized clinical cases, but demonstrate insufficient differential diagnosis comprehensiveness. Given that diagnostic accuracy and differential diagnosis comprehensiveness are fundamental elements of clinical reasoning education, these findings suggest that the quality of LLMs-generated outputs is not yet sufficient for direct application in teaching scenarios and should be used with caution under teacher supervision.