通用大语言模型临床推理输出在医学教育中的适用性评估与比较

Applicability Assessment and Comparison of Clinical Reasoning Outputs from General-Purpose Large Language Models in Medical Education

  • 摘要: 目的 比较6种通用大语言模型(large language models,LLMs)在标准化临床病例中的诊断准确性及鉴别诊断全面性,评估其在临床推理教学中的适用性。方法 纳入2023年3月—2025年2月发表于《英国医学杂志》(BMJ)“Endgames”专栏的74例标准化临床病例,以纯文本病例信息作为提示词,通过官方应用程序编程接口(application programming interface,API)调用6种通用LLMs (ChatGPT-4、ChatGPT-4o、DeepSeek-V3、DeepSeek-R1、Gemini-2.0、Kimi-k1.5),对模型的诊断准确率和鉴别诊断全面性进行评估与比较。诊断准确率的组间比较采用Cochran's Q检验,事后两两比较采用McNemar检验;鉴别诊断全面性的组间比较采用Friedman检验,事后两两比较采用配对Wilcoxon符号秩检验,P值均经Bonferroni法校正;等级分布差异采用Pearson卡方检验。结果 6种模型的诊断准确率介于62.2%~78.4%之间,模型间总体差异具有统计学意义(P=0.022),但任意两模型间差异均无统计学意义(P均>0.05)。分层分析显示,无图病例中模型的准确率(63.3%~81.6%)高于含图病例(52.0%~72.0%),各模型间差异均无统计学意义(P均>0.05)。6种模型中,DeepSeek-R1的鉴别诊断平均覆盖率最高(61.2%),ChatGPT-4o最低(50.3%),模型间总体差异具有统计学意义(χ2=16.6,P=0.005)。ChatGPT-4o的鉴别诊断全面性显著低于Gemini-2.0(校正后P=0.009,r=0.436),且低于DeepSeek-R1呈边缘显著(校正后P=0.052,r=0.451)。等级分布方面,约70%~80%的病例仅达“有限”或“中等”全面性,模型间差异无统计学意义(χ2=17.90,P=0.268)。结论 当前通用LLMs在标准化临床病例中的诊断准确率可达中等水平,但鉴别诊断全面性整体不足。由于诊断准确率与鉴别诊断全面性是临床推理教学的基础要素,提示其输出质量尚不支持直接用于教学场景,需经教师审核后谨慎使用。

     

    Abstract: Objective To compare the diagnostic accuracy and differential diagnosis comprehensiveness of six general-purpose large language models (LLMs) on standardized clinical cases, and to assess their applicability for clinical reasoning education. Methods A total of 74 standardized clinical cases published in the BMJ "Endgames" column between March 2023 and February 2025 were included. Six general-purpose LLMs (ChatGPT-4, ChatGPT4o, DeepSeek-V3, DeepSeek-R1, Gemini2.0, and Kimi-k1.5) were prompted with identical text-only case information via their application programming interfaces (APIs). Diagnostic accuracy and differential diagnosis comprehensiveness were evaluated and compared across models. For statistical analysis, Cochran's Q test was used for overall comparisons of diagnostic accuracy, with post-hoc pairwise comparisons performed using McNemar tests; the Friedman test was used for overall comparisons of differential diagnosis comprehensiveness, with post-hoc pairwise comparisons performed using Wilcoxon signed-rank tests; all P-values were adjusted using the Bonferroni method. Grade distribution differences were analyzed using Pearson's chi-square test. Results Diagnostic accuracy ranged from 62.2% to 78.4% across the six models, with a statistically significant overall difference among models (P=0.022); however, no pairwise comparison remained significant after Bonferroni correction (all P>0.05). Stratified analysis showed that accuracy was higher for cases without images (63.3%-81.6%) than for those with images (52.0%-72.0%), with no significant differences among models in either subgroup (all P>0.05). Regarding differential diagnosis comprehensiveness, DeepSeek-R1 achieved the highest mean coverage rate (61.2%), while ChatGPT-4o scored the lowest (50.3%), with a statistically significant overall difference among models (χ2=16.6, P=0.005). Post-hoc pairwise comparisons revealed that ChatGPT-4o was significantly inferior to Gemini-2.0 (adjusted P=0.009, r=0.436) and showed a borderline inferiority to DeepSeek-R1 (adjusted P=0.052, r=0.451). Regarding grade distribution, approximately 70% -80% of cases achieved only "moderate" or "limited" comprehensiveness, with no significant difference among models (χ2=17.90, P=0.268). Conclusions Current general-purpose LLMs can achieve a moderate level of diagnostic accuracy on standardized clinical cases, but demonstrate insufficient differential diagnosis comprehensiveness. Given that diagnostic accuracy and differential diagnosis comprehensiveness are fundamental elements of clinical reasoning education, these findings suggest that the quality of LLMs-generated outputs is not yet sufficient for direct application in teaching scenarios and should be used with caution under teacher supervision.

     

/

返回文章
返回