人工智能驱动的全链路临床诊断思维教学模式研发与评价

Development and Evaluation of an Artificial Intelligence-Driven Full-Link Teaching Model for Clinical Diagnostic Thinking

  • 摘要: 目的 回顾性总结本团队自研人工智能(artificial intelligence, AI)全链路临床诊断思维训练教学智能体的临床教学应用情况,观察不同教学模式下医学生的教学成效差异,初步分析计划-执行-检查-处理(Plan-Do-Check-Act, PDCA)管理机制在AI辅助临床教学中的应用趋势与潜在价值,以期为“AI+临床医学教学”的优化改良提供参考性实践数据与经验借鉴。方法 本研究为回顾性观察性研究,选取2024年9月—2025年9月在上海交通大学医学院完成《诊断学基础》课程学习的临床医学专业学生为观察对象。基于既定的教学改革实施方案,将研究对象分为传统教学组(A组)、混合式AI教学组(B组)、PDCA循环优化的混合式AI教学组(C组)。3组教学大纲、核心教学内容及总学时(32学时)保持基本统一,其中A为传统教学组,采用常规课堂授课与病区小组式带教;B组采用传统+AI辅助教学模式,使用AI全链路临床诊断思维训练系统辅助训练,但未实施PDCA管理;C组为PDCA管理下的传统+AI辅助教学模式,即在传统授课与AI训练基础上叠加PDCA全周期教学质量管理闭环。收集3组课程开始前的基线测评资料、课程结束后《诊断学基础》期末理论成绩及客观结构化临床考试(objective structured clinical examination,OSCE)临床技能考核数据,比较不同教学模式的教学效果差异。结果 共纳入180名临床医学专业学生,A组、B组、C组各60人。3组年龄、男性比例、既往平均学分绩点、见习前诊断学理论测验成绩、医患沟通测评成绩等基线资料差异均无统计学意义(P均>0.05)。课程结束后,A组、B组、C组《诊断学基础》期末理论成绩分别为(83.1±5.9)分、(83.4±6.0)分、(82.6±5.4)分,组间差异无统计学意义(P=0.739)。在OSCE临床技能考核成绩方面,B组、C组“病史问诊”(16.0±0.7)分、(17.2±0.9)分与“病例分析”(14.9±0.7)分、(16.0±0.7)分两个模块的评分高于A组(14.4±0.8)分、(13.0±0.8)分,且C组优于B组,组间差异均具有统计学意义(P均<0.05)。结论 相较于传统教学模式,叠加自研AI全链路教学工具的教学方案,与PDCA闭环管理结合的复合教学模式,可能有助于改善医学生临床诊断思维相关实践能力,且未对其理论知识的掌握产生负面干预。AI赋能联合PDCA闭环优化的教学形式,或存在提升临床教学效果稳定性、同质化水平的潜在作用,可作为临床医学诊断思维教学的可选改良方案,但其长期效果、推广价值仍需进一步验证。

     

    Abstract: Objective To retrospectively summarize the clinical teaching application of our team's self-developed artificial intelligence (AI) full-chain clinical diagnostic reasoning training teaching agent, observe the differences in teaching outcomes among medical students under different teaching models, and preliminarily analyze the application trends and potential value of the Plan-Do-Check-Act (PDCA) management mechanism in AI-assisted clinical teaching, so as to provide reference practical data and experiential insights for the optimization and improvement of "AI + clinical medical teaching". Methods This study was a retrospective observational study that enrolled clinical medicine students who had completed the Fundamentals of Diagnostics course at Shanghai Jiao Tong University School of Medicine from September 2024 to September 2025 as the study subjects. Based on the established teaching reform implementation plan, the participants were divided into three groups:traditional teaching group (Group A), blended AI teaching group (Group B), and PDCA cycle-optimized blended AI teaching group (Group C). The three groups maintained essential consistency in teaching syllabi, core content, and total class hours (32 credit hours). Group A received conventional lecture-based classroom instruction and ward-based small-group mentoring. Group B received traditional teaching plus AI-assisted instruction using the AI full-chain clinical diagnostic reasoning training system, but without implementing PDCA management. Group C received traditional teaching plus AI-assisted instruction with PDCA management, i.e., a full-cycle teaching quality management closed-loop system based on conventional teaching and AI training. Baseline assessment data before the course, final theoretical examination scores in Basic Diagnostics, and objective structured clinical examination (OSCE) clinical skills assessment data after course completion were collected across the three groups to compare the differences in teaching effectiveness among different teaching models. Results A total of 180 clinical medicine students were enrolled, with 60 students in each of Groups A, B, and C. No statistically significant differences were observed among the three groups in baseline data, including age, male proportion, previous grade point average, pre-internship diagnostic theory test scores, and doctor-patient communication assessment scores (all P > 0.05), indicating comparability of baseline data between groups. Post-course assessments showed that the final theoretical examination scores in Basic Diagnostics for Groups A, B, and C were (83.1 ± 5.9), (83.4 ± 6.0), and (82.6 ± 5.4), respectively, with no statistically significant difference between groups (P=0.739). In terms of OSCE clinical skills assessment, Groups B and C scored higher than Group A in the modules of "history taking"(14.4 ± 0.8), (16.0 ± 0.7), and (17.2 ± 0.9), respectively and "case analysis"(13.0 ± 0.8), (14.9 ± 0.7), and (16.0 ± 0.7), respectively, and Group C outperformed Group B, with all between-group differences reaching statistical significance (all P< 0.05). Conclusion Compared with the traditional teaching model, the teaching protocol incorporating our self-developed AI full-chain teaching tool, combined with the PDCA closed-loop management composite teaching model, may help improve medical students' practical abilities in clinical diagnostic reasoning, without exerting negative interference on their mastery of theoretical knowledge. The teaching format integrating AI empowerment with PDCA closed-loop optimization may have potential effects in enhancing the stability and homogenization level of clinical teaching outcomes. It may serve as an optional improvement strategy for clinical diagnostic reasoning teaching. However, its long-term effectiveness and broader applicability still require further validation. In terms of OSCE clinical skills assessment, Groups B and C scored higher than Group A in both the "history taking" module(16.0±0.7) and (17.2±0.9) vs. (14.4±0.8), respectively and the "case analysis" module(14.9±0.7) and (16.0±0.7) vs. (13.0±0.8), respectively. Furthermore, Group C outperformed Group B in both modules. All between-group differences were statistically significant (all P < 0.05).

     

/

返回文章
返回