Abstract:
Objective To retrospectively summarize the clinical teaching application of our team's self-developed artificial intelligence (AI) full-chain clinical diagnostic reasoning training teaching agent, observe the differences in teaching outcomes among medical students under different teaching models, and preliminarily analyze the application trends and potential value of the Plan-Do-Check-Act (PDCA) management mechanism in AI-assisted clinical teaching, so as to provide reference practical data and experiential insights for the optimization and improvement of "AI + clinical medical teaching".
Methods This study was a retrospective observational study that enrolled clinical medicine students who had completed the Fundamentals of Diagnostics course at Shanghai Jiao Tong University School of Medicine from September 2024 to September 2025 as the study subjects. Based on the established teaching reform implementation plan, the participants were divided into three groups:traditional teaching group (Group A), blended AI teaching group (Group B), and PDCA cycle-optimized blended AI teaching group (Group C). The three groups maintained essential consistency in teaching syllabi, core content, and total class hours (32 credit hours). Group A received conventional lecture-based classroom instruction and ward-based small-group mentoring. Group B received traditional teaching plus AI-assisted instruction using the AI full-chain clinical diagnostic reasoning training system, but without implementing PDCA management. Group C received traditional teaching plus AI-assisted instruction with PDCA management, i.e., a full-cycle teaching quality management closed-loop system based on conventional teaching and AI training. Baseline assessment data before the course, final theoretical examination scores in Basic Diagnostics, and objective structured clinical examination (OSCE) clinical skills assessment data after course completion were collected across the three groups to compare the differences in teaching effectiveness among different teaching models.
Results A total of 180 clinical medicine students were enrolled, with 60 students in each of Groups A, B, and C. No statistically significant differences were observed among the three groups in baseline data, including age, male proportion, previous grade point average, pre-internship diagnostic theory test scores, and doctor-patient communication assessment scores (all
P > 0.05), indicating comparability of baseline data between groups. Post-course assessments showed that the final theoretical examination scores in Basic Diagnostics for Groups A, B, and C were (83.1 ± 5.9), (83.4 ± 6.0), and (82.6 ± 5.4), respectively, with no statistically significant difference between groups (
P=0.739). In terms of OSCE clinical skills assessment, Groups B and C scored higher than Group A in the modules of "history taking"(14.4 ± 0.8), (16.0 ± 0.7), and (17.2 ± 0.9), respectively and "case analysis"(13.0 ± 0.8), (14.9 ± 0.7), and (16.0 ± 0.7), respectively, and Group C outperformed Group B, with all between-group differences reaching statistical significance (all
P< 0.05).
Conclusion Compared with the traditional teaching model, the teaching protocol incorporating our self-developed AI full-chain teaching tool, combined with the PDCA closed-loop management composite teaching model, may help improve medical students' practical abilities in clinical diagnostic reasoning, without exerting negative interference on their mastery of theoretical knowledge. The teaching format integrating AI empowerment with PDCA closed-loop optimization may have potential effects in enhancing the stability and homogenization level of clinical teaching outcomes. It may serve as an optional improvement strategy for clinical diagnostic reasoning teaching. However, its long-term effectiveness and broader applicability still require further validation. In terms of OSCE clinical skills assessment, Groups B and C scored higher than Group A in both the "history taking" module(16.0±0.7) and (17.2±0.9)
vs. (14.4±0.8), respectively and the "case analysis" module(14.9±0.7) and (16.0±0.7)
vs. (13.0±0.8), respectively. Furthermore, Group C outperformed Group B in both modules. All between-group differences were statistically significant (all
P < 0.05).