Abstract:
Existing medical large language models are largely limited to text-based recommendations and lack the ability to execute tasks and verify safety in physical environments. To bridge the gap between semantic reasoning and the manipulation of embodied entities, this paper proposes a parallel orthopedic foundation model (ParaOrthoVLA) designed for executable surgical planning and closed-loop rehabilitation. This model deeply integrates the vision-language-action (VLA) architecture with the theory of parallel systems (ACP). Methodologically, we construct a multimodal alignment encoder and design a hierarchical VLA action layer to translate clinical intentions into structured action control scripts for scenarios such as pre-hospital emergency care, intraoperative robotic arm motion, and postoperative rehabilitation, thereby establishing a conceptual mechanism for physical manipulation. To ensure the safety of high-risk medical procedures, we construct a parallel orthopedic world model that performs “verify-then-execute” closed-loop testing of action scripts through multi-agent semantic games and high-fidelity physical simulation. Additionally, we introduce Retrieval-Augmented Generation (RAG) and Human-in-the-Loop (HITL) mechanisms to translate clinical guidelines into hard physical boundaries and soft compliance constraints. Theoretical workflow deductions indicate that this system architecture has the potential to screen for physical collisions and logical violations, providing a forward-looking verification framework for establishing a closed-loop decision-making and control system across the entire orthopedic workflow, while ensuring clinical safety and traceable accountability.