基于视觉基础模型的自监督儿童腹部CT多器官分割方法

Self-Supervised Multi-Organ Segmentation in Pediatric Abdominal CT Based on Vision Foundation Models

  • 摘要: 目的 针对儿童腹部CT影像标注数据稀缺及现有模型泛化能力不足的问题,基于视觉基础模型DINOv3构建面向儿童CT域适应的自监督预训练架构,并验证其在儿童腹部多器官分割任务中的性能。方法 基于大规模成人CT数据集CT-3M构建通用放射学解剖表征,并引入Gram锚定机制,将冻结的成人预训练模型作为结构教师,引导模型在无标注儿童CT数据上进行局部拓扑结构的域对齐。结合多尺度特征聚合策略和轻量级Primus解码器,在公开儿童CT数据集上进行下游分割任务测试。基于逐病例配对结果,采用Wilcoxon符号秩检验比较本研究模型与基线模型nnU-Net的平均Dice相似系数(Dice similarity coefficient,DSC)和平均交并比(intersection over union,IoU),并计算相对性能提升幅度。结果 共收集867例腹部CT影像数据,构建了包含367 588张二维CT切片的预训练数据集。在公开Pediatric-CT-SEG数据集(359例)上,本研究模型的平均DSC为(71.38±1.08)%、平均IoU为(63.73±1.01)%,较基准模型nnU-Net分别提升3.22%和3.59%,差异具有统计学意义(P<0.05)。且在十二指肠(5.78%)、胰腺(4.69%)及左(2.89%)、右肾上腺(1.12%)、胆囊(1.86%)等小体积或边界模糊器官上的平均DSC均有稳定提升。消融实验表明,成人预训练、儿童域适配、高分辨率适配及多尺度特征聚合后的DSC分别提升1.44%、3.66%、5.78%和6.45%。结论 本研究提出的自监督预训练框架能有效缓解成人与儿童腹部CT影像间的域偏移问题,显著提升儿童腹部多器官(尤其是小器官及复杂边界结构)的分割精度,为标注数据稀缺场景下的儿科影像智能分析提供了可靠的技术方案。

     

    Abstract: Objective To address the scarcity of annotated data for pediatric abdominal CT imaging and the insufficient generalization capability of existing models,we constructed a self-supervised pretraining architecture tailored for pediatric CT domain adaptation based on the visual foundation model DINOv3,and validated its performance in the task of pediatric abdominal multi-organ segmentation. Methods We built a general-purpose radiological visual representation using the large-scale adult CT dataset CT-3M,and introduced a Gram-anchoring mechanism that employs a frozen adult pretrained model as a structural teacher to guide domain alignment of local topological structures on unlabeled pediatric CT data. Combined with a multi-scale feature aggregation strategy and a lightweight Primus decoder,downstream segmentation tasks were evaluated on a public pediatric CT dataset. Based on case-wise paired results,we compared the mean Dice similarity coefficient (DSC) and mean intersection over union (IoU) between our model and the baseline nnU-Net using the Wilcoxon signed-rank test,and computed the relative performance improvements. Results A total of 867 abdominal CT imaging cases were collected,constituting a pretraining dataset comprising 367 588 two-dimensional CT slices. On the public Pediatric-CT-SEG dataset (359 cases),our model achieved a mean DSC of (71.38± 1.08)% and a mean IoU of (63.73±1.01)%,representing improvements of 3.22% and 3.59% over the baseline nnU-Net,respectively,with statistically significant differences (P<0.05). Stable improvements in mean DSC were also observed for small-volume or boundary-ambiguous organs,including the duodenum (5.78%),pancreas (4.69%),left adrenal gland (2.89%),right adrenal gland (1.12%),and gallblad-der (1.86%). Ablation experiments demonstrated that DSC improved by 1.44%,3.66%,5.78%,and 6.45% following adult pretraining,pediatric domain adaptation,high-resolution adaptation,and multi-scale feature aggregation,respectively. Conclusions The self-supervised pretraining framework proposed in this study effectively alleviates the domain shift between adult and pediatric abdominal CT images,significantly enhances segmentation accuracy for pediatric abdominal multi-organs-particularly small organs and structures with complex boundaries-and provides a reliable technical solution for intelligent pediatric imaging analysis in scenarios with limited annotated data.

     

/

返回文章
返回