WANG Qian, GU Taifu, HAN Qinghe, et al
Journal of Clinical Radiology. 2026, 45(9): 1496-1508.
Objective This study aims to systematically evaluate and compare the diagnostic performance of various artificial intelligence approaches, including traditional radiomics, ResNet50, morphological features, Vision Transformer (ViT), and a ViT model integrated with morphological features (ViT+shape), in differentiating orbital solitary fibrous tumor (SFT) from two common benign focal masses (cavernous hemangioma and schwannoma). The study seeks to identify the model with optimal generalizability and, based on this model, to further optimize the combination of MRI sequences, while evaluating the added value of contrast-enhanced imaging in this differential diagnosis. Methods This study retrospectively included 520 patients with surgically and pathologically confirmed focal orbital masses, comprising an internal cohort of 327 patients (69 SFT, 156 cavernous hemangiomas, and 102 schwannomas) and an external cohort of 193 patients (25 SFT, 113 cavernous hemangiomas, and 55 schwannomas). The study cohort was divided into a training set, an internal test set, and an external test set. We constructed five classification models based on conventional radiomics, ResNet50, shape features (Shape), ViT, and ViT+Shape, and comprehensively evaluated their generalization performance on the internal test set and the independent external test set. Each model was trained and evaluated for three-class classification under five sequence combinations: T1WI; T2WI; the non-contrast combination of T1WI and T2WI; contrast-enhanced T1WI (CE-T1WI); and the full-sequence combination of T1WI, T2WI, and CE-T1WI. Model performance was comprehensively assessed using macro-average AUC (AUC_macro), micro-average AUC (AUC_micro), per-class precision/recall/F1-score, Cohen's Kappa coefficient, and confusion matrices, with the DeLong test combined with Bonferroni correction used for statistical comparison of AUC between models. Results In the multi-sequence combined analysis of the internal test set, the Shape model achieved the highest AUC_macro (0.856), followed by ViT+Shape (0.849), while ResNet50 had the lowest (0.560). After Bonferroni correction, ResNet50 demonstrated significantly lower AUC for SFT compared with Shape, ViT, and ViT+Shape (DeLong test, Bonferroni correction, all P < 0.05). In the external test set, the performance of the radiomics model declined substantially, with the AUC_macro dropping to 0.495 and the Kappa coefficient to only 0.044. The recall for cavernous hemangioma was merely 13.3%, and the confusion matrix revealed that 57.5% of cases were misclassified as SFT. In contrast, the ViT+Shape model performed robustly, achieving an AUC_macro of 0.838 and a Kappa coefficient of 0.433. The addition of contrast-enhanced imaging to the triple-sequence combination showed only marginal enhancement in diagnostic performance compared with the non-contrast dual-sequence combination. In the external test set, the ViT+Shape model demonstrated the most significant improvement, with the macro-AUC rising from 0.810 to 0.838 under the triple-sequence combination. Nevertheless, this difference did not reach statistical significance after Bonferroni correction (P > 0.05). Conclusion The ViT+Shape model, constructed with prior morphological knowledge, demonstrates high accuracy and robustness in differentiating SFT, cavernous hemangioma, and schwannoma, with stable and reliable performance across external validation. Notably, the model achieves comparable diagnostic performance and clinical utility using only non-contrast T1WI and T2WI sequences, comparable to the full-sequence protocol that includes contrast-enhanced imaging. This provides a feasible technical approach to reduce unnecessary contrast-enhanced examinations.