Performance evaluation of three multimodal large language models for pediatric profile-based orthodontic screening
Xiangyu Ge, Jingcheng Chen, Chenyang Yuan, Zhenghan Chu, Xiangyu Li, Xi Zhang, Yanan Chen, WeiYing Zheng, Chunqin Miao
This study aimed to compare the performance of three multimodal large language models (ChatGPT, DeepSeek, and Gemini) in analyzing pediatric profile photographs and providing early orthodontic intervention recommendations, thereby assessing their clinical feasibility and reliability. In this cross-sectional study, 100 children aged 5–12 years who attended Jiaxing Second Hospital between January and June 2025 were enrolled. Standardized profile photographs were obtained and processed uniformly before being analyzed by the three models using identical prompts. Model outputs were anonymized and independently evaluated under single-blind, randomized conditions by orthodontic experts and parents. An eight-dimension weighted scoring system was applied, encompassing professionalism, Clinical Plausibility, completeness, individualization, safety, comprehensibility, empathy, and readability. Statistical analyses included the Friedman test, Wilcoxon signed-rank test, and Kendall’s W effect size. All three models achieved high overall scores, ranging from 3.9 to 4.2. ChatGPT consistently produced slightly higher mean scores (4.07–4.15), while DeepSeek and Gemini showed comparable performance (3.91–4.09). Inter-model differences were not statistically significant (all q > 0.05), and effect sizes were uniformly negligible (Kendall’s W = 0.003–0.029). ChatGPT, DeepSeek, and Gemini demonstrated comparable and overall reliable performance in pediatric orthodontic screening based on profile photographs, with ChatGPT showing a slight but nonsignificant advantage. At present, LLMs may serve as supportive tools for early orthodontic assessment but cannot substitute for clinical expertise. This study verifies that multimodal large language models can function as lightweight auxiliary tools for preliminary pediatric orthodontic screening, markedly improving the accessibility and clinical efficiency of early maxillofacial deformity screening. The tool delivers differentiated clinical value for three distinct user groups: Primary general and pediatric physicians can rapidly stratify risks of craniofacial developmental abnormalities using standardized pediatric profile photographs, enabling precise triage to orthodontic specialists and reducing unnecessary consultations and missed diagnoses.Caregivers of pediatric patients may perform self-administered preliminary assessments by capturing standardized facial images at home to detect craniofacial developmental irregularities at an early stage, seize the critical 5–12-year window for growth modification, and prevent delayed intervention.Orthodontic specialists may only utilize model outputs as supporting references for patient education and rapid initial screening; such outputs cannot replace comprehensive clinical examinations and formal diagnoses.This category of AI tools exclusively performs predictive screening functions and must not be regarded as clinical diagnostic modalities. Future research may integrate multimodal imaging datasets including cephalometric radiographs, CBCT scans, and intraoral scans to conduct large-sample, multicenter clinical validation. Further refinement of the model’s analytical performance will facilitate the deployment of this screening protocol in primary oral healthcare settings and home-based health management scenarios.