Torun B A, Torun A, Çağlayan M S. Interdisciplinary Evaluation of Urogenital Radiological Findings: Human Expert vs Multimodal Generative AI Classification Performance. Med J Islam Repub Iran 2026; 40 (1) :639-644
URL:
http://mjiri.iums.ac.ir/article-1-10301-en.html
Department of Gynecologic Oncology, Adana City Training and Research Hospital, University of Health Sciences, Adana, Turkey , b.asena.torun@gmail.com
Abstract: (144 Views)
Background: Multimodal generative artificial intelligence (AI) systems are increasingly evaluated for medical image classification, but their performance in interdisciplinary urogenital radiological assessment remains uncertain. This study aimed to compare the overall and domain-specific classification performance and uncertainty patterns of two domain-specific human specialists and three multimodal generative AI systems in a standardized single-image urogenital radiology task.
Methods: This exploratory pilot comparative study evaluated 60 publicly available anonymized radiological cases selected from Radiopaedia, including 20 urological anomaly/variation cases, 20 gynecological anomaly/variation cases, and 20 normal anatomy cases. Two human specialists and three multimodal AI systems (GPT-4o, Gemini 1.5 Pro, and Microsoft Copilot) independently classified randomized single images into predefined categories under blinded conditions. Overall accuracy was summarized with 95% confidence intervals, while domain-specific accuracy and uncertain/unable-to-classify responses were analyzed descriptively.
Results: Overall accuracy was highest for Gemini 1.5 Pro (36/60, 60.0%; 95% CI, 47.4-71.4), followed by the urologist (34/60, 56.7%; 95% CI, 44.1-68.4), GPT-4o (31/60, 51.7%; 95% CI, 39.3-63.8), the gynecologic oncologist (29/60, 48.3%; 95% CI, 36.2-60.7), and Copilot (19/60, 31.7%; 95% CI, 21.3-44.2). Human specialists performed best within their own domains. Gemini performed strongly in urological and gynecological categories but had low normal-category accuracy (2/20, 10.0%), suggesting overclassification.
Conclusion: In this standardized single-image pilot task, multimodal AI systems showed potential as supportive classification tools but also demonstrated clinically relevant limitations, particularly false-positive interpretation of normal anatomy. These findings should be interpreted as hypothesis-generating and not as evidence of autonomous diagnostic capability.