SCENARIO-BASED PREHOSPITAL TRIAGE OF CARBON MONOXIDE POISONING: COMPARING FIVE LARGE LANGUAGE MODELS WITH EMERGENCY MEDICAL SERVICES PERSONNEL IN A PROSPECTIVE STUDY Karbonmonoksit Zehirlenmesinde Senaryo Tabanlı Hastane Öncesi Triyaj: Beş Büyük Dil Modeli ile Acil Sağlık Hizmetleri Personelinin Prospektif Karşılaştırılması
Journal of Kirikkale University Faculty of Medicine, cilt.28, sa.2, ss.314-320, 2026 (Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 28 Sayı: 2
- Basım Tarihi: 2026
- Doi Numarası: 10.24938/kutfd.1922944
- Dergi Adı: Journal of Kirikkale University Faculty of Medicine
- Derginin Tarandığı İndeksler: Scopus, TR DİZİN (ULAKBİM)
- Sayfa Sayıları: ss.314-320
- Anahtar Kelimeler: Artificial intelligence, carbon monoxide, Emergency medical service, poisoning, triage
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Karadeniz Teknik Üniversitesi Adresli: Evet
Özet
Objective: Carbon monoxide (CO) poisoning requires rapid identification and timely decisions regarding the need for hyperbaric oxygen therapy (HBOT) to improve clinical outcomes. This study aimed to compare the decision-making performance of emergency medical services (EMS) personnel and large language models (LLMs) in accurately determining the need for HBOT during the prehospital phase of CO poisoning cases, and to explore the potential implementation of LLMs as decision-support tools for patient triage and referral. Material and Methods: In this prospective, simulation-based diagnostic accuracy study, 128 standardized scenario-based clinical cases (64 requiring HBOT, 64 not requiring HBOT) were developed based on established indications. Sixty EMS personnel and five LLMs (GPT-4o, GPT-4.5, GPT-o3, Gemini 2.5 Pro, and DeepSeek-R1) evaluated each scenario. Their responses were compared with the gold standard answers established by consensus between a medical toxicologist and a hyperbaric medicine specialist. Results: The overall accuracy of the EMS personnel was 65.1%, whereas the highest accuracy was observed with GPT-o3 (96.9%). All LLMs demonstrated 100% sensitivity, although the specificity varied, ranging from 26.6% (GPT-4.5) to 93.8% (GPT-o3). The accuracy of GPT-o3 (96.9%, p<0.001), GPT-4o (78.9%, p=0.001), and Gemini 2.5 Pro (75.8%, p=0.012) was significantly higher than that of the EMS personnel. GPT-o3 also significantly outperformed all other models (all p<0.001). Conclusion: Our study suggests that LLMs can support prehospital decision-making, facilitating timely referral to appropriate centers, improving early access to HBOT, and potentially enhancing outcomes. However, these findings are based on standardized simulation scenarios, and real-world clinical validation is necessary before widespread implementation.