Uncovering Language Disparity of ChatGPT on Retinal Vascular Disease Classification: Cross-Sectional Study.

Liu, Xiaocong; Wu, Jiageng; Shao, An; Shen, Wenyue; Ye, Panpan; Wang, Yao; Ye, Juan; Jin, Kai; Yang, Jie

Liu, Xiaocong; Wu, Jiageng; Shao, An; Shen, Wenyue; Ye, Panpan; Wang, Yao; Ye, Juan; Jin, Kai; Yang, Jie.

Afiliación

Liu X; Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
Wu J; School of Public Health, Zhejiang University School of Medicine, Zhejiang, China.
Shao A; School of Public Health, Zhejiang University School of Medicine, Zhejiang, China.
Shen W; Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
Ye P; Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
Wang Y; Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
Ye J; Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
Jin K; Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.
Yang J; Eye Center, The Second Affiliated Hospital, Zhejiang University, Zhejiang, China.

J Med Internet Res ; 26: e51926, 2024 Jan 22.

Article en En | MEDLINE | ID: mdl-38252483

ABSTRACT

ABSTRACT

BACKGROUND:

Benefiting from rich knowledge and the exceptional ability to understand text, large language models like ChatGPT have shown great potential in English clinical environments. However, the performance of ChatGPT in non-English clinical settings, as well as its reasoning, have not been explored in depth.

OBJECTIVE:

This study aimed to evaluate ChatGPT's diagnostic performance and inference abilities for retinal vascular diseases in a non-English clinical environment.

METHODS:

In this cross-sectional study, we collected 1226 fundus fluorescein angiography reports and corresponding diagnoses written in Chinese and tested ChatGPT with 4 prompting strategies (direct diagnosis or diagnosis with a step-by-step reasoning process and in Chinese or English).

RESULTS:

Compared with ChatGPT using Chinese prompts for direct diagnosis that achieved an F1-score of 70.47%, ChatGPT using English prompts for direct diagnosis achieved the best diagnostic performance (80.05%), which was inferior to ophthalmologists (89.35%) but close to ophthalmologist interns (82.69%). As for its inference abilities, although ChatGPT can derive a reasoning process with a low error rate (0.4 per report) for both Chinese and English prompts, ophthalmologists identified that the latter brought more reasoning steps with less incompleteness (44.31%), misinformation (1.96%), and hallucinations (0.59%) (all P<.001). Also, analysis of the robustness of ChatGPT with different language prompts indicated significant differences in the recall (P=.03) and F1-score (P=.04) between Chinese and English prompts. In short, when prompted in English, ChatGPT exhibited enhanced diagnostic and inference capabilities for retinal vascular disease classification based on Chinese fundus fluorescein angiography reports.

CONCLUSIONS:

ChatGPT can serve as a helpful medical assistant to provide diagnosis in non-English clinical environments, but there are still performance gaps, language disparities, and errors compared to professionals, which demonstrate the potential limitations and the need to continually explore more robust large language models in ophthalmology practice.

Asunto(s)

Inteligencia Artificial; Errores Diagnósticos; Angiografía con Fluoresceína; Lenguaje; Enfermedades de la Retina; Enfermedades Vasculares; Humanos; Estudios Transversales; Enfermedades Vasculares/clasificación; Enfermedades Vasculares/diagnóstico; Enfermedades Vasculares/diagnóstico por imagen; Enfermedades de la Retina/clasificación; Enfermedades de la Retina/diagnóstico; Enfermedades de la Retina/diagnóstico por imagen

Palabras clave

ChatGPT; artificial intelligence; clinical decision support; large language models; retinal vascular disease

Texto completo

Añadir a Mi BVS

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Colección: 01-internacional Base de datos: MEDLINE Contexto en salud: 1_ASSA2030 Problema de salud: 1_acesso_equitativo_servicos Asunto principal: Enfermedades de la Retina / Enfermedades Vasculares / Inteligencia Artificial / Angiografía con Fluoresceína / Errores Diagnósticos / Lenguaje Tipo de estudio: Observational_studies / Prevalence_studies / Prognostic_studies / Risk_factors_studies Límite: Humans Idioma: En Revista: J Med Internet Res Asunto de la revista: INFORMATICA MEDICA Año: 2024 Tipo del documento: Article País de afiliación: China

Texto completo

Añadir a Mi BVS

Imprimir

XML

PubMed Links

Buscar en Google