The utility of ChatGPT as a generative medical translator.

Grimm, David R; Lee, Yu-Jin; Hu, Katherine; Liu, Longsha; Garcia, Omar; Balakrishnan, Karthik; Ayoub, Noel F

Grimm, David R; Lee, Yu-Jin; Hu, Katherine; Liu, Longsha; Garcia, Omar; Balakrishnan, Karthik; Ayoub, Noel F.

Afiliación

Grimm DR; Division of Pediatric Otolaryngology, Department of Otolaryngology-Head and Neck Surgery, Stanford University School of Medicine, Stanford, CA, 94305, USA.
Lee YJ; Division of Pediatric Otolaryngology, Department of Otolaryngology-Head and Neck Surgery, Stanford University School of Medicine, Stanford, CA, 94305, USA.
Hu K; Division of Pediatric Otolaryngology, Department of Otolaryngology-Head and Neck Surgery, Stanford University School of Medicine, Stanford, CA, 94305, USA.
Liu L; Division of Pediatric Otolaryngology, Department of Otolaryngology-Head and Neck Surgery, Stanford University School of Medicine, Stanford, CA, 94305, USA.
Garcia O; Division of Pediatric Otolaryngology, Department of Otolaryngology-Head and Neck Surgery, Stanford University School of Medicine, Stanford, CA, 94305, USA.
Balakrishnan K; Division of Pediatric Otolaryngology, Department of Otolaryngology-Head and Neck Surgery, Stanford University School of Medicine, Stanford, CA, 94305, USA.
Ayoub NF; Division of Pediatric Otolaryngology, Department of Otolaryngology-Head and Neck Surgery, Stanford University School of Medicine, Stanford, CA, 94305, USA. noelayoub@gmail.com.

Eur Arch Otorhinolaryngol ; 2024 May 05.

Article en En | MEDLINE | ID: mdl-38705894

ABSTRACT

ABSTRACT

PURPOSE:

Large language models continue to dramatically change the medical landscape. We aimed to explore the utility of ChatGPT in providing accurate, actionable, and understandable generative medical translations in English, Spanish, and Mandarin pertaining to Otolaryngology.

METHODS:

Responses of GPT-4 to commonly asked patient questions listed on official otolaryngology clinical practice guidelines (CPG) were evaluated with the Patient Education materials Assessment Tool-printable (PEMAT-P.) Additional critical elements were identified a priori to evaluate ChatGPT's accuracy and thoroughness in its responses. Multiple fluent speakers of English, Mandarin, and Spanish evaluated each response generated by ChatGPT.

RESULTS:

Total PEMAT-P scores differed between English, Mandarin, and Spanish GPT-4 generated responses depicting a moderate effect size of language, Eta-Square 0.07 with scores ranging from 73 to 77 (P-value = 0.03). Overall understandability scores did not differ between English, Mandarin, and Spanish depicting a small effect size of language, Eta-Square 0.02 scores ranging from 76 to 79 (P-value = 0.17), nor did overall actionability scores Eta-Square 0 score ranging 66-73 (P-value = 0.44). Overall a priori procedure-specific responses similarly did not differ between English, Spanish, and Mandarin Eta-Square 0.02 scores ranging 61-78 (P-value = 0.22).

CONCLUSION:

GPT-4 produces accurate, understandable, and actionable outputs in English, Spanish, and Mandarin. Responses generated by GPT-4 in Spanish and Mandarin are comparable to English counterparts indicating a novel use for these models within Otolaryngology, and implications for bridging healthcare access and literacy gaps. LEVEL OF EVIDENCE IV.

Palabras clave

Artificial intelligence; Disparities; Large language models

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Banco de datos: MEDLINE Idioma: En Revista: Eur Arch Otorhinolaryngol Asunto de la revista: OTORRINOLARINGOLOGIA Año: 2024 Tipo del documento: Article País de afiliación: Estados Unidos

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google