Comparative Analysis of Artificial Intelligence Virtual Assistant and Large Language Models in Post-Operative Care.

Borna, Sahar; Gomez-Cabello, Cesar A; Pressman, Sophia M; Haider, Syed Ali; Sehgal, Ajai; Leibovich, Bradley C; Cole, Dave; Forte, Antonio Jorge

Borna, Sahar; Gomez-Cabello, Cesar A; Pressman, Sophia M; Haider, Syed Ali; Sehgal, Ajai; Leibovich, Bradley C; Cole, Dave; Forte, Antonio Jorge.

Afiliação

Borna S; Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA.
Gomez-Cabello CA; Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA.
Pressman SM; Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA.
Haider SA; Division of Plastic Surgery, Mayo Clinic, Jacksonville, FL 32224, USA.
Sehgal A; Center for Digital Health, Mayo Clinic, Rochester, MN 55905, USA.
Leibovich BC; Center for Digital Health, Mayo Clinic, Rochester, MN 55905, USA.
Cole D; Department of Urology, Mayo Clinic, Rochester, MN 55905, USA.
Forte AJ; Center for Digital Health, Mayo Clinic, Rochester, MN 55905, USA.

Eur J Investig Health Psychol Educ ; 14(5): 1413-1424, 2024 May 15.

Article em En | MEDLINE | ID: mdl-38785591

ABSTRACT

ABSTRACT

In postoperative care, patient education and follow-up are pivotal for enhancing the quality of care and satisfaction. Artificial intelligence virtual assistants (AIVA) and large language models (LLMs) like Google BARD and ChatGPT-4 offer avenues for addressing patient queries using natural language processing (NLP) techniques. However, the accuracy and appropriateness of the information vary across these platforms, necessitating a comparative study to evaluate their efficacy in this domain. We conducted a study comparing AIVA (using Google Dialogflow) with ChatGPT-4 and Google BARD, assessing the accuracy, knowledge gap, and response appropriateness. AIVA demonstrated superior performance, with significantly higher accuracy (mean 0.9) and lower knowledge gap (mean 0.1) compared to BARD and ChatGPT-4. Additionally, AIVA's responses received higher Likert scores for appropriateness. Our findings suggest that specialized AI tools like AIVA are more effective in delivering precise and contextually relevant information for postoperative care compared to general-purpose LLMs. While ChatGPT-4 shows promise, its performance varies, particularly in verbal interactions. This underscores the importance of tailored AI solutions in healthcare, where accuracy and clarity are paramount. Our study highlights the necessity for further research and the development of customized AI solutions to address specific medical contexts and improve patient outcomes.

Palavras-chave

Bard; ChatGPT; artificial intelligence; large language model; machine learning; natural language processing

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Idioma: En Revista: Eur J Investig Health Psychol Educ Ano de publicação: 2024 Tipo de documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google