Utility and Comparative Performance of Current Artificial Intelligence Large Language Models as Postoperative Medical Support Chatbots in Aesthetic Surgery.

Abi-Rafeh, Jad; Henry, Nader; Xu, Hong Hao; Bassiri-Tehrani, Brian; Arezki, Adel; Kazan, Roy; Gilardino, Mirko S; Nahai, Foad

Abi-Rafeh, Jad; Henry, Nader; Xu, Hong Hao; Bassiri-Tehrani, Brian; Arezki, Adel; Kazan, Roy; Gilardino, Mirko S; Nahai, Foad.

Aesthet Surg J ; 44(8): 889-896, 2024 Jul 15.

Article in En | MEDLINE | ID: mdl-38318684

ABSTRACT

ABSTRACT

BACKGROUND:

Large language models (LLMs) have revolutionized the way plastic surgeons and their patients can access and leverage artificial intelligence (AI).

OBJECTIVES:

The present study aims to compare the performance of 2 current publicly available and patient-accessible LLMs in the potential application of AI as postoperative medical support chatbots in an aesthetic surgeon's practice.

METHODS:

Twenty-two simulated postoperative patient presentations following aesthetic breast plastic surgery were devised and expert-validated. Complications varied in their latency within the postoperative period, as well as urgency of required medical attention. In response to each patient-reported presentation, Open AI's ChatGPT and Google's Bard, in their unmodified and freely available versions, were objectively assessed for their comparative accuracy in generating an appropriate differential diagnosis, most-likely diagnosis, suggested medical disposition, treatments or interventions to begin from home, and/or red flag signs/symptoms indicating deterioration.

RESULTS:

ChatGPT cumulatively and significantly outperformed Bard across all objective assessment metrics examined (66% vs 55%, respectively; P < .05). Accuracy in generating an appropriate differential diagnosis was 61% for ChatGPT vs 57% for Bard (P = .45). ChatGPT asked an average of 9.2 questions on history vs Bard's 6.8 questions (P < .001), with accuracies of 91% vs 68% reporting the most-likely diagnosis, respectively (P < .01). Appropriate medical dispositions were suggested with accuracies of 50% by ChatGPT vs 41% by Bard (P = .40); appropriate home interventions/treatments with accuracies of 59% vs 55% (P = .94), and red flag signs/symptoms with accuracies of 79% vs 54% (P < .01), respectively. Detailed and comparative performance breakdowns according to complication latency and urgency are presented.

CONCLUSIONS:

ChatGPT represents the superior LLM for the potential application of AI technology in postoperative medical support chatbots. Imperfect performance and limitations discussed may guide the necessary refinement to facilitate adoption.

Subject(s)

Artificial Intelligence; Postoperative Complications; Humans; Female; Postoperative Complications/etiology; Postoperative Complications/diagnosis; Postoperative Complications/therapy; Mammaplasty/methods; Mammaplasty/adverse effects; Adult; Diagnosis, Differential

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google

Full text: 1 Collection: 01-internacional Database: MEDLINE Main subject: Postoperative Complications / Artificial Intelligence Type of study: Prognostic_studies Limits: Adult / Female / Humans Language: En Journal: Aesthet Surg J Year: 2024 Document type: Article Publication country: ENGLAND / ESCOCIA / GB / GREAT BRITAIN / INGLATERRA / REINO UNIDO / SCOTLAND / UK / UNITED KINGDOM

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google