ChatGPT Conquers the Saudi Medical Licensing Exam: Exploring the Accuracy of Artificial Intelligence in Medical Knowledge Assessment and Implications for Modern Medical Education.

Aljindan, Fahad K; Al Qurashi, Abdullah A; Albalawi, Ibrahim Abdullah S; Alanazi, Abeer Mohammed M; Aljuhani, Hussam Abdulkhaliq M; Falah Almutairi, Faisal; Aldamigh, Omar A; Halawani, Ibrahim R; K Zino Alarki, Subhi M

Aljindan, Fahad K; Al Qurashi, Abdullah A; Albalawi, Ibrahim Abdullah S; Alanazi, Abeer Mohammed M; Aljuhani, Hussam Abdulkhaliq M; Falah Almutairi, Faisal; Aldamigh, Omar A; Halawani, Ibrahim R; K Zino Alarki, Subhi M.

Affiliation

Aljindan FK; Department of Plastic Surgery, King Abdullah Medical City, Makkah, SAU.
Al Qurashi AA; College of Medicine, King Saud Bin Abdulaziz University for Health Sciences, Jeddah, SAU.
Albalawi IAS; College of Medicine, Tabuk University for Health Sciences, Tabuk, SAU.
Alanazi AMM; Department of Pediatrics, Faculty of Medicine, University of Tabuk, Tabuk, SAU.
Aljuhani HAM; Faculty of Medicine, Ibn Sina National College, Jeddah, SAU.
Falah Almutairi F; College of Medicine, Unaizah College of Medicine and Medical Sciences, Qassim University, Unaizah, SAU.
Aldamigh OA; College of Medicine, King Faisal University, Dammam, SAU.
Halawani IR; Faculty of Medicine, King Abdulaziz University, Jeddah, SAU.
K Zino Alarki SM; Department of Surgery, King Khalid University Hospital, Riyadh, SAU.

Cureus ; 15(9): e45043, 2023 Sep.

Article in En | MEDLINE | ID: mdl-37829968

ABSTRACT

Background The application of artificial intelligence (AI) in education is undergoing rapid advancements, with models such as ChatGPT-4 showing potential in medical education. This study aims to evaluate the proficiency of ChatGPT-4 in answering Saudi Medical Licensing Exam (SMLE) questions. Methodology A dataset of 220 questions across four medical disciplines was used. The model was trained using a specific code to answer the questions accurately, and its performance was assessed using key performance indicators, difficulty level, and exam sections. Results ChatGPT-4 demonstrated an overall accuracy of 88.6%. It showed high proficiency with Easy and Average questions, but accuracy decreased for Hard questions. Performance was consistent across all disciplines, indicating a broad knowledge base. However, an error analysis revealed areas for further refinement, particularly with category (Option) A questions across all sections. Conclusions This study underscores the potential of ChatGPT-4 as an AI-assisted tool in medical education, demonstrating high proficiency in answering SMLE questions. Future research is recommended to expand the scope of training and evaluation as well as to enhance the model's performance on complex clinical questions.

Key words

ai in healthcare; artificial intelligence; chatgpt-4; medical education; saudi medical licensing exam; smle; standardized medical exam

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google

Full text: 1 Collection: 01-internacional Database: MEDLINE Language: En Journal: Cureus Year: 2023 Document type: Article Country of publication: United States

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google

Full text: 1 Collection: 01-internacional Database: MEDLINE Language: En Journal: Cureus Year: 2023 Document type: Article Country of publication: United States