AlpaPICO: Extraction of PICO frames from clinical trial documents using LLMs.

Ghosh, Madhusudan; Mukherjee, Shrimon; Ganguly, Asmit; Basuchowdhuri, Partha; Naskar, Sudip Kumar; Ganguly, Debasis

Ghosh, Madhusudan; Mukherjee, Shrimon; Ganguly, Asmit; Basuchowdhuri, Partha; Naskar, Sudip Kumar; Ganguly, Debasis.

Afiliação

Ghosh M; School of Mathematical and Computational Sciences, Indian Association for the Cultivation of Science, India. Electronic address: madhusuda.iacs@gmail.com.
Mukherjee S; School of Mathematical and Computational Sciences, Indian Association for the Cultivation of Science, India. Electronic address: iacsshrimon@gmail.com.
Ganguly A; Computer Science and Engineering, Indian Institute of Technology, Patna, India. Electronic address: asmitganguly.personal@gmail.com.
Basuchowdhuri P; School of Mathematical and Computational Sciences, Indian Association for the Cultivation of Science, India. Electronic address: partha.basuchowdhuri@iacs.res.in.
Naskar SK; Department of Computer Science and Engineering, Jadavpur University, India. Electronic address: sudip.naskar@gmail.com.
Ganguly D; School of Computing Science, University of Glasgow, United Kingdom of Great Britain and Northern Ireland. Electronic address: Debasis.Ganguly@glasgow.ac.uk.

Methods ; 226: 78-88, 2024 Jun.

Article em En | MEDLINE | ID: mdl-38643910

ABSTRACT

ABSTRACT

In recent years, there has been a surge in the publication of clinical trial reports, making it challenging to conduct systematic reviews. Automatically extracting Population, Intervention, Comparator, and Outcome (PICO) from clinical trial studies can alleviate the traditionally time-consuming process of manually scrutinizing systematic reviews. Existing approaches of PICO frame extraction involves supervised approach that relies on the existence of manually annotated data points in the form of BIO label tagging. Recent approaches, such as In-Context Learning (ICL), which has been shown to be effective for a number of downstream NLP tasks, require the use of labeled examples. In this work, we adopt ICL strategy by employing the pretrained knowledge of Large Language Models (LLMs), gathered during the pretraining phase of an LLM, to automatically extract the PICO-related terminologies from clinical trial documents in unsupervised set up to bypass the availability of large number of annotated data instances. Additionally, to showcase the highest effectiveness of LLM in oracle scenario where large number of annotated samples are available, we adopt the instruction tuning strategy by employing Low Rank Adaptation (LORA) to conduct the training of gigantic model in low resource environment for the PICO frame extraction task. More specifically, both of the proposed frameworks utilize AlpaCare as base LLM which employs both few-shot in-context learning and instruction tuning techniques to extract PICO-related terms from the clinical trial reports. We applied these approaches to the widely used coarse-grained datasets such as EBM-NLP, EBM-COMET and fine-grained datasets such as EBM-NLPrev and EBM-NLPh. Our empirical results show that our proposed ICL-based framework produces comparable results on all the version of EBM-NLP datasets and the proposed instruction tuned version of our framework produces state-of-the-art results on all the different EBM-NLP datasets. Our project is available at https//github.com/shrimonmuke0202/AlpaPICO.git.

Assuntos

Ensaios Clínicos como Assunto; Processamento de Linguagem Natural; Humanos; Ensaios Clínicos como Assunto/métodos; Mineração de Dados/métodos; Aprendizado de Máquina

Palavras-chave

Bio-medical NER; In-context learning; Instruction tuning; LLM; Llama; PICO frame extraction

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Assunto principal: Processamento de Linguagem Natural / Ensaios Clínicos como Assunto Limite: Humans Idioma: En Revista: Methods Assunto da revista: BIOQUIMICA Ano de publicação: 2024 Tipo de documento: Article

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google