Your browser doesn't support javascript.
loading
MeSH: a window into full text for document summarization.
Bhattacharya, Sanmitra; Ha-Thuc, Viet; Srinivasan, Padmini.
Affiliation
  • Bhattacharya S; Department of Computer Science, The University of Iowa, Iowa City, IA 52242, USA. sanmitra-bhattacharya@uiowa.edu
Bioinformatics ; 27(13): i120-8, 2011 Jul 01.
Article in En | MEDLINE | ID: mdl-21685060
ABSTRACT
MOTIVATION Previous research in the biomedical text-mining domain has historically been limited to titles, abstracts and metadata available in MEDLINE records. Recent research initiatives such as TREC Genomics and BioCreAtIvE strongly point to the merits of moving beyond abstracts and into the realm of full texts. Full texts are, however, more expensive to process not only in terms of resources needed but also in terms of accuracy. Since full texts contain embellishments that elaborate, contextualize, contrast, supplement, etc., there is greater risk for false positives. Motivated by this, we explore an approach that offers a compromise between the extremes of abstracts and full texts. Specifically, we create reduced versions of full text documents that contain only important portions. In the long-term, our goal is to explore the use of such summaries for functions such as document retrieval and information extraction. Here, we focus on designing summarization strategies. In particular, we explore the use of MeSH terms, manually assigned to documents by trained annotators, as clues to select important text segments from the full text documents.

RESULTS:

Our experiments confirm the ability of our approach to pick the important text portions. Using the ROUGE measures for evaluation, we were able to achieve maximum ROUGE-1, ROUGE-2 and ROUGE-SU4 F-scores of 0.4150, 0.1435 and 0.1782, respectively, for our MeSH term-based method versus the maximum baseline scores of 0.3815, 0.1353 and 0.1428, respectively. Using a MeSH profile-based strategy, we were able to achieve maximum ROUGE F-scores of 0.4320, 0.1497 and 0.1887, respectively. Human evaluation of the baselines and our proposed strategies further corroborates the ability of our method to select important sentences from the full texts. CONTACT sanmitra-bhattacharya@uiowa.edu; padmini-srinivasan@uiowa.edu.
Subject(s)

Full text: 1 Collection: 01-internacional Database: MEDLINE Main subject: Information Storage and Retrieval / Medical Subject Headings Country/Region as subject: America do norte Language: En Journal: Bioinformatics Journal subject: INFORMATICA MEDICA Year: 2011 Type: Article Affiliation country: United States

Full text: 1 Collection: 01-internacional Database: MEDLINE Main subject: Information Storage and Retrieval / Medical Subject Headings Country/Region as subject: America do norte Language: En Journal: Bioinformatics Journal subject: INFORMATICA MEDICA Year: 2011 Type: Article Affiliation country: United States