A biomedically oriented automatically annotated Twitter COVID-19 dataset.

Hernandez, Luis Alberto Robles; Callahan, Tiffany J; Banda, Juan M

Hernandez, Luis Alberto Robles; Callahan, Tiffany J; Banda, Juan M.

Afiliação

Hernandez LAR; Department of Computer Science, Georgia State University, Atlanta, GA 30303, USA.
Callahan TJ; Computational Bioscience Program, University of Colorado Anschutz Medical Campus, Aurora, CO 80045, USA.
Banda JM; Department of Computer Science, Georgia State University, Atlanta, GA 30303, USA.

Genomics Inform ; 19(3): e21, 2021 Sep.

Article em En | MEDLINE | ID: mdl-34638168

RESUMO

The use of social media data, like Twitter, for biomedical research has been gradually increasing over the years. With the coronavirus disease 2019 (COVID-19) pandemic, researchers have turned to more non-traditional sources of clinical data to characterize the disease in near-real time, study the societal implications of interventions, as well as the sequelae that recovered COVID-19 cases present. However, manually curated social media datasets are difficult to come by due to the expensive costs of manual annotation and the efforts needed to identify the correct texts. When datasets are available, they are usually very small and their annotations don't generalize well over time or to larger sets of documents. As part of the 2021 Biomedical Linked Annotation Hackathon, we release our dataset of over 120 million automatically annotated tweets for biomedical research purposes. Incorporating best-practices, we identify tweets with potentially high clinical relevance. We evaluated our work by comparing several SpaCy-based annotation frameworks against a manually annotated gold-standard dataset. Selecting the best method to use for automatic annotation, we then annotated 120 million tweets and released them publicly for future downstream usage within the biomedical domain.

Palavras-chave

COVID-19; biomedical annotations; datasets; social media data

Texto completo

Adicionar na Minha BVS

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Tipo de estudo: Guideline Idioma: En Revista: Genomics Inform Ano de publicação: 2021 Tipo de documento: Article País de afiliação: Estados Unidos País de publicação: Coréia do Sul

Texto completo

Adicionar na Minha BVS

Imprimir

XML

PubMed Links

Buscar no Google