Your browser doesn't support javascript.
loading
A detailed workflow to develop QIIME2-formatted reference databases for taxonomic analysis of DNA metabarcoding data.
Dubois, Benjamin; Debode, Frédéric; Hautier, Louis; Hulin, Julie; Martin, Gilles San; Delvaux, Alain; Janssen, Eric; Mingeot, Dominique.
Afiliación
  • Dubois B; Life Sciences Department, Bioengineering Unit, Walloon Agricultural Research Center, Chaussée de Charleroi 234, 5030, Gembloux, Belgium. b.dubois@cra.wallonie.be.
  • Debode F; Life Sciences Department, Bioengineering Unit, Walloon Agricultural Research Center, Chaussée de Charleroi 234, 5030, Gembloux, Belgium.
  • Hautier L; Life Sciences Department, Plant and Forest Health Unit, Walloon Agricultural Research Center, Rue de Liroux 2, 5030, Gembloux, Belgium.
  • Hulin J; Knowledge and Valorization of Agricultural Products Department, Quality and Authentication Unit, Walloon Agricultural Research Center, Chaussée de Namur 24, 5030, Gembloux, Belgium.
  • Martin GS; Life Sciences Department, Plant and Forest Health Unit, Walloon Agricultural Research Center, Rue de Liroux 2, 5030, Gembloux, Belgium.
  • Delvaux A; Knowledge and Valorization of Agricultural Products Department, Protection, Control Products and Residues Unit, Walloon Agricultural Research Center, Rue du Bordia 11, 5030, Gembloux, Belgium.
  • Janssen E; Knowledge and Valorization of Agricultural Products Department, Quality and Authentication Unit, Walloon Agricultural Research Center, Chaussée de Namur 24, 5030, Gembloux, Belgium.
  • Mingeot D; Life Sciences Department, Bioengineering Unit, Walloon Agricultural Research Center, Chaussée de Charleroi 234, 5030, Gembloux, Belgium.
BMC Genom Data ; 23(1): 53, 2022 07 08.
Article en En | MEDLINE | ID: mdl-35804326
ABSTRACT

BACKGROUND:

The DNA metabarcoding approach has become one of the most used techniques to study the taxa composition of various sample types. To deal with the high amount of data generated by the high-throughput sequencing process, a bioinformatics workflow is required and the QIIME2 platform has emerged as one of the most reliable and commonly used. However, only some pre-formatted reference databases dedicated to a few barcode sequences are available to assign taxonomy. If users want to develop a new custom reference database, several bottlenecks still need to be addressed and a detailed procedure explaining how to develop and format such a database is currently missing. In consequence, this work is aimed at presenting a detailed workflow explaining from start to finish how to develop such a curated reference database for any barcode sequence.

RESULTS:

We developed DB4Q2, a detailed workflow that allowed development of plant reference databases dedicated to ITS2 and rbcL, two commonly used barcode sequences in plant metabarcoding studies. This workflow addresses several of the main bottlenecks connected with the development of a curated reference database. The detailed and commented structure of DB4Q2 offers the possibility of developing reference databases even without extensive bioinformatics skills, and avoids 'black box' systems that are sometimes encountered. Some filtering steps have been included to discard presumably fungal and misidentified sequences. The flexible character of DB4Q2 allows several key sequence processing steps to be included or not, and downloading issues can be avoided. Benchmarking the databases developed using DB4Q2 revealed that they performed well compared to previously published reference datasets.

CONCLUSION:

This study presents DB4Q2, a detailed procedure to develop custom reference databases in order to carry out taxonomic analyses with QIIME2, but also with other bioinformatics platforms if desired. This work also provides ready-to-use plant ITS2 and rbcL databases for which the prediction accuracy has been assessed and compared to that of other published databases.
Asunto(s)
Palabras clave

Texto completo: 1 Colección: 01-internacional Banco de datos: MEDLINE Asunto principal: ADN / Código de Barras del ADN Taxonómico Tipo de estudio: Prognostic_studies Idioma: En Revista: BMC Genom Data Año: 2022 Tipo del documento: Article País de afiliación: Bélgica

Texto completo: 1 Colección: 01-internacional Banco de datos: MEDLINE Asunto principal: ADN / Código de Barras del ADN Taxonómico Tipo de estudio: Prognostic_studies Idioma: En Revista: BMC Genom Data Año: 2022 Tipo del documento: Article País de afiliación: Bélgica