Search | VHL Regional Portal

GenPipes: an open-source framework for distributed and scalable genomic analyses.

Bourgey, Mathieu; Dali, Rola; Eveleigh, Robert; Chen, Kuang Chung; Letourneau, Louis; Fillon, Joel; Michaud, Marc; Caron, Maxime; Sandoval, Johanna; Lefebvre, Francois; Leveque, Gary; Mercier, Eloi; Bujold, David; Marquis, Pascale; Van, Patrick Tran; Anderson de Lima Morais, David; Tremblay, Julien; Shao, Xiaojian; Henrion, Edouard; Gonzalez, Emmanuel; Quirion, Pierre-Olivier; Caron, Bryan; Bourque, Guillaume.

Gigascience ; 8(6)2019 06 01.

Article in English | MEDLINE | ID: mdl-31185495

ABSTRACT

BACKGROUND: With the decreasing cost of sequencing and the rapid developments in genomics technologies and protocols, the need for validated bioinformatics software that enables efficient large-scale data processing is growing. FINDINGS: Here we present GenPipes, a flexible Python-based framework that facilitates the development and deployment of multi-step workflows optimized for high-performance computing clusters and the cloud. GenPipes already implements 12 validated and scalable pipelines for various genomics applications, including RNA sequencing, chromatin immunoprecipitation sequencing, DNA sequencing, methylation sequencing, Hi-C, capture Hi-C, metagenomics, and Pacific Biosciences long-read assembly. The software is available under a GPLv3 open source license and is continuously updated to follow recent advances in genomics and bioinformatics. The framework has already been configured on several servers, and a Docker image is also available to facilitate additional installations. CONCLUSIONS: GenPipes offers genomics researchers a simple method to analyze different types of data, customizable to their needs and resources, as well as the flexibility to create their own workflows.

Subject(s)

Genomics/methods , Software , DNA Methylation , Epigenomics/methods , Humans , Metagenomics/methods , Sequence Analysis, DNA/methods , Sequence Analysis, RNA/methods

Design of a data model for developing laboratory information management and analysis systems for protein production.

Pajon, Anne; Ionides, John; Diprose, Jon; Fillon, Joël; Fogh, Rasmus; Ashton, Alun W; Berman, Helen; Boucher, Wayne; Cygler, Miroslaw; Deleury, Emeline; Esnouf, Robert; Janin, Joël; Kim, Rosalind; Krimm, Isabelle; Lawson, Catherine L; Oeuillet, Eric; Poupon, Anne; Raymond, Stéphane; Stevens, Tim; van Tilbeurgh, Herman; Westbrook, John; Wood, Peter; Ulrich, Eldon; Vranken, Wim; Xueli, Li; Laue, Ernest; Stuart, David I; Henrick, Kim.

Proteins ; 58(2): 278-84, 2005 Feb 01.

Article in English | MEDLINE | ID: mdl-15562521

ABSTRACT

Data management has emerged as one of the central issues in the high-throughput processes of taking a protein target sequence through to a protein sample. To simplify this task, and following extensive consultation with the international structural genomics community, we describe here a model of the data related to protein production. The model is suitable for both large and small facilities for use in tracking samples, experiments, and results through the many procedures involved. The model is described in Unified Modeling Language (UML). In addition, we present relational database schemas derived from the UML. These relational schemas are already in use in a number of data management projects.

Subject(s)

Genomics/methods , Protein Engineering/methods , Proteins/chemistry , Proteomics/methods , Algorithms , Amino Acid Sequence , Data Interpretation, Statistical , Databases, Protein , Internet , Models, Biological , Programming Languages , Research , Software , Software Design , Systems Biology , Unified Medical Language System

ABSTRACT

Subject(s)

ABSTRACT

Subject(s)

SEND TO:

SELECTION OF CITATIONS

SEARCH DETAIL