Búsqueda | BVS Bolivia

Correction to: Recommendations for performance optimizations when using GATK3.8 and GATK4.

Heldenbrand, Jacob R; Baheti, Saurabh; Bockol, Matthew A; Drucker, Travis M; Hart, Steven N; Hudson, Matthew E; Iyer, Ravishankar K; Kalmbach, Michael T; Kendig, Katherine I; Klee, Eric W; Mattson, Nathan R; Wieben, Eric D; Wiepert, Mathieu; Wildman, Derek E; Mainzer, Liudmila S.

BMC Bioinformatics ; 20(1): 722, 2019 12 17.

Artículo en Inglés | MEDLINE | ID: mdl-31847808

RESUMEN

Following publication of the original article [1], the author explained that Table 2 is displayed incorrectly. The correct Table 2 is given below. The original article has been corrected.

Recommendations for performance optimizations when using GATK3.8 and GATK4.

BMC Bioinformatics ; 20(1): 557, 2019 Nov 08.

Artículo en Inglés | MEDLINE | ID: mdl-31703611

RESUMEN

BACKGROUND: Use of the Genome Analysis Toolkit (GATK) continues to be the standard practice in genomic variant calling in both research and the clinic. Recently the toolkit has been rapidly evolving. Significant computational performance improvements have been introduced in GATK3.8 through collaboration with Intel in 2017. The first release of GATK4 in early 2018 revealed rewrites in the code base, as the stepping stone toward a Spark implementation. As the software continues to be a moving target for optimal deployment in highly productive environments, we present a detailed analysis of these improvements, to help the community stay abreast with changes in performance. RESULTS: We re-evaluated multiple options, such as threading, parallel garbage collection, I/O options and data-level parallelization. Additionally, we considered the trade-offs of using GATK3.8 and GATK4. We found optimized parameter values that reduce the time of executing the best practices variant calling procedure by 29.3% for GATK3.8 and 16.9% for GATK4. Further speedups can be accomplished by splitting data for parallel analysis, resulting in run time of only a few hours on whole human genome sequenced to the depth of 20X, for both versions of GATK. Nonetheless, GATK4 is already much more cost-effective than GATK3.8. Thanks to significant rewrites of the algorithms, the same analysis can be run largely in a single-threaded fashion, allowing users to process multiple samples on the same CPU. CONCLUSIONS: In time-sensitive situations, when a patient has a critical or rapidly developing condition, it is useful to minimize the time to process a single sample. In such cases we recommend using GATK3.8 by splitting the sample into chunks and computing across multiple nodes. The resultant walltime will be nnn.4 hours at the cost of $41.60 on 4 c5.18xlarge instances of Amazon Cloud. For cost-effectiveness of routine analyses or for large population studies, it is useful to maximize the number of samples processed per unit time. Thus we recommend GATK4, running multiple samples on one node. The total walltime will be â¼34.1 hours on 40 samples, with 1.18 samples processed per hour at the cost of $2.60 per sample on c5.18xlarge instance of Amazon Cloud.

Asunto(s)

Genómica/métodos , Programas Informáticos , Algoritmos , Cromosomas Humanos/genética , Genoma Humano , Haplotipos/genética , Secuenciación de Nucleótidos de Alto Rendimiento , Humanos

Sentieon DNASeq Variant Calling Workflow Demonstrates Strong Computational Performance and Accuracy.

Kendig, Katherine I; Baheti, Saurabh; Bockol, Matthew A; Drucker, Travis M; Hart, Steven N; Heldenbrand, Jacob R; Hernaez, Mikel; Hudson, Matthew E; Kalmbach, Michael T; Klee, Eric W; Mattson, Nathan R; Ross, Christian A; Taschuk, Morgan; Wieben, Eric D; Wiepert, Mathieu; Wildman, Derek E; Mainzer, Liudmila S.

Front Genet ; 10: 736, 2019.

Artículo en Inglés | MEDLINE | ID: mdl-31481971

RESUMEN

As reliable, efficient genome sequencing becomes ubiquitous, the need for similarly reliable and efficient variant calling becomes increasingly important. The Genome Analysis Toolkit (GATK), maintained by the Broad Institute, is currently the widely accepted standard for variant calling software. However, alternative solutions may provide faster variant calling without sacrificing accuracy. One such alternative is Sentieon DNASeq, a toolkit analogous to GATK but built on a highly optimized backend. We conducted an independent evaluation of the DNASeq single-sample variant calling pipeline in comparison to that of GATK. Our results support the near-identical accuracy of the two software packages, showcase optimal scalability and great speed from Sentieon, and describe computational performance considerations for the deployment of DNASeq.

BBBomics-Human Blood Brain Barrier Transcriptomics Hub.

Kalari, Krishna R; Thompson, Kevin J; Nair, Asha A; Tang, Xiaojia; Bockol, Matthew A; Jhawar, Navya; Swaminathan, Suresh K; Lowe, Val J; Kandimalla, Karunya K.

Front Neurosci ; 10: 71, 2016.

Artículo en Inglés | MEDLINE | ID: mdl-26973449

MAP-RSeq: Mayo Analysis Pipeline for RNA sequencing.

Kalari, Krishna R; Nair, Asha A; Bhavsar, Jaysheel D; O'Brien, Daniel R; Davila, Jaime I; Bockol, Matthew A; Nie, Jinfu; Tang, Xiaojia; Baheti, Saurabh; Doughty, Jay B; Middha, Sumit; Sicotte, Hugues; Thompson, Aubrey E; Asmann, Yan W; Kocher, Jean-Pierre A.

BMC Bioinformatics ; 15: 224, 2014 Jun 27.

Artículo en Inglés | MEDLINE | ID: mdl-24972667

RESUMEN

BACKGROUND: Although the costs of next generation sequencing technology have decreased over the past years, there is still a lack of simple-to-use applications, for a comprehensive analysis of RNA sequencing data. There is no one-stop shop for transcriptomic genomics. We have developed MAP-RSeq, a comprehensive computational workflow that can be used for obtaining genomic features from transcriptomic sequencing data, for any genome. RESULTS: For optimization of tools and parameters, MAP-RSeq was validated using both simulated and real datasets. MAP-RSeq workflow consists of six major modules such as alignment of reads, quality assessment of reads, gene expression assessment and exon read counting, identification of expressed single nucleotide variants (SNVs), detection of fusion transcripts, summarization of transcriptomics data and final report. This workflow is available for Human transcriptome analysis and can be easily adapted and used for other genomes. Several clinical and research projects at the Mayo Clinic have applied the MAP-RSeq workflow for RNA-Seq studies. The results from MAP-RSeq have thus far enabled clinicians and researchers to understand the transcriptomic landscape of diseases for better diagnosis and treatment of patients. CONCLUSIONS: Our software provides gene counts, exon counts, fusion candidates, expressed single nucleotide variants, mapping statistics, visualizations, and a detailed research data report for RNA-Seq. The workflow can be executed on a standalone virtual machine or on a parallel Sun Grid Engine cluster. The software can be downloaded from http://bioinformaticstools.mayo.edu/research/maprseq/.

Asunto(s)

Perfilación de la Expresión Génica , Genómica/métodos , Instituciones de Salud , Secuenciación de Nucleótidos de Alto Rendimiento/métodos , Análisis de Secuencia de ARN/métodos , Programas Informáticos , Secuencia de Bases , Exones/genética , Humanos

RESUMEN

RESUMEN

Asunto(s)

RESUMEN

RESUMEN

Asunto(s)

ENVIAR RESULTADO:

SELECCIÓN DE REFERENCIAS

DETALLE DE LA BÚSQUEDA