Your browser doesn't support javascript.
loading
Mostrar: 20 | 50 | 100
Resultados 1 - 2 de 2
Filtrar
Mais filtros








Base de dados
Intervalo de ano de publicação
1.
bioRxiv ; 2024 Jun 12.
Artigo em Inglês | MEDLINE | ID: mdl-38915693

RESUMO

Background: Variant Call Format (VCF) is the standard file format for interchanging genetic variation data and associated quality control metrics. The usual row-wise encoding of the VCF data model (either as text or packed binary) emphasises efficient retrieval of all data for a given variant, but accessing data on a field or sample basis is inefficient. Biobank scale datasets currently available consist of hundreds of thousands of whole genomes and hundreds of terabytes of compressed VCF. Row-wise data storage is fundamentally unsuitable and a more scalable approach is needed. Results: We present the VCF Zarr specification, an encoding of the VCF data model using Zarr which makes retrieving subsets of the data much more efficient. Zarr is a cloud-native format for storing multi-dimensional data, widely used in scientific computing. We show how this format is far more efficient than standard VCF based approaches, and competitive with specialised methods for storing genotype data in terms of compression ratios and calculation performance. We demonstrate the VCF Zarr format (and the vcf2zarr conversion utility) on a subset of the Genomics England aggV2 dataset comprising 78,195 samples and 59,880,903 variants, with a 5X reduction in storage and greater than 300X reduction in CPU usage in some representative benchmarks. Conclusions: Large row-encoded VCF files are a major bottleneck for current research, and storing and processing these files incurs a substantial cost. The VCF Zarr specification, building on widely-used, open-source technologies has the potential to greatly reduce these costs, and may enable a diverse ecosystem of next-generation tools for analysing genetic variation data directly from cloud-based object stores.

2.
Planta ; 259(3): 61, 2024 Feb 06.
Artigo em Inglês | MEDLINE | ID: mdl-38319406

RESUMO

MAIN CONCLUSION: Agrobacterium-mediated transformation of Nicotiana tabacum, using an intragenic T-DNA region derived entirely from the N. tabacum genome, results in the equivalence of micro-translocations within genomes. Intragenic Agrobacterium-mediated gene transfer was achieved in Nicotiana tabacum using a T-DNA composed entirely of N. tabacum DNA, including T-DNA borders and the acetohydroxyacid synthase gene conferring resistance to sulfonylurea herbicides. Genomic analysis of a resulting plant, with single locus inheritance of herbicide resistance, identified a single insertion of the intragenic T-DNA on chromosome 5. The insertion event was composed of three N. tabacum DNA fragments from other chromosomes, as assembled on the T-DNA vector. This validates that intragenic transformation of plants can mimic micro-translocations within genomes, with the absence of foreign DNA.


Assuntos
Acetolactato Sintase , Rearranjo Gênico , Translocação Genética , DNA , Agrobacterium/genética , Nicotiana/genética
SELEÇÃO DE REFERÊNCIAS
DETALHE DA PESQUISA