Pesquisa | Biblioteca Virtual em Saúde

1.

Reactivation of a transplant recipient's inherited human herpesvirus 6 and implications to the graft.

Hannolainen, Leo; Pyöriä, Lari; Pratas, Diogo; Lohi, Jouko; Skuja, Sandra; Rasa-Dzelzkaleja, Santa; Murovska, Modra; Hedman, Klaus; Jahnukainen, Timo; Perdomo, Maria Fernanda.

J Infect Dis ; 2024 May 20.

Artigo em Inglês | MEDLINE | ID: mdl-38768311

RESUMO

BACKGROUND: The implications of inherited chromosomally integrated human herpesvirus 6 (iciHHV-6) in solid organ transplantation remain uncertain. Although this trait has been linked to unfavorable clinical outcomes, an association between viral reactivation and complications has only been conclusively established in a few cases. In contrast to these studies, which followed donor-derived transmission, our investigation is the first to examine the pathogenicity of a recipient´s iciHHV-6B and its impact on the graft. METHODS: We used hybrid capture sequencing for in-depth analysis of the viral sequences reconstructed from sequential liver biopsies. Moreover, we investigated viral replication through in situ hybridization (U38-U94 genes), real-time PCR (U89/U90 genes), immunohistochemistry, and immunofluorescence (against viral lysate). We also performed whole transcriptome sequencing of the liver biopsies to profile the host immune response. RESULTS: We report a case of reactivation of a recipient´s iciHHV-6B and subsequent infection of the graft. Using a novel approach integrating the analysis of viral and mitochondrial DNAs, we located the iciHHV-6B intra-graft. We demonstrated active replication via the emergence of viral minor variants across time points, in addition to positive viral mRNAs and antigen stainings in tissue sections. Furthermore, we detected significant upregulation of cell surface molecules, transcription factors, and cytokines associated with antiviral immune responses, arguing against immunotolerance. CONCLUSIONS: Our analysis underscores the potential pathological impact of iciHHV-6B, emphasizing the need for close monitoring of reactivation in transplant recipients. Most crucially, it highlights the critical role that the host's virome can play in shaping the outcome of transplantation, urging further investigations.

2.

Mapping human pathogens in wastewater using a metatranscriptomic approach.

Carneiro, João; Pascoal, Francisco; Semedo, Miguel; Pratas, Diogo; Tomasino, Maria Paola; Rego, Adriana; Carvalho, Maria de Fátima; Mucha, Ana Paula; Magalhães, Catarina.

Environ Res ; 231(Pt 1): 116040, 2023 Aug 15.

Artigo em Inglês | MEDLINE | ID: mdl-37150387

RESUMO

The monitoring of cities' wastewaters for the detection of potentially pathogenic viruses and bacteria has been considered a priority during the COVID-19 pandemic to monitor public health in urban environments. The methodological approaches frequently used for this purpose include deoxyribonucleic acid (DNA)/Ribonucleic acid (RNA) isolation followed by quantitative polymerase chain reaction (qPCR) and reverse transcription (RT)âqPCR targeting pathogenic genes. More recently, the application of metatranscriptomic has opened opportunities to develop broad pathogenic monitoring workflows covering the entire pathogenic community within the sample. Nevertheless, the high amount of data generated in the process requires an appropriate analysis to detect the pathogenic community from the entire dataset. Here, an implementation of a bioinformatic workflow was developed to produce a map of the detected pathogenic bacteria and viruses in wastewater samples by analysing metatranscriptomic data. The main objectives of this work was the development of a computational methodology that can accurately detect both human pathogenic virus and bacteria in wastewater samples. This workflow can be easily reproducible with open-source software and uses efficient computational resources. The results showed that the used algorithms can predict potential human pathogens presence in the tested samples and that active forms of both bacteria and virus can be identified. By comparing the computational method implemented in this study to other state-of-the-art workflows, the implementation analysis was faster, while providing higher accuracy and sensitivity. Considering these results, the processes and methods to monitor wastewater for potential human pathogens can become faster and more accurate. The proposed workflow is available at https://github.com/waterpt/watermonitor and can be implemented in currently wastewater monitoring programs to ascertain the presence of potential human pathogenic species.

Assuntos

COVID-19 , Vírus , Humanos , Águas Residuárias , Pandemias , Vírus/genética , Bactérias/genética

3.

TargIDe: a machine-learning workflow for target identification of molecules with antibiofilm activity against Pseudomonas aeruginosa.

Carneiro, João; Magalhães, Rita P; de la Oliva Roque, Victor M; Simões, Manuel; Pratas, Diogo; Sousa, Sérgio F.

J Comput Aided Mol Des ; 37(5-6): 265-278, 2023 06.

Artigo em Inglês | MEDLINE | ID: mdl-37085636

RESUMO

Bacterial biofilms are a source of infectious human diseases and are heavily linked to antibiotic resistance. Pseudomonas aeruginosa is a multidrug-resistant bacterium widely present and implicated in several hospital-acquired infections. Over the last years, the development of new drugs able to inhibit Pseudomonas aeruginosa by interfering with its ability to form biofilms has become a promising strategy in drug discovery. Identifying molecules able to interfere with biofilm formation is difficult, but further developing these molecules by rationally improving their activity is particularly challenging, as it requires knowledge of the specific protein target that is inhibited. This work describes the development of a machine learning multitechnique consensus workflow to predict the protein targets of molecules with confirmed inhibitory activity against biofilm formation by Pseudomonas aeruginosa. It uses a specialized database containing all the known targets implicated in biofilm formation by Pseudomonas aeruginosa. The experimentally confirmed inhibitors available on ChEMBL, together with chemical descriptors, were used as the input features for a combination of nine different classification models, yielding a consensus method to predict the most likely target of a ligand. The implemented algorithm is freely available at https://github.com/BioSIM-Research-Group/TargIDe under licence GNU General Public Licence (GPL) version 3 and can easily be improved as more data become available.

Assuntos

Antibacterianos , Pseudomonas aeruginosa , Humanos , Antibacterianos/farmacologia , Fluxo de Trabalho , Biofilmes , Aprendizado de Máquina , Testes de Sensibilidade Microbiana

4.

Herpesviruses, polyomaviruses, parvoviruses, papillomaviruses, and anelloviruses in vestibular schwannoma.

Jauhiainen, Maria K; Mohanraj, Ushanandini; Lehecka, Martin; Niemelä, Mika; Hirvonen, Timo P; Pratas, Diogo; Perdomo, Maria F; Söderlund-Venermo, Maria; Mäkitie, Antti A; Sinkkonen, Saku T.

J Neurovirol ; 29(2): 226-231, 2023 04.

Artigo em Inglês | MEDLINE | ID: mdl-36857017

RESUMO

Etiology of vestibular schwannoma (VS) is unknown. Viruses can infect and reside in neural tissues for decades, and new viruses with unknown tumorigenic potential have been discovered. The presence of herpesvirus, polyomavirus, parvovirus, and anellovirus DNA was analyzed by quantitative PCR in 46 formalin-fixed paraffin-embedded VS samples. Five samples were analyzed by targeted next-generation sequencing. Viral DNA was detected altogether in 24/46 (52%) tumor samples, mostly representing anelloviruses (46%). Our findings show frequent persistence of anelloviruses, considered normal virome, in VS. None of the other viruses showed an extensive presence, thereby suggesting insignificant role in VS.

Assuntos

Anelloviridae , Herpesviridae , Neuroma Acústico , Parvovirus , Polyomavirus , Humanos , Polyomavirus/genética , Anelloviridae/genética , Neuroma Acústico/genética , Herpesviridae/genética , Parvovirus/genética , DNA Viral/genética

5.

Unmasking the tissue-resident eukaryotic DNA virome in humans.

Pyöriä, Lari; Pratas, Diogo; Toppinen, Mari; Hedman, Klaus; Sajantila, Antti; Perdomo, Maria F.

Nucleic Acids Res ; 51(7): 3223-3239, 2023 04 24.

Artigo em Inglês | MEDLINE | ID: mdl-36951096

RESUMO

Little is known on the landscape of viruses that reside within our cells, nor on the interplay with the host imperative for their persistence. Yet, a lifetime of interactions conceivably have an imprint on our physiology and immune phenotype. In this work, we revealed the genetic make-up and unique composition of the known eukaryotic human DNA virome in nine organs (colon, liver, lung, heart, brain, kidney, skin, blood, hair) of 31 Finnish individuals. By integration of quantitative (qPCR) and qualitative (hybrid-capture sequencing) analysis, we identified the DNAs of 17 species, primarily herpes-, parvo-, papilloma- and anello-viruses (>80% prevalence), typically persisting in low copies (mean 540 copies/ million cells). We assembled in total 70 viral genomes (>90% breadth coverage), distinct in each of the individuals, and identified high sequence homology across the organs. Moreover, we detected variations in virome composition in two individuals with underlying malignant conditions. Our findings reveal unprecedented prevalences of viral DNAs in human organs and provide a fundamental ground for the investigation of disease correlates. Our results from post-mortem tissues call for investigation of the crosstalk between human DNA viruses, the host, and other microbes, as it predictably has a significant impact on our health.

Assuntos

DNA Viral , Genoma Humano , Vírus , Humanos , DNA Viral/genética , DNA Viral/análise , Eucariotos/genética , Viroma , Vírus/genética , Especificidade de Órgãos

6.

The complexity landscape of viral genomes.

Silva, Jorge Miguel; Pratas, Diogo; Caetano, Tânia; Matos, Sérgio.

Gigascience ; 112022 08 11.

Artigo em Inglês | MEDLINE | ID: mdl-35950839

RESUMO

BACKGROUND: Viruses are among the shortest yet highly abundant species that harbor minimal instructions to infect cells, adapt, multiply, and exist. However, with the current substantial availability of viral genome sequences, the scientific repertory lacks a complexity landscape that automatically enlights viral genomes' organization, relation, and fundamental characteristics. RESULTS: This work provides a comprehensive landscape of the viral genome's complexity (or quantity of information), identifying the most redundant and complex groups regarding their genome sequence while providing their distribution and characteristics at a large and local scale. Moreover, we identify and quantify inverted repeats abundance in viral genomes. For this purpose, we measure the sequence complexity of each available viral genome using data compression, demonstrating that adequate data compressors can efficiently quantify the complexity of viral genome sequences, including subsequences better represented by algorithmic sources (e.g., inverted repeats). Using a state-of-the-art genomic compressor on an extensive viral genomes database, we show that double-stranded DNA viruses are, on average, the most redundant viruses while single-stranded DNA viruses are the least. Contrarily, double-stranded RNA viruses show a lower redundancy relative to single-stranded RNA. Furthermore, we extend the ability of data compressors to quantify local complexity (or information content) in viral genomes using complexity profiles, unprecedently providing a direct complexity analysis of human herpesviruses. We also conceive a features-based classification methodology that can accurately distinguish viral genomes at different taxonomic levels without direct comparisons between sequences. This methodology combines data compression with simple measures such as GC-content percentage and sequence length, followed by machine learning classifiers. CONCLUSIONS: This article presents methodologies and findings that are highly relevant for understanding the patterns of similarity and singularity between viral groups, opening new frontiers for studying viral genomes' organization while depicting the complexity trends and classification components of these genomes at different taxonomic levels. The whole study is supported by an extensive website (https://asilab.github.io/canvas/) for comprehending the viral genome characterization using dynamic and interactive approaches.

Assuntos

Genoma Viral , Vírus , Composição de Bases , Genômica/métodos , Humanos , Vírus/genética

7.

The haplotype-resolved chromosome pairs of a heterozygous diploid African cassava cultivar reveal novel pan-genome and allele-specific transcriptome features.

Qi, Weihong; Lim, Yi-Wen; Patrignani, Andrea; Schläpfer, Pascal; Bratus-Neuenschwander, Anna; Grüter, Simon; Chanez, Christelle; Rodde, Nathalie; Prat, Elisa; Vautrin, Sonia; Fustier, Margaux-Alison; Pratas, Diogo; Schlapbach, Ralph; Gruissem, Wilhelm.

Gigascience ; 112022 03 24.

Artigo em Inglês | MEDLINE | ID: mdl-35333302

RESUMO

BACKGROUND: Cassava (Manihot esculenta) is an important clonally propagated food crop in tropical and subtropical regions worldwide. Genetic gain by molecular breeding has been limited, partially because cassava is a highly heterozygous crop with a repetitive and difficult-to-assemble genome. FINDINGS: Here we demonstrate that Pacific Biosciences high-fidelity (HiFi) sequencing reads, in combination with the assembler hifiasm, produced genome assemblies at near complete haplotype resolution with higher continuity and accuracy compared to conventional long sequencing reads. We present 2 chromosome-scale haploid genomes phased with Hi-C technology for the diploid African cassava variety TME204. With consensus accuracy >QV46, contig N50 >18 Mb, BUSCO completeness of 99%, and 35k phased gene loci, it is the most accurate, continuous, complete, and haplotype-resolved cassava genome assembly so far. Ab initio gene prediction with RNA-seq data and Iso-Seq transcripts identified abundant novel gene loci, with enriched functionality related to chromatin organization, meristem development, and cell responses. During tissue development, differentially expressed transcripts of different haplotype origins were enriched for different functionality. In each tissue, 20-30% of transcripts showed allele-specific expression (ASE) differences. ASE bias was often tissue specific and inconsistent across different tissues. Direction-shifting was observed in <2% of the ASE transcripts. Despite high gene synteny, the HiFi genome assembly revealed extensive chromosome rearrangements and abundant intra-genomic and inter-genomic divergent sequences, with large structural variations mostly related to LTR retrotransposons. We use the reference-quality assemblies to build a cassava pan-genome and demonstrate its importance in representing the genetic diversity of cassava for downstream reference-guided omics analysis and breeding. CONCLUSIONS: The phased and annotated chromosome pairs allow a systematic view of the heterozygous diploid genome organization in cassava with improved accuracy, completeness, and haplotype resolution. They will be a valuable resource for cassava breeding and research. Our study may also provide insights into developing cost-effective and efficient strategies for resolving complex genomes with high resolution, accuracy, and continuity.

Assuntos

Manihot , Alelos , Cromossomos , Diploide , Haplótipos , Manihot/genética , Melhoramento Vegetal , Análise de Sequência de DNA , Transcriptoma

8.

Detection of Low-Copy Human Virus DNA upon Prolonged Formalin Fixation.

Mielonen, Outi I; Pratas, Diogo; Hedman, Klaus; Sajantila, Antti; Perdomo, Maria F.

Viruses ; 14(1)2022 01 12.

Artigo em Inglês | MEDLINE | ID: mdl-35062338

RESUMO

Formalin fixation, albeit an outstanding method for morphological and molecular preservation, induces DNA damage and cross-linking, which can hinder nucleic acid screening. This is of particular concern in the detection of low-abundance targets, such as persistent DNA viruses. In the present study, we evaluated the analytical sensitivity of viral detection in lung, liver, and kidney specimens from four deceased individuals. The samples were either frozen or incubated in formalin (±paraffin embedding) for up to 10 days. We tested two DNA extraction protocols for the control of efficient yields and viral detections. We used short-amplicon qPCRs (63-159 nucleotides) to detect 11 DNA viruses, as well as hybridization capture of these plus 27 additional ones, followed by deep sequencing. We observed marginally higher ratios of amplifiable DNA and scantly higher viral genoprevalences in the samples extracted with the FFPE dedicated protocol. Based on the findings in the frozen samples, most viruses were detected regardless of the extended fixation times. False-negative calls, particularly by qPCR, correlated with low levels of viral DNA (<250 copies/million cells) and longer PCR amplicons (>150 base pairs). Our data suggest that low-copy viral DNAs can be satisfactorily investigated from FFPE specimens, and encourages further examination of historical materials.

Assuntos

DNA Viral/isolamento & purificação , Formaldeído , Técnicas de Diagnóstico Molecular/métodos , Fixação de Tecidos/métodos , Humanos , Rim , Fígado , Pulmão , Hibridização de Ácido Nucleico , Reação em Cadeia da Polimerase em Tempo Real , Análise de Sequência de DNA

9.

AlcoR: alignment-free simulation, mapping, and visualization of low-complexity regions in biological data.

Silva, Jorge M; Qi, Weihong; Pinho, Armando J; Pratas, Diogo.

Gigascience ; 122022 Dec 28.

Artigo em Inglês | MEDLINE | ID: mdl-38091509

RESUMO

BACKGROUND: Low-complexity data analysis is the area that addresses the search and quantification of regions in sequences of elements that contain low-complexity or repetitive elements. For example, these can be tandem repeats, inverted repeats, homopolymer tails, GC-biased regions, similar genes, and hairpins, among many others. Identifying these regions is crucial because of their association with regulatory and structural characteristics. Moreover, their identification provides positional and quantity information where standard assembly methodologies face significant difficulties because of substantial higher depth coverage (mountains), ambiguous read mapping, or where sequencing or reconstruction defects may occur. However, the capability to distinguish low-complexity regions (LCRs) in genomic and proteomic sequences is a challenge that depends on the model's ability to find them automatically. Low-complexity patterns can be implicit through specific or combined sources, such as algorithmic or probabilistic, and recurring to different spatial distances-namely, local, medium, or distant associations. FINDINGS: This article addresses the challenge of automatically modeling and distinguishing LCRs, providing a new method and tool (AlcoR) for efficient and accurate segmentation and visualization of these regions in genomic and proteomic sequences. The method enables the use of models with different memories, providing the ability to distinguish local from distant low-complexity patterns. The method is reference and alignment free, providing additional methodologies for testing, including a highly flexible simulation method for generating biological sequences (DNA or protein) with different complexity levels, sequence masking, and a visualization tool for automatic computation of the LCR maps into an ideogram style. We provide illustrative demonstrations using synthetic, nearly synthetic, and natural sequences showing the high efficiency and accuracy of AlcoR. As large-scale results, we use AlcoR to unprecedentedly provide a whole-chromosome low-complexity map of a recent complete human genome and the haplotype-resolved chromosome pairs of a heterozygous diploid African cassava cultivar. CONCLUSIONS: The AlcoR method provides the ability of fast sequence characterization through data complexity analysis, ideally for scenarios entangling the presence of new or unknown sequences. AlcoR is implemented in C language using multithreading to increase the computational speed, is flexible for multiple applications, and does not contain external dependencies. The tool accepts any sequence in FASTA format. The source code is freely provided at https://github.com/cobilab/alcor.

10.

The Human Bone Marrow Is Host to the DNAs of Several Viruses.

Toppinen, Mari; Sajantila, Antti; Pratas, Diogo; Hedman, Klaus; Perdomo, Maria F.

Front Cell Infect Microbiol ; 11: 657245, 2021.

Artigo em Inglês | MEDLINE | ID: mdl-33968803

RESUMO

The long-term impact of viruses residing in the human bone marrow (BM) remains unexplored. However, chronic inflammatory processes driven by single or multiple viruses could significantly alter hematopoiesis and immune function. We performed a systematic analysis of the DNAs of 38 viruses in the BM. We detected, by quantitative PCRs and next-generation sequencing, viral DNA in 88.9% of the samples, up to five viruses in one individual. Included were, among others, several herpesviruses, hepatitis B virus, Merkel cell polyomavirus and, unprecedentedly, human papillomavirus 31. Given the reactivation and/or oncogenic potential of these viruses, their repercussion on hematopoietic and malignant disorders calls for careful examination. Furthermore, the implications of persistent infections on the engraftment, regenerative capacity, and outcomes of bone marrow transplantation deserve in-depth evaluation.

Assuntos

Transplante de Medula Óssea , Medula Óssea , DNA Viral , Vírus da Hepatite B , Sequenciamento de Nucleotídeos em Larga Escala , Humanos

11.

AC2: An Efficient Protein Sequence Compression Tool Using Artificial Neural Networks and Cache-Hash Models.

Silva, Milton; Pratas, Diogo; Pinho, Armando J.

Entropy (Basel) ; 23(5)2021 Apr 26.

Artigo em Inglês | MEDLINE | ID: mdl-33925812

RESUMO

Recently, the scientific community has witnessed a substantial increase in the generation of protein sequence data, triggering emergent challenges of increasing importance, namely efficient storage and improved data analysis. For both applications, data compression is a straightforward solution. However, in the literature, the number of specific protein sequence compressors is relatively low. Moreover, these specialized compressors marginally improve the compression ratio over the best general-purpose compressors. In this paper, we present AC2, a new lossless data compressor for protein (or amino acid) sequences. AC2 uses a neural network to mix experts with a stacked generalization approach and individual cache-hash memory models to the highest-context orders. Compared to the previous compressor (AC), we show gains of 2-9% and 6-7% in reference-free and reference-based modes, respectively. These gains come at the cost of three times slower computations. AC2 also improves memory usage against AC, with requirements about seven times lower, without being affected by the sequences' input size. As an analysis application, we use AC2 to measure the similarity between each SARS-CoV-2 protein sequence with each viral protein sequence from the whole UniProt database. The results consistently show higher similarity to the pangolin coronavirus, followed by the bat and human coronaviruses, contributing with critical results to a current controversial subject. AC2 is available for free download under GPLv3 license.

12.

A semi-automatic methodology for analysing distributed and private biobanks.

Almeida, João Rafael; Pratas, Diogo; Oliveira, José Luís.

Comput Biol Med ; 130: 104180, 2021 03.

Artigo em Inglês | MEDLINE | ID: mdl-33360272

RESUMO

Privacy issues limit the analysis and cross-exploration of most distributed and private biobanks, often raised by the multiple dimensionality and sensitivity of the data associated with access restrictions and policies. These characteristics prevent collaboration between entities, constituting a barrier to emergent personalized and public health challenges, namely the discovery of new druggable targets, identification of disease-causing genetic variants, or the study of rare diseases. In this paper, we propose a semi-automatic methodology for the analysis of distributed and private biobanks. The strategies involved in the proposed methodology efficiently enable the creation and execution of unified genomic studies using distributed repositories, without compromising the information present in the datasets. We apply the methodology to a case study in the current Covid-19, ensuring the combination of the diagnostics from multiple entities while maintaining privacy through a completely identical procedure. Moreover, we show that the methodology follows a simple, intuitive, and practical scheme.

Assuntos

Bancos de Espécimes Biológicos , COVID-19 , Saúde Pública , SARS-CoV-2 , Humanos

13.

Persistent minimal sequences of SARS-CoV-2.

Pratas, Diogo; Silva, Jorge M.

Bioinformatics ; 36(21): 5129-5132, 2021 01 29.

Artigo em Inglês | MEDLINE | ID: mdl-32730589

RESUMO

MOTIVATION: Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has caused more than 14 million cases and more than half million deaths. Given the absence of implemented therapies, new analysis, diagnosis and therapeutics are of great importance. RESULTS: Analysis of SARS-CoV-2 genomes from the current outbreak reveals the presence of short persistent DNA/RNA sequences that are absent from the human genome and transcriptome (PmRAWs). For the PmRAWs with length 12, only four exist at the same location in all SARS-CoV-2. At the gene level, we found one PmRAW of size 13 at the Spike glycoprotein coding sequence. This protein is fundamental for binding in human ACE2 and further use as an entry receptor to invade target cells. Applying protein structural prediction, we localized this PmRAW at the surface of the Spike protein, providing a potential targeted vector for diagnostics and therapeutics. In addition, we show a new pattern of relative absent words (RAWs), characterized by the progressive increase of GC content (Guanine and Cytosine) according to the decrease of RAWs length, contrarily to the virus and host genome distributions. New analysis shows the same property during the Ebola virus outbreak. At a computational level, we improved the alignment-free method to identify pathogen-specific signatures in balance with GC measures and removed previous size limitations. AVAILABILITY AND IMPLEMENTATION: https://github.com/cobilab/eagle. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Assuntos

COVID-19 , Glicoproteína da Espícula de Coronavírus , Humanos , Ligação Proteica , SARS-CoV-2

14.

Statistical Complexity Analysis of Turing Machine tapes with Fixed Algorithmic Complexity Using the Best-Order Markov Model.

Silva, Jorge M; Pinho, Eduardo; Matos, Sérgio; Pratas, Diogo.

Entropy (Basel) ; 22(1)2020 Jan 16.

Artigo em Inglês | MEDLINE | ID: mdl-33285880

RESUMO

Sources that generate symbolic sequences with algorithmic nature may differ in statistical complexity because they create structures that follow algorithmic schemes, rather than generating symbols from a probabilistic function assuming independence. In the case of Turing machines, this means that machines with the same algorithmic complexity can create tapes with different statistical complexity. In this paper, we use a compression-based approach to measure global and local statistical complexity of specific Turing machine tapes with the same number of states and alphabet. Both measures are estimated using the best-order Markov model. For the global measure, we use the Normalized Compression (NC), while, for the local measures, we define and use normal and dynamic complexity profiles to quantify and localize lower and higher regions of statistical complexity. We assessed the validity of our methodology on synthetic and real genomic data showing that it is tolerant to increasing rates of editions and block permutations. Regarding the analysis of the tapes, we localize patterns of higher statistical complexity in two regions, for a different number of machine states. We show that these patterns are generated by a decrease of the tape's amplitude, given the setting of small rule cycles. Additionally, we performed a comparison with a measure that uses both algorithmic and statistical approaches (BDM) for analysis of the tapes. Naturally, BDM is efficient given the algorithmic nature of the tapes. However, for a higher number of states, BDM is progressively approximated by our methodology. Finally, we provide a simple algorithm to increase the statistical complexity of a Turing machine tape while retaining the same algorithmic complexity. We supply a publicly available implementation of the algorithm in C++ language under the GPLv3 license. All results can be reproduced in full with scripts provided at the repository.

15.

Efficient DNA sequence compression with neural networks.

Silva, Milton; Pratas, Diogo; Pinho, Armando J.

Gigascience ; 9(11)2020 11 11.

Artigo em Inglês | MEDLINE | ID: mdl-33179040

RESUMO

BACKGROUND: The increasing production of genomic data has led to an intensified need for models that can cope efficiently with the lossless compression of DNA sequences. Important applications include long-term storage and compression-based data analysis. In the literature, only a few recent articles propose the use of neural networks for DNA sequence compression. However, they fall short when compared with specific DNA compression tools, such as GeCo2. This limitation is due to the absence of models specifically designed for DNA sequences. In this work, we combine the power of neural networks with specific DNA models. For this purpose, we created GeCo3, a new genomic sequence compressor that uses neural networks for mixing multiple context and substitution-tolerant context models. FINDINGS: We benchmark GeCo3 as a reference-free DNA compressor in 5 datasets, including a balanced and comprehensive dataset of DNA sequences, the Y-chromosome and human mitogenome, 2 compilations of archaeal and virus genomes, 4 whole genomes, and 2 collections of FASTQ data of a human virome and ancient DNA. GeCo3 achieves a solid improvement in compression over the previous version (GeCo2) of $2.4\%$, $7.1\%$, $6.1\%$, $5.8\%$, and $6.0\%$, respectively. To test its performance as a reference-based DNA compressor, we benchmark GeCo3 in 4 datasets constituted by the pairwise compression of the chromosomes of the genomes of several primates. GeCo3 improves the compression in $12.4\%$, $11.7\%$, $10.8\%$, and $10.1\%$ over the state of the art. The cost of this compression improvement is some additional computational time (1.7-3 times slower than GeCo2). The RAM use is constant, and the tool scales efficiently, independently of the sequence size. Overall, these values outperform the state of the art. CONCLUSIONS: GeCo3 is a genomic sequence compressor with a neural network mixing approach that provides additional gains over top specific genomic compressors. The proposed mixing method is portable, requiring only the probabilities of the models as inputs, providing easy adaptation to other data compressors or compression-based data analysis tools. GeCo3 is released under GPLv3 and is available for free download at https://github.com/cobilab/geco3.

Assuntos

Sequenciamento de Nucleotídeos em Larga Escala , Software , Algoritmos , Sequência de Bases , Redes Neurais de Computação , Análise de Sequência de DNA

16.

A hybrid pipeline for reconstruction and analysis of viral genomes at multi-organ level.

Pratas, Diogo; Toppinen, Mari; Pyöriä, Lari; Hedman, Klaus; Sajantila, Antti; Perdomo, Maria F.

Gigascience ; 9(8)2020 08 01.

Artigo em Inglês | MEDLINE | ID: mdl-32815536

RESUMO

BACKGROUND: Advances in sequencing technologies have enabled the characterization of multiple microbial and host genomes, opening new frontiers of knowledge while kindling novel applications and research perspectives. Among these is the investigation of the viral communities residing in the human body and their impact on health and disease. To this end, the study of samples from multiple tissues is critical, yet, the complexity of such analysis calls for a dedicated pipeline. We provide an automatic and efficient pipeline for identification, assembly, and analysis of viral genomes that combines the DNA sequence data from multiple organs. TRACESPipe relies on cooperation among 3 modalities: compression-based prediction, sequence alignment, and de novo assembly. The pipeline is ultra-fast and provides, additionally, secure transmission and storage of sensitive data. FINDINGS: TRACESPipe performed outstandingly when tested on synthetic and ex vivo datasets, identifying and reconstructing all the viral genomes, including those with high levels of single-nucleotide polymorphisms. It also detected minimal levels of genomic variation between different organs. CONCLUSIONS: TRACESPipe's unique ability to simultaneously process and analyze samples from different sources enables the evaluation of within-host variability. This opens up the possibility to investigate viral tissue tropism, evolution, fitness, and disease associations. Moreover, additional features such as DNA damage estimation and mitochondrial DNA reconstruction and analysis, as well as exogenous-source controls, expand the utility of this pipeline to other fields such as forensics and ancient DNA studies. TRACESPipe is released under GPLv3 and is available for free download at https://github.com/viromelab/tracespipe.

Assuntos

Genoma Viral , Software , Sequência de Bases , Genômica , Humanos , Alinhamento de Sequência

17.

The landscape of persistent human DNA viruses in femoral bone.

Toppinen, Mari; Pratas, Diogo; Väisänen, Elina; Söderlund-Venermo, Maria; Hedman, Klaus; Perdomo, Maria F; Sajantila, Antti.

Forensic Sci Int Genet ; 48: 102353, 2020 09.

Artigo em Inglês | MEDLINE | ID: mdl-32668397

RESUMO

The imprints left by persistent DNA viruses in the tissues can testify to the changes driving virus evolution as well as provide clues on the provenance of modern and ancient humans. However, the history hidden in skeletal remains is practically unknown, as only parvovirus B19 and hepatitis B virus DNA have been detected in hard tissues so far. Here, we investigated the DNA prevalences of 38 viruses in femoral bone of recently deceased individuals. To this end, we used quantitative PCRs and a custom viral targeted enrichment followed by next-generation sequencing. The data was analyzed with a tailor-made bioinformatics pipeline. Our findings revealed bone to be a much richer source of persistent DNA viruses than earlier perceived, discovering ten additional ones, including several members of the herpes- and polyomavirus families, as well as human papillomavirus 31 and torque teno virus. Remarkably, many of the viruses found have oncogenic potential and/or may reactivate in the elderly and immunosuppressed individuals. Thus, their persistence warrants careful evaluation of their clinical significance and impact on bone biology. Our findings open new frontiers for the study of virus evolution from ancient relics as well as provide new tools for the investigation of human skeletal remains in forensic and archaeological contexts.

Assuntos

DNA Viral/análise , Fêmur/química , Genética Forense , Adulto , Idoso , Idoso de 80 Anos ou mais , Feminino , Fêmur/virologia , Genótipo , Sequenciamento de Nucleotídeos em Larga Escala , Humanos , Masculino , Pessoa de Meia-Idade , Parvovirus B19 Humano/genética , Reação em Cadeia da Polimerase , Análise de Sequência de DNA

18.

Smash++: an alignment-free and memory-efficient tool to find genomic rearrangements.

Hosseini, Morteza; Pratas, Diogo; Morgenstern, Burkhard; Pinho, Armando J.

Gigascience ; 9(5)2020 05 01.

Artigo em Inglês | MEDLINE | ID: mdl-32432328

RESUMO

BACKGROUND: The development of high-throughput sequencing technologies and, as its result, the production of huge volumes of genomic data, has accelerated biological and medical research and discovery. Study on genomic rearrangements is crucial owing to their role in chromosomal evolution, genetic disorders, and cancer. RESULTS: We present Smash++, an alignment-free and memory-efficient tool to find and visualize small- and large-scale genomic rearrangements between 2 DNA sequences. This computational solution extracts information contents of the 2 sequences, exploiting a data compression technique to find rearrangements. We also present Smash++ visualizer, a tool that allows the visualization of the detected rearrangements along with their self- and relative complexity, by generating an SVG (Scalable Vector Graphics) image. CONCLUSIONS: Tested on several synthetic and real DNA sequences from bacteria, fungi, Aves, and Mammalia, the proposed tool was able to accurately find genomic rearrangements. The detected regions were in accordance with previous studies, which took alignment-based approaches or performed FISH (fluorescence in situ hybridization) analysis. The maximum peak memory usage among all experiments was â¼1 GB, which makes Smash++ feasible to run on present-day standard computers.

Assuntos

Biologia Computacional/métodos , Genômica/métodos , Software , Algoritmos , Rearranjo Gênico , Genoma , Sequenciamento de Nucleotídeos em Larga Escala , Análise de Sequência de DNA/métodos

19.

AC: A Compression Tool for Amino Acid Sequences.

Hosseini, Morteza; Pratas, Diogo; Pinho, Armando J.

Interdiscip Sci ; 11(1): 68-76, 2019 Mar.

Artigo em Inglês | MEDLINE | ID: mdl-30721401

RESUMO

Advancement of protein sequencing technologies has led to the production of a huge volume of data that needs to be stored and transmitted. This challenge can be tackled by compression. In this paper, we propose AC, a state-of-the-art method for lossless compression of amino acid sequences. The proposed method works based on the cooperation between finite-context models and substitutional tolerant Markov models. Compared to several general-purpose and specific-purpose protein compressors, AC provides the best bit-rates. This method can also compress the sequences nine times faster than its competitor, paq8l. In addition, employing AC, we analyze the compressibility of a large number of sequences from different domains. The results show that viruses are the most difficult sequences to be compressed. Archaea and bacteria are the second most difficult ones, and eukaryota are the easiest sequences to be compressed.

Assuntos

Algoritmos , Compressão de Dados , Sequenciamento de Nucleotídeos em Larga Escala/métodos , Software , Sequência de Aminoácidos , Cadeias de Markov

20.

Cryfa: a secure encryption tool for genomic data.

Hosseini, Morteza; Pratas, Diogo; Pinho, Armando J.

Bioinformatics ; 35(1): 146-148, 2019 01 01.

Artigo em Inglês | MEDLINE | ID: mdl-30020420

RESUMO

Summary: The ever-increasing growth of high-throughput sequencing technologies has led to a great acceleration of medical and biological research and discovery. As these platforms advance, the amount of information for diverse genomes increases at unprecedented rates. Confidentiality, integrity and authenticity of such genomic information should be ensured due to its extremely sensitive nature. In this paper, we propose Cryfa, a fast secure encryption tool for genomic data, namely in Fasta, Fastq, VCF, SAM and BAM formats, which is also capable of reducing the storage size of Fasta and Fastq files. Cryfa uses advanced encryption standard (AES) encryption combined with a shuffling mechanism, which leads to a substantial enhancement of the security against low data complexity attacks. Compared to AES Crypt, a general-purpose encryption tool, Cryfa is an industry-oriented tool, which is able to provide confidentiality, integrity and authenticity of data at four times more speed; in addition, it can reduce the file sizes to 1/3. Due to the absence of a method similar to Cryfa, we have simulated its behavior with a combination of encryption and compression tools, for comparison purpose. For instance, our tool is nine times faster than its fastest competitor in Fasta files. Also, Cryfa has a very low memory usage (only a few megabytes), which makes it feasible to run on any computer. Availability and implementation: Source codes and binaries are available, under GPLv3, at https://github.com/pratas/cryfa. Supplementary information: Supplementary data are available at Bioinformatics online.

Assuntos

Compressão de Dados , Genômica , Sequenciamento de Nucleotídeos em Larga Escala , Software , Biologia Computacional

RESUMO

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

RESUMO

Assuntos

RESUMO

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

ENVIAR RESULTADO:

SELEÇÃO DE REFERÊNCIAS

DETALHE DA PESQUISA