Búsqueda | Portal Regional de la BVS

Species-aware DNA language models capture regulatory elements and their evolution.

Karollus, Alexander; Hingerl, Johannes; Gankin, Dennis; Grosshauser, Martin; Klemon, Kristian; Gagneur, Julien.

Genome Biol ; 25(1): 83, 2024 Apr 02.

Artículo en Inglés | MEDLINE | ID: mdl-38566111

RESUMEN

BACKGROUND: The rise of large-scale multi-species genome sequencing projects promises to shed new light on how genomes encode gene regulatory instructions. To this end, new algorithms are needed that can leverage conservation to capture regulatory elements while accounting for their evolution. RESULTS: Here, we introduce species-aware DNA language models, which we trained on more than 800 species spanning over 500 million years of evolution. Investigating their ability to predict masked nucleotides from context, we show that DNA language models distinguish transcription factor and RNA-binding protein motifs from background non-coding sequence. Owing to their flexibility, DNA language models capture conserved regulatory elements over much further evolutionary distances than sequence alignment would allow. Remarkably, DNA language models reconstruct motif instances bound in vivo better than unbound ones and account for the evolution of motif sequences and their positional constraints, showing that these models capture functional high-order sequence and evolutionary context. We further show that species-aware training yields improved sequence representations for endogenous and MPRA-based gene expression prediction, as well as motif discovery. CONCLUSIONS: Collectively, these results demonstrate that species-aware DNA language models are a powerful, flexible, and scalable tool to integrate information from large compendia of highly diverged genomes.

Asunto(s)

ADN , Secuencias Reguladoras de Ácidos Nucleicos , Sitios de Unión , Alineación de Secuencia , Algoritmos , Secuencia Conservada/genética , Evolución Molecular

Deep learning-driven fragment ion series classification enables highly precise and sensitive de novo peptide sequencing.

Klaproth-Andrade, Daniela; Hingerl, Johannes; Bruns, Yanik; Smith, Nicholas H; Träuble, Jakob; Wilhelm, Mathias; Gagneur, Julien.

Nat Commun ; 15(1): 151, 2024 Jan 02.

Artículo en Inglés | MEDLINE | ID: mdl-38167372

RESUMEN

Unlike for DNA and RNA, accurate and high-throughput sequencing methods for proteins are lacking, hindering the utility of proteomics in applications where the sequences are unknown including variant calling, neoepitope identification, and metaproteomics. We introduce Spectralis, a de novo peptide sequencing method for tandem mass spectrometry. Spectralis leverages several innovations including a convolutional neural network layer connecting peaks in spectra spaced by amino acid masses, proposing fragment ion series classification as a pivotal task for de novo peptide sequencing, and a peptide-spectrum confidence score. On spectra for which database search provided a ground truth, Spectralis surpassed 40% sensitivity at 90% precision, nearly doubling state-of-the-art sensitivity. Application to unidentified spectra confirmed its superiority and showcased its applicability to variant calling. Altogether, these algorithmic innovations and the substantial sensitivity increase in the high-precision range constitute an important step toward broadly applicable peptide sequencing.

Asunto(s)

Aprendizaje Profundo , Algoritmos , Análisis de Secuencia de Proteína/métodos , Péptidos/química , Secuencia de Aminoácidos

RESUMEN

Asunto(s)

RESUMEN

Asunto(s)

ENVIAR RESULTADO:

SELECCIÓN DE REFERENCIAS

DETALLE DE LA BÚSQUEDA