Your browser doesn't support javascript.
loading
An intrinsically interpretable neural network architecture for sequence-to-function learning.
Balci, Ali Tugrul; Ebeid, Mark Maher; Benos, Panayiotis V; Kostka, Dennis; Chikina, Maria.
Afiliação
  • Balci AT; Joint Carnegie Mellon University-University of Pittsburgh Program in Computational Biology, Pittsburgh, PA 15213, United States.
  • Ebeid MM; Department of Computational and Systems Biology, University of Pittsburgh, Pittsburgh, PA 15213, United States.
  • Benos PV; Joint Carnegie Mellon University-University of Pittsburgh Program in Computational Biology, Pittsburgh, PA 15213, United States.
  • Kostka D; Department of Computational and Systems Biology, University of Pittsburgh, Pittsburgh, PA 15213, United States.
  • Chikina M; Department of Epidemiology, University of Florida, Gainesville, FL 32610, United States.
Bioinformatics ; 39(39 Suppl 1): i413-i422, 2023 06 30.
Article em En | MEDLINE | ID: mdl-37387140
ABSTRACT
MOTIVATION Sequence-based deep learning approaches have been shown to predict a multitude of functional genomic readouts, including regions of open chromatin and RNA expression of genes. However, a major limitation of current methods is that model interpretation relies on computationally demanding post hoc analyses, and even then, one can often not explain the internal mechanics of highly parameterized models. Here, we introduce a deep learning architecture called totally interpretable sequence-to-function model (tiSFM). tiSFM improves upon the performance of standard multilayer convolutional models while using fewer parameters. Additionally, while tiSFM is itself technically a multilayer neural network, internal model parameters are intrinsically interpretable in terms of relevant sequence motifs.

RESULTS:

We analyze published open chromatin measurements across hematopoietic lineage cell-types and demonstrate that tiSFM outperforms a state-of-the-art convolutional neural network model custom-tailored to this dataset. We also show that it correctly identifies context-specific activities of transcription factors with known roles in hematopoietic differentiation, including Pax5 and Ebf1 for B-cells, and Rorc for innate lymphoid cells. tiSFM's model parameters have biologically meaningful interpretations, and we show the utility of our approach on a complex task of predicting the change in epigenetic state as a function of developmental transition. AVAILABILITY AND IMPLEMENTATION The source code, including scripts for the analysis of key findings, can be found at https//github.com/boooooogey/ATAConv, implemented in Python.
Assuntos

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Linfócitos / Imunidade Inata Idioma: En Ano de publicação: 2023 Tipo de documento: Article

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Linfócitos / Imunidade Inata Idioma: En Ano de publicação: 2023 Tipo de documento: Article