Automatic recognition of complementary strands: lessons regarding machine learning abilities in RNA folding.

Chasles, Simon; Major, François

Chasles, Simon; Major, François.

Afiliação

Chasles S; Institute for Research in Immunology and Cancer, Montréal, QC, Canada.
Major F; Department of Computer Science and Operations Research, Université de Montréal, Montréal, QC, Canada.

Front Genet ; 14: 1254226, 2023.

Article em En | MEDLINE | ID: mdl-37732325

RESUMO

Introduction: Prediction of RNA secondary structure from single sequences still needs substantial improvements. The application of machine learning (ML) to this problem has become increasingly popular. However, ML algorithms are prone to overfitting, limiting the ability to learn more about the inherent mechanisms governing RNA folding. It is natural to use high-capacity models when solving such a difficult task, but poor generalization is expected when too few examples are available. Methods: Here, we report the relation between capacity and performance on a fundamental related problem: determining whether two sequences are fully complementary. Our analysis focused on the impact of model architecture and capacity as well as dataset size and nature on classification accuracy. Results: We observed that low-capacity models are better suited for learning with mislabelled training examples, while large capacities improve the ability to generalize to structurally dissimilar data. It turns out that neural networks struggle to grasp the fundamental concept of base complementarity, especially in lengthwise extrapolation context. Discussion: Given a more complex task like RNA folding, it comes as no surprise that the scarcity of useable examples hurdles the applicability of machine learning techniques to this field.

Palavras-chave

RNA folding; artificial data; base complementarity; binary classification; machine learning; neural networks

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Tipo de estudo: Prognostic_studies Idioma: En Revista: Front Genet Ano de publicação: 2023 Tipo de documento: Article País de afiliação: Canadá

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google