Your browser doesn't support javascript.
loading
Mostrar: 20 | 50 | 100
Resultados 1 - 1 de 1
Filtrar
Más filtros

Banco de datos
Tipo del documento
Asunto de la revista
País de afiliación
Intervalo de año de publicación
1.
Bioinformatics ; 33(14): i83-i91, 2017 Jul 15.
Artículo en Inglés | MEDLINE | ID: mdl-28881966

RESUMEN

MOTIVATION: Moonlighting proteins (MPs) are an important class of proteins that perform more than one independent cellular function. MPs are gaining more attention in recent years as they are found to play important roles in various systems including disease developments. MPs also have a significant impact in computational function prediction and annotation in databases. Currently MPs are not labeled as such in biological databases even in cases where multiple distinct functions are known for the proteins. In this work, we propose a novel method named DextMP, which predicts whether a protein is a MP or not based on its textual features extracted from scientific literature and the UniProt database. RESULTS: DextMP extracts three categories of textual information for a protein: titles, abstracts from literature, and function description in UniProt. Three language models were applied and compared: a state-of-the-art deep unsupervised learning algorithm along with two other language models of different types, Term Frequency-Inverse Document Frequency in the bag-of-words and Latent Dirichlet Allocation in the topic modeling category. Cross-validation results on a dataset of known MPs and non-MPs showed that DextMP successfully predicted MPs with over 91% accuracy with significant improvement over existing MP prediction methods. Lastly, we ran DextMP with the best performing language models and text-based feature combinations on three genomes, human, yeast and Xenopus laevis , and found that about 2.5-35% of the proteomes are potential MPs. AVAILABILITY AND IMPLEMENTATION: Code available at http://kiharalab.org/DextMP . CONTACT: dkihara@purdue.edu.


Asunto(s)
Minería de Datos/métodos , Proteómica/métodos , Programas Informáticos , Aprendizaje Automático no Supervisado , Animales , Bases de Datos Factuales , Humanos , Modelos Biológicos , Anotación de Secuencia Molecular , Saccharomyces cerevisiae/genética , Saccharomyces cerevisiae/metabolismo , Xenopus laevis/genética , Xenopus laevis/metabolismo
SELECCIÓN DE REFERENCIAS
DETALLE DE LA BÚSQUEDA