Architectures and accuracy of artificial neural network for disease classification from omics data.

Yu, Hui; Samuels, David C; Zhao, Ying-Yong; Guo, Yan

Yu, Hui; Samuels, David C; Zhao, Ying-Yong; Guo, Yan.

Affiliation

Yu H; Department of Internal Medicine, University of New Mexico, Albuquerque, NM, 87131, USA.
Samuels DC; Vanderbilt Genetics Institute, Department of Molecular Physiology and Biophysics, Vanderbilt University Medical School, Nashville, TN, 37232, USA.
Zhao YY; Key Laboratory of Resource Biology and Biotechnology in Western China, School of Life Sciences, Northwest University, Xi'an, 710069, Shaanxi, China.
Guo Y; Department of Internal Medicine, University of New Mexico, Albuquerque, NM, 87131, USA. yanguo1978@gmail.com.

BMC Genomics ; 20(1): 167, 2019 Mar 04.

Article in En | MEDLINE | ID: mdl-30832569

ABSTRACT

BACKGROUND: Deep learning has made tremendous successes in numerous artificial intelligence applications and is unsurprisingly penetrating into various biomedical domains. High-throughput omics data in the form of molecular profile matrices, such as transcriptomes and metabolomes, have long existed as a valuable resource for facilitating diagnosis of patient statuses/stages. It is timely imperative to compare deep learning neural networks against classical machine learning methods in the setting of matrix-formed omics data in terms of classification accuracy and robustness. RESULTS: Using 37 high throughput omics datasets, covering transcriptomes and metabolomes, we evaluated the classification power of deep learning compared to traditional machine learning methods. Representative deep learning methods, Multi-Layer Perceptrons (MLP) and Convolutional Neural Networks (CNN), were deployed and explored in seeking optimal architectures for the best classification performance. Together with five classical supervised classification methods (Linear Discriminant Analysis, Multinomial Logistic Regression, Naïve Bayes, Random Forest, Support Vector Machine), MLP and CNN were comparatively tested on the 37 datasets to predict disease stages or to discriminate diseased samples from normal samples. MLPs achieved the highest overall accuracy among all methods tested. More thorough analyses revealed that single hidden layer MLPs with ample hidden units outperformed deeper MLPs. Furthermore, MLP was one of the most robust methods against imbalanced class composition and inaccurate class labels. CONCLUSION: Our results concluded that shallow MLPs (of one or two hidden layers) with ample hidden neurons are sufficient to achieve superior and robust classification performance in exploiting numerical matrix-formed omics data for diagnosis purpose. Specific observations regarding optimal network width, class imbalance tolerance, and inaccurate labeling tolerance will inform future improvement of neural network applications on functional genomics data.

Subject(s)

Deep Learning/trends; Gene Expression Profiling/statistics & numerical data; Machine Learning/trends; Neural Networks, Computer; Algorithms; Artificial Intelligence/statistics & numerical data; Bayes Theorem; Deep Learning/statistics & numerical data; Gene Expression Profiling/methods; Humans; Logistic Models; Machine Learning/statistics & numerical data; Metabolome/genetics; Support Vector Machine/statistics & numerical data; Support Vector Machine/trends

Key words

Artificial neural network; Cancer diagnosis; Deep learning; Omics; Supervised classification

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google

Full text: 1 Collection: 01-internacional Database: MEDLINE Main subject: Neural Networks, Computer / Gene Expression Profiling / Machine Learning / Deep Learning Type of study: Prognostic_studies / Risk_factors_studies Limits: Humans Language: En Journal: BMC Genomics Journal subject: GENETICA Year: 2019 Document type: Article Affiliation country: United States Country of publication: United kingdom

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google