Predicting peptide presentation by major histocompatibility complex class I: an improved machine learning approach to the immunopeptidome.

Boehm, Kevin Michael; Bhinder, Bhavneet; Raja, Vijay Joseph; Dephoure, Noah; Elemento, Olivier

Boehm, Kevin Michael; Bhinder, Bhavneet; Raja, Vijay Joseph; Dephoure, Noah; Elemento, Olivier.

Afiliación

Boehm KM; Weill Cornell/Rockefeller/Sloan Kettering Tri-Institutional MD-PhD Program, 1300 York Avenue, New York, NY, USA. kmb2012@med.cornell.edu.
Bhinder B; Caryl and Israel Englander Institute for Precision Medicine, Weill Cornell Medical College, 413 East 69th Street, New York, NY, USA.
Raja VJ; Institute for Computational Biomedicine, Weill Cornell Medical College, 1305 York Avenue, New York, NY, USA.
Dephoure N; Department of Biochemistry, Weill Cornell Medical College, 1300 York Avenue, New York, NY, USA.
Elemento O; Department of Biochemistry, Weill Cornell Medical College, 1300 York Avenue, New York, NY, USA.

BMC Bioinformatics ; 20(1): 7, 2019 Jan 05.

Article en En | MEDLINE | ID: mdl-30611210

ABSTRACT

ABSTRACT

BACKGROUND:

To further our understanding of immunopeptidomics, improved tools are needed to identify peptides presented by major histocompatibility complex class I (MHC-I). Many existing tools are limited by their reliance upon chemical affinity data, which is less biologically relevant than sampling by mass spectrometry, and other tools are limited by incomplete exploration of machine learning approaches. Herein, we assemble publicly available data describing human peptides discovered by sampling the MHC-I immunopeptidome with mass spectrometry and use this database to train random forest classifiers (ForestMHC) to predict presentation by MHC-I.

RESULTS:

As measured by precision in the top 1% of predictions, our method outperforms NetMHC and NetMHCpan on test sets, and it outperforms both these methods and MixMHCpred on new data from an ovarian carcinoma cell line. We also find that random forest scores correlate monotonically, but not linearly, with known chemical binding affinities, and an information-based analysis of classifier features shows the importance of anchor positions for our classification. The random-forest approach also outperforms a deep neural network and a convolutional neural network trained on identical data. Finally, we use our large database to confirm that gene expression partially determines peptide presentation.

CONCLUSIONS:

ForestMHC is a promising method to identify peptides bound by MHC-I. We have demonstrated the utility of random forest-based approaches in predicting peptide presentation by MHC-I, assembled the largest known database of MS binding data, and mined this database to show the effect of gene expression on peptide presentation. ForestMHC has potential applicability to basic immunology, rational vaccine design, and neoantigen binding prediction for cancer immunotherapy. This method is publicly available for applications and further validation.

Asunto(s)

Antígenos de Histocompatibilidad Clase I/metabolismo; Aprendizaje Automático; Péptidos/inmunología; Proteoma/metabolismo; Algoritmos; Línea Celular Tumoral; Bases de Datos de Proteínas; Regulación de la Expresión Génica; Humanos; Péptidos/química; Reproducibilidad de los Resultados

Palabras clave

Antigen presentation; Immunopeptidomics; MHC-I; Machine learning; Random forest

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Colección: 01-internacional Banco de datos: MEDLINE Asunto principal: Péptidos / Antígenos de Histocompatibilidad Clase I / Proteoma / Aprendizaje Automático Tipo de estudio: Prognostic_studies / Risk_factors_studies Límite: Humans Idioma: En Revista: BMC Bioinformatics Asunto de la revista: INFORMATICA MEDICA Año: 2019 Tipo del documento: Article País de afiliación: Estados Unidos

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google