Quantification of biases in predictions of protein-protein binding affinity changes upon mutations.

Tsishyn, Matsvei; Pucci, Fabrizio; Rooman, Marianne

Tsishyn, Matsvei; Pucci, Fabrizio; Rooman, Marianne.

Afiliación

Tsishyn M; Computational Biology and Bioinformatics, Université Libre de Bruxelles, Roosevelt Ave, 1050, Brussels, Belgium.
Pucci F; Interuniversity Institute of Bioinformatics in Brussels, Brussels, Belgium.
Rooman M; Computational Biology and Bioinformatics, Université Libre de Bruxelles, Roosevelt Ave, 1050, Brussels, Belgium.

Brief Bioinform ; 25(1)2023 11 22.

Article en En | MEDLINE | ID: mdl-38197311

ABSTRACT

ABSTRACT

Understanding the impact of mutations on protein-protein binding affinity is a key objective for a wide range of biotechnological applications and for shedding light on disease-causing mutations, which are often located at protein-protein interfaces. Over the past decade, many computational methods using physics-based and/or machine learning approaches have been developed to predict how protein binding affinity changes upon mutations. They all claim to achieve astonishing accuracy on both training and test sets, with performances on standard benchmarks such as SKEMPI 2.0 that seem overly optimistic. Here we benchmarked eight well-known and well-used predictors and identified their biases and dataset dependencies, using not only SKEMPI 2.0 as a test set but also deep mutagenesis data on the severe acute respiratory syndrome coronavirus 2 spike protein in complex with the human angiotensin-converting enzyme 2. We showed that, even though most of the tested methods reach a significant degree of robustness and accuracy, they suffer from limited generalizability properties and struggle to predict unseen mutations. Interestingly, the generalizability problems are more severe for pure machine learning approaches, while physics-based methods are less affected by this issue. Moreover, undesirable prediction biases toward specific mutation properties, the most marked being toward destabilizing mutations, are also observed and should be carefully considered by method developers. We conclude from our analyses that there is room for improvement in the prediction models and suggest ways to check, assess and improve their generalizability and robustness.

Asunto(s)

Glicoproteína de la Espiga del Coronavirus; Humanos; Unión Proteica; Mutación; Sesgo

Palabras clave

machine learning; prediction biases; protein complex structure; proteinprotein binding affinity; proteinprotein interactions; symmetry principle

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google

Texto completo: 1 Colección: 01-internacional Banco de datos: MEDLINE Asunto principal: Glicoproteína de la Espiga del Coronavirus Tipo de estudio: Prognostic_studies / Risk_factors_studies Límite: Humans Idioma: En Revista: Brief Bioinform Asunto de la revista: BIOLOGIA / INFORMATICA MEDICA Año: 2023 Tipo del documento: Article País de afiliación: Bélgica

Texto completo

Imprimir

XML

PubMed Links

Buscar en Google