Your browser doesn't support javascript.
loading
PON-tstab: Protein Variant Stability Predictor. Importance of Training Data Quality.
Yang, Yang; Urolagin, Siddhaling; Niroula, Abhishek; Ding, Xuesong; Shen, Bairong; Vihinen, Mauno.
Afiliação
  • Yang Y; School of Computer Science and Technology, Soochow University, No. 1. Shizi Street, Suzhou 215006, China. yyang@suda.edu.cn.
  • Urolagin S; Department of Experimental Medical Science, BMC B13, Lund University, SE-22 184 Lund, Sweden. yyang@suda.edu.cn.
  • Niroula A; Center for Systems Biology, Soochow University, No. 1. Shizi Street, Suzhou 215006, China. yyang@suda.edu.cn.
  • Ding X; Department of Experimental Medical Science, BMC B13, Lund University, SE-22 184 Lund, Sweden. siddhaling@dubai.bits-pilani.ac.in.
  • Shen B; Department of Experimental Medical Science, BMC B13, Lund University, SE-22 184 Lund, Sweden. abhishek.niroula@med.lu.se.
  • Vihinen M; School of Computer Science and Technology, Soochow University, No. 1. Shizi Street, Suzhou 215006, China. 20154227019@stu.suda.edu.cn.
Int J Mol Sci ; 19(4)2018 Mar 28.
Article em En | MEDLINE | ID: mdl-29597263
ABSTRACT
Several methods have been developed to predict effects of amino acid substitutions on protein stability. Benchmark datasets are essential for method training and testing and have numerous requirements including that the data is representative for the investigated phenomenon. Available machine learning algorithms for variant stability have all been trained with ProTherm data. We noticed a number of issues with the contents, quality and relevance of the database. There were errors, but also features that had not been clearly communicated. Consequently, all machine learning variant stability predictors have been trained on biased and incorrect data. We obtained a corrected dataset and trained a random forests-based tool, PON-tstab, applicable to variants in any organism. Our results highlight the importance of the benchmark quality, suitability and appropriateness. Predictions are provided for three categories stability decreasing, increasing and those not affecting stability.
Assuntos
Palavras-chave

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Proteínas / Modelos Moleculares / Bases de Dados de Proteínas / Aprendizado de Máquina Tipo de estudo: Prognostic_studies / Risk_factors_studies Idioma: En Ano de publicação: 2018 Tipo de documento: Article

Texto completo: 1 Base de dados: MEDLINE Assunto principal: Proteínas / Modelos Moleculares / Bases de Dados de Proteínas / Aprendizado de Máquina Tipo de estudo: Prognostic_studies / Risk_factors_studies Idioma: En Ano de publicação: 2018 Tipo de documento: Article