XGBoost odor prediction model: finding the structure-odor relationship of odorant molecules using the extreme gradient boosting algorithm.

Tyagi, Pankaj; Sharma, Anju; Semwal, Rahul; Tiwary, Uma Shanker; Varadwaj, Pritish Kumar

Tyagi, Pankaj; Sharma, Anju; Semwal, Rahul; Tiwary, Uma Shanker; Varadwaj, Pritish Kumar.

Afiliação

Tyagi P; Department of Information Technology, Indian Institute of Information Technology Allahabad, Allahabad, India.
Sharma A; Department of Pharmacoinformatics, National Institute of Pharmaceutical Education and Research (NIPER), Mohali, India.
Semwal R; Department of Computer Sciences & Engineering, Indian Institute of Information Technology Nagpur, Nagpur, India.
Tiwary US; Department of Information Technology, Indian Institute of Information Technology Allahabad, Allahabad, India.
Varadwaj PK; Department of Bioinformatics and Applied Sciences, Indian Institute of Information Technology Allahabad, Allahabad, India.

J Biomol Struct Dyn ; : 1-12, 2023 Sep 18.

Article em En | MEDLINE | ID: mdl-37723894

RESUMO

Determining the structure-odor relationship has always been a very challenging task. The main challenge in investigating the correlation between the molecular structure and its associated odor is the ambiguous and obscure nature of verbally defined odor descriptors, particularly when the odorant molecules are from different sources. With the recent developments in machine learning (ML) technology, ML and data analytic techniques are significantly being used for quantitative structure-activity relationship (QSAR) in the chemistry domain toward knowledge discovery where the traditional Edisonian methods have not been useful. The smell perception of odorant molecules is one of the aforementioned tasks, as olfaction is one of the least understood senses as compared to other senses. In this study, the XGBoost odor prediction model was generated to classify smells of odorant molecules from their SMILES strings. We first collected the dataset of 1278 odorant molecules with seven basic odor descriptors, and then 1875 physicochemical properties of odorant molecules were calculated. To obtain relevant physicochemical features, a feature reduction algorithm called PCA was also employed. The ML model developed in this study was able to predict all seven basic smells with high precision (>99%) and high sensitivity (>99%) when tested on an independent test dataset. The results of the proposed study were also compared with three recently conducted studies. The results indicate that the XGBoost-PCA model performed better than the other models for predicting common odor descriptors. The methodology and ML model developed in this study may be helpful in understanding the structure-odor relationship.Communicated by Ramaswamy H. Sarma.

Palavras-chave

Olfaction; XGBoost; machine learning; odor perception; structure odor relationship

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Base de dados: MEDLINE Idioma: En Ano de publicação: 2023 Tipo de documento: Article

Texto completo

Imprimir

XML

PubMed Links