Your browser doesn't support javascript.
loading
PreDBP-PLMs: Prediction of DNA-binding proteins based on pre-trained protein language models and convolutional neural networks.
Qi, Dawei; Song, Chen; Liu, Taigang.
Affiliation
  • Qi D; College of Information Technology, Shanghai Ocean University, Shanghai, 201306, China.
  • Song C; College of Information Technology, Shanghai Ocean University, Shanghai, 201306, China.
  • Liu T; College of Information Technology, Shanghai Ocean University, Shanghai, 201306, China. Electronic address: tgliu@shou.edu.cn.
Anal Biochem ; 694: 115603, 2024 Jul 08.
Article in En | MEDLINE | ID: mdl-38986796
ABSTRACT
The recognition of DNA-binding proteins (DBPs) is the crucial step to understanding their roles in various biological processes such as genetic regulation, gene expression, cell cycle control, DNA repair, and replication within cells. However, conventional experimental methods for identifying DBPs are usually time-consuming and expensive. Therefore, there is an urgent need to develop rapid and efficient computational methods for the prediction of DBPs. In this study, we proposed a novel predictor named PreDBP-PLMs to further improve the identification accuracy of DBPs by fusing the pre-trained protein language model (PLM) ProtT5 embedding with evolutionary features as input to the classic convolutional neural network (CNN) model. Firstly, the ProtT5 embedding was combined with different evolutionary features derived from the position-specific scoring matrix (PSSM) to represent protein sequences. Then, the optimal feature combination was selected and input to the CNN classifier for the prediction of DBPs. Finally, the 5-fold cross-validation (CV), the leave-one-out CV (LOOCV), and the independent set test were adopted to examine the performance of PreDBP-PLMs on the benchmark datasets. Compared to the existing state-of-the-art predictors, PreDBP-PLMs exhibits an accuracy improvement of 0.5 % and 5.2 % on the PDB186 and PDB2272 datasets, respectively. It demonstrated that the proposed method could serve as a useful tool for the recognition of DBPs.
Key words

Full text: 1 Collection: 01-internacional Database: MEDLINE Language: En Journal: Anal Biochem Year: 2024 Document type: Article Affiliation country: China

Full text: 1 Collection: 01-internacional Database: MEDLINE Language: En Journal: Anal Biochem Year: 2024 Document type: Article Affiliation country: China