Prediction of mono- and di-nucleotide-specific DNA-binding sites in proteins using neural networks.

Andrabi, Munazah; Mizuguchi, Kenji; Sarai, Akinori; Ahmad, Shandar

Andrabi, Munazah; Mizuguchi, Kenji; Sarai, Akinori; Ahmad, Shandar.

Afiliação

Andrabi M; National Institute of Biomedical Innovation, Ibaraki-shi, Osaka, Japan. munazah@nibio.go.jp

BMC Struct Biol ; 9: 30, 2009 May 13.

Article em En | MEDLINE | ID: mdl-19439068

ABSTRACT

ABSTRACT

BACKGROUND:

DNA recognition by proteins is one of the most important processes in living systems. Therefore, understanding the recognition process in general, and identifying mutual recognition sites in proteins and DNA in particular, carries great significance. The sequence and structural dependence of DNA-binding sites in proteins has led to the development of successful machine learning methods for their prediction. However, all existing machine learning methods predict DNA-binding sites, irrespective of their target sequence and hence, none of them is helpful in identifying specific protein-DNA contacts. In this work, we formulate the problem of predicting specific DNA-binding sites in terms of contacts between the residue environments of proteins and the identity of a mononucleotide or a dinucleotide step in DNA. The aim of this work is to take a protein sequence or structural features as inputs and predict for each amino acid residue if it binds to DNA at locations identified by one of the four possible mononucleotides or one of the 10 unique dinucleotide steps. Contact predictions are made at various levels of resolution viz. in terms of side chain, backbone and major or minor groove atoms of DNA.

RESULTS:

Significant differences in residue preferences for specific contacts are observed, which combined with other features, lead to promising levels of prediction. In general, PSSM-based predictions, supported by secondary structure and solvent accessibility, achieve a good predictability of approximately 70-80%, measured by the area under the curve (AUC) of ROC graphs. The major and minor groove contact predictions stood out in terms of their poor predictability from sequences or PSSM, which was very strongly (>20 percentage points) compensated by the addition of secondary structure and solvent accessibility information, revealing a predominant role of local protein structure in the major/minor groove DNA-recognition. Following a detailed analysis of results, a web server to predict mononucleotide and dinucleotide-step contacts using PSSM was developed and made available at http//sdcpred.netasa.org/ or http//tardis.nibio.go.jp/netasa/sdcpred/.

CONCLUSION:

Most residue-nucleotide contacts can be predicted with high accuracy using only sequence and evolutionary information. Major and minor groove contacts, however, depend profoundly on the local structure. Overall, this study takes us a step closer to the ultimate goal of predicting mutual recognition sites in protein and DNA sequences.

Assuntos

Biologia Computacional/métodos; Proteínas de Ligação a DNA/química; DNA/química; Redes Neurais de Computação; Análise de Sequência de Proteína/métodos; Sequência de Aminoácidos; Aminoácidos/química; Sítios de Ligação; Nucleotídeos/química

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google

Texto completo: 1 Coleções: 01-internacional Base de dados: MEDLINE Assunto principal: DNA / Redes Neurais de Computação / Biologia Computacional / Análise de Sequência de Proteína / Proteínas de Ligação a DNA Tipo de estudo: Prognostic_studies / Risk_factors_studies Idioma: En Revista: BMC Struct Biol Assunto da revista: BIOLOGIA Ano de publicação: 2009 Tipo de documento: Article País de afiliação: Japão

Texto completo

Imprimir

XML

PubMed Links

Buscar no Google