Search PubMed⌕ Search

Biomedical subjects

Yael Mandel-Gutfreund

Publications and source records attributed to Yael Mandel-Gutfreund.

3 recordsLinked to original sources

Hidden Markov models that use predicted local structure for fold recognition: alphabets of backbone geometry.

An important problem in computational biology is predicting the structure of the large number of putative proteins discovered by genome sequencing projects. Fold-recognition methods attempt to solve the problem by relating the target proteins to known structures, searching for template proteins homologous to the target. Remote homologs that may have significant structural similarity are often not detectable by sequence similarities alone. To address this, we incorporated predicted local structure, a generalization of secondary structure, into two-track profile hidden Markov models (HMMs). We did not rely on a simple helix-strand-coil definition of secondary structure, but experimented with a variety of local structure descriptions, following a principled protocol to establish which descriptions are most useful for improving fold recognition and alignment quality. On a test set of 1298 nonhomologous proteins, HMMs incorporating a 3-letter STRIDE alphabet improved fold recognition accuracy by 15% over amino-acid-only HMMs and 23% over PSI-BLAST, measured by ROC-65 numbers. We compared two-track HMMs to amino-acid-only HMMs on a difficult alignment test set of 200 protein pairs (structurally similar with 3-24% sequence identity). HMMs with a 6-letter STRIDE secondary track improved alignment quality by 62%, relative to DALI structural alignments, while HMMs with an STR track (an expanded DSSP alphabet that subdivides strands into six states) improved by 40% relative to CE.

Algorithms↗

Annotating nucleic acid-binding function based on protein structure.

Many of the targets of structural genomics will be proteins with little or no structural similarity to those currently in the database. Therefore, novel function prediction methods that do not rely on sequence or fold similarity to other known proteins are needed. We present an automated approach to predict nucleic-acid-binding (NA-binding) proteins, specifically DNA-binding proteins. The method is based on characterizing the structural and sequence properties of large, positively charged electrostatic patches on DNA-binding protein surfaces, which typically coincide with the DNA-binding-sites. Using an ensemble of features extracted from these electrostatic patches, we predict DNA-binding proteins with high accuracy. We show that our method does not rely on sequence or structure homology and is capable of predicting proteins of novel-binding motifs and protein structures solved in an unbound state. Our method can also distinguish NA-binding proteins from other proteins that have similar, large positive electrostatic patches on their surfaces, but that do not bind nucleic acids.

Amino Acid Motifs↗

On the significance of alternating patterns of polar and non-polar residues in beta-strands.

A common assumption about protein sequences in beta-strands is that they have alternating patterns of polar and non-polar residues. It is thought that such patterns reflect the interior/exterior geometry of amino acid residue side-chains on a beta-sheet. Here we study the prevalence of simple hydrophobicity patterns in parallel and antiparallel beta-sheets in proteins of known structure and in the sequences of amyloidogenic proteins. The occurrence of 32 possible pentapeptide binary patterns (polar (P)/non-polar (N)) is computed in 1911 non-homologous protein structures. Despite their tendency to aggregate in experimentally designed proteins, the purely alternating hydrophobic/polar patterns (PNPNP and NPNPN) are most frequent in beta-sheets, typically occurring in antiparallel strands. The overall distribution of the pentapeptide binary patterns is significantly different in strands within parallel and antiparallel sheets. In both types of sheets, complementary patterns (where the hydrophobic and polar residues pair with one another) associate preferentially. We do not find alternating patterns to be common in amyloidogenic proteins or in short fragments involved directly in amyloid formation. However, we do note some similarities between patterns present in amyloidogenic sequences and those in parallel strands.

Amino Acid Sequence↗