Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Secondary structure”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Integrating protein secondary structure prediction and multiple sequence alignment.

Modern protein secondary structure prediction methods are based on exploiting evolutionary information contained in multiple sequence alignments. Critical steps in the secondary structure prediction process are (i) the selection of a set of sequences that are homologous to a given query sequence, (ii) the choice of the multiple sequence alignment method, and (iii) the choice of the secondary structure prediction method. Because of the close relationship between these three steps and their critical influence on the prediction results, secondary structure prediction has received increased attention from the bioinformatics community over the last few years. In this treatise, we discuss recent developments in computational methods for protein secondary structure prediction and multiple sequence alignment, focus on the integration of these methods, and provide some recommendations for state-of-the-art secondary structure prediction in practice.

Amino Acid Sequence↗

Tree graphs of RNA secondary structures and their comparisons.

To facilitate comparison of RNA secondary structures each structure is represented as an ordered labeled tree. Several alternate secondary structures yielding a set of trees can be computed for any given RNA molecule (sequence). Frequently recurring subtrees are searched in this set of trees. The consensus structure motifs are then selected and used to construct a secondary structure model of the RNA. Given the difficulties involved in RNA secondary structure calculations, this procedure may significantly improve our predictive capabilities. In addition, the change of secondary structures between two different RNA sequences is described as a transformation of ordered trees. The transferable ratio of tree A from tree B is defined as a proportion of the largest common subtrees in trees A and B occurring in tree A. The method is applied to the study of the mechanism of human alpha 1 globin pre-mRNA splicing. In the study, two tentative splicing mechanisms, A and B, with different orders of intron excision from alpha 1 globin pre-mRNA have been stimulated. A possible relationship between the structural features of the secondary structures and the order of intron excision in the pathway of precursor splicing of human alpha 1 globin is discussed.

Base Sequence↗

Mapping in solution shows the peach latent mosaic viroid to possess a new pseudoknot in a complex, branched secondary structure.

We have investigated the secondary structure of peach latent mosaic viroid (PLMVd) in solution, and we present here the first description of the structure of a branched viroid in solution. Different PLMVd transcripts of plus polarity were produced by using the circularly permuted RNA method and the exploitation of RNA internal secondary structure to position the 5' and 3' termini and studied by nuclease mapping and binding shift assays using DNA and RNA oligonucleotides. We show that PLMVd folds into a complex, branched secondary structure. In general, this structure is similar to that reported previously, which was based on sequence comparison and computer modelling. The structural microheterogeneity is apparently limited to only some small domains. More importantly, this structure includes a novel pseudoknot that is conserved in all PLMVd isolates and seems to allow folding into a very compact form. This pseudoknot is also found in chrysanthemum chlorotic mottle viroid, suggesting that it is a unique feature of the viroid members of the PLMVd subgroup.

Base Sequence↗

Protein secondary structure prediction in different structural classes.

Information about the secondary structure of a protein can be helpful in understanding its native folded state. In previous work, it was shown that the medium-range interactions predominate in all-alpha class and the long-range interactions predominate in all-beta class proteins. Based on this, in this work the performance of several structure prediction methods in different structural classes of globular proteins was analyzed. It was found that all the methods predict the secondary structures of all-alpha proteins more accurately than other classes.

Protein Structure, Secondary↗

New joint prediction algorithm (Q7-JASEP) improves the prediction of protein secondary structure.

The classical problem of secondary structure prediction is approached by a new joint algorithm (Q7-JASEP) that combines the best aspects of six different methods. The algorithm includes the statistical methods of Chou-Fasman, Nagano, and Burgess-Ponnuswamy-Scheraga, the homology method of Nishikawa, the information theory method of Garnier-Osgurthope-Robson, and the artificial neural network approach of Qian-Sejnowski. Steps in the algorithm are (i) optimizing each individual method with respect to its correlation coefficient (Q7) for assigning a structural type from the predictive score of the method, (ii) weighting each method, (iii) combining the scores from different methods, and (iv) comparing the scores for alpha-helix, beta-strand, and coil conformational states to assign the secondary structure at each residue position. The present application to 45 globular proteins demonstrates good predictive power in cross-validation testing (with average correlation coefficients per test protein of Q7, alpha = 0.41, Q7, beta = 0.47, Q7,c = 0.41 for alpha-helix, beta-strand, and coil conformations). By the criterion of correlation coefficient (Q7) for each type of secondary structure, Q7-JASEP performs better than any of the component methods. When all protein classes are included for training and testing (by cross-validation), the results here equal the best in the literature, by the Q7 criterion. More generally, the basic algorithm can be applied to any protein class and to any type of structure/sequence or function/sequence correlation for which multiple predictive methods exist.

Algorithms↗

Denatured states of ribonuclease A have compact dimensions and residual secondary structure.

Using small-angle X-ray scattering and Fourier transform infrared spectroscopy, we have determined that the thermally denatured state of native ribonuclease A is on average a compact structure having residual secondary structure. Under strongly reducing conditions, the protein further unfolds into a looser structure with larger dimensions but still retains a comparable amount of secondary structure. The dimensions of the thermally and chemically denatured states of the reduced protein are different but both are more compact than is predicted for a random coil of the same length. These results demonstrate that thermal denaturation in ribonuclease A is not a simple two-state transition from a native to a completely disordered random coil state.

Animals↗

Folding type specific secondary structure propensities of synonymous codons.

We have proposed new amino acid secondary structure propensities in proteins with different folding types based on synonymous codons. They have been derived from 200 all alpha, all beta, alpha/beta, and alpha + beta proteins of known structures and their coding genes. The secondary structure propensities of the same codon in gene coding for different folding type proteins are not the same. For instance, amino acid Ile coded by AUU is indifferent to form the alpha unit in the alpha + beta protein class, but it is a former and a breaker for the alpha unit in the all alpha protein class and the alpha/beta class, respectively. On the other hand, the secondary structure propensities of different synonymous codons in the coding genes with the same folding type are also not all the same. As an example, CGU, CGG, and AGA, which are synonymous codons of Arg, are preferential to form the alpha unit in all alpha proteins, while CGA is an alpha unit breaker and the other two synonymous codons, CGC and AGG, are indifferent to form or break the alpha unit. As a result, protein secondary structure information contained both in mRNA sequences and in amino acid sequences has been introduced in these codon-based amino acid secondary structure propensities. These codon-based amino acid secondary structure propensities are helpful to in vitro protein design and protein secondary structure prediction.

Amino Acid Sequence↗

Prediction of protein secondary structures by a neural network.

We have studied the prediction of globular protein secondary structures by neural networks. Protein secondary structures are allocated to amino acid residues using Kabsch and Sander's dictionary of protein secondary structures and the neural network is taught the protein secondary structures. The input layer of the neural network allows sequences of residues including 20 amino acids, chain break, B, X and Z. We consider classifying secondary structures into groups of 3, 4 and 8. In each case, we calculate the percentage of correct predictions. We discuss the effect of overlearning on the protein secondary structure prediction. In addition, we include the application of a neural network with a modular architecture to prediction of protein secondary structures. We compare the results from neural networks with a modular architecture and with a simple three-layer structure.

Algorithms↗

A physical basis for protein secondary structure.

A physical theory of protein secondary structure is proposed and tested by performing exceedingly simple Monte Carlo simulations. In essence, secondary structure propensities are predominantly a consequence of two competing local effects, one favoring hydrogen bond formation in helices and turns, the other opposing the attendant reduction in sidechain conformational entropy on helix and turn formation. These sequence specific biases are densely dispersed throughout the unfolded polypeptide chain, where they serve to preorganize the folding process and largely, but imperfectly, anticipate the native secondary structure.

Protein Conformation↗

A program for predicting significant RNA secondary structures.

We describe a program for the analysis of RNA secondary structure. There are two new features in this program. (i) To get vector speeds on a vector pipeline machine (such as Cray X-MP/24) we have vectorized the secondary structure dynamic algorithm. (ii) The statistical significance of a locally 'optimal' secondary structure is assessed by a Monte Carlo method. The results can be depicted graphically including profiles of the stability of local secondary structures and the distribution of the potentially significant secondary structures in the RNA molecules. Interesting regions where both the potentially significant secondary structures and 'open' structures (single-stranded coils) occur can be identified by the plots mentioned above. Furthermore, the speed of the vectorized code allows repeated Monte Carlo simulations with different overlapping window sizes. Thus, the optimal size of the significant secondary structure occurring in the interesting region can be assessed by repeating the Monte Carlo simulation. The power of the program is demonstrated in the analysis of local secondary structures of human T-cell lymphotrophic virus type III (HIV).

Algorithms↗

Two-stage multi-class support vector machines to protein secondary structure prediction.

Bioinformatics techniques to protein secondary structure (PSS) prediction are mostly single-stage approaches in the sense that they predict secondary structures of proteins by taking into account only the contextual information in amino acid sequences. In this paper, we propose two-stage Multi-class Support Vector Machine (MSVM) approach where a MSVM predictor is introduced to the output of the first stage MSVM to capture the sequential relationship among secondary structure elements for the prediction. By using position specific scoring matrices, generated by PSI-BLAST, the two-stage MSVM approach achieves Q3 accuracies of 78.0% and 76.3% on the RS126 dataset of 126 nonhomologous globular proteins and the CB396 dataset of 396 nonhomologous proteins, respectively, which are better than the highest scores published on both datasets to date.

Computational Biology↗

Secondary structure characteristics of proenkephalin peptides E, B, and F.

The conformations of three adrenal medullary enkephalin containing polypeptides (ECPs) were investigated to gain an understanding of their potential structure-activity relationships. Secondary structure characteristics of peptides E, B, and F were examined by circular dichrosim (CD) under conditions designed to mimic both the soluble state and the anisotropic environment which exists at the biological effector site. Conformational differences between the three peptides were further examined by Fourier Transform Infrared Spectroscopy (FTIR) and by empirical predictions for conformation and hydrophobic periodicity. Although all three peptides have a similar structure, existing in random configurations in aqueous solutions, they do exhibit unique individual potentials to assume secondary structure in less polar environments. These conformational differences may be important factors in determining their unique individual biological activities.

Adrenal Medulla↗

Spectroscopic methods for analysis of protein secondary structure.

Several methods for determination of the secondary structure of proteins by spectroscopic measurements are reviewed. Circular dichroism (CD) spectroscopy provides rapid determinations of protein secondary structure with dilute solutions and a way to rapidly assess conformational changes resulting from addition of ligands. Both CD and Raman spectroscopies are particularly useful for measurements over a range of temperatures. Infrared (IR) and Raman spectroscopy require only small volumes of protein solution. The frequencies of amide bands are analyzed to determine the distribution of secondary structures in proteins. NMR chemical shifts may also be used to determine the positions of secondary structure within the primary sequence of a protein. However, the chemical shifts must first be assigned to particular residues, making the technique considerably slower than the optical methods. These data, together with sophisticated molecular modeling techniques, allow for refinement of protein structural models as well as rapid assessment of conformational changes resulting from ligand binding or macromolecular interactions. A selected number of examples are given to illustrate the power of the techniques in applications of biological interest.

Animals↗

Use of amino acid environment-dependent substitution tables and conformational propensities in structure prediction from aligned sequences of homologous proteins. II. Secondary structures.

A three-step method is presented to predict secondary structures of proteins, by utilizing aligned sequences of homologous proteins. Mean propensities and amino acid substitution patterns at a given site in the aligned sequences are first evaluated for four conformational states (i.e. alpha-helix, beta-strand, buried coil and exposed coil). Capping rules are applied in order to define boundaries of the secondary-structure segments more precisely. In the second step beta-strand is predicted by searching regions predicted as coil for the two patterns characteristic of alternating and fully buried beta-strands. The complete sequences of the solvent-accessibility classes predicted by substitution tables and propensities are also searched using Fourier transform methods for alpha-helical periodicity. After applying capping rules, the alpha-helices and beta-strands predicted in the second step replace, where appropriate, the conformational states predicted in the first step. Finally, in the third step, if one of the four conformational states is assigned to the residues at an equivalent site of aligned sequences in more than a given fraction of the proteins, such a state is reassigned to all the residues at that site. The method is applied to 13 protein families, which contain four folding types, alpha, beta, alpha/beta and alpha + beta. The accuracy of the prediction ranges from 60 to 79% (mean percentage over the 13 families is 69%). For comparison the Garnier-Osguthorpe-Robson (GOR) method is also applied to them. Although the mean prediction accuracy for the GOR method, 58%, can be improved to 63% by applying the second and third steps in this method, there remain four families with less than 55% accuracy. The mean accuracy is relatively higher and poor predictions are reduced in this method.

Amino Acid Sequence↗

Recognition of super-secondary structure in proteins.

A procedure to recognize super-secondary structure in protein sequences is described. An idealized template, derived from known super-secondary structures, is used to locate probable sites by matching with secondary structure probability profiles. We applied the method to the identification of beta alpha beta units in beta/alpha type proteins with 75% accuracy. The location of super-secondary structure was then used to refine the original (Garnier et al., 1978) secondary structure prediction resulting in an 8.8% improvement, which correctly assigned 83% of secondary structure elements in 14 proteins. Slight modifications to the Garnier et al. method are suggested, producing a more accurate identification of protein class and a better prediction for beta/alpha type proteins. A method for the incorporation of hydrophobic information into the prediction is also described.

Amino Acid Sequence↗

When awaiting 'Bio' Champollion: dynamic programming regularization of the protein secondary structure predictions.

Predictions of protein secondary structure using current methods are often unrealistic, i.e. the predicted alpha-helices or beta-strands are too short. To improve the realism, various heuristic 'filtering' or 'smoothing' methods are used. They are more or less intuitive and are based on ad hoc corrections. We present a regularization method to obtain a realistic secondary structure from predicted propensities. It is based on the known dynamic programming algorithm and is quite objective. It can be used with any prediction method which yields propensities. The regularized predictions conserve well the overall prediction accuracy and improve the 'protein-likeness' of the prediction.

Algorithms↗

Secondary structural wobble: the limits of protein prediction accuracy.

At present, accuracies of secondary structural prediction scarcely go beyond 70-75%. Secondary structural comparison is carried out among sequence-identified proteins. The results show natural wobble between different secondary structural types is possible in homologous families, and the best prediction accuracy will rarely be 100%. Besides shortcoming of the prediction approaches, secondary structural wobble is found to be responsible for nearly all secondary structural prediction limits. Only average 73.2% of amino acid residue is conserved in secondary structural types. The wobble allows alpha-class/coil and beta-class/coil transitions but not direct alpha-class/beta-class transition. Propensity values representing the statistical occurrence of 20 amino acid residues in secondary structural wobbles are given.

Animals↗

Conservation of substructures in proteins: interfaces of secondary structural elements in proteasomal subunits.

It is observed that during divergent evolution of two proteins with a common phylogenetic origin, the structural similarity of their backbones is often preserved even when the sequence similarity between them decreases to a virtually undetectable level. Here we analyzed, whether the conservation of structure along evolution involves also the local atomic structures in the interfaces between secondary structural elements. We have used as study case one protein family, the proteasomal subunits, for which 17 crystal structures are known. These include 14 different subunits of Saccharomyces cerevisiae, 2 subunits of Thermoplasma acidophilum and one subunit of Escherichia coli. The structural core of the 17 proteasomal subunits has 23 secondary structural elements. Any two adjacent secondary structural elements form a molecular interface consisting of two molecular patches. We found 61 interfaces that occurred in all 17 subunits. The 3D shape of equivalent molecular patches from different proteasomal subunits were compared by superposition. Our results demonstrate that pairs of equivalent molecular patches show an RMSD which is lower than that of randomly chosen patches from unrelated proteins. This is true even when patch comparisons with identical residues were excluded from the analysis. Furthermore it is known that the sequential dissimilarity is correlated to the RMSD between the backbones of the members of protein families. The question arises whether this is also true for local atomic structures. The results show that the correlation of individual patch RMSD values and local sequence dissimilarities is low and has a wide range from 0 to 0.41, however, it is surprising that there is a good correlation between the average RMSD of all corresponding patches and the global sequence dissimilarity. This average patch RMSD correlates slightly stronger than the C(alpha)-trace RMSD to the global sequence dissimilarity.

Algorithms↗