Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

A proteolytically sensitive region common to several rat liver cytochromes P450: effect of cleavage on substrate binding.

Limited proteolysis of rat liver microsomes was used to probe the topography and structure of cytochrome P450 bound to the endoplasmic reticulum. Three cytochromes P450 from two families were examined. Monoclonal antibodies to cytochrome P450 forms 1A1, 2B1, and 2E1 were used to immunopurify these proteolyzed cytochromes P450 from microsomes from rats treated with 3-methylcholanthrene, phenobarbital, and acetone, respectively. Electrophoretic and immunoblot analysis of tryptic fragments revealed a highly sensitive cleavage site in all three cytochromes P450. N-Terminal sequencing was performed on the fragments after transfer onto poly(vinylidene difluoride) membranes and showed that this preferential cleavage site is at amino acid position 298 of P450 1A1, position 277 of P450 2B1, and position 278 of P450 2E1. Multiple sequence alignment revealed that these positions are at the amino terminal of a highly conserved region of these cytochromes P450. The important functional role implied by primary sequence conservation along with the proteolytic sensitivity at its amino terminal suggests that this region is a protein domain. Comparison with the known structure of the bacterial cytochrome P450cam predicts that this proteolytically sensitive site is within an interhelical turn region connected to the distal helix that partially encompasses the heme-containing active site. Substrate binding to the cleaved cytochromes P450 was examined in order to determine whether the newly added conformational freedom near the cleavage site functionally altered these cytochromes P450. Cleavage of P450 2B1 abolished benzphetamine binding, which indicates that the cleavage site contains an important structural determinant for binding this substrate. However, cleavage did not affect benzo[a]pyrene binding to P450 1A1.

Amino Acid Sequence

Identification of Cys-150 in the active site of phosphomannose isomerase from Candida albicans.

Candida albicans phosphomannose isomerase (PMI) (EC 5.3.1.8) has been recently cloned and overexpressed in Escherichia coli. The enzyme can be irreversibly inactivated by iodoacetate in 50 mM borate buffer, pH 9.0, in a time-dependent manner at a rate of 4.2 +/- 0.03 min-1 M-1. This inhibition can be prevented by the substrate mannose 6-phosphate with a Ks of 0.22 +/- 0.05 mM, slightly lower than its Km value. However, metals such as zinc and cadmium, which are reversible, competitive inhibitors for PMI, do not protect the enzyme against modification. The protein has been labeled by using [2-14C]iodoacetate, in the presence or absence of substrate, and the protein is fully inactivated when 1.0 thiol group is modified per molecule of enzyme. Tryptic maps of the modified protein have been produced. The protected peptide has been identified and sequenced, and the phenylthiohydantoin amino acids have been collected. The modified amino acid is Cys-150. This cysteine residue is conserved in mammalian and yeast phosphomannose isomerases, but not in bacterial species where it is replaced with asparagine. We therefore purified PMI from E. coli and showed that this enzyme is not sensitive to inactivation by iodoacetate. The iodoacetate is presumably inhibiting PMI by sterically blocking the mannose 6-phosphate binding site. Multiple sequence alignment procedures were used to try to identify potential ligands of the zinc atom that is essential for enzyme activity and thus to delineate the active site region.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

Photoaffinity labeling with UMP of lysine 992 of carbamyl phosphate synthetase from Escherichia coli allows identification of the binding site for the pyrimidine inhibitor.

UMP is a highly specific reagent for photoaffinity labeling of the allosteric inhibitor site of carbamyl phosphate synthetase (CPS) from Escherichia coli and has been found to be photoincorporated in the COOH-terminal domain of the large subunit [Rubio et al. (1991) Biochemistry 30, 1068-1075]. In the present work we identify lysine 992 as the residue that is covalently attached to UMP. This identification is based on two lines of evidence. First, [14C]UMP is found to be incorporated between residues 939 and 1006, as shown by peptide mapping and by mass estimates of [14C]UMP-peptides generated by chemical and enzymatic cleavage of CPS. Secondly, we have purified two radioactive peptides derived exclusively from those enzyme molecules (approximately 5% of the total enzyme) that had incorporated [14C]-UMP. Edman analyses show the sequences of the labeled peptides (989)LVNXVHEGRPHIQD and (989)LVNXVHE to be overlapping. Since neither a phenylthiohydantoin (Pth) derivative (in cycle 4) nor any radioactivity is released from the membrane during sequencing, we can conclude that Lys992 and [14C]-UMP form a covalent adduct that remains bound to the membrane. Formation of this adduct agrees with all of the evidence and with the finding that UMP labeling prevents trypsin cleavage at Lys992. Lysine 992 is invariant in those CPSs that are inhibited by UMP, and is located 30 residues upstream of the site whose phosphorylation in hamster CAD reduces inhibition of CAD by UTP. Multiple sequence alignment of the residues surrounding Lys992 of the E. coli enzyme and the corresponding residues of the yeast and animal enzymes supports the existence of a uridine nucleotide binding fold in this region of the protein. We conclude that sequence changes in the binding fold provide a structural basis for the different regulatory properties found among CPSs I, II, and III.

Affinity Labels

Single amino acid substitutions disrupt tetramer formation in the dihydroneopterin aldolase enzyme of Pneumocystis carinii.

In the opportunistic pathogen Pneumocystis carinii, dihydroneopterin aldolase function is expressed as the N-terminal portion of the multifunctional folic acid synthesis protein (Fas). This region encompasses two domains, FasA and FasB, which are 27% amino acid identical. FasA and FasB also share significant amino acid sequence similarity with bacterial dihydroneopterin aldolases. In the present study, this enzyme function has been overproduced as an independent monofunctional activity in Escherichia coli. Recombinant FasAB-Met23 (amino acids 23-290 of the predicted open reading frame) was purified and shown to contain dihydroneopterin aldolase activity. The native FasAB-Met23 is a tetramer of the 30-kDa subunit, demonstrating characteristics of an associating-dissociating equilibrium system in which only the multimeric form of the enzyme is active. Multiple sequence alignment of FasA and FasB with other dihydroneopterin aldolases highlights only three positions where the amino acid is invariable between all the predicted proteins. The role of these conserved amino acid residues in enzyme function was investigated using site-directed mutagenesis. Mutant FasAB-Met23 species were overproduced and purified to near homogeneity. Three FasA domain mutants and two FasB domain mutants had little or no detectable dihydroneopterin aldolase activity, implicating both FasA and FasB in the catalytic mechanism. We show that each mutant protein containing an inactivating amino acid substitution has lost its ability to form stable tetramers.

Aldehyde-Lyases

Predicting differential antigen-antibody contact regions based on solvent accessibility.

A novel computational approach was examined for predicting epitopes from primary structures of the seven immunologically distinct botulinum neurotoxins (BoNT/A-G) and tetanus toxin (TeTX). An artificial neural network [Rost and Sander (1994), Proteins 20, 216] was used to estimate residue solvent accessibilities in multiple aligned sequences. A similar network trained to predict secondary structures was also used to examine this protein family, whose tertiary fold is presently unknown. The algorithm was validated by showing that it was 80% accurate in determining the secondary structure of avian egg-white lysozyme and that it correctly identified highly solvent-exposed residues that correspond to the major contact regions of lysozyme-antibody cocrystals. When sequences of the heavy (H) chains of TeTX and BoNT/A-G were analyzed, this algorithm predicted that the most highly exposed regions were clustered at the sequentially nonconserved N- and C-termini [Lebeda and Olson (1994), Proteins 20, 293]. The secondary structures and the remaining highly solvent-accessible regions were, in contrast, predicted to be conserved. In experiments reported by others, H-chain fragments that induced immunological protection against BoNT/A overlap with these predicted most highly exposed regions. It is also known that the C-terminal halves of the TeTX and BoNT/A H-chains interfere with holotoxin binding to ectoacceptors on nerve endings. Thus, the present results provide a theoretical framework for predicting the sites that could assist in the development of genetically engineered vaccines and that could interact with neurally located toxin ectoacceptors. Finally, because the most highly solvent-exposed regions were not well conserved, it is hypothesized that nonconserved, potential contact sites partially account for the existence of different dominant binding regions for type-specific neutralizing antibodies.

Algorithms

On the classification and evolution of protein modules.

Our efforts to classify the functional units of many proteins, the modules, are reviewed. The data from the sequencing projects for various model organisms are extremely helpful in deducing the evolution of proteins and modules. For example, a dramatic increase of modular proteins can be observed from yeast to C. elegans in accordance with new protein functions that had to be introduced in multicellular organisms. Our sequence characterization of modules relies on sensitive similarity search algorithms and the collection of multiple sequence alignments for each module. To trace the evolution of modules and to further automate the classification, we have developed a sequence and a module alerting system that checks newly arriving sequence data for the presence of already classified modules. Using these systems, we were able to identify an unexpected similarity between extracellular C1Q modules with bacterial proteins.

Amino Acid Sequence

An investigation of the role of Glu-842, Glu-844 and His-846 in the function of the cytoplasmic domain of the epidermal growth factor receptor.

Activation of several protein kinases is mediated, at least in part, by phosphorylation of conserved Thr or Tyr residues located in a variable loop region, near the active site. In certain kinases, this activation loop also controls access of peptide substrates to the active site. In the corresponding region of the epidermal growth factor (EGF) receptor, a potential phosphorylation site, Tyr-845, does not appear to have a major regulatory role. In order to find out whether this variable loop can modulate the peptide phosphorylation and self-phosphorylation activities of the EGF receptor kinase, we investigated the role of residues around Tyr-845, using site-directed mutagenesis. Multiple sequence alignment showed that residues Glu-842, Glu-844 and His-846 are conserved or nearly conserved in eight members of the EGF receptor family. Mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala were expressed in the baculovirus/insect cell system, purified to near-homogeneity and characterized with respect to their peptide phosphorylation and self-phosphorylation activities. All three mutants were active, and these changes did not affect ATP binding directly. However, all mutations increased the Km(app.) for peptide substrates and MnATP in peptide phosphorylation reactions. The Vmax. for the phosphorylation of peptide RREELQDDYEDD was unaltered, but the Vmax. for self-phosphorylation (with variable [MnATP]) decreased 4-, 2- and 7-fold for mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala respectively, compared with the wild-type. These results suggest that binding of this peptide restored an optimal conformation at the active site that might be impaired by the mutations. A study of the dependence of initial rates of self-phosphorylation on cytoplasmic domain concentration showed that the order of reaction increased with the progress of self-phosphorylation. Both pre-phosphorylation and high concentrations of ammonium sulphate restored maximal or near-maximal levels of self-phosphorylation in the mutants, possibly through compensating conformational changes. A plausible homology model, based on the cyclic AMP-dependent protein kinase catalytic subunit, accommodated the sequence Glu-841-Glu-Lys-Glu as an insertion in the peptide binding loop at the edge of the active site cleft. The model suggests that Glu-844 and His-846 may participate in H-bonding interactions, thus stabilizing the active site region, while Glu-842 does not appear to interact with regions of the catalytic core.

Amino Acid Sequence

Prediction from sequence comparisons of residues of factor H involved in the interaction with complement component C3b.

The amino acid sequence of the region of bovine factor H containing the C3b binding site has been derived from sequencing overlapping cDNA clones. A cDNA sequence encoding 669 amino acids was obtained. Like human and mouse factor H the sequence can be arranged into a number of internally homologous units (CPs), each of which is about 60 amino acids long and is based on a framework of four conserved cysteine residues. Bovine factor H is of the same molecular mass as human and mouse factor H, and is therefore likely to be composed of 20 contiguous CPs. Comparisons with human and mouse factor H indicate that the partial bovine sequence encodes CPs 2-12 inclusive of bovine factor H. Bovine factor H binds to human ammonia-treated C3 (causing thiolester cleavage) [C3(NH3)] and promotes the cleavage of human C3(NH3) in the presence of bovine factor I. Other studies indicate that CPs 2-5 of human factor H encompass the C3b binding and factor I cofactor activity site. Multiple sequence alignments of human factor H, mouse factor H (which also interacts with human C3b) and bovine factor H with CP modules whose structures have been determined experimentally, have been used to predict residues in the hypervariable loops of CPs 2-5 and to identify residues of potential importance in human C3 binding and factor I cofactor activity. Leu-17 and Gly-20 of CP 2, Ser-17, Ala-19, Glu-21, Asp-23 and Glu-25 of CP 3 and Lys-18 of CP 4 are all conserved between the three species. It may be that CPs 3 and 4 interact with C3(NH3) directly, whilst CPs 2 and 5 maintain the correct orientation for CPs 3 and 4 to interact.

Amino Acid Sequence

Identification of sequence similarity between 60 kDa and 70 kDa molecular chaperones: evidence for a common evolutionary background?

Recent findings support the premise that chaperonins (60 kDa stress-proteins) and alpha-subunits of F-type ATPases (alpha-ATPase) are evolutionary related protein families. Two-dimensional gel patterns of synthesized proteins in unstressed and heat-shocked embryonic Drosophila melanogaster SL2 cells revealed that antibodies raised against the alpha-subunit of the F1-ATPase complex from rat liver recognize an inducible p71 member of the 70 kDa stress-responsive protein family. Molecular recognition of this stress-responsive 70 kDa protein by antibodies raised against the F1-ATPase alpha-subunit suggests the possibility of partial sequence similarity within these ATP-binding protein families. A multiple sequence alignment between alpha-ATPases and 60 kDa and 70 kDa molecular chaperones is presented. Statistical evaluation of sequence similarity reveals a significant degree of sequence conservation within the three protein families. The finding suggests a common evolutionary origin for the ATPases and molecular chaperone protein families of 60 kDa and 70 kDa, despite the lack of obvious structural resemblance between them.

Amino Acid Sequence

Comparative anatomy of the aldo-keto reductase superfamily.

The aldo-keto reductases metabolize a wide range of substrates and are potential drug targets. This protein superfamily includes aldose reductases, aldehyde reductases, hydroxysteroid dehydrogenases and dihydrodiol dehydrogenases. By combining multiple sequence alignments with known three-dimensional structures and the results of site-directed mutagenesis studies, we have developed a structure/function analysis of this superfamily. Our studies suggest that the (alpha/beta)8-barrel fold provides a common scaffold for an NAD(P)(H)-dependent catalytic activity, with substrate specificity determined by variation of loops on the C-terminal side of the barrel. All the aldo-keto reductases are dependent on nicotinamide cofactors for catalysis and retain a similar cofactor binding site, even among proteins with less than 30% amino acid sequence identity. Likewise, the aldo-keto reductase active site is highly conserved. However, our alignments indicate that variation ofa single residue in the active site may alter the reaction mechanism from carbonyl oxidoreduction to carbon-carbon double-bond reduction, as in the 3-oxo-5beta-steroid 4-dehydrogenases (Delta4-3-ketosteroid 5beta-reductases) of the superfamily. Comparison of the proposed substrate binding pocket suggests residues 54 and 118, near the active site, as possible discriminators between sugar and steroid substrates. In addition, sequence alignment and subsequent homology modelling of mouse liver 17beta-hydroxysteroid dehydrogenase and rat ovary 20alpha-hydroxysteroid dehydrogenase indicate that three loops on the C-terminal side of the barrel play potential roles in determining the positional and stereo-specificity of the hydroxysteroid dehydrogenases. Finally, we propose that the aldo-keto reductase superfamily may represent an example of divergent evolution from an ancestral multifunctional oxidoreductase and an example of convergent evolution to the same active-site constellation as the short-chain dehydrogenase/reductase superfamily.

Alcohol Oxidoreductases

Cloning and sequencing of four new mammalian monocarboxylate transporter (MCT) homologues confirms the existence of a transporter family with an ancient past.

Measurement of monocarboxylate transport kinetics in a range of cell types has provided strong circumstantial evidence for a family of monocarboxylate transporters (MCTs). Two mammalian MCT isoforms (MCT1 and MCT2) and a chicken isoform (REMP or MCT3) have already been cloned, sequenced and expressed, and another MCT-like sequence (XPCT) has been identified. Here we report the identification of new human MCT homologues in the database of expression sequence tags and the cloning and sequencing of four new full-length MCT-like sequences from human cDNA libraries, which we have denoted MCT3, MCT4, MCT5 and MCT6. Northern blotting revealed a unique tissue distribution for the expression of mRNA for each of the seven putative MCT isoforms (MCT1-MCT6 and XPCT). All sequences were predicted to have 12 transmembrane (TM) helical domains with a large intracellular loop between TM6 and TM7. Multiple sequence alignments showed identities ranging from 20% to 55%, with the greatest conservation in the predicted TM regions and more variation in the C-terminal than the N-terminal region. Searching of additional sequence databases identified candidate MCT homologues from the yeast Saccharomyces cerevisiae, the nematode worm Caenorhabditis elegans and the archaebacterium Sulfolobus solfataricus. Together these sequences constitute a new family of transporters with some strongly conserved sequence motifs, the possible functions of which are discussed.

Amino Acid Sequence

The profilin multigene family of maize: differential expression of three isoforms.

Profilin is a small (12-15 kDa) actin- and phospholipid-binding protein previously known only from studies on animals and lower eukaryotes but recently identified as a birch pollen allergen. Here we have identified and characterized three members of the profilin multigene family from the plant Zea mays. Two cDNAs isolated from a maize pollen library (ZmPRO 1 and ZmPRO 3) each have a single, large open reading frame encoding a putative polypeptide 131 amino acids long with a predicted molecular weight of approximately 14 kDa. A third maize pollen cDNA (ZmPRO 2) has two in-frame translation initiation codons. Use of the first ATG would result in a polypeptide 137 amino acids long with a molecular weight of 14.8 kDa. The three maize profilins are highly homologous to each other (> 90% nucleotide and amino acid sequence identity) as well as other plant profilins but show far less similarity (30-40% amino acid sequence identity) to animal and lower eukaryote profilins. Multiple sequence alignments indicate that only nine residues are shared by all eukaryotic profilins examined. However, limited comparisons reveal domains in the NH2 and COOH termini that have a high degree of similarity suggesting functional conservation. The maize gene family size is estimated to contain three to six members based on Southern blot experiments with gene-specific and coding region probes. Northern blot analysis demonstrates that the three maize profilin cDNAs characterized here are utilized in a tissue-specific manner and are anther or pollen specific.

Actins

Equus caballus gelsolin--cDNA sequence and protein structural implications.

We have generated and characterized the cDNA from equine smooth muscle that encodes gelsolin, an actin-modulating protein. Overlapping cDNA clones synthesized by the reverse transcriptase/polymerase chain reaction and clones isolated from a horse genomic library provided the complete primary structure for the intracellular isoform of gelsolin, while cDNA complemented with protein sequence data produced the full-length primary transcript of the gelsolin isoform found circulating in equine plasma. The deduced amino acid sequences of the intracellular and secreted versions of equine gelsolin infer polypeptides of 731 and 755 residues with apparent molecular masses of 80.7 kDa and 83.2 kDa, respectively. Multiple sequence alignment analysis of equine, human, porcine, and murine orthologs of gelsolin demonstrates prominent similarities among all of these proteins, with the horse and human molecules exhibiting the largest degree of likeness with respect to polypeptide length and overall sequence composition. Both horse and human plasma gelsolins are comprised of 755 amino acids with 94% of the residues identical, while the degree of sequence identity in the shorter (731 residues) cytoplasmic gelsolins is 95%. Analysis of the sequences and structures of the six related domains that comprise gelsolin emphasizes the strong correlation that exists between primary structural conservation among mammalian gelsolins and maintenance of the three-dimensional domain fold characteristic of members of this protein family.

Amino Acid Sequence

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking," Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana, contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a swapped domain topology; the other has no disulfide bonds and possesses two distinct, tandem domains. To explore the structural evolution of bicycle proteins, we attempted to predict bicycle protein structures with Alphafold2 (AF2) and other deep learning programs. While AF2 did not recover the two experimental structures using existing databases, it succeeded when provided with multiple sequence alignments (MSAs) of protein sequences from newly sequenced closely related species. Using this approach, we generated 2,400 high-confidence bicycle protein predictions from seven aphid species. While all aphid bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance. Our extension of AF2 with custom MSAs of proteins from closely related species provides a generalizable, powerful approach for predicting structures of rapidly evolving protein families.

Animals

Efficient algorithms for molecular sequence analysis.

Efficient (linear time) algorithms are described for identifying global molecular sequence features allowing for errors including repeats, matches between sequences, dyad symmetry pairings, and other sequence patterns. A multiple sequence alignment algorithm is also described. Specific applications are given to hepatitis B viruses and the J5-C (J, joining; C, constant) region of the immunoglobulin kappa gene.

Algorithms

Structural model of the nucleotide-binding conserved component of periplasmic permeases.

The amino acid sequences of 17 bacterial membrane proteins that are components of periplasmic permeases and function in the uptake of a variety of small molecules and ions are highly homologous to each other and contain sequence motifs characteristic of nucleotide-binding proteins. These proteins are known to bind ATP and are postulated to be the energy-coupling components of the permeases. Several medically important eukaryotic proteins, including the multidrug-resistance transporters and the protein encoded by the cystic fibrosis gene, are also homologous to this family. By multiple sequence alignment of these 17 proteins, the consensus sequence, secondary structure, and surface exposure were predicted. The secondary structural motifs that are conserved among nucleotide-binding proteins were identified in adenylate kinase, p21ras, and elongation factor Tu by superposition of their known tertiary structures. The equivalent secondary structural elements in the predicted conserved component were located. These, together with sequence information, served as guides for alignment with adenylate kinase. A model for the structure of the ATP-binding domain of the permease proteins is proposed by analogy to the adenylate kinase structure. The characteristics of several permease mutations and biochemical data lend support to the model.

Adenylate Kinase

GTPase domains of ras p21 oncogene protein and elongation factor Tu: analysis of three-dimensional structures, sequence families, and functional sites.

GTPase domains are functional and structural units employed as molecular switches in a variety of important cellular functions, such as growth control, protein biosynthesis, and membrane traffic. Amino acid sequences of more than 100 members of different subfamilies are known, but crystal structures of only mammalian ras p21 and bacterial elongation factor Tu have been determined. After optimal superposition of these remarkably similar structures, careful multiple sequence alignment, and calculation of residue-residue interactions, we analyzed the two subfamilies in terms of structural conservation, sequence conservation, and residue contact strength. There are three main results. (i) A structure-based alignment of p21 and elongation factor Tu. (ii) The definition of a common conserved structural core that may be useful as the basis of model building by homology of the three-dimensional structure of any GTPase domain. (iii) Identification of sequence regions, other than the effector loop and the nucleotide binding site, that may be involved in the functional cycle: they are loop L4, known to change conformation after GTP hydrolysis; helix alpha 2, especially Arg-73 and Met-67 in ras p21; loops L8 and L10, including ras p21 Arg-123, Lys-147, and Leu-120; and residues located spatially near the N and C termini. These regions are candidate sites for interaction either with the GTP/GDP exchange factor, with a GTPase-affected function, or with a molecule delivered to a destination site with the aid of the GTPase domain.

Amino Acid Sequence

Improved prediction of protein secondary structure by use of sequence profiles and neural networks.

The explosive accumulation of protein sequences in the wake of large-scale sequencing projects is in stark contrast to the much slower experimental determination of protein structures. Improved methods of structure prediction from the gene sequence alone are therefore needed. Here, we report a substantial increase in both the accuracy and quality of secondary-structure predictions, using a neural-network algorithm. The main improvements come from the use of multiple sequence alignments (better overall accuracy), from "balanced training" (better prediction of beta-strands), and from "structure context training" (better prediction of helix and strand lengths). This method, cross-validated on seven different test sets purged of sequence similarity to learning sets, achieves a three-state prediction accuracy of 69.7%, significantly better than previous methods. In addition, the predicted structures have a more realistic distribution of helix and strand segments. The predictions may be suitable for use in practice as a first estimate of the structural type of newly sequenced proteins.

Amino Acid Sequence