Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Structure-guided analysis reveals nine sequence motifs conserved among DNA amino-methyltransferases, and suggests a catalytic mechanism for these enzymes.

Previous X-ray crystallographic studies have revealed that the catalytic domain of a DNA methyltransferase (Mtase) generating C5-methylcytosine bears a striking structural similarity to that of a Mtase generating N6-methyladenine. Guided by this common structure, we performed a multiple sequence alignment of 42 amino-Mtases (N6-adenine and N4-cytosine). This comparison revealed nine conserved motifs, corresponding to the motifs I to VIII and X previously defined in C5-cytosine Mtases. The amino and C5-cytosine Mtases thus appear to be more closely related than has been appreciated. The amino Mtases could be divided into three groups, based on the sequential order of motifs, and this variation in order may explain why only two motifs were previously recognized in the amino Mtases. The Mtases grouped in this way show several other group-specific properties, including differences in amino acid sequence, molecular mass and DNA sequence specificity. Surprisingly, the N4-cytosine and N6-adenine Mtases do not form separate groups. These results have implications for the catalytic mechanisms, evolution and diversification of this family of enzymes. Furthermore, a comparative analysis of the S-adenosyl-L-methionine and adenine/cytosine binding pockets suggests that, structurally and functionally, they are remarkably similar to one another.

Amino Acid Sequence↗

Thermodynamic prediction of conserved secondary structure: application to the RRE element of HIV, the tRNA-like element of CMV and the mRNA of prion protein.

An algorithm for prediction of conserved secondary structure of single-stranded RNA is presented. For each RNA of a set of homologous RNAs optimal and suboptimal secondary structures are calculated and stored in a base-pair probability matrix. A multiple sequence alignment is performed for the set of RNAs. The resulting gaps are introduced into the individual probability matrices. These homologous probability matrices are summed to give a consensus probability matrix emphasizing the conserved secondary structure elements of the RNA set. Thus the algorithm combines the advantages of thermodynamic structure prediction by energy minimization with the information obtained from phylogenetic alignment of sequences. The algorithm is applied to three examples. The REV-responsive element of HIV, the structure of which is well known from the literature, was chosen to test the algorithm. The second example is the 3' terminal segment of genomic single-stranded RNAs of cucumber mosaic viruses; a structure similar to that of the related brome mosaic virus was expected and was confirmed. The third example is the prion-protein mRNA from different organisms; the structure of this mRNA is not known. By application of the algorithm highly conserved hairpins were found in the prion-protein mRNA.

Algorithms↗

Using evolutionary trees in protein secondary structure prediction and other comparative sequence analyses.

Previously proposed methods for protein secondary structure prediction from multiple sequence alignments do not efficiently extract the evolutionary information that these alignments contain. The predictions of these methods are less accurate than they could be, because of their failure to consider explicitly the phylogenetic tree that relates aligned protein sequences. As an alternative, we present a hidden Markov model approach to secondary structure prediction that more fully uses the evolutionary information contained in protein sequence alignments. A representative example is presented, and three experiments are performed that illustrate how the appropriate representation of evolutionary relatedness can improve inferences. We explain why similar improvement can be expected in other secondary structure prediction methods and indeed any comparative sequence analysis method.

Amino Acid Sequence↗

Correlated mutations contain information about protein-protein interaction.

Many proteins have evolved to form specific molecular complexes and the specificity of this interaction is essential for their function. The network of the necessary inter-residue contacts must consequently constrain the protein sequences to some extent. In other words, the sequence of an interacting protein must reflect the consequence of this process of adaptation. It is reasonable to assume that the sequence changes accumulated during the evolution of one of the interacting proteins must be compensated by changes in the other. Here we apply a method for detecting correlated changes in multiple sequence alignments to a set of interacting protein domains and show that positions where changes occur in a correlated fashion in the two interacting molecules tend to be close to the protein-protein interfaces. This leads to the possibility of developing a method for predicting contacting pairs of residues from the sequence alone. Such a method would not need the knowledge of the structure of the interacting proteins, and hence would be both radically different and more widely applicable than traditional docking methods. We indeed demonstrate here that the information about correlated sequence changes is sufficient to single out the right inter-domain docking solution amongst many wrong alternatives of two-domain proteins. The same approach is also used here in one case (haemoglobin) where we attempt to predict the interface of two different proteins rather than two protein domains. Finally, we report here a prediction about the inter-domain contact regions of the heat- shock protein Hsc70 based only on sequence information.

Amino Acid Sequence↗

Detection of protein three-dimensional side-chain patterns: new examples of convergent evolution.

Detection of recurring three-dimensional side-chain patterns is a potential means of inferring protein function. This paper presents a new method for detecting such patterns and discusses various implications. The method allows detection of side-chain patterns without any prior knowledge of function, requiring only protein structure data and associated multiple sequence alignments. A recursive, depth-first search algorithm finds all possible groups of identical amino acids common to two protein structures independent of sequence order. The search is highly constrained by distance constraints, and by ignoring amino acids unlikely to be involved in protein function. A weighted root-mean-square deviation (RMSD) between equivalenced groups of amino acids is used as a measure of similarity. The statistical significance of any RMSD is assigned by reference to a distribution fitted to simulated data. Searches with the Ser/His/Asp catalytic triad, a His/His porphyrin binding pattern, and the zinc-finger Cys/Cys/His/His pattern are performed to test the method on known examples. An all-against-all comparison of representatives from the structural classification of proteins (SCOP) is performed, revealing several new examples of evolutionary convergence to common patterns of side-chains within different tertiary folds and in different orders along the sequence. These include a di-zinc binding Asp/Asp/His/His/Ser pattern common to alkaline phosphatase/bacterial aminopeptidase, and an Asp/Glu/His/His/Asn/Asn pattern common to the active sites of DNase I and endocellulase E1. Implications for protein evolution, function prediction and the rational design of functional regulators are discussed.

Acetylglucosaminidase↗

Evolutionary conserved rigid module-domain interactions can be detected at the sequence level: the examples of complement and blood coagulation proteases.

Several extracellular modular proteins, including proteases of the complement and blood coagulation cascades, are shown here to exhibit conserved sequence patterns specific for a particular module-domain association. This was detected by comparative analysis of sequence variability in different multiple sequence alignments, which provides a new tool to investigate the evolution of modular proteins. A first example deals with the proteins featuring a common complement control protein (CCP) module-serine protease (SP) domain pattern at their C-terminal end, defined here as the CCP-SP sub-family. These proteins include the complement proteases C1r, C1s and MASPs, the Limulus clotting factor C, and the proteins of the haptoglobin family. A second example deals with blood coagulation factors VII, IX and X and protein C, all featuring a common epidermal growth factor (EGF)-SP C-terminal assembly. Highly specific motifs are found at the connection between the CCP or EGF module and the activation peptide of the SP domain: [P/A]-x-C-x-[P/A]-[I/V]-C-G-x-[P/S/K] in the case of the CCP-SP proteins, and C-x-[P/S]-x-x-x-[Y/F]-P-C-G in the case of the EGF-SP proteins. Each motif is strictly conserved in the whole sub-family and it is detected in no more than one other known protein sequence. Strikingly, most of the conserved residues specific to each sub-family appear to be clustered at the interface between the SP domain and the CCP or EGF module. We propose that a rigid module-domain interaction occurs in these proteins and has been conserved through evolution. The functional implications of these assemblies, underlined by such evolutionary constraints, are discussed.

Amino Acid Sequence↗

Novel molecular architecture of the multimeric archaeal PEP-synthase homologue (MAPS) from Staphylothermus marinus.

The phosphoenolpyruvate (PEP)-synthases belong to the family of structurally and functionally related PEP-utilizing enzymes. The only archaeal member of this family characterized thus far is the Multimeric Archaeal PEP-Synthase homologue from Staphylothermus marinus (MAPS). This protein complex differs from the bacterial and eukaryotic representatives characterized to date in its homomultimeric, as opposed to dimeric or tetrameric, structure. We have probed the molecular architecture of MAPS using limited proteolytic digestion in conjunction with electron microscopic, biochemical, and biophysical techniques. The 2.2 MDa particle was found to be organized in a concentric fashion. The 93.7 kDa monomers possess a pronounced tripartite domain structure and are arranged such that the N-terminal domains form an outer shell, the intermediate domains form an inner shell, and the C-terminal domains form a core structure responsible for the assembly into a multimeric complex. The core domain was shown to be capable of assembling into the native multimer by recombinant expression in Escherichia coli. Deletion mutants as well as a synthetic peptide were investigated for their state of oligomerization using native polyacrylamide gel electrophoresis, molecular sieve chromatography, analytical ultracentrifugation, circular dichroism (CD) spectroscopy, and chemical cross-linking. Our data confirmed the existence of a short C-terminal, alpha-helical oligomerization motif that had been suggested by multiple sequence alignments and secondary structure predictions. We propose that this motif bundles the monomers into six groups of four. An additional formation of 12 dimers between globular domains from different bundles leads to the multimeric assembly. According to our model, each of the six bundles of globular domains is positioned at the corners of an imaginary octahedron, and the helical C-terminal segments are oriented towards the centre of the particle. The edges of the octahedron represent the dimeric contacts. Phylogenetic analysis suggests that the ancient predecessor of this family of enzymes contained the C-terminal oligomerization motif as a feature that was preserved in some hyperthermophiles.

Amino Acid Sequence↗

The binding site for UCH-L3 on ubiquitin: mutagenesis and NMR studies on the complex between ubiquitin and UCH-L3.

The ubiquitin fold is a versatile and widely used targeting signal that is added post-translationally to a variety of proteins. Covalent attachment of one or more ubiquitin domains results in localization of the target protein to the proteasome, the nucleus, the cytoskeleton or the endocytotic machinery. Recognition of the ubiquitin domain by a variety of enzymes and receptors is vital to the targeting function of ubiquitin. Several parallel pathways exist and these must be able to distinguish among ubiquitin, several different types of polymeric ubiquitin, and the various ubiquitin-like domains. Here we report the first molecular description of the binding site on ubiquitin for ubiquitin C-terminal hydrolase L3 (UCH-L3). The site on ubiquitin was experimentally determined using solution NMR, and site-directed mutagenesis. The site on UCH-L3 was modeled based on X-ray crystallography, multiple sequence alignments, and computer-aided docking. Basic residues located on ubiquitin (K6, K11, R72, and R74) are postulated to contact acidic residues on UCH-L3 (E10, E14, D33, E219). These putative interactions are testable and fully explain the selectivity of ubiquitin domain binding to this enzyme.

Allosteric Site↗

Effective use of sequence correlation and conservation in fold recognition.

Protein families are a rich source of information; sequence conservation and sequence correlation are two of the main properties that can be derived from the analysis of multiple sequence alignments. Sequence conservation is related to the direct evolutionary pressure to retain the chemical characteristics of some positions in order to maintain a given function. Sequence correlation is attributed to the small sequence adjustments needed to maintain protein stability against constant mutational drift. Here, we showed that sequence conservation and correlation were each frequently informative enough to detect incorrectly folded proteins. Furthermore, combining conservation, correlation, and polarity, we achieved an almost perfect discrimination between native and incorrectly folded proteins. Thus, we made use of this information for threading by evaluating the models suggested by a threading method according to the degree of proximity of the corresponding correlated, conserved, and apolar residues. The results showed that the fold recognition capacity of a given threading approach could be improved almost fourfold by selecting the alignments that score best under the three different sequence-based approaches.

Amino Acid Sequence↗

Structural clues in the sequences of the aquaporins.

The large number of sequences available for the aquaporin family represents a valuable source of information to incorporate into three-dimensional structure determination. Phylogenetic analysis was used to define type sequences to avoid extreme over-representation of some subfamilies, and as a measure of the quality of multiple sequence alignment. Inspection of the sequence alignment suggested eight conserved segments that define the core architecture of six transmembrane helices and two functional loops, B and E, projecting into the plane of the membrane. The sum of the core segments and the minimum lengths of the interlinking loops constitute the 208 residues necessary to satisfy the aquaporin architecture. Analysis of hydrophobic and conservation periodicity and of correlated mutations across the alignment indicated the likely assignment and orientation of the helices in the bilayer. This assignment is examined with respect to the structure of the erythrocyte aquaporin 1 determined by electron crystallography. The aquaporin 1 tetramer is described as three rings of helices, each ring with a different exposure to the lipid environment. The sequence analysis clearly suggests that two helices are exposed along their whole lengths, two helices are exposed only at their N termini, and two helices are not exposed to lipid. It is further proposed that, besides loops B and E, the highly conserved motifs on helices 1 and 4, ExxxTxxF/L, could line the water channel.

Amino Acid Sequence↗

Co-evolution of proteins with their interaction partners.

The divergent evolution of proteins in cellular signaling pathways requires ligands and their receptors to co-evolve, creating new pathways when a new receptor is activated by a new ligand. However, information about the evolution of binding specificity in ligand-receptor systems is difficult to glean from sequences alone. We have used phosphoglycerate kinase (PGK), an enzyme that forms its active site between its two domains, to develop a standard for measuring the co-evolution of interacting proteins. The N-terminal and C-terminal domains of PGK form the active site at their interface and are covalently linked. Therefore, they must have co-evolved to preserve enzyme function. By building two phylogenetic trees from multiple sequence alignments of each of the two domains of PGK, we have calculated a correlation coefficient for the two trees that quantifies the co-evolution of the two domains. The correlation coefficient for the trees of the two domains of PGK is 0. 79, which establishes an upper bound for the co-evolution of a protein domain with its binding partner. The analysis is extended to ligands and their receptors, using the chemokines as a model. We show that the correlation between the chemokine ligand and receptor trees' distances is 0.57. The chemokine family of protein ligands and their G-protein coupled receptors have co-evolved so that each subgroup of chemokine ligands has a matching subgroup of chemokine receptors. The matching subfamilies of ligands and their receptors create a framework within which the ligands of orphan chemokine receptors can be more easily determined. This approach can be applied to a variety of ligand and receptor systems.

Chemokines↗

An aspartic acid residue in TPR-1, a specific region of protein-priming DNA polymerases, is required for the functional interaction with primer terminal protein.

A multiple sequence alignment of eukaryotic-type DNA polymerases led to the identification of two regions of amino acid residues that are only present in the group of DNA polymerases that make use of terminal proteins. (TPs) as primers to initiate DNA replication of linear genomes. These amino acid regions (named terminal region (TPR protein-1 and TPR-2) are inserted between the generally conserved motifs Dx(2)SLYP and Kx(3)NSxYG (TPR-1) and motifs Kx(3)NSxYG and YxDTDS (TPR-2) of the eukaryotic-type family of DNA polymerases. We carried out site-directed mutagenesis in two of the most conserved residues of phi29 DNA polymerase TPR-1 to study the possible role of this specific region. Two mutant DNA polymerases, in conserved residues AsP332 and Leu342, were purified and subjected to a detailed biochemical analysis of their enzymatic activities. Both mutant DNA polymerases were essentially normal when assayed for synthetic activities in DNA-primed reactions. However, mutant D332Y was drastically affected in phi29 TP-DNA replication as a consequence of a large reduction in the catalytic efficiency of the protein-primed reactions. The molecular basis of this defect is a non-functional interaction with TP that strongly reduces the activity of the DNA polymerase/TP heterodimer.

Amino Acid Motifs↗

ConSurf: an algorithmic tool for the identification of functional regions in proteins by surface mapping of phylogenetic information.

Experimental approaches for the identification of functionally important regions on the surface of a protein involve mutagenesis, in which exposed residues are replaced one after another while the change in binding to other proteins or changes in activity are recorded. However, practical considerations limit the use of these methods to small-scale studies, precluding a full mapping of all the functionally important residues on the surface of a protein. We present here an alternative approach involving the use of evolutionary data in the form of multiple-sequence alignment for a protein family to identify hot spots and surface patches that are likely to be in contact with other proteins, domains, peptides, DNA, RNA or ligands. The underlying assumption in this approach is that key residues that are important for binding should be conserved throughout evolution, just like residues that are crucial for maintaining the protein fold, i.e. buried residues. A main limitation in the implementation of this approach is that the sequence space of a protein family may be unevenly sampled, e.g. mammals may be overly represented. Thus, a seemingly conserved position in the alignment may reflect a taxonomically uneven sampling, rather than being indicative of structural or functional importance. To avoid this problem, we present here a novel methodology based on evolutionary relations among proteins as revealed by inferred phylogenetic trees, and demonstrate its capabilities for mapping binding sites in SH2 and PTB signaling domains. A computer program that implements these ideas is available freely at: http://ashtoret.tau.ac.il/ approximately rony

Algorithms↗

Three-dimensional cluster analysis identifies interfaces and functional residue clusters in proteins.

Three-dimensional cluster analysis offers a method for the prediction of functional residue clusters in proteins. This method requires a representative structure and a multiple sequence alignment as input data. Individual residues are represented in terms of regional alignments that reflect both their structural environment and their evolutionary variation, as defined by the alignment of homologous sequences. From the overall (global) and the residue-specific (regional) alignments, we calculate the global and regional similarity matrices, containing scores for all pairwise sequence comparisons in the respective alignments. Comparing the matrices yields two scores for each residue. The regional conservation score (C(R)(x)) defines the conservation of each residue x and its neighbors in 3D space relative to the protein as a whole. The similarity deviation score (S(x)) detects residue clusters with sequence similarities that deviate from the similarities suggested by the full-length sequences. We evaluated 3D cluster analysis on a set of 35 families of proteins with available cocrystal structures, showing small ligand interfaces, nucleic acid interfaces and two types of protein-protein interfaces (transient and stable). We present two examples in detail: fructose-1,6-bisphosphate aldolase and the mitogen-activated protein kinase ERK2. We found that the regional conservation score (C(R)(x)) identifies functional residue clusters better than a scoring scheme that does not take 3D information into account. C(R)(x) is particularly useful for the prediction of poorly conserved, transient protein-protein interfaces. Many of the proteins studied contained residue clusters with elevated similarity deviation scores. These residue clusters correlate with specificity-conferring regions: 3D cluster analysis therefore represents an easily applied method for the prediction of functionally relevant spatial clusters of residues in proteins.

Adenosine Triphosphate↗

A novel method of protein secondary structure prediction with high segment overlap measure: support vector machine approach.

We have introduced a new method of protein secondary structure prediction which is based on the theory of support vector machine (SVM). SVM represents a new approach to supervised pattern classification which has been successfully applied to a wide range of pattern recognition problems, including object recognition, speaker identification, gene function prediction with microarray expression profile, etc. In these cases, the performance of SVM either matches or is significantly better than that of traditional machine learning approaches, including neural networks.The first use of the SVM approach to predict protein secondary structure is described here. Unlike the previous studies, we first constructed several binary classifiers, then assembled a tertiary classifier for three secondary structure states (helix, sheet and coil) based on these binary classifiers. The SVM method achieved a good performance of segment overlap accuracy SOV=76.2 % through sevenfold cross validation on a database of 513 non-homologous protein chains with multiple sequence alignments, which out-performs existing methods. Meanwhile three-state overall per-residue accuracy Q(3) achieved 73.5 %, which is at least comparable to existing single prediction methods. Furthermore a useful "reliability index" for the predictions was developed. In addition, SVM has many attractive features, including effective avoidance of overfitting, the ability to handle large feature spaces, information condensing of the given data set, etc. The SVM method is conveniently applied to many other pattern classification tasks in biology.

Computer Simulation↗

Automated structure-based prediction of functional sites in proteins: applications to assessing the validity of inheriting protein function from homology in genome annotation and to protein docking.

A major problem in genome annotation is whether it is valid to transfer the function from a characterised protein to a homologue of unknown activity. Here, we show that one can employ a strategy that uses a structure-based prediction of protein functional sites to assess the reliability of functional inheritance. We have automated and benchmarked a method based on the evolutionary trace approach. Using a multiple sequence alignment, we identified invariant polar residues, which were then mapped onto the protein structure. Spatial clusters of these invariant residues formed the predicted functional site. For 68 of 86 proteins examined, the method yielded information about the observed functional site. This algorithm for functional site prediction was then used to assess the validity of transferring the function between homologues. This procedure was tested on 18 pairs of homologous proteins with unrelated function and 70 pairs of proteins with related function, and was shown to be 94 % accurate. This automated method could be linked to schemes for genome annotation. Finally, we examined the use of functional site prediction in protein-protein and protein-DNA docking. The use of predicted functional sites was shown to filter putative docked complexes with a discrimination similar to that obtained by manually including biological information about active sites or DNA-binding residues.

Algorithms↗

Crystal structure of Rv2118c: an AdoMet-dependent methyltransferase from Mycobacterium tuberculosis H37Rv.

Rv2118c belongs to the class of conserved hypothetical proteins from Mycobacterium tuberculosis H37Rv. The crystal structure of Rv2118c in complex with S-adenosyl-l-methionine (AdoMet) has been determined at 1.98 A resolution. The crystallographic asymmetric unit consists of a monomer, but symmetry-related subunits interact extensively, leading to a tetrameric structure. The structure of the monomer can be divided functionally into two domains: the larger catalytic C-terminal domain that binds the cofactor AdoMet and is involved in the transfer of methyl group from AdoMet to the substrate and a smaller N-terminal domain. The structure of the catalytic domain is very similar to that of other AdoMet-dependent methyltransferases. The N-terminal domain is primarily a beta-structure with a fold not found in other methyltransferases of known structure. Database searches reveal a conserved family of Rv2118c-like proteins from various organisms. Multiple sequence alignments show several regions of high sequence similarity (motifs) in this family of proteins. Structure analysis and homology to yeast Gcd14p suggest that Rv2118c could be an RNA methyltransferase, but further studies are required to establish its functional role conclusively.

Amino Acid Motifs↗

Characterization of disease-associated single amino acid polymorphisms in terms of sequence and structure properties.

In the present work, we use structural information to characterize a set of disease-associated single amino acid polymorphisms exhaustively. The analysis of different properties, such as substitution matrix elements, secondary structure, accessibility, free energies of transfer from water to octanol, amino acid volume, etc., suggests that many disease-causing mutations are associated with extreme changes in the value of parameters relating to protein stability. Overall, our results indicate that, while knowledge of protein structure clearly helps in understanding these mutations, a finer understanding can come only from a quantitative knowledge of protein stability and of the protein environment in the cell. Interestingly, use of evolutionary information from multiple sequence alignments can be used to increase our knowledge of disease-associated mutations.

Computational Biology↗