Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Phylogeny for the faint of heart: a tutorial.

Phylogenetic trees seem to be finding ever broader applications, and researchers from very different backgrounds are becoming interested in what they might have to say. This tutorial aims to introduce the basics of building and interpreting phylogenetic trees. It is intended for those wanting to understand better what they are looking at when they look at someone else's trees or to begin learning how to build their own. Topics covered include: how to read a tree, assembling a dataset, multiple sequence alignment (how it works and when it does not), phylogenetic methods, bootstrap analysis and long-branch artefacts, and software and resources.

Amino Acid Sequence↗

Knowledge-based grouping of modeled HLA peptide complexes.

Human leukocyte antigens are the most polymorphic of human genes and multiple sequence alignment shows that such polymorphisms are clustered in the functional peptide binding domains. Because of such polymorphism among the peptide binding residues, the prediction of peptides that bind to specific HLA molecules is very difficult. In recent years two different types of computer based prediction methods have been developed and both the methods have their own advantages and disadvantages. The nonavailability of allele specific binding data restricts the use of knowledge-based prediction methods for a wide range of HLA alleles. Alternatively, the modeling scheme appears to be a promising predictive tool for the selection of peptides that bind to specific HLA molecules. The scoring of the modeled HLA-peptide complexes is a major concern. The use of knowledge based rules (van der Waals clashes and solvent exposed hydrophobic residues) to distinguish binders from nonbinders is applied in the present study. The rules based on (1) number of observed atomic clashes between the modeled peptide and the HLA structure, and (2) number of solvent exposed hydrophobic residues on the modeled peptide effectively discriminate experimentally known binders from poor/nonbinders. Solved crystal complexes show no vdW Clash (vdWC) in 95% cases and no solvent exposed hydrophobic peptide residues (SEHPR) were seen in 86% cases. In our attempt to compare experimental binding data with the predicted scores by this scoring scheme, 77% of the peptides are correctly grouped as good binders with a sensitivity of 71%.

Alleles↗

Molecular evolution of thyroid peroxidase.

Thyroid peroxidase is a member of a family of mammalian peroxidases that includes myeloperoxidase, lactoperoxidase, eosinophil peroxidase, and salivary peroxidase. Protein sequences showing a high degree of sequence similarity with mammalian peroxidases have recently been observed in several invertebrate species. A multiple sequence alignment prepared with five mammalian and six invertebrate peroxidases shows complete conservation of amino acid residues considered to be important in the formation of peroxidase compound 1. These include the distal and proximal histidines, a catalytic arginine residue, and an asparagine residue hydrogen bonded to the proximal histidine. TPO-2, an alternatively spliced form of TPO, lacks the essential asparagine (Asn 579). It is now possible to speak more broadly of the family of animal peroxidases, rather than mammalian peroxidases. The animal peroxidases comprise a group of homologous proteins that differ markedly from the plant/fungal/bacterial peroxidases in primary, secondary and tertiary structure, but which share with them a common function. Animal peroxidases probably arose independently of the plant/fungal/bacterial peroxidase superfamily and most likely belong to a different gene family. The relationship between animal and non-animal peroxidases probably represents an example of convergent evolution to a common enzymatic mechanism.

Amino Acid Sequence↗

Annealing function of GroEL: structural and bioinformatic analysis.

The Escherichia coli chaperonin system, GroEL-GroES, facilitates folding of substrate proteins (SPs) that are otherwise destined to aggregate. The iterative annealing mechanism suggests that the allostery-driven GroEL transitions leading to changes in the microenvironment of the SP constitutes the annealing action of chaperonins. To describe the molecular basis for the changes in the nature of SP-GroEL interactions we use the crystal structures of GroEL (T state), GroEL-ATP (R state) and the GroEL-GroES-(ADP)(7) (R" state) complex to determine the residue-specific changes in the accessible surface area and the number of tertiary contacts as a result of the T-->R-->R" transitions. We find large changes in the accessible area in many residues in the apical domain, but relatively smaller changes are associated with residues in the equatorial domain. In the course of the T-->R transition the microenvironment of the SP changes which suggests that GroEL is an annealing machine even without GroES. This is reflected in the exposure of Glu386 which loses six contacts in the T-->R transition. We also evaluate the conservation of residues that participate in the various chaperonin functions. Multiple sequence alignments and chemical sequence entropy calculations reveal that, to a large extent, only the chemical identities and not the residues themselves important for the nominal functions (peptide binding, nucleotide binding, GroES and substrate protein release) are strongly conserved. Using chemical sequence entropy, which is computed by classifying aminoacids into four types (hydrophobic, polar, positively charged and negatively charged) we make several new predictions that are relevant for peptide binding and annealing function of GroEL. We identify a number of conserved peptide binding sites in the apical domain which coincide with those found in the 1.7 A crystal structure of 'mini-chaperone' complexed with the N-terminal tag. Correlated mutations in the HSP60 family, that might control allostery in GroEL, are also strongly conserved. Most importantly, we find that charged solvent-exposed residues in the T state (Lys 226, Glu 252 and Asp 253) are strongly conserved. This leads to the prediction that mutating these residues, that control the annealing function of the SP, can decrease the efficacy of the chaperonin function.

Algorithms↗

Oxygen transport proteins: III. Structural studies of the scorpion (Buthus sindicus) hemocyanin, partial primary structure of its subunit Bsin1.

The hemocyanin (Hc) from Buthus sindicus, studied in the native state, demonstrated to be an aggregate of eight different types of subunits arranged in four cubic hexamers. Both, the 'top' and the 'side' views of the native molecule have been identified from the negatively stained specimens using transmission electron microscopy. Out of these, eight different polypeptide chains, the partial primary structure (68%) of a subunit Bsin1 (Mr = 72422.7 Da) was established using a combination of automated Edman degradation and mass spectrometry. A multiple sequence alignment with other closely related cheliceratan Hc subunits revealed average identities of ca. 60%. Most of the structurally important residues, i.e. copper and calcium-binding ligands, as well as the residues involved in the presumed oxygen entrance pathway, proved to be strictly conserved in Bsin1. Sequence variations have been observed around the functionally important chloride-binding site, not only for the B. sindicus subunit Bsin1, but also for the subunit Aaus-6 of the scorpion A. australis and the subunit Ecal-a from the spider Eurypelma californicum Hcs. Deviation in the primary structure related to the chloride-binding site suggest that the effect of chloride ions may vary in different hemocyanins. Furthermore, the secondary structural contents of the Hc subunit Bsin1 were determined by circular dichroism revealing ca. 33% alpha-helix, 18%, beta-sheet, 19% beta-turn, and 30% random coil composition. These values are in good agreement with the crystal structure of the closely related Hc subunit Lpol-II from horseshoe crab L. polyphemus. Electron microscopic studies of the purified Hc subunit under native conditions revealed that Bsin1 has self aggregation properties. Results of these studies are discussed.

Amino Acid Sequence↗

Structural similarities and evolutionary relationships in chloride-dependent alpha-amylases.

The alpha-amylase sequences contained in databanks were screened for the presence of amino acid residues Arg195, Asn298 and Arg/Lys337 forming the chloride-binding site of several specialized alpha-amylases allosterically activated by this anion. This search provides 38 alpha-amylases potentially binding a chloride ion. All belong to animals, including mammals, birds, insects, acari, nematodes, molluscs, crustaceans and are also found in three extremophilic Gram-negative bacteria. An evolutionary distance tree based on complete amino acid sequences was constructed, revealing four distinct clusters of species. On the basis of multiple sequence alignment and homology modeling, invariable structural elements were defined, corresponding to the active site, the substrate binding site, the accessory binding sites, the Ca(2+) and Cl(-) binding sites, a protease-like catalytic triad and disulfide bonds. The sequence variations within functional elements allowed engineering strategies to be proposed, aimed at identifying and modifying the specificity, activity and stability of chloride-dependent alpha-amylases.

Amino Acid Sequence↗

Structural/functional assignment of unknown bacteriophage T4 proteins by iterative database searches.

Among the total of 274 orfs within bacteriophage T4, only half have been reasonably well characterized, and the functions of the rest have remained obscure. In order to predict the molecular functions of the orfs, a position-specific iterated (PSI)-BLAST search of bacteriophage T4 against the sequence database of known 3D structures was carried out. PSI-BLAST is one of the most powerful iterative sequence search methods using multiple sequence alignment, with the ability to detect many more proteins with distant homology than standard pairwise methods. The 3D structures of proteins are considered to be better preserved than the sequences, and the detected distantly homologous proteins are likely to possess highly similar 3D structures. Thirteen orfs of phage T4, whose homologues were not detected by standard pairwise methods, were found to have significantly homologous counterparts by this method. The plausibility of the results was confirmed by checking whether important residues at substrate/ligand-binding sites were conserved. Among them, two orfs, vs.1 and e.1, which are similar to Escherichia coli lytic enzyme and MutT protein, respectively, had not been studied previously. Also, gp rIIA, a rapid lysis protein, whose gene structure had been intensively studied during the development of molecular biology in the 1950s and yet whose molecular function remains unknown, has an N-terminal domain that is significantly similar to the N-terminal region of the heat shock protein Hsp90.

Amino Acid Sequence↗

Cloning and expression of a nuclear encoded plastid specific 33 kDa ribonucleoprotein gene (33RNP) from pea that is light stimulated.

We report the cloning and sequencing of both cDNA and genomic DNA of a 33 kDa chloroplast ribonucleoprotein (33RNP) from pea. The analysis of the predicted amino acid sequence of the cDNA clone revealed that the encoded protein contains two RNA binding domains, including the conserved consensus ribonucleoprotein sequences CS-RNP1 and CS-RNP2, on the C-terminus half and the presence of a putative transit peptide sequence in the N-terminus region. The phylogenetic and multiple sequence alignment analysis of pea chloroplast RNP along with RNPs reported from the other plant sources revealed that the pea 33RNP is very closely related to Nicotiana sylvestris 31RNP and 28RNP and also to 31RNP and 28RNP of Arabidopsis and spinach, respectively. The pea 33RNP was expressed in Escherichia coli and purified to homogeneity. The in vitro import of precursor protein into chloroplasts confirmed that the N-terminus putative transit peptide is a bona fide transit peptide and 33RNP is localized in the chloroplast. The nucleic acid-binding properties of the recombinant protein, as revealed by South-Western analysis, showed that 33RNP has higher binding affinity for poly (U) and oligo dT than for ssDNA and dsDNA. The steady state transcript level was higher in leaves than in roots and the expression of this gene is light stimulated. Sequence analysis of the genomic clone revealed that the gene contains four exons and three introns. We have also isolated and analyzed the 5' flanking region of the pea 33RNP gene.

Amino Acid Sequence↗

Characterization of a Drosophila melanogaster orthologue of muskelin.

Muskelin was identified in vertebrates as a novel, intracellular, kelch repeat protein that is needed in cell-spreading responses to the matrix adhesion molecule, thrombospondin-1. The identification and characterization of an orthologue of muskelin in Drosophila melanogaster is now reported. The Drosophila muskelin gene, located on chromosome 2R, is encoded in ten exons. Drosophila muskelin is expressed in embryos, larvae and adult flies. The protein has 45% sequence identity to vertebrate muskelins, with highest sequence identity in an amino-terminal domain and the six kelch repeats that form a beta-propeller structure. Multiple sequence alignment of human, mouse, rat and Drosophila muskelins and protein database searches revealed a novel highly conserved motif within the amino-terminal domain, lissencephaly homology motif (LisH) and C-terminal to LisH motifs in the central region of the molecule, and several conserved consensus motifs for phosphorylation by protein kinase C and casein kinase II. These findings provide new information on the modular structure of muskelin and indicate potential for conserved mechanisms of function.

Amino Acid Sequence↗

Serine proteases and their homologs in the Drosophila melanogaster genome: an initial analysis of sequence conservation and phylogenetic relationships.

Serine proteases (SPs) and serine protease homologs (SPHs) constitute the second largest family of genes in the Drosophila melanogaster genome. Eighty-four SPs comprise less than 300 amino acid residues, and a significant portion of them are probably digestive enzymes. Some larger SPs may contain one or more regions important for protein-protein interactions, including clip domains, low-density lipoprotein receptor class A repeats, and scavenger receptor cysteine-rich domains. We identified 37 clusters of SP or SPH genes, which probably evolved from relatively recent gene duplication and sequence divergence. A majority of the SPs may be trypsin-like and activated by cleavage after a specific arginine or lysine residue. Among the 147 SPs and 57 SPHs studied, 24 SPs and 13 SPHs contain at least one regulatory clip domain. A multiple sequence alignment of the clip domains provided further information on structural conservation of these regulatory modules. Detailed sequence comparison led to an improved classification system for SPs containing clip domains. These analyses have established a framework of information about evolutionary relationships among the Drosophila SPs and SPHs, which may facilitate research on these proteins as well as homologous molecules from other invertebrate species.

Amino Acid Sequence↗

Identification of a novel member of the TGF-beta superfamily highly expressed in human placenta.

While conducting a gene discovery effort targeted to transcripts of the prevalent and intermediate frequency classes in placenta throughout gestation, we identified a novel member of the TGF-beta superfamily that is expressed at high levels in human placenta. Hence, we named this factor 'Placental Transforming Growth Factor Beta' (PTGFB). The full-length sequence of the 1.2-kb PTGFB mRNA has the potential of encoding a putative pre-pro-PTGFB protein of 295 amino acids and a putative mature PTGFB protein of 112 amino acids. Multiple sequence alignments of PTGFB and representative members of all TGF-beta subfamilies evidenced a number of conserved residues, including the seven cysteines that are almost invariant in all members of the TGF-beta superfamily. The single-copy PTGFB gene was shown to be composed of only two exons of 309 bp and 891 bp, separated by a 2.9-kb intron. The gene was localized to chromosome 19p12-13.1 by fluorescence in-situ hybridization. Northern analyses revealed a complex tissue-specific pattern of expression and a second transcript of 1.9 kb that is predominant in adult skeletal muscle. Most importantly, the 1.2-kb PTGFB transcript was shown to be expressed in placenta at much higher levels than in any other human fetal or adult tissue surveyed.

Adult↗

Evolution of the proximal promoter region of the mammalian growth hormone gene.

The evolutionary relationship between the proximal growth hormone (GH) gene promoter sequences of 12 mammalian species was explored by comparison of their trinucleotide composition and by multiple sequence alignment. Both approaches yielded results that were consistent with the known fossil record-based phylogeny of the analysed sequences, suggesting that the two methods of tree reconstruction might be equally efficient and reliable. The pattern of evolution inferred for the mammalian GH gene promoters was found to vary both temporally and spatially. Thus, two distinct regions devoid of any evolutionary changes exist in primates, but only one of these 'gaps' is also observed in rodents, and neither is seen in ruminants. Furthermore, different evolutionary rates must have prevailed during different periods of evolutionary time and in different lineages, with a dramatic increase in evolutionary rate apparent in primates. Since a similar pattern of discontinuity has been previously noted for the evolution of the GH-coding regions, it may reflect the action of positive selection operating upon the GH gene as a single cohesive unit. Strong evidence for the action of gene conversion between primate GH gene promoters is provided by the fact that the human GH1 and GH2 sequences, which are thought to have diverged before the divergence of Old World monkeys from great apes, are more similar to one another than either is to the rhesus monkey GH2 promoter. Finally, it was noted that a number of nucleotide positions in the GH1 gene promoter that are polymorphic in humans appear to be highly conserved in mammals. This apparent conundrum, which could represent a caveat for the interpretation of phylogenetic footprinting studies, is potentially explicable in terms either of reduced genetic diversity in highly inbred animal species or insufficient population data from non-human species.

Animals↗

Molecular phylogenetic analysis of felid herpesvirus 1.

The position of felid herpesvirus 1 within the alphaherpesvirus subfamily was investigated using molecular phylogenetic techniques applied to multiple sequence alignments of recently reported FHV-1 gene homologs (glycoprotein B, ribonucleotide reductase and DNA polymerase). FHV-1 was most closely related to other carnivore alphaherpesviruses, (phocid herpesvirus 1 and canid herpesvirus 1) and to the equid herpesviruses 1 and 4.

Alphaherpesvirinae↗

The identification of a conserved domain in both spartin and spastin, mutated in hereditary spastic paraplegia.

Multiple sequence alignment has revealed the presence of a sequence domain of approximately 80 amino acids in two molecules, spartin and spastin, mutated in hereditary spastic paraplegia. The domain, which corresponds to a slightly extended version of the recently described ESP domain of unknown function, was also identified in VPS4, SKD1, RPK118, and SNX15, all of which have a well established and consistent role in endosomal trafficking. Recent functional information indicates that spastin is likely to be involved in microtubule interaction. With this new information relating to its likely function, we propose the more descriptive name 'MIT' (contained within microtubule-interacting and trafficking molecules) for the domain and predict endosomal trafficking as the principal functionality of all molecules in which it is present.

Adenosine Triphosphatases↗

Conserved motifs in T-cell receptor CDR1 and CDR2: implications for ligand and CD8 co-receptor binding.

Recent X-ray crystallographic structures of the T-cell receptor (TCR) alpha and beta chains, as well as their trimolecular complexes with peptide-MHC ligand, have established their structural similarity with the immunoglobulin molecules. The complementarity-determining region (CDR1) and CDR2 encoded within the TCR germline variable (V) sequence genes are well conserved across different TCR V alpha and V beta subfamilies. Multiple sequence alignments have been made based on structural information; they indicate that there will be only a limited number of canonical conformations for the first and second CDR loops. The limited diversity shown by CDRs 1 and 2 contrasts with the extreme junctional CDR3 diversity. Furthermore, CDR2 alignments have revealed conservation of a positive net charge in V alpha subfamilies. A model has been proposed for a direct interaction of the lateral part of CDR2 alpha with the negatively charged membrane-proximal 'stalk' region of the CD8 molecule.

Amino Acid Sequence↗

Pathway evolution, structurally speaking.

Small-molecule metabolism forms the core of the metabolic processes of all living organisms. As early as 1945, possible mechanisms for the evolution of such a complex metabolic system were considered. The problem is to explain the appearance and development of a highly regulated complex network of interacting proteins and substrates from a limited structural and functional repertoire. By permitting the co-analysis of phylogeny and metabolism, the combined exploitation of pathway and structural databases, as well as the use of multiple-sequence alignment search algorithms, sheds light on this problem. Much of the current research suggests a chemistry-driven 'patchwork' model of pathway evolution, but other mechanisms may play a role. In the future, as metabolic structure and sequence space are further explored, it should become easier to trace the finer details of pathway development and understand how complexity has evolved.

Amino Acid Sequence↗

Protein Explorer: easy yet powerful macromolecular visualization.

Protein Explorer (PE, http://www.proteinexplorer.org) enables students, educators and other nonspecialists to visualize macromolecular structures easily. It also offers several advanced capabilities useful to protein structure specialists. Great attention has been given to making PE easy to use. Explanations, color keys and troubleshooting information are displayed automatically. There are also 'Frequently Asked Questions', a one-hour 'Quick-Tour', an alphabetical 'Help/Index/Glossary', and a detailed 'Tutorial'; all making PE much easier to use than either Chime or RasMol. Moreover, it is much more powerful; in addition to basic macromolecular visualization capabilities common to most similar programs, it offers one-click visualization of interfaces between moieties ('contacts'), cation-pi interactions and salt bridges, as well as easy-to-use routines to visualize regions of conservation in three-dimensional protein structures based on multiple sequence alignments.

Computational Biology↗

Interaction of pyrimethamine, cycloguanil, WR99210 and their analogues with Plasmodium falciparum dihydrofolate reductase: structural basis of antifolate resistance.

The nature of the interactions between Plasmodium falciparum dihydrofolate reductase (pfDHFR) and antimalarial antifolates, i.e., pyrimethamine (Pyr), cycloguanil (Cyc) and WR99210 including some of their analogues, was investigated by molecular modeling in conjunction with the determination of the inhibition constants (Ki). A three-dimensional structural model of pfDHFR was constructed using multiple sequence alignment and homology modeling procedures, followed by extensive molecular dynamics calculations. Mutations at amino acid residues 16 and 108 known to be associated with antifolate resistance were introduced into the structure, and the interactions of the inhibitors with the enzymes were assessed by docking and molecular dynamics for both wild-type and mutant DHFRs. The Ki values of a number of analogues tested support the validity of the model. A 'steric constraint' hypothesis is proposed to explain the structural basis of the antifolate resistance.

Amino Acid Sequence↗