Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Primary structure and unusual carbohydrate moiety of functional unit 2-c of keyhole limpet hemocyanin (KLH).

The complete amino acid sequence of the Megathura crenulata hemocyanin functional unit KLH2-c was determined by direct sequencing and matrix-assisted laser desorption ionization mass spectrometry of the protein, and of peptides obtained by cleavage with EndoLysC proteinase, chymotrypsin and cyanogen bromide. This is the first complete primary structure of a functional unit c from a gastropod hemocyanin. KLH2-c consists of 420 amino acid residues. Circular dichroism spectra indicated approx. 31% beta-sheet and 29% alpha-helix contents. A multiple sequence alignment with other molluscan hemocyanin functional units revealed average identities between 41 and 49%, but 55% in case of Octopus hemocyanin functional unit c which is the structural equivalent to KLH2-c. KLH2-c has a molecular mass of approx. 48 kDa as calculated from its sequence and a measured mass of approx. 56 kDa; the mass difference is attributed to the sugar side chains usually decorating molluscan hemocyanin. However, inspection of the sequence of KLH2-c revealed no potential N-linked carbohydrate attachment sites, and this was supported by its inability to bind concanavalin A. Also KLH1-c was unreactive, whereas most, if not all, other functional units of KLH1 and KLH2 reacted positively to this lectin. On the other hand, peanut agglutinin specifically binds KLH2-c, indicating the presence of O-glycosidically linked carbohydrates in this functional unit. This contrasts to all other KLH functional units (including KLH1-c), which lack O-linked glycosides. The present results are discussed in view of the recent X-ray structure of the functional unit g from Octopus hemocyanin, and a published record of the Thomsen Friedenreich tumor antigenic epitope in KLH.

Amino Acid Sequence↗

PCR primers and functional probes for amplification and detection of bacterial genes for extracellular peptidases in single strains and in soil.

A set of primers and functional probes was developed for the detection of peptidase gene fragments of proteolytic bacteria. Based on DNA sequence data, degenerate PCR primers and internal DIG-labeled probes specific for genes encoding alkaline metallopeptidases (apr) (E.3.4.24), neutral metallopeptidases (npr) (E.3.4.24) and serine peptidases (sub) (E.3.4.21) were derived by multiple sequence alignments. Type strains with known peptidase genes and proteolytic bacteria from a grassland rhizosphere soil, a garden soil and an arable field were investigated for their genotypic proteolytic potential. For 52 out of 53 proteolytic bacterial isolates, at least one of the three peptidase classes could be identified by this approach. The amplified gene fragments were of the expected sizes with each of the three primer sets. The functional probes APR, NPR and SUB have been shown to hybridize specifically to the corresponding gene fragments. sub and npr genes were mainly found in Bacillus species. apr genes were only found in the Pseudomonas fluorescens biotypes and in two morphologically identical Flavobacterium-Cytophaga strains from two different sites. In most of the Bacillus spp., both sub and the npr and in the Flavobacterium-Cytophaga strains even all the three genes could be detected. PCR with DNA isolated from soil led to one main product of the expected size with each primer pair whose identity was additionally confirmed by Southern blot hybridization with the corresponding probes.

Bacteria↗

Evidence for the assignment of two strains of SPLV to the genus Potyvirus based on coat protein and 3' non-coding region sequence data.

The use of potyvirus-specific primers and subsequent application of the RACE procedure allowed the cloning of the 3' terminal 1088 nucleotides of the genomic RNA of the Taiwan isolate of sweetpotato latent virus (SPLV-T) and the 3' genomic 1085 nucleotides of a SPLV-like virus from China (SPLV-CH). The sequence of an internal part of the presumptive nuclear inclusion b gene was also determined for both isolates. Detailed sequence analyses revealed the presence of consensus motifs which indicated that SPLV-CH and SPLV-T should be regarded as members of the genus Potyvirus. Multiple sequence alignments and phylogenetic analyses were also performed and unambiguously assessed these isolates as strains of a distinct Potyvirus. SPLV was not related to other potyviruses infecting sweetpotato nor to any other sequenced virus. From the presence of the DAG box, SPLV-CH is expected to be a typical aphid transmitted Potyvirus whereas a conceivable explanation is proposed for the non-aphid transmission of SPLV-T.

Amino Acid Sequence↗

Use of a neural network secondary structure prediction to define targets for mutagenesis of herpes simplex virus glycoprotein B.

Herpes simplex virus glycoprotein B (HSV gB) is essential for penetration of virus into cells, for cell-to-cell spread of virus, and for cell-cell fusion. Every member of the family Herpesviridae has a gB homolog, underlining its importance. The antigenic structure of gB has been studied extensively, but little is known about which regions of the protein are important for its roles in virus entry and spread. In contrast to successes with other HSV glycoproteins, attempts to map functional domains of gB by insertion mutagenesis have been largely frustrated by the misfolding of most mutants. The present study shows that this problem can be overcome by targeting mutations to the loop regions that connect alpha-helices and beta-strands, avoiding the helices and strands themselves. The positions of loops in the primary sequence were predicted by the PHD neural network procedure, using a multiple sequence alignment of 19 alphaherpesvirus gB sequences as input. Comparison of the prediction with a panel of insertion mutants showed that all mutants with insertions in predicted alpha-helices or beta-strands failed to fold correctly and consequently had no activity in virus entry; in contrast, half the mutants with insertions in predicted loops were able to fold correctly. There are 27 predicted loops of four or more residues in gB; targeting of mutations to these regions will minimize the number of misfolded mutants and maximize the likelihood of identifying functional domains of the protein.

Amino Acid Sequence↗

Knowledge-based grouping of modeled HLA peptide complexes.

Human leukocyte antigens are the most polymorphic of human genes and multiple sequence alignment shows that such polymorphisms are clustered in the functional peptide binding domains. Because of such polymorphism among the peptide binding residues, the prediction of peptides that bind to specific HLA molecules is very difficult. In recent years two different types of computer based prediction methods have been developed and both the methods have their own advantages and disadvantages. The nonavailability of allele specific binding data restricts the use of knowledge-based prediction methods for a wide range of HLA alleles. Alternatively, the modeling scheme appears to be a promising predictive tool for the selection of peptides that bind to specific HLA molecules. The scoring of the modeled HLA-peptide complexes is a major concern. The use of knowledge based rules (van der Waals clashes and solvent exposed hydrophobic residues) to distinguish binders from nonbinders is applied in the present study. The rules based on (1) number of observed atomic clashes between the modeled peptide and the HLA structure, and (2) number of solvent exposed hydrophobic residues on the modeled peptide effectively discriminate experimentally known binders from poor/nonbinders. Solved crystal complexes show no vdW Clash (vdWC) in 95% cases and no solvent exposed hydrophobic peptide residues (SEHPR) were seen in 86% cases. In our attempt to compare experimental binding data with the predicted scores by this scoring scheme, 77% of the peptides are correctly grouped as good binders with a sensitivity of 71%.

Alleles↗

Molecular evolution of thyroid peroxidase.

Thyroid peroxidase is a member of a family of mammalian peroxidases that includes myeloperoxidase, lactoperoxidase, eosinophil peroxidase, and salivary peroxidase. Protein sequences showing a high degree of sequence similarity with mammalian peroxidases have recently been observed in several invertebrate species. A multiple sequence alignment prepared with five mammalian and six invertebrate peroxidases shows complete conservation of amino acid residues considered to be important in the formation of peroxidase compound 1. These include the distal and proximal histidines, a catalytic arginine residue, and an asparagine residue hydrogen bonded to the proximal histidine. TPO-2, an alternatively spliced form of TPO, lacks the essential asparagine (Asn 579). It is now possible to speak more broadly of the family of animal peroxidases, rather than mammalian peroxidases. The animal peroxidases comprise a group of homologous proteins that differ markedly from the plant/fungal/bacterial peroxidases in primary, secondary and tertiary structure, but which share with them a common function. Animal peroxidases probably arose independently of the plant/fungal/bacterial peroxidase superfamily and most likely belong to a different gene family. The relationship between animal and non-animal peroxidases probably represents an example of convergent evolution to a common enzymatic mechanism.

Amino Acid Sequence↗

Oxygen transport proteins: III. Structural studies of the scorpion (Buthus sindicus) hemocyanin, partial primary structure of its subunit Bsin1.

The hemocyanin (Hc) from Buthus sindicus, studied in the native state, demonstrated to be an aggregate of eight different types of subunits arranged in four cubic hexamers. Both, the 'top' and the 'side' views of the native molecule have been identified from the negatively stained specimens using transmission electron microscopy. Out of these, eight different polypeptide chains, the partial primary structure (68%) of a subunit Bsin1 (Mr = 72422.7 Da) was established using a combination of automated Edman degradation and mass spectrometry. A multiple sequence alignment with other closely related cheliceratan Hc subunits revealed average identities of ca. 60%. Most of the structurally important residues, i.e. copper and calcium-binding ligands, as well as the residues involved in the presumed oxygen entrance pathway, proved to be strictly conserved in Bsin1. Sequence variations have been observed around the functionally important chloride-binding site, not only for the B. sindicus subunit Bsin1, but also for the subunit Aaus-6 of the scorpion A. australis and the subunit Ecal-a from the spider Eurypelma californicum Hcs. Deviation in the primary structure related to the chloride-binding site suggest that the effect of chloride ions may vary in different hemocyanins. Furthermore, the secondary structural contents of the Hc subunit Bsin1 were determined by circular dichroism revealing ca. 33% alpha-helix, 18%, beta-sheet, 19% beta-turn, and 30% random coil composition. These values are in good agreement with the crystal structure of the closely related Hc subunit Lpol-II from horseshoe crab L. polyphemus. Electron microscopic studies of the purified Hc subunit under native conditions revealed that Bsin1 has self aggregation properties. Results of these studies are discussed.

Amino Acid Sequence↗

Structural similarities and evolutionary relationships in chloride-dependent alpha-amylases.

The alpha-amylase sequences contained in databanks were screened for the presence of amino acid residues Arg195, Asn298 and Arg/Lys337 forming the chloride-binding site of several specialized alpha-amylases allosterically activated by this anion. This search provides 38 alpha-amylases potentially binding a chloride ion. All belong to animals, including mammals, birds, insects, acari, nematodes, molluscs, crustaceans and are also found in three extremophilic Gram-negative bacteria. An evolutionary distance tree based on complete amino acid sequences was constructed, revealing four distinct clusters of species. On the basis of multiple sequence alignment and homology modeling, invariable structural elements were defined, corresponding to the active site, the substrate binding site, the accessory binding sites, the Ca(2+) and Cl(-) binding sites, a protease-like catalytic triad and disulfide bonds. The sequence variations within functional elements allowed engineering strategies to be proposed, aimed at identifying and modifying the specificity, activity and stability of chloride-dependent alpha-amylases.

Amino Acid Sequence↗

Structural/functional assignment of unknown bacteriophage T4 proteins by iterative database searches.

Among the total of 274 orfs within bacteriophage T4, only half have been reasonably well characterized, and the functions of the rest have remained obscure. In order to predict the molecular functions of the orfs, a position-specific iterated (PSI)-BLAST search of bacteriophage T4 against the sequence database of known 3D structures was carried out. PSI-BLAST is one of the most powerful iterative sequence search methods using multiple sequence alignment, with the ability to detect many more proteins with distant homology than standard pairwise methods. The 3D structures of proteins are considered to be better preserved than the sequences, and the detected distantly homologous proteins are likely to possess highly similar 3D structures. Thirteen orfs of phage T4, whose homologues were not detected by standard pairwise methods, were found to have significantly homologous counterparts by this method. The plausibility of the results was confirmed by checking whether important residues at substrate/ligand-binding sites were conserved. Among them, two orfs, vs.1 and e.1, which are similar to Escherichia coli lytic enzyme and MutT protein, respectively, had not been studied previously. Also, gp rIIA, a rapid lysis protein, whose gene structure had been intensively studied during the development of molecular biology in the 1950s and yet whose molecular function remains unknown, has an N-terminal domain that is significantly similar to the N-terminal region of the heat shock protein Hsp90.

Amino Acid Sequence↗

Cloning and expression of a nuclear encoded plastid specific 33 kDa ribonucleoprotein gene (33RNP) from pea that is light stimulated.

We report the cloning and sequencing of both cDNA and genomic DNA of a 33 kDa chloroplast ribonucleoprotein (33RNP) from pea. The analysis of the predicted amino acid sequence of the cDNA clone revealed that the encoded protein contains two RNA binding domains, including the conserved consensus ribonucleoprotein sequences CS-RNP1 and CS-RNP2, on the C-terminus half and the presence of a putative transit peptide sequence in the N-terminus region. The phylogenetic and multiple sequence alignment analysis of pea chloroplast RNP along with RNPs reported from the other plant sources revealed that the pea 33RNP is very closely related to Nicotiana sylvestris 31RNP and 28RNP and also to 31RNP and 28RNP of Arabidopsis and spinach, respectively. The pea 33RNP was expressed in Escherichia coli and purified to homogeneity. The in vitro import of precursor protein into chloroplasts confirmed that the N-terminus putative transit peptide is a bona fide transit peptide and 33RNP is localized in the chloroplast. The nucleic acid-binding properties of the recombinant protein, as revealed by South-Western analysis, showed that 33RNP has higher binding affinity for poly (U) and oligo dT than for ssDNA and dsDNA. The steady state transcript level was higher in leaves than in roots and the expression of this gene is light stimulated. Sequence analysis of the genomic clone revealed that the gene contains four exons and three introns. We have also isolated and analyzed the 5' flanking region of the pea 33RNP gene.

Amino Acid Sequence↗

Identification of a novel member of the TGF-beta superfamily highly expressed in human placenta.

While conducting a gene discovery effort targeted to transcripts of the prevalent and intermediate frequency classes in placenta throughout gestation, we identified a novel member of the TGF-beta superfamily that is expressed at high levels in human placenta. Hence, we named this factor 'Placental Transforming Growth Factor Beta' (PTGFB). The full-length sequence of the 1.2-kb PTGFB mRNA has the potential of encoding a putative pre-pro-PTGFB protein of 295 amino acids and a putative mature PTGFB protein of 112 amino acids. Multiple sequence alignments of PTGFB and representative members of all TGF-beta subfamilies evidenced a number of conserved residues, including the seven cysteines that are almost invariant in all members of the TGF-beta superfamily. The single-copy PTGFB gene was shown to be composed of only two exons of 309 bp and 891 bp, separated by a 2.9-kb intron. The gene was localized to chromosome 19p12-13.1 by fluorescence in-situ hybridization. Northern analyses revealed a complex tissue-specific pattern of expression and a second transcript of 1.9 kb that is predominant in adult skeletal muscle. Most importantly, the 1.2-kb PTGFB transcript was shown to be expressed in placenta at much higher levels than in any other human fetal or adult tissue surveyed.

Adult↗

Evolution of the proximal promoter region of the mammalian growth hormone gene.

The evolutionary relationship between the proximal growth hormone (GH) gene promoter sequences of 12 mammalian species was explored by comparison of their trinucleotide composition and by multiple sequence alignment. Both approaches yielded results that were consistent with the known fossil record-based phylogeny of the analysed sequences, suggesting that the two methods of tree reconstruction might be equally efficient and reliable. The pattern of evolution inferred for the mammalian GH gene promoters was found to vary both temporally and spatially. Thus, two distinct regions devoid of any evolutionary changes exist in primates, but only one of these 'gaps' is also observed in rodents, and neither is seen in ruminants. Furthermore, different evolutionary rates must have prevailed during different periods of evolutionary time and in different lineages, with a dramatic increase in evolutionary rate apparent in primates. Since a similar pattern of discontinuity has been previously noted for the evolution of the GH-coding regions, it may reflect the action of positive selection operating upon the GH gene as a single cohesive unit. Strong evidence for the action of gene conversion between primate GH gene promoters is provided by the fact that the human GH1 and GH2 sequences, which are thought to have diverged before the divergence of Old World monkeys from great apes, are more similar to one another than either is to the rhesus monkey GH2 promoter. Finally, it was noted that a number of nucleotide positions in the GH1 gene promoter that are polymorphic in humans appear to be highly conserved in mammals. This apparent conundrum, which could represent a caveat for the interpretation of phylogenetic footprinting studies, is potentially explicable in terms either of reduced genetic diversity in highly inbred animal species or insufficient population data from non-human species.

Animals↗

Molecular phylogenetic analysis of felid herpesvirus 1.

The position of felid herpesvirus 1 within the alphaherpesvirus subfamily was investigated using molecular phylogenetic techniques applied to multiple sequence alignments of recently reported FHV-1 gene homologs (glycoprotein B, ribonucleotide reductase and DNA polymerase). FHV-1 was most closely related to other carnivore alphaherpesviruses, (phocid herpesvirus 1 and canid herpesvirus 1) and to the equid herpesviruses 1 and 4.

Alphaherpesvirinae↗

Conserved motifs in T-cell receptor CDR1 and CDR2: implications for ligand and CD8 co-receptor binding.

Recent X-ray crystallographic structures of the T-cell receptor (TCR) alpha and beta chains, as well as their trimolecular complexes with peptide-MHC ligand, have established their structural similarity with the immunoglobulin molecules. The complementarity-determining region (CDR1) and CDR2 encoded within the TCR germline variable (V) sequence genes are well conserved across different TCR V alpha and V beta subfamilies. Multiple sequence alignments have been made based on structural information; they indicate that there will be only a limited number of canonical conformations for the first and second CDR loops. The limited diversity shown by CDRs 1 and 2 contrasts with the extreme junctional CDR3 diversity. Furthermore, CDR2 alignments have revealed conservation of a positive net charge in V alpha subfamilies. A model has been proposed for a direct interaction of the lateral part of CDR2 alpha with the negatively charged membrane-proximal 'stalk' region of the CD8 molecule.

Amino Acid Sequence↗

Protein Explorer: easy yet powerful macromolecular visualization.

Protein Explorer (PE, http://www.proteinexplorer.org) enables students, educators and other nonspecialists to visualize macromolecular structures easily. It also offers several advanced capabilities useful to protein structure specialists. Great attention has been given to making PE easy to use. Explanations, color keys and troubleshooting information are displayed automatically. There are also 'Frequently Asked Questions', a one-hour 'Quick-Tour', an alphabetical 'Help/Index/Glossary', and a detailed 'Tutorial'; all making PE much easier to use than either Chime or RasMol. Moreover, it is much more powerful; in addition to basic macromolecular visualization capabilities common to most similar programs, it offers one-click visualization of interfaces between moieties ('contacts'), cation-pi interactions and salt bridges, as well as easy-to-use routines to visualize regions of conservation in three-dimensional protein structures based on multiple sequence alignments.

Computational Biology↗

Interaction of pyrimethamine, cycloguanil, WR99210 and their analogues with Plasmodium falciparum dihydrofolate reductase: structural basis of antifolate resistance.

The nature of the interactions between Plasmodium falciparum dihydrofolate reductase (pfDHFR) and antimalarial antifolates, i.e., pyrimethamine (Pyr), cycloguanil (Cyc) and WR99210 including some of their analogues, was investigated by molecular modeling in conjunction with the determination of the inhibition constants (Ki). A three-dimensional structural model of pfDHFR was constructed using multiple sequence alignment and homology modeling procedures, followed by extensive molecular dynamics calculations. Mutations at amino acid residues 16 and 108 known to be associated with antifolate resistance were introduced into the structure, and the interactions of the inhibitors with the enzymes were assessed by docking and molecular dynamics for both wild-type and mutant DHFRs. The Ki values of a number of analogues tested support the validity of the model. A 'steric constraint' hypothesis is proposed to explain the structural basis of the antifolate resistance.

Amino Acid Sequence↗

The crystal structure of vascular endothelial growth factor (VEGF) refined to 1.93 A resolution: multiple copy flexibility and receptor binding.

BACKGROUND: Vascular endothelial growth factor (VEGF) is an endothelial cell-specific angiogenic and vasculogenic mitogen. VEGF also plays a role in pathogenic vascularization which is associated with a number of clinical disorders, including cancer and rheumatoid arthritis. The development of VEGF antagonists, which prevent the interaction of VEGF with its receptor, may be important for the treatment of such disorders. VEGF is a homodimeric member of the cystine knot growth factor superfamily, showing greatest similarity to platelet-derived growth factor (PDGF). VEGF binds to two different tyrosine kinase receptors, kinase domain receptor (KDR) and Fms-like tyrosine kinase 1 (Flt-1), and a number of VEGF homologs are known with distinct patterns of specificity for these same receptors. The structure of VEGF will help define the location of the receptor-binding site, and shed light on the differences in specificity and cross-reactivity among the VEGF homologs. RESULTS: We have determined the crystal structure of the receptor-binding domain of VEGF at 1.93 A resolution in a triclinic space group containing eight monomers in the asymmetric unit. Superposition of the eight copies of VEGF shows that the beta-sheet core regions of the monomers are very similar, with slightly greater differences in most loop regions. For one loop, the different copies represent different snapshots of a concerted motion. Mutagenesis mapping shows that this loop is part of the receptor-binding site of VEGF. CONCLUSIONS: A comparison of the eight independent copies of VEGF in the asymmetric unit indicates the conformational space sampled by the protein in solution; the root mean square differences observed are similar to those seen in ensembles of the highest precision NMR structures. Mapping the receptor-binding determinants on a multiple sequence alignment of VEGF homologs, suggests the differences in specificity towards KDR and Flt-1 may derive from both sequence variation and changes in the flexibility of binding loops. The structure can also be used to predict possible receptor-binding determinants for related cystine knot growth factors, such as PDGF.

Amino Acid Sequence↗

The structure of a Staphylococcus aureus leucocidin component (LukF-PV) reveals the fold of the water-soluble species of a family of transmembrane pore-forming toxins.

BACKGROUND: Leucocidins and gamma-hemolysins are bi-component toxins secreted by Staphylococcus aureus. These toxins activate responses of specific cells and form lethal transmembrane pores. Their leucotoxic and hemolytic activities involve the sequential binding and the synergistic association of a class S and a class F component, which form hetero-oligomeric complexes. The components of each protein class are produced as non-associated, water-soluble proteins that undergo conformational changes and oligomerization after recognition of their cell targets. RESULTS: The crystal structure of the monomeric water-soluble form of the F component of Panton-Valentine leucocidin (LukF-PV) has been solved by the multiwavelength anomalous dispersion (MAD) method and refined at 2.0 A resolution. The core of this three-domain protein is similar to that of alpha-hemolysin, but significant differences occur in regions that may be involved in the mechanism of pore formation. The glycine-rich stem, which undergoes a major rearrangement in this process, forms an additional domain in LukF-PV. The fold of this domain is similar to that of the neurotoxins and cardiotoxins from snake venom. CONCLUSIONS: The structure analysis and a multiple sequence alignment of all toxic components, suggest that LukF-PV represents the fold of any water-soluble secreted protein in this family of transmembrane pore-forming toxins. The comparison of the structures of LukF-PV and alpha-hemolysin provides some insights into the mechanism of transmembrane pore formation for the bi-component toxins, which may diverge from that of the alpha-hemolysin heptamer.

Amino Acid Sequence↗