Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Comparison of conservation within and between the Ser/Thr and Tyr protein kinase family: proposed model for the catalytic domain of the epidermal growth factor receptor.

The protein kinase family can be subdivided into two main groups based on their ability to phosphorylate Ser/Thr or Tyr substrates. In order to understand the basis of this functional difference, we have carried out a comparative analysis of sequence conservation within and between the Ser/Thr and Tyr protein kinases. A multiple sequence alignment of 86 protein kinase sequences was generated. For each position in the alignment we have computed the conservation of residue type in the Ser/Thr, in the Tyr and in both of the kinase subfamilies. To understand the structural and/or functional basis for the conservation, we have mapped these conservation properties onto the backbone of the recently determined structure of the cAMP-dependent Ser/Thr kinase. The results show that the kinase structure can be roughly segregated, based upon conservation, into three zones. The inner zone contains residues highly conserved in all the kinase family and describes the hydrophobic core of the enzyme together with residues essential for substrate and ATP binding and catalysis. The outer zone contains residues highly variable in all kinases and represents the solvent-exposed surface of the protein. The third zone is comprised of residues conserved in either the Ser/Thr or Tyr kinases or in both, but which are not conserved between them. These are sandwiched between the hydrophobic core and the solvent-exposed surface. In addition to analyzing overall conservation in the kinase family, we have also looked at conservation of its substrate and ATP binding sites. The ATP site is highly conserved throughout the kinases, whereas the substrate binding site is more variable. The active site contains several positions which differ between the Ser/Thr and Tyr kinases and may be responsible for discriminating between hydroxyl bearing side chains. Using this information we propose a model for Tyr substrate binding to the catalytic domain of the epidermal growth factor receptor (EGFR).

Amino Acid Sequence

Structural and evolutionary relationships among the immunophilins: two ubiquitous families of peptidyl-prolyl cis-trans isomerases.

The immunophilins, protein receptors for the immunosuppressing drugs cyclosporin A and FK506 and related proteins from plants, fungi, and bacteria, have been analyzed structurally and evolutionarily. The cyclosporin A binding proteins (cyclophilins) represent one ubiquitous family of homologous proteins, and the FK506- and rapamycin-binding proteins (FKBPs) constitute a second, unrelated family. Multiple sequence alignments of members of each of these two protein families define the highly conserved residues that are likely to play important structural and functional roles, and mutations in representative members of these two families that abolish or alter function have been evaluated. FKBPs have undergone greater evolutionary divergence than the cyclophilins. Evolutionary trees were constructed using two distinct programs, and these trees establish the structural relationships that allow division of each of these families into subgroups. The results lead to the suggestion that several genes encoding isozymic forms of the FKBPs and possibly also of the cyclophilins existed in prokaryotes before the emergence of eukaryotes on earth and that representatives of these genes were transmitted to both kingdoms to give rise to current subfamilies of these proteins. By contrast, compartmentalization of both classes of immunophilins appears to have arisen independently in prokaryotes and eukaryotes, late in evolutionary history.

Amino Acid Isomerases

Genomic divergence of an HIV-2 from a German AIDS patient probably infected in Mali.

The complete nucleotide sequence of an HIV-2 isolate derived from a German AIDS patient with predominantly neurological symptoms is reported. The HIV-2BEN sequence is highly divergent from those of previously described HIV-2 and SIV strains. Evolutionary tree analysis of eight HIV-2 sequences reveals the existence of three HIV-2 groups. HIV-2BEN belongs to a group with two isolates from Ghana and The Gambia. Based on a comparison of HIV-2BEN with six HIV-2 isolates, SIVsmm and SIVmac, the variability of the structural env and gag proteins is similar within the HIV-2/SIVsmm/mac and HIV-1 groups. In contrast, the regulatory HIV-1 proteins are more highly conserved than those from HIV-2 strains. Multiple sequence alignments reveal that some domains of the envelope and regulatory proteins are well conserved among HIV-1, HIV-2/SIVsmm/mac, SIVagm and SIVmnd. The identification of conserved domains within the external glycoprotein could help to develop broadly active vaccines.

Acquired Immunodeficiency Syndrome

Site-directed mutagenesis of the lipoate acetyltransferase of Escherichia coli.

Remote but significant similarities between the primary and predicted secondary structures of the chloramphenicol acetyltransferases (CAT) and lipoate acyltransferase subunits (LAT, E2) of the 2-oxo acid dehydrogenase complexes, have suggested that both types of enzyme may use similar catalytic mechanisms. Multiple sequence alignments for CAT and LAT have highlighted two conserved motifs that contain the active-site histidine and serine residues of CAT. Site-directed replacement of Ser550 in the E2p subunit (LAT) of the pyruvate dehydrogenase complex of Escherichia coli, deemed to be equivalent to the active-site Ser148 of CAT, supported the CAT-based model of LAT catalysis. The effects of other substitutions were also consistent with the predicted similarity in catalytic mechanism although specific details of active-site geometry may not be conserved.

Acetyltransferases

Identification of a gag protein epitope conserved among all four groups of primate immunodeficiency viruses by using monoclonal antibodies.

Five monoclonal antibodies (MAbs) were raised against the gag proteins of simian immunodeficiency virus (SIV) from African green monkey (SIVagmTYO-7). Two MAbs reacted with the matrix protein p17 and the other three with the core protein p24. Studies on the cross-reactivity of the MAbs revealed that the anti-p24 MAbs detected an epitope shared by the viruses belonging to the human immunodeficiency virus type 2 (HIV-2)/SIVmac group and SIVagmTYO-7 and SIVagmTYO-5. The anti-p17 MAbs recognized an epitope present on all these viruses and on SIVagmTYO-1, HIV-1 and SIVmnd. This finding demonstrates for the first time that the matrix protein, p17 or p18, respectively, of all nine HIV and SIV isolates tested in this study expresses at least one conserved immunogenic epitope recognized serologically. By using synthetic peptides, this epitope was identified at the N terminus of p17. Furthermore, this epitope was analysed by multiple sequence alignments of the peptide with homologous sequences of HIV and SIV p17.

Amino Acid Sequence

Sequence and structural analysis of murine adenovirus type 1 hexon.

The genomic region encoding the major capsid protein (hexon) of murine adenovirus type 1 (MAV-1) has been isolated and sequenced. The sequence predicts a 908 residue MAV-1 hexon protein and is flanked by a portion of the upstream pVI gene and the downstream endoproteinase gene. The order of these genes and their location in the middle of the genome are the same as those found in other adenoviruses sequenced to date. Multiple sequence alignment with the other five known hexon protein sequences reveals an overall residue identity of 51% and residue conservation of 66%. In comparison with human adenovirus type 2 (Ad2), MAV-1 hexon has major deletions between residues 141 to 170, 270 to 284 and 446 to 455. Since these regions in the Ad2 hexon are partially exposed on the outer surface of the virion, they may represent type-specific antigenic determinants. The MAV-1 hexon sequence has been modelled using the known three-dimensional structure of the Ad2 hexon. The variable regions in which the mutations, deletions and insertions occur are located in the l1 and l2 loops of the molecule that form the protruding hexon towers on the external surface of the virion.

Amino Acid Sequence

Citrate synthase from the thermophilic archaebacterium Thermoplasma acidophilium. Cloning and sequencing of the gene.

The gene encoding the citric acid cycle enzyme, citrate synthase, has been cloned from the thermoacidophilic archaebacterium, Thermoplasma acidophilum. We report the sequencing of this gene and its flanking regions, and the derived amino acid sequence of the enzyme is compared by multiple-sequence alignment analysis with those of citrate synthases from eubacterial and eukaryotic organisms. The similarity is less than 30% between the archaebacterial and non-archaebacterial sequences, although the majority of residues implicated in the catalytic action of the enzyme have been conserved across all three kingdoms. The cloned archaebacterial gene has been expressed in Escherichia coli to produce catalytically active citrate synthase. This is the first reported sequence of citrate synthase from the archaebacteria.

Amino Acid Sequence

The purification, characterization and analysis of primary and secondary-structure of prolyl oligopeptidase from human lymphocytes. Evidence that the enzyme belongs to the alpha/beta hydrolase fold family.

Prolyl oligopeptidase was isolated and purified to homogeneity from human lymphocytes, yielding a specific activity of 7780 mU/mg. The molecular mass using size-exclusion chromatography matches the 76 kDa obtained by SDS/PAGE. This provides evidence that prolyl oligopeptidase is a monomer. The isoelectric point is 4.8 as judged by isoelectric focusing in free solution. Di-isopropyl fluorophosphate and phenylmethylsulphonyl fluoride completely abolish the activity, classifying the enzyme as a serine proteinase. The inhibition by p-chloromercuribenzoic acid indicates the importance of a free sulfhydryl group near the active-site. alpha 1-Casein and ornithine decarboxylase, two proteins containing a PEST sequence, inhibit prolyl oligopeptidase, but were not hydrolyzed. This demonstrates that prolyl oligopeptidase is not participating in the metabolism of proteins according to a PEST-dependent pathway. alpha 1-Antitrypsin partially inhibits the enzyme but in contrast, aprotinin does not. Its inability to cleave corticotropin-releasing factor, ubiquitin, albumin and aprotinin, together with the hydrolysis of bradykinin between Pro7-Arg8 confirms the affinity of prolyl oligopeptidase for small peptides. Multiple sequence alignment does not reveal any similarity with proteases of known tertiary structure. Secondary-structure prediction displays striking similarity with dipeptidyl peptidase IV and acylaminoacyl peptidase. Two characteristic features of the members of the prolyl oligopeptidase family of serine proteases are high-lighted: the linear arrangement of the catalytic triad is nucleophile-acid-base and the proteolytic cleavage releasing the catalytically active C-terminal region of around 500 amino acids from the N-terminal sequence. Secondary structure prediction and comparison of the active-site of serine proteinases with known three-dimensional coordinates prove that Asp641 is the third member of the catalytic triad. The secondary structural organization of the protease domain of prolyl oligopeptidase is in accordance with the alpha/beta hydrolase fold.

Amino Acid Sequence

Amino acid transporters of lower eukaryotes: regulation, structure and topogenesis.

Lower eukaryotes such as the yeast Saccharomyces cerevisiae and the filamentous fungus Aspergillus nidulans possess a multiplicity of amino acid transporters or permeases which exhibit different properties with respect to substrate affinity, specificity, capacity and regulation. Regulation of amino acid uptake in response to physiological conditions of growth is achieved principally by a dual mechanism; control of gene expression, mediated by a complex interplay of pathway-specific and wide-domain transcription regulatory proteins, and control of transport activities, mediated by a series of protein factors, including a kinase, and possibly, by amino acids. All fungal and a number of bacterial amino acid permeases show significant sequence similarities (33-62% identity scores in binary comparisons), revealing a unique transporter family conserved across the prokaryotic-eukaryotic boundary. Prediction of the topology of this transporter family utilizing a multiple sequence alignment strongly suggests the presence of a common structural motif consisting of 12 alpha-helical putative transmembrane segments and cytoplasmically located N- and C-terminal hydrophilic regions. Interestingly, recent genetic and molecular results strongly suggest that yeast amino acid permeases are integrated into the plasma membrane through a specific intracellular translocation system. Finally, speculating on their predicted structure and on amino acid sequence similarities conserved within this family of permeases reveals regions of putative importance in amino acid transporter structure, function, post-translational regulation or biogenesis.

Amino Acid Transport Systems

Insecticidal properties of a crystal protein gene product isolated from Bacillus thuringiensis subsp. kenyae.

A protoxin gene, localized to a high-molecular-weight plasmid from Bacillus thuringiensis subsp. kenyae, was cloned on a 19-kb BamHI DNA fragment into Escherichia coli. Characterization of the gene revealed it to be a member of the CryIE toxin subclass which has been reported to be as toxic as the CryIC subclass to larvae from Spodoptera exigua in assays with crude E. coli extracts. To directly test the purified recombinant gene product, the gene was subcloned as a 4.8-kb fragment into an expression vector resulting in the overexpression of a 134-kDa protein in the form of phase-bright inclusions in E. coli. Treatment of solubilized inclusion bodies with either trypsin or gut juice from the silkworm Bombyx mori resulted in the appearance of a protease-resistant 65-kDa protein. In force-feeding bioassays, the purified activated protein was highly toxic to larvae of B. mori but not to larvae of Choristoneura fumiferana. In diet bioassays with larvae from S. exigua, the purified protoxin was nontoxic. However, prior activation of the protoxin by tryptic digestion resulted in the appearance of some toxic activity. These results demonstrate that this new subclass of protein toxin may not be useful for the control of Spodoptera species as previously reported. Hierarchical clustering of the nine known lepidopteran-specific CryI toxin subclasses through multiple sequence alignment suggests that the toxins fall into four possible subgroups or clusters.

Animals

Membrane topology of Escherichia coli diacylglycerol kinase.

The topology of Escherichia coli diacylglycerol kinase (DAGK) within the cytoplasmic membrane was elucidated by a combined approach involving both multiple aligned sequence analysis and fusion protein experiments. Hydropathy plots of the five prokaryotic DAGK sequences available were uniform in their prediction of three transmembrane segments. The hydropathy predictions were experimentally tested genetically by fusing C-terminal deletion derivatives of DAGK to beta-lactamase and beta-galactosidase. Following expression, the enzymatic activities of the chimeric proteins were measured and used to determine the cellular location of the fusion junction. These studies confirmed the hydropathy predictions for DAGK with respect to the number and approximate sequence locations of the transmembrane segments. Further analysis of the aligned DAGK sequences detected probable alpha-helical N-terminal capping motifs and two amphipathic alpha-helices within the enzyme. The combined fusion and sequence data indicate that DAGK is a polytopic integral membrane protein with three transmembrane segments with the N terminus of the protein in the cytoplasm, the C terminus in the periplasmic space, and two amphipathic helices near the cytoplasmic surface.

Amino Acid Sequence

Biochemical characterization and sequence analysis of the gluconate:NADP 5-oxidoreductase gene from Gluconobacter oxydans.

Gluconate:NADP 5-oxidoreductase (GNO) from the acetic acid bacterium Gluconobacter oxydans subsp. oxydans DSM3503 was purified to homogeneity. This enzyme is involved in the nonphosphorylative, ketogenic oxidation of glucose and oxidizes gluconate to 5-ketogluconate. GNO was localized in the cytoplasm, had an isoelectric point of 4.3, and showed an apparent molecular weight of 75,000. In sodium dodecyl sulfate gel electrophoresis, a single band appeared corresponding to a molecular weight of 33,000, which indicated that the enzyme was composed of two identical subunits. The pH optimum of gluconate oxidation was pH 10, and apparent Km values were 20.6 mM for the substrate gluconate and 73 microM for the cosubstrate NADP. The enzyme was almost inactive with NAD as a cofactor and was very specific for the substrates gluconate and 5-ketogluconate. D-Glucose, D-sorbitol, and D-mannitol were not oxidized, and 2-ketogluconate and L-sorbose were not reduced. Only D-fructose was accepted, with a rate that was 10% of the rate of 5-ketogluconate reduction. The gno gene encoding GNO was identified by hybridization with a gene probe complementary to the DNA sequence encoding the first 20 N-terminal amino acids of the enzyme. The gno gene was cloned on a 3.4-kb DNA fragment and expressed in Escherichia coli. Sequencing of the gene revealed an open reading frame of 771 bp, encoding a protein of 257 amino acids with a predicted relative molecular mass of 27.3 kDa. Plasmid-encoded gno was functionally expressed, with 6.04 U/mg of cell-free protein in E. coli and with 6.80 U/mg of cell-free protein in G. oxydans, which corresponded to 85-fold overexpression of the G. oxydans wild-type GNO activity. Multiple sequence alignments showed that GNO was affiliated with the group II alcohol dehydrogenases, or short-chain dehydrogenases, which display a typical pattern of six strictly conserved amino acid residues.

Amino Acid Sequence

Structure-function relationship of bacterial prolipoprotein diacylglyceryl transferase: functionally significant conserved regions.

The structure-function relationship of bacterial prolipoprotein diacylgyceryl transferase (LGT) Has been investigated by a comparison of the primary structures of this enzyme in phylogenetically distant bacterial species, analysis of the sequences of mutant enzymes, and specific chemical modification of the Escherichia coli enzyme. A clone containing the gene for LGT, lgt, of the gram-positive species Staphylococcus aureus was isolated by complementation of the temperature-sensitive lgt mutant of E. coli (strain SK634) defective in LGT activity. In vivo and in vitro assays for prolipoprotein diacylglyceryl modification activity indicated that the complementing clone restored the prolipoprotein modification activity in the mutant strain. Sequence determination of the insert DNA revealed an open reading frame of 837 bp encoding a protein of 279 amino acids with a calculated molecular mass of 31.6 kDa. S. aureus LGT showed 24% identity and 47% similarity with E. coli, Salmonella typhimurium, and Haemophilus influenzae LGT.S. aureus LGT, while 12 amino acids shorter than the E. coli enzyme, had a hydropathic profile and a predicted pI (10.4) similar to those of the E. coli enzyme. Multiple sequence alignment among E. coli, S. typhimurium, H. influenzae, and S. aureus LGT proteins revealed regions of highly conserved amino acid sequences throughout the molecule. Three independent lgt mutant alleles from E. coli SK634, SK635, and SK636 and one lgt allele from S. typhimurium SE5221, all defective in LGT activity at the nonpermissive temperature, were cloned by PCR and sequenced. The mutant alleles were found to contain a single base alteration resulting in the substitution of a conserved amino acid. The longest set of identical amino acids without any gap was H-103-GGLIG-108 in LGT from these four microorganisms. In E. coli lgt mutant SK634, Gly-104 in this region was mutated to Ser, and the mutant organism was temperature sensitive in growth and exhibited low LGT activity in vitro. Diethylpyrocarbonate inactivated the E. coli LGT with a second-order rate constant of 18.6 M-1S-1, and the inactivation of LGT activity was reversed by hydroxylamine at pH 7. The inactivation kinetics were consistent with the modification of a single residue, His or Tyr, essential for LGT activity.

Amino Acid Sequence

The Alcaligenes eutrophus protein HoxN mediates nickel transport in Escherichia coli.

HoxN, an integral membrane protein with seven transmembrane helices and a molecular mass of 33.1 kDa, is involved in high-affinity nickel transport in Alcaligenes eutrophus H16. From genetic analyses, it has been concluded that HoxN is a single-component ion carrier. To investigate this assumption, hoxN was introduced into Escherichia coli. The recombinant strain showed significantly enhanced nickel uptake in a short-interval assay. Likewise, growth in the presence of 63NiCl2 yielded a more than 15-fold-increased cellular nickel content. The HoxN-based nickel transport activity could also be demonstrated in a physiological assay: an E. coli strain coexpressing hoxN and the urease operon of Klebsiella aerogenes exhibited urease activity 10-fold greater than that in the strain lacking a functional hoxN. These results strongly suggest that HoxN is sufficient to operate as a nickel permease. Multiple sequence alignment of HoxN and four other bacterial membrane proteins implicated in nickel metabolism revealed two conserved signatures which may play a role in the nickel translocation process.

Alcaligenes

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking", Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana (Witch Hazel), contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a helix swap; the other has no disulfide bonds and possesses two tandem domains. To explore the structural evolution of bicycle proteins, we predicted bicycle protein structures with Alphafold2 (AF2). While AF2 did not recover the two experimental structures using existing databases, it succeeded after we provided multiple sequence alignments (MSAs) containing protein sequences encoded in new genome sequences from closely related aphid species. Using this customized approach at scale, we generated 2400 high-confidence predictions for bicycle proteins from seven aphid species. This dataset revealed that bicycle proteins without cysteines are outliers in fold space and appear to have evolved from ancestral proteins with disulfide-bonded saposin-like folds. While all bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance.

AlphaFold predictions

The VPH1 gene encodes a 95-kDa integral membrane polypeptide required for in vivo assembly and activity of the yeast vacuolar H(+)-ATPase.

Yeast vacuolar acidification-defective (vph) mutants were identified using the pH-sensitive fluorescence of 6-carboxyfluorescein diacetate (Preston, R. A., Murphy, R. F., and Jones, E. W. (1989) Proc. Natl. Acad. Sci. U.S.A. 86, 7027-7031). Vacuoles purified from yeast bearing the vph1-1 mutation had no detectable bafilomycin-sensitive ATPase activity or ATP-dependent proton pumping. The peripherally bound nucleotide-binding subunits of the vacuolar H(+)-ATPase (60 and 69 kDa) were no longer associated with vacuolar membranes yet were present in wild type levels in yeast whole cell extracts. The VPH1 gene was cloned by complementation of the vph1-1 mutation and independently cloned by screening a lambda gt11 expression library with antibodies directed against a 95-kDa vacuolar integral membrane protein. Deletion disruption of the VPH1 gene revealed that the VPH1 gene is not essential for viability but is required for vacuolar H(+)-ATPase assembly and vacuolar acidification. VPH1 encodes a predicted polypeptide of 840 amino acid residues (molecular mass 95.6 kDa) and contains six putative membrane-spanning regions. Cell fractionation and immunodetection demonstrate that Vph1p is a vacuolar integral membrane protein that co-purifies with vacuolar H(+)-ATPase activity. Multiple sequence alignments show extensive homology over the entire lengths of the following four polypeptides: Vph1p, the 116-kDa polypeptide of the rat clathrin-coated vesicles/synaptic vesicle proton pump, the predicted polypeptide encoded by the yeast gene STV1 (Similar To VPH1, identified as an open reading frame next to the BUB2 gene), and the TJ6 mouse immune suppressor factor.

Amino Acid Sequence

Computational sequence analysis revisited: new databases, software tools, and the research opportunities they engender.

The increasing quantity and complexity of sequences and structural data for proteins and nucleic acids create both problems and opportunities for biomedical researchers. Fortunately, a new generation of practical computer tools for data analysis and integrated information retrieval is emerging. Recent developments in fast database searching, multiple sequence alignment, and molecular modeling are discussed and windows-based, mouse-driven software for CD-ROM and network information retrieval are described. Each method is illustrated with a practical example pertinent to lipid research. In particular, the connection among cholesteryl ester transfer protein, bactericidal permeability-increasing protein, and lipopolysaccharide-binding proteins is determined; novel repetitive sequence motifs in mammalian farnesyltransferase subunits and related yeast prenyltransferases are derived; biochemical insights from a three-dimensional model of human apolipoprotein D based on two insect lipocalins are discussed; the relationship between apolipoprotein D and gross cystic disease fluid protein from human breast is reviewed; and prospects for modeling apolipoprotein E-related proteins are described. In addition, information on a number of general and special-purpose sequence, motif, and structural databases is included.

Amino Acid Sequence

Multiple DNA and protein sequence alignment on a workstation and a supercomputer.

This paper describes a multiple alignment method using a workstation and supercomputer. The method is based on the alignment of a set of aligned sequences with the new sequence, and uses a recursive procedure of such alignment. The alignment is executed in a reasonable computation time on diverse levels from a workstation to a supercomputer, from the viewpoint of alignment results and computational speed by parallel processing. The application of the algorithm is illustrated by several examples of multiple alignment of 12 amino acid and DNA sequences of HIV (human immunodeficiency virus) env genes. Colour graphic programs on a workstation and parallel processing on a supercomputer are discussed.

Algorithms