Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

An investigation of the role of Glu-842, Glu-844 and His-846 in the function of the cytoplasmic domain of the epidermal growth factor receptor.

Activation of several protein kinases is mediated, at least in part, by phosphorylation of conserved Thr or Tyr residues located in a variable loop region, near the active site. In certain kinases, this activation loop also controls access of peptide substrates to the active site. In the corresponding region of the epidermal growth factor (EGF) receptor, a potential phosphorylation site, Tyr-845, does not appear to have a major regulatory role. In order to find out whether this variable loop can modulate the peptide phosphorylation and self-phosphorylation activities of the EGF receptor kinase, we investigated the role of residues around Tyr-845, using site-directed mutagenesis. Multiple sequence alignment showed that residues Glu-842, Glu-844 and His-846 are conserved or nearly conserved in eight members of the EGF receptor family. Mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala were expressed in the baculovirus/insect cell system, purified to near-homogeneity and characterized with respect to their peptide phosphorylation and self-phosphorylation activities. All three mutants were active, and these changes did not affect ATP binding directly. However, all mutations increased the Km(app.) for peptide substrates and MnATP in peptide phosphorylation reactions. The Vmax. for the phosphorylation of peptide RREELQDDYEDD was unaltered, but the Vmax. for self-phosphorylation (with variable [MnATP]) decreased 4-, 2- and 7-fold for mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala respectively, compared with the wild-type. These results suggest that binding of this peptide restored an optimal conformation at the active site that might be impaired by the mutations. A study of the dependence of initial rates of self-phosphorylation on cytoplasmic domain concentration showed that the order of reaction increased with the progress of self-phosphorylation. Both pre-phosphorylation and high concentrations of ammonium sulphate restored maximal or near-maximal levels of self-phosphorylation in the mutants, possibly through compensating conformational changes. A plausible homology model, based on the cyclic AMP-dependent protein kinase catalytic subunit, accommodated the sequence Glu-841-Glu-Lys-Glu as an insertion in the peptide binding loop at the edge of the active site cleft. The model suggests that Glu-844 and His-846 may participate in H-bonding interactions, thus stabilizing the active site region, while Glu-842 does not appear to interact with regions of the catalytic core.

Amino Acid Sequence

The profilin multigene family of maize: differential expression of three isoforms.

Profilin is a small (12-15 kDa) actin- and phospholipid-binding protein previously known only from studies on animals and lower eukaryotes but recently identified as a birch pollen allergen. Here we have identified and characterized three members of the profilin multigene family from the plant Zea mays. Two cDNAs isolated from a maize pollen library (ZmPRO 1 and ZmPRO 3) each have a single, large open reading frame encoding a putative polypeptide 131 amino acids long with a predicted molecular weight of approximately 14 kDa. A third maize pollen cDNA (ZmPRO 2) has two in-frame translation initiation codons. Use of the first ATG would result in a polypeptide 137 amino acids long with a molecular weight of 14.8 kDa. The three maize profilins are highly homologous to each other (> 90% nucleotide and amino acid sequence identity) as well as other plant profilins but show far less similarity (30-40% amino acid sequence identity) to animal and lower eukaryote profilins. Multiple sequence alignments indicate that only nine residues are shared by all eukaryotic profilins examined. However, limited comparisons reveal domains in the NH2 and COOH termini that have a high degree of similarity suggesting functional conservation. The maize gene family size is estimated to contain three to six members based on Southern blot experiments with gene-specific and coding region probes. Northern blot analysis demonstrates that the three maize profilin cDNAs characterized here are utilized in a tissue-specific manner and are anther or pollen specific.

Actins

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking," Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana, contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a swapped domain topology; the other has no disulfide bonds and possesses two distinct, tandem domains. To explore the structural evolution of bicycle proteins, we attempted to predict bicycle protein structures with Alphafold2 (AF2) and other deep learning programs. While AF2 did not recover the two experimental structures using existing databases, it succeeded when provided with multiple sequence alignments (MSAs) of protein sequences from newly sequenced closely related species. Using this approach, we generated 2,400 high-confidence bicycle protein predictions from seven aphid species. While all aphid bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance. Our extension of AF2 with custom MSAs of proteins from closely related species provides a generalizable, powerful approach for predicting structures of rapidly evolving protein families.

Animals

Efficient algorithms for molecular sequence analysis.

Efficient (linear time) algorithms are described for identifying global molecular sequence features allowing for errors including repeats, matches between sequences, dyad symmetry pairings, and other sequence patterns. A multiple sequence alignment algorithm is also described. Specific applications are given to hepatitis B viruses and the J5-C (J, joining; C, constant) region of the immunoglobulin kappa gene.

Algorithms

Structural model of the nucleotide-binding conserved component of periplasmic permeases.

The amino acid sequences of 17 bacterial membrane proteins that are components of periplasmic permeases and function in the uptake of a variety of small molecules and ions are highly homologous to each other and contain sequence motifs characteristic of nucleotide-binding proteins. These proteins are known to bind ATP and are postulated to be the energy-coupling components of the permeases. Several medically important eukaryotic proteins, including the multidrug-resistance transporters and the protein encoded by the cystic fibrosis gene, are also homologous to this family. By multiple sequence alignment of these 17 proteins, the consensus sequence, secondary structure, and surface exposure were predicted. The secondary structural motifs that are conserved among nucleotide-binding proteins were identified in adenylate kinase, p21ras, and elongation factor Tu by superposition of their known tertiary structures. The equivalent secondary structural elements in the predicted conserved component were located. These, together with sequence information, served as guides for alignment with adenylate kinase. A model for the structure of the ATP-binding domain of the permease proteins is proposed by analogy to the adenylate kinase structure. The characteristics of several permease mutations and biochemical data lend support to the model.

Adenylate Kinase

GTPase domains of ras p21 oncogene protein and elongation factor Tu: analysis of three-dimensional structures, sequence families, and functional sites.

GTPase domains are functional and structural units employed as molecular switches in a variety of important cellular functions, such as growth control, protein biosynthesis, and membrane traffic. Amino acid sequences of more than 100 members of different subfamilies are known, but crystal structures of only mammalian ras p21 and bacterial elongation factor Tu have been determined. After optimal superposition of these remarkably similar structures, careful multiple sequence alignment, and calculation of residue-residue interactions, we analyzed the two subfamilies in terms of structural conservation, sequence conservation, and residue contact strength. There are three main results. (i) A structure-based alignment of p21 and elongation factor Tu. (ii) The definition of a common conserved structural core that may be useful as the basis of model building by homology of the three-dimensional structure of any GTPase domain. (iii) Identification of sequence regions, other than the effector loop and the nucleotide binding site, that may be involved in the functional cycle: they are loop L4, known to change conformation after GTP hydrolysis; helix alpha 2, especially Arg-73 and Met-67 in ras p21; loops L8 and L10, including ras p21 Arg-123, Lys-147, and Leu-120; and residues located spatially near the N and C termini. These regions are candidate sites for interaction either with the GTP/GDP exchange factor, with a GTPase-affected function, or with a molecule delivered to a destination site with the aid of the GTPase domain.

Amino Acid Sequence

Improved prediction of protein secondary structure by use of sequence profiles and neural networks.

The explosive accumulation of protein sequences in the wake of large-scale sequencing projects is in stark contrast to the much slower experimental determination of protein structures. Improved methods of structure prediction from the gene sequence alone are therefore needed. Here, we report a substantial increase in both the accuracy and quality of secondary-structure predictions, using a neural-network algorithm. The main improvements come from the use of multiple sequence alignments (better overall accuracy), from "balanced training" (better prediction of beta-strands), and from "structure context training" (better prediction of helix and strand lengths). This method, cross-validated on seven different test sets purged of sequence similarity to learning sets, achieves a three-state prediction accuracy of 69.7%, significantly better than previous methods. In addition, the predicted structures have a more realistic distribution of helix and strand segments. The predictions may be suitable for use in practice as a first estimate of the structural type of newly sequenced proteins.

Amino Acid Sequence

Identification of active-site residues of the adenovirus endopeptidase.

Multiple sequence alignment of the 12 adenovirus endopeptidases known to date identified a number of conserved residues which might be important for enzyme activity. Eleven mutants were created in the cloned gene by site-directed mutagenesis to identify the active site of this thiol endopeptidase. Analysis of the proteolytic activity in a crude system using viral precursor proteins, as well as in a purified system with activated proteinases using a new chromophoric octapeptide substrate, yielded results consistent with Cys-104 and His-54 being two members of the active site. This result was confirmed by the carboxymethylation of the reactive Cys-104 and its prevention by the active-thiol-specific agent E64. Although Cys-122 and Cys-126 were also reactive cysteines, mutation of these residues did not affect enzyme activity. Replacement of the active-site Cys-104 by serine converted the enzyme into a serine-like proteinase, sensitive to serine proteinase inhibitors. The absence of homology to other proteinases, particularly at the active-site cysteine, coupled with the requirement for activation by a substrate cleavage fragment, indicates that the adenovirus endoproteinase may represent a new subclass of cysteine proteinases.

Adenoviridae

Molecular cloning of the isoquinoline 1-oxidoreductase genes from Pseudomonas diminuta 7, structural analysis of iorA and iorB, and sequence comparisons with other molybdenum-containing hydroxylases.

The iorA and iorB genes from the isoquinoline-degrading bacterium Pseudomonas diminuta 7, encoding the heterodimeric molybdo-iron-sulfur-protein isoquinoline 1-oxidoreductase, were cloned and sequenced. The deduced amino acid sequences IorA and IorB showed homologies (i) to the small (gamma) and large (alpha) subunits of complex molybdenum-containing hydroxylases (alpha beta gamma/alpha 2 beta 2 gamma 2) possessing a pterin molybdenum cofactor with a monooxo-monosulfido-type molybdenum center, (ii) to the N- and C-terminal regions of aldehyde oxidoreductase from Desulfovibrio gigas, and (iii) to the N- and C-terminal domains of eucaryotic xanthine dehydrogenases, respectively. The closest similarity to IorB was shown by aldehyde dehydrogenase (Adh) from the acetic acid bacterium Acetobacter polyoxogenes. Five conserved domains of IorB were identified by multiple sequence alignments. Whereas IorB and Adh showed an identical sequential arrangement of these conserved domains, in all other molybdenum-containing hydroxylases the relative position of "domain A" differed. IorA contained eight conserved cysteine residues. The amino acid pattern harboring the four cysteine residues proposed to ligate the Fe/S I cluster was homologous to the consensus binding site of bacterial and chloroplast-type [2Fe-2S] ferredoxins, whereas the pattern including the four cysteines assumed to ligate the Fe/S II center showed no similarities to any described [2Fe-2S] binding motif. The N-terminal region of IorB comprised a putative signal peptide similar to typical leader peptides, indicating that isoquinoline 1-oxidoreductase is associated with the cell membrane.

Amino Acid Sequence

Cloning of a novel family of mammalian GTP-binding proteins (RagA, RagBs, RagB1) with remote similarity to the Ras-related GTPases.

cDNA clones of two novel Ras-related GTP-binding proteins (RagA and RagB) were isolated from rat and human cDNA libraries. Their deduced amino acid sequences comprise four of the six known conserved GTP-binding motifs (PM1, -2, -3, G1), the remaining two (G2, G3) being strikingly different from those of the Ras family, and an unusually large C-terminal domain (100 amino acids) presumably unrelated to GTP binding. RagA and RagB differ by seven conservative amino acid substitutions (98% identity), and by 33 additional residues at the N terminus of RagB. In addition, two isoforms of RagB (RagBs and RagB1) were found that differed only by an insertion of 28 codons between the GTP-binding motifs PM2 and PM3, apparently generated by alternative mRNA splicing. Polymerase chain reaction amplification with specific primers indicated that both long and short form of RagB transcripts were present in adrenal gland, thymus, spleen, and kidney, whereas in brain, only the long form RagB1 was detected. A long splicing variant of RagA was not detected. Recombinant glutathione S-transferase (GST) fusion proteins of RagA and RagBs bound large amounts of radiolabeled GTP gamma S in a specific and saturable manner. In contrast, GTP gamma S binding of GST-RagB1 hardly exceeded that of recombinant GST. GTP gamma S bound to recombinant RagA, and RagBs was rapidly exchangeable for GTP, whereas no intrinsic GTPase activity was detected. A multiple sequence alignment indicated that RagA and RagB cannot be assigned to any of the known subfamilies of Ras-related GTPases but exhibit a 52% identity with a yeast protein (Gtr1) presumably involved in phosphate transport and/or cell growth. It is suggested that RagA and RagB are the mammalian homologues of Gtr1 and that they represent a novel subfamily of Ras-homologous GTP binding proteins.

Amino Acid Sequence

PHD--an automatic mail server for protein secondary structure prediction.

By the middle of 1993, > 30,000 protein sequences has been listed. For 1000 of these, the three-dimensional (tertiary) structure has been experimentally solved. Another 7000 can be modelled by homology. For the remaining 21,000 sequences, secondary structure prediction provides a rough estimate of structural features. Predictions in three states range between 35% (random) and 88% (homology modelling) overall accuracy. Using information about evolutionary conservation as contained in multiple sequence alignments, the secondary structure of 4700 protein sequences was predicted by the automatic e-mail server PHD. For proteins with at least one known homologue, the method has an expected overall three-state accuracy of 71.4% for proteins with at least one known homologue (evaluated on 126 unique protein chains).

Algorithms

SEQSEE: a comprehensive program suite for protein sequence analysis.

SEQSEE (SEQuence SEEker) is a multi-purpose, menu-driven suite of programs designed to provide a fully integrated, state-of-the-art package for the analysis and display of protein sequences and protein databases. It is currently configured to run on most UNIX-based machines including Sun, SGI and NeXT workstations with conversion to other architectures (e.g. Vax or Cray) being a relatively simple task. SEQSEE is capable of performing nearly all of the analytical and comparative tasks found in most comprehensive commercially available software packages. These include sequence/database searching, sequence retrieval, sequence entry and editing, statistical sequence analysis, multiple sequence alignment, flexible pattern matching, and secondary structure prediction. SEQSEE also integrates a number of unique databases which allow it to perform many additional functions such as structure-based sequence alignments and homology-based secondary structure prediction. Additional enhancements to many previously published algorithms have substantially improved the performance of SEQSEE over that found for most other commercial products. The source code, the documentation and all of the required databases for SEQSEE are freely available and may be obtained by anonymous ftp.

Algorithms

ENVIRON: a software package to compare protein three-dimensional structures with homologous sequences using local structural motifs.

This work presents a method to compare local clusters of interacting residues as observed in a known three-dimensional protein structure with corresponding clusters inferred from homologous protein sequences, assuming conserved protein folding. For this purpose the local environment of a selected residue in a known protein structure is defined as the ensemble of amino acids in contact with it in the folded state. Using a multiple sequence alignment to identify corresponding residues in homologous proteins, a detailed comparison can be performed between the local environment of a selected amino acid in the template protein structure and the expected local environments at the sets of equivalent residues, derived from the aligned protein sequences. The comparison makes it possible to detect conserved local features such as hydrogen bonding or complementarity in residue substitution. A global measure of environmental similarity is also defined, to search for conserved amino acid clusters subject to functional or structural constraints. The proposed approach is useful for investigating protein function as well as for site-directed mutagenesis experiments, where appropriate amino acid substitutions can be suggested by observing naturally occurring protein variants.

Algorithms

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Inference of Cytochrome P450 Evolutionary History Using Structural and Physicochemical Metrics.

Cytochrome P450s are a superfamily of heme-binding monooxygenases involved with the detoxification of intrinsic and extrinsic toxins. They are near ubiquitous within biological domains and are found in all domains. Members of families within the superfamily are defined based on amino acid identity thresholds, with thresholds as low as 40% in some families. Relationships among Cytochrome P450 families have proven elusive due to sub-Twilight Zone interfamily identities (<30%) that result in poor multiple sequence alignment quality and thus low levels of support for downstream phylogenetic reconstructions. Despite the low identities, Cytochrome P450 structures are remarkably well conserved both within and among families. In such cases, structural phylogenetics has the potential to unveil elusive relationships because the selectively favored physicochemical properties giving rise to the structure and function of the proteins persist despite sequence-level divergence. Recently, in two separate publications, we demonstrated that by utilizing physicochemical vectors, dynamic time warping, and hierarchical clustering (PCDTW), large swaths of protein domain families and betacoronavirus receptor-binding domain clades were congruent with validated functional/structural relationships. These were important findings because anomalous sequence alignment-based maximum likelihood phylogenetic findings, which were not congruent with the known functional relationships, were resolved. That also validated the use of physicochemical vectors in making inferences about structural/functional homology. Additionally, it illuminated that the same methods might be applied to other protein families with relationships that are difficult to resolve from sequence data alone. Herein, we used Molecular Weight and Hydrophobicity Physicochemical Dynamic Time Warping (MWHP PCDTW) along with structural and sequence alignment-based phylogenetic methodologies to analyze all of the Cytochrome P450s found both in the high-fidelity Structural Classificaction of Proteins (SCOP) database and the reviewed sequences with both experimentally resolved and de novo predicted structures in the Protein Data Bank and the AlphaFold (AF) Protein Structure Database, respectively. We compared the resulting phylogenetic topologies and found that in some cases, structure-based methods may be less able to resolve random/convergent similarity than physicochemical and sequence-based methodologies. This finding agrees with previous findings that demonstrate the usefulness of physicochemical properties in resolving both random structural similarity and potentially convergent relationships.

Cytochrome P-450 Enzyme System

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Analysis of genetic diversity in cytotoxin-producing and non-cytotoxin-producing Helicobacter pylori strains.

Analysis of 32 Helicobacter pylori strains indicated a strong association between the presence of the cagA gene and a specific type of vacA allele found predominantly in cytotoxin-producing strains (P < .001). To determine whether tox+/CagA+ and tox-/CagA- strains constituted two separate noncombining lineages, sequences of the H. pylori ureC gene, cysS homologue, and the intergenic region between cysS and vacA were determined for multiple strains. The mean levels of nucleotide identity in the three regions were 96.7% +/- 0.5%, 95.0% +/- 1.0%, and 89.0% +/- 2.9%, respectively. Multiple sequence alignments and dendrograms based on these three regions failed to identify two clonal populations of organisms for which cagA and vacA genotypes were markers. The presence of a 63- to 64-bp insertion in the cysS-vacA intergenic region was unrelated to the vacA genotype of the strains. These data suggest that recombination between Helicobacter genomes may occur in vivo.

Base Sequence

Molecular cloning of the cDNA for the catalytic subunit of human DNA polymerase delta.

The cDNA of human DNA polymerase delta was cloned. The cDNA had a length of 3.5 kb and encoded a protein of 1107 amino acid residues with a calculated molecular mass of 124 kDa. Northern blot analysis showed that the cDNA hybridized to a mRNA of 3.4 kb. Monoclonal and polyclonal antibodies to the C-terminal 20 residues specifically immunoblotted the human pol delta catalytic polypeptide. A multiple sequence alignment was constructed. This showed that human pol delta is closely related to yeast pol delta and the herpes virus DNA polymerases. The levels of pol delta message were found to be induced concomitantly with DNA pol delta activity and DNA synthesis in serum restimulated proliferating IMR90 cultured cells. The human pol delta gene was localized to chromosome 19 by Southern blotting of EcoRI digested DNA from a panel of rodent/human cell hybrids.

Amino Acid Sequence