Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

The transthyretin-related protein family.

A number of proteins related to the homotetrameric transport protein transthyretin (TTR) forms a highly conserved protein family, which we present in an integrated analysis of data from different sources combined with an initial biochemical characterization. Homologues of the transthyretin-related protein (TRP) can be found in a wide range of species including bacteria, plants and animals, whereas transthyretins have so far only been identified in vertebrates. A multiple sequence alignment of 49 TRP sequences from 47 species to TTR suggests that the tertiary and quaternary features of the three-dimensional structure are most likely preserved. Interestingly, while some of the TRP orthologues show as little as 30% identity, the residues at the putative ligand-binding site are almost entirely conserved. RT/PCR analysis in Caenorhabditis elegans confirms that one TRP gene is transcribed, spliced and predominantly expressed in the worm, which suggests that at least one of the two C. elegans TRP genes encodes a functional protein. We used double-stranded RNA-mediated interference techniques in order to determine the loss-of-function phenotype for the two TRP genes in C. elegans but detected no apparent phenotype. The cloning and initial characterization of purified TRP from Escherichia coli reveals that, while still forming a homotetramer, this protein does not recognize thyroid hormones that are the natural ligands of TTR. The ligand for TRP is not known; however, genomic data support a functional role involving purine catabolism especially linked to urate oxidase (uricase) activity.

Amino Acid Sequence↗

Site-directed mutagenesis of a loop at the active site of E1 (alpha2beta2) of the pyruvate dehydrogenase complex. A possible common sequence motif.

Limited proteolysis of the pyruvate decarboxylase (E1, alpha2beta2) component of the pyruvate dehydrogenase (PDH) multienzyme complex of Bacillus stearothermophilus has indicated the importance for catalysis of a site (Tyr281-Arg282) in the E1alpha subunit (Chauhan, H.J., Domingo, G.J., Jung, H.-I. & Perham, R.N. (2000) Eur. J. Biochem. 267, 7158-7169). This site appears to be conserved in the alpha-subunit of heterotetrameric E1s and multiple sequence alignments suggest that there are additional conserved amino-acid residues in this region, part of a common pattern with the consensus sequence -YR-H-D-YR-DE-. This region lies about 50 amino acids on the C-terminal side of a 30-residue motif previously recognized as involved in binding thiamin diphosphate (ThDP) in all ThDP-dependent enzymes. The role of individual residues in this set of conserved amino acids in the E1alpha chain was investigated by means of site-directed mutagenesis. We propose that particular residues are involved in: (a) binding the 2-oxo acid substrate, (b) decarboxylation of the 2-oxo acid and reductive acetylation of the tethered lipoyl domain in the PDH complex, (c) an "open-close" mechanism of the active site, and (d) phosphorylation by the E1-specific kinase (in eukaryotic PDH and branched chain 2-oxo acid dehydrogenase complexes).

Amino Acid Sequence↗

Identification and functional expression of a second human beta-galactoside alpha2,6-sialyltransferase, ST6Gal II.

BLAST analysis of the human and mouse genome sequence databases using the sequence of the human CMP-sialic acid:beta-galactoside alpha-2,6-sialyltransferase cDNA (hST6Gal I, EC2.4.99.1) as a probe allowed us to identify a putative sialyltransferase gene on chromosome 2. The sequence of the corresponding cDNA was also found as an expressed sequence tag of human brain. This gene contained a 1590 bp open reading frame divided in five exons and the deduced amino-acid sequence didn't correspond to any sialyltransferase already known in other species. Multiple sequence alignment and subsequent phylogenic analysis showed that this new enzyme belonged to the ST6Gal subfamily and shared 48% identity with hST6Gal-I. Consequently, we named this new sialyltransferase ST6Gal II. A construction in pFlag vector transfected in COS-7 cells gave raise to a soluble active form of ST6Gal II. Enzymatic assays indicate that the best acceptor substrate of ST6Gal II was the free disaccharide Galbeta1-4GlcNAc structure whereas ST6Gal I preferred Galbeta1-4GlcNAc-R disaccharide sequence linked to a protein. The alpha2,6-linkage was confirmed by the increase of Sambucus nigra agglutinin-lectin binding to the cell surface of CHO transfected with the cDNA encoding ST6Gal II and by specific sialidases treatment. In addition, the ST6Gal II gene showed a very tissue specific pattern of expression because it was found essentially in brain whereas ST6Gal I gene is ubiquitously expressed.

Amino Acid Sequence↗

Tracking interactions that stabilize the dimer structure of starch phosphorylase from Corynebacterium callunae. Roles of Arg234 and Arg242 revealed by sequence analysis and site-directed mutagenesis.

Glycogen phosphorylases (GPs) constitute a family of widely spread catabolic alpha1,4-glucosyltransferases that are active as dimers of two identical, pyridoxal 5'-phosphate-containing subunits. In GP from Corynebacterium callunae, physiological concentrations of phosphate are required to inhibit dissociation of protomers and cause a 100-fold increase in kinetic stability of the functional quarternary structure. To examine interactions involved in this large stabilization, we have cloned and sequenced the coding gene and have expressed fully active C. callunae GP in Escherichia coli. By comparing multiple sequence alignment to structure-function assignments for regulated and nonregulated GPs that are stable in the absence of phosphate, we have scrutinized the primary structure of C. callunae enzyme for sequence changes possibly related to phosphate-dependent dimer stability. Location of Arg234, Arg236, and Arg242 within the predicted subunit-to-subunit contact region made these residues primary candidates for site-directed mutagenesis. Individual Arg-->Ala mutants were purified and characterized using time-dependent denaturation assays in urea and at 45 degrees C. R234A and R242A are enzymatically active dimers and in the absence of added phosphate, they display a sixfold and fourfold greater kinetic stability of quarternary interactions than the wild-type, respectively. The stabilization by 10 mm of phosphate was, however, up to 20-fold greater in the wild-type than in the two mutants. The replacement of Arg236 by Ala was functionally silent under all conditions tested. Arg234 and Arg242 thus partially destabilize the C. callunae GP dimer structure, and phosphate binding causes a change of their tertiary or quartenary contacts, likely by an allosteric mechanism, which contributes to a reduced protomer dissociation rate.

Allosteric Site↗

Equus caballus gelsolin--cDNA sequence and protein structural implications.

We have generated and characterized the cDNA from equine smooth muscle that encodes gelsolin, an actin-modulating protein. Overlapping cDNA clones synthesized by the reverse transcriptase/polymerase chain reaction and clones isolated from a horse genomic library provided the complete primary structure for the intracellular isoform of gelsolin, while cDNA complemented with protein sequence data produced the full-length primary transcript of the gelsolin isoform found circulating in equine plasma. The deduced amino acid sequences of the intracellular and secreted versions of equine gelsolin infer polypeptides of 731 and 755 residues with apparent molecular masses of 80.7 kDa and 83.2 kDa, respectively. Multiple sequence alignment analysis of equine, human, porcine, and murine orthologs of gelsolin demonstrates prominent similarities among all of these proteins, with the horse and human molecules exhibiting the largest degree of likeness with respect to polypeptide length and overall sequence composition. Both horse and human plasma gelsolins are comprised of 755 amino acids with 94% of the residues identical, while the degree of sequence identity in the shorter (731 residues) cytoplasmic gelsolins is 95%. Analysis of the sequences and structures of the six related domains that comprise gelsolin emphasizes the strong correlation that exists between primary structural conservation among mammalian gelsolins and maintenance of the three-dimensional domain fold characteristic of members of this protein family.

Amino Acid Sequence↗

Subunit organization of the abalone Haliotis tuberculata hemocyanin type 2 (HtH2), and the cDNA sequence encoding its functional units d, e, f, g and h.

We have developed a HPLC procedure to isolate the two different hemocyanin types (HtH1 and HtH2) of the European abalone Haliotis tuberculata. On the basis of limited proteolytic cleavage, two-dimensional immunoelectrophoresis, PAGE, N-terminal protein sequencing and cDNA sequencing, we have identified eight different 40-60-kDa functional units (FUs) in HtH2, termed HtH2-a to HtH2-h, and determined their linear arrangement within the elongated 400-kDa subunit. From a Haliotis cDNA library, we have isolated and sequenced a cDNA clone which encodes the five C-terminal FUs d, e, f, g and h of HtH2. As shown by multiple sequence alignments, defg of HtH2 correspond structurally to defg from Octopus dofleini hemocyanin. HtH2-e is the first FU of a gastropod hemocyanin to be sequenced. The new Haliotis hemocyanin sequences are compared to their counterparts in Octopus, Helix pomatia and HtH1 (from the latter, the sequences of FU-f, FU-g and FU-h have recently been determined) and discussed in relation to the recent 2.3 A X-ray structure of FU-g from Octopus hemocyanin and the 15 A three-dimensional reconstruction of the Megathura crenulata hemocyanin didecamer from electron micrographs. This data allows, for the first time, an insight into the evolution of the two functionally different hemocyanin isoforms found in marine gastropods. It appears that they evolved several hundred million years ago within the Prosobranchia, after separation of the latter from the branch leading to the Pulmonata. Moreover, as a structural explanation for the inefficiency of the type 1 hemocyanin to form multidecamers in vivo, the additional N-glycosylation sites in HtH1 compared to HtH2 are discussed.

Amino Acid Sequence↗

Mutagenesis and modelling of linoleate-binding to pea seed lipoxygenase.

We have produced a model to define the linoleate-binding pocket of pea 9/13-lipoxygenase and have validated it by the construction and characterization of eight point mutants. Three of the mutations reduced, to varying degrees, the catalytic centre activity (kcat) of the enzyme with linoleate. In two of the mutants, reductions in turnover were associated with changes in iron-coordination. Multiple sequence alignments of recombinant plant and mammalian lipoxygenases of known positional specificity, and the results from numerous other mutagenesis and modelling studies, have been combined to discuss the possible role of the mutated residues in pea 9/13-lipoxygenase catalysis. A new nomenclature for recombinant plant lipoxygenases based on positional specificity has subsequently been proposed. The null-effect of mutating pea 9/13-lipoxygenase at the equivalent residue to that which controlled dual positional specificity in cucumber 13/9-lipoxygenase, strongly suggests that the mechanisms controlling dual positional specificity in pea 9/13-lipoxygenase and cucumber 13/9-lipoxygenase are different. This was supported from modelling of another isoform of pea lipoxygenase, pea 13/9-lipoxygenase. Dual positional specificity in pea lipoxygenases is more likely to be determined by the degree of penetration of the methyl terminus of linoleate and the volume of the linoleate-binding pocket rather than substrate orientation. A single model for positional specificity, that has proved to be inappropriate for arachidonate-binding to mammalian 5-, 12- and 15-lipoxygenases, would appear to be true also for linoleate-binding to plant 9- and 13-lipoxygenases.

Amino Acid Sequence↗

Overview on the sub-grouping of the crustacean hyperglycemic hormone family.

The Crustacean hyperglycemic hormones (CHHs) are an ever extending family of crustacean hormones mainly involved in carbohydrate metabolism, molt and reproduction. In this paper, we drew together 32 available CHH sequences, and applied the techniques of multiple sequence alignment, motif searching and amino acid conservation analysis to the characterization of the molecules independently of their biological function. The analysis clearly showed that the proteins clustered into two groups (CHH and VIH). Amino acid conservation analysis also subdivided the VIH group into sequences involved in reproduction (RIH) or in molt (MIH). Motif searching identified five motifs in each group of mature hormones. Motifs A2 and A3 were conserved in all sequences while motifs A1 and A1' were specific of the CHH and VIH groups respectively. This approach demonstrated the S. gregaria ion transport peptides as true members of the CHH group. The two main groups, CHH and VIH, are also discussed in terms of functional homogeneity.

Animals↗

Crystal structure of gamma-glutamylcysteine synthetase: insights into the mechanism of catalysis by a key enzyme for glutathione homeostasis.

Gamma-glutamylcysteine synthetase (gammaGCS), a rate-limiting enzyme in glutathione biosynthesis, plays a central role in glutathione homeostasis and is a target for development of potential therapeutic agents against parasites and cancer. We have determined the crystal structures of Escherichia coli gammaGCS unliganded and complexed with a sulfoximine-based transition-state analog inhibitor at resolutions of 2.5 and 2.1 A, respectively. In the crystal structure of the complex, the bound inhibitor is phosphorylated at the sulfoximido nitrogen and is coordinated to three Mg2+ ions. The cysteine-binding site was identified; it is formed inductively at the transition state. In the unliganded structure, an open space exists around the representative cysteine-binding site and is probably responsible for the competitive binding of glutathione. Upon inhibitor binding, the side chains of Tyr-241 and Tyr-300 turn, forming a hydrogen-bonding triad with the carboxyl group of the inhibitor's cysteine moiety, allowing this moiety to fit tightly into the cysteine-binding site with concomitant accommodation of its side chain into a shallow pocket. This movement is caused by a conformational change of a switch loop (residues 240-249). Based on this crystal structure, the cysteine-binding sites of mammalian and parasitic gammaGCSs were predicted by multiple sequence alignment, although no significant sequence identity exists between the E. coli gammaGCS and its eukaryotic homologues. The identification of this cysteine-binding site provides important information for the rational design of novel gammaGCS inhibitors.

Amino Acid Sequence↗

Computational prediction of native protein ligand-binding and enzyme active site sequences.

Recent studies reveal that the core sequences of many proteins were nearly optimized for stability by natural evolution. Surface residues, by contrast, are not so optimized, presumably because protein function is mediated through surface interactions with other molecules. Here, we sought to determine the extent to which the sequences of protein ligand-binding and enzyme active sites could be predicted by optimization of scoring functions based on protein ligand-binding affinity rather than structural stability. Optimization of binding affinity under constraints on the folding free energy correctly predicted 83% of amino acid residues (94% similar) in the binding sites of two model receptor-ligand complexes, streptavidin-biotin and glucose-binding protein. To explore the applicability of this methodology to enzymes, we applied an identical algorithm to the active sites of diverse enzymes from the peptidase, beta-gal, and nucleotide synthase families. Although simple optimization of binding affinity reproduced the sequences of some enzyme active sites with high precision, imposition of additional, geometric constraints on side-chain conformations based on the catalytic mechanism was required in other cases. With these modifications, our sequence optimization algorithm correctly predicted 78% of residues from all of the enzymes, with 83% similar to native (90% correct, with 95% similar, excluding residues with high variability in multiple sequence alignments). Furthermore, the conformations of the selected side chains were often correctly predicted within crystallographic error. These findings suggest that simple selection pressures may have played a predominant role in determining the sequences of ligand-binding and active sites in proteins.

Algorithms↗

Low-frequency normal modes that describe allosteric transitions in biological nanomachines are robust to sequence variations.

By representing the high-resolution crystal structures of a number of enzymes using the elastic network model, it has been shown that only a few low-frequency normal modes are needed to describe the large-scale domain movements that are triggered by ligand binding. Here we explore a link between the nearly invariant nature of the modes that describe functional dynamics at the mesoscopic level and the large evolutionary sequence variations at the residue level. By using a structural perturbation method (SPM), which probes the residue-specific response to perturbations (or mutations), we identify a sparse network of strongly conserved residues that transmit allosteric signals in three structurally unrelated biological nanomachines, namely, DNA polymerase, myosin motor, and the Escherichia coli chaperonin. Based on the response of every mode to perturbations, which are generated by interchanging specific sequence pairs in a multiple sequence alignment, we show that the functionally relevant low-frequency modes are most robust to sequence variations. Our work shows that robustness of dynamical modes at the mesoscopic level is encoded in the structure through a sparse network of residues that transmit allosteric signals.

Allosteric Regulation↗

Identification of the binding sites of regulatory proteins in bacterial genomes.

We present an algorithm that extracts the binding sites (represented by position-specific weight matrices) for many different transcription factors from the regulatory regions of a genome, without the need for delineating groups of coregulated genes. The algorithm uses the fact that many DNA-binding proteins in bacteria bind to a bipartite motif with two short segments more conserved than the intervening region. It identifies all statistically significant patterns of the form W(1)N(x)W(2), where W(1) and W(2) are two short oligonucleotides separated by x arbitrary bases, and groups them into clusters of similar patterns. These clusters are then used to derive quantitative recognition profiles of putative regulatory proteins. For a given cluster, the algorithm finds the matching sequences plus the flanking regions in the genome and performs a multiple sequence alignment to derive position-specific weight matrices. We have analyzed the Escherichia coli genome with this algorithm and found approximately 1,500 significant patterns, which give rise to approximately 160 distinct position-specific weight matrices. A fraction of these matrices match the binding sites of one-third of the approximately 60 characterized transcription factors with high statistical significance. Many of the remaining matrices are likely to describe binding sites and regulons of uncharacterized transcription factors. The significance of these matrices was evaluated by their specificity, the location of the predicted sites, and the biological functions of the corresponding regulons, allowing us to suggest putative regulatory functions. The algorithm is efficient for analyzing newly sequenced bacterial genomes for which little is known about transcriptional regulation.

Algorithms↗

Crystal structure of conserved hypothetical protein Aq1575 from Aquifex aeolicus.

The crystal structure of a conserved hypothetical protein, Aq1575, from Aquifex aeolicus has been determined by using x-ray crystallography. The protein belongs to the domain of unknown function DUF28 in the Pfam and PALI databases for which there was no structural information available until now. A structural homology search with the DALI algorithm indicates that this protein has a new fold with no obvious similarity to those of other proteins of known three-dimensional structure. The protein reveals a monomer consisting of three domains arranged along a pseudo threefold symmetry axis. There is a large cleft with approximate dimensions of 10 A x 10 A x 20 A in the center of the three domains along the symmetry axis. Two possible active sites are suggested based on the structure and multiple sequence alignment. There are several highly conserved residues in these putative active sites. The structure based molecular properties and thermostability of the protein are discussed.

Amino Acid Sequence↗

Determining the basis of channel-tetramerization specificity by x-ray crystallography and a sequence-comparison algorithm: Family Values (FamVal).

We have developed a semiempirical algorithm called Family Values (FamVal), which identifies residues that encode functional specificity in a protein sequence. Given a multiple sequence alignment (MSA) grouped into functionally distinct subfamilies, FamVal calculates a specificity score for each subfamily at every amino acid position of an MSA. This algorithm was used to predict specificity-encoding positions within the tetramerization assembly (T1) domain of voltage-gated potassium (Kv) channel subfamilies Kv3 and Kv4. The importance of one such position (Arg to Ala at MSA position 93) was confirmed by in vitro pull-down assays. The structural basis of this assembly discrimination was elucidated by determining the crystal structure of the Kv4 T1 domain and comparing it to the Kv3 T1 domain.

Algorithms↗

Coupled prediction of protein secondary and tertiary structure.

The strong coupling between secondary and tertiary structure formation in protein folding is neglected in most structure prediction methods. In this work we investigate the extent to which nonlocal interactions in predicted tertiary structures can be used to improve secondary structure prediction. The architecture of a neural network for secondary structure prediction that utilizes multiple sequence alignments was extended to accept low-resolution nonlocal tertiary structure information as an additional input. By using this modified network, together with tertiary structure information from native structures, the Q3-prediction accuracy is increased by 7-10% on average and by up to 35% in individual cases for independent test data. By using tertiary structure information from models generated with the ROSETTA de novo tertiary structure prediction method, the Q3-prediction accuracy is improved by 4-5% on average for small and medium-sized single-domain proteins. Analysis of proteins with particularly large improvements in secondary structure prediction using tertiary structure information provides insight into the feedback from tertiary to secondary structure.

Computer Simulation↗

Coding potential of laboratory and clinical strains of human cytomegalovirus.

Six strains of human cytomegalovirus have been sequenced, including two laboratory strains (AD169 and Towne) that have been extensively passaged in fibroblasts and four clinical isolates that have been passaged to a limited extent in the laboratory (Toledo, FIX, PH, and TR). All of the sequenced viral genomes have been cloned as infectious bacterial artificial chromosomes. A total of 252 ORFs with the potential to encode proteins have been identified that are conserved in all four clinical isolates of the virus. Multiple sequence alignments revealed substantial variation in the amino acid sequences encoded by many of the conserved ORFs.

Chromosomes, Artificial, Bacterial↗

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking," Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana, contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a swapped domain topology; the other has no disulfide bonds and possesses two distinct, tandem domains. To explore the structural evolution of bicycle proteins, we attempted to predict bicycle protein structures with Alphafold2 (AF2) and other deep learning programs. While AF2 did not recover the two experimental structures using existing databases, it succeeded when provided with multiple sequence alignments (MSAs) of protein sequences from newly sequenced closely related species. Using this approach, we generated 2,400 high-confidence bicycle protein predictions from seven aphid species. While all aphid bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance. Our extension of AF2 with custom MSAs of proteins from closely related species provides a generalizable, powerful approach for predicting structures of rapidly evolving protein families.

Animals↗

Efficient algorithms for molecular sequence analysis.

Efficient (linear time) algorithms are described for identifying global molecular sequence features allowing for errors including repeats, matches between sequences, dyad symmetry pairings, and other sequence patterns. A multiple sequence alignment algorithm is also described. Specific applications are given to hepatitis B viruses and the J5-C (J, joining; C, constant) region of the immunoglobulin kappa gene.

Algorithms↗