Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Immunoglobulin-binding FcrA and Enn proteins and M proteins of group A streptococci evolved independently from a common ancestral protein.

Significant sequence homology between M proteins and immunoglobulin (Ig)-binding proteins of group A streptococci suggests that these proteins arose by gene duplication followed by the development of functional diversity due to mutations and intragenic recombinations. The deduced sequence of multiple Ig-binding proteins and M proteins were compared to distinguish between two evolutionary models. Did these functionally distinct genes originate in the distant past from duplication of a common ancestral gene and then functionally evolve independently or did they evolve more recently, one from the other by duplication of a fixed gene? Multiple alignments of conserved sequences of these proteins are consistent with the former hypothesis. Comparison of N termini of Ig-binding proteins revealed less diversity than that of the M proteins' N termini, suggesting that these proteins are under less selective pressure to change.

Amino Acid Sequence

Herpesviral deoxythymidine kinases contain a site analogous to the phosphoryl-binding arginine-rich region of porcine adenylate kinase; comparison of secondary structure predictions and conservation.

Twelve herpesviral deoxythymidine kinases were examined for regions of sequence similarity by multiple alignment. Six highly conserved sites were observed. Site 1 corresponded to a glycine-rich loop that forms part of the ATP-binding pocket in porcine adenylate kinase (PAK), and site 5 corresponded to a region in PAK, located on one lobe of the cleft, that contains arginine residues that bind substrate phosphoryl groups. Site 3, consisting of the motif -DRH-, is thought to be involved in thymine/deoxythymidine recognition; site 4, which is nearby, probably participates in this function as well. The functions of sites 2 and 6 have not been identified. Secondary structure predictions were made by the Garnier method and averaged for each position in the multiple alignment. The structure predicted for all six sites was typically a short flexible region (turn or coil) at or adjacent to the site, flanked by rigid structures (helix or sheet) on either side.

Adenylate Kinase

Prediction of an rRNA methyltransferase domain in human tumor-specific nucleolar protein P120.

Using computer methods for identification of amino acid motifs in sequence databases and multiple alignment, it is shown that human proliferation-associated nucleolar protein P120 contains a putative methyltransferase domain that is conserved in a group of bacterial proteins. It is hypothesized that P120 and the related prokaryotic proteins are rRNA methylases required for division of all types of cells.

Amino Acid Sequence

Novel GACG-hairpin pair motif in the 5' untranslated region of type C retroviruses related to murine leukemia virus.

We searched for the presence of common RNA structural motifs in mammalian type C retroviruses related to murine leukemia viruses and the closely related avian spleen necrosis virus. A novel motif consisting of a pair of hairpins, called hairpin pair motif, was detected in the 5' untranslated regions of the genomes of these retroviruses. A combination of computational analyses that included the assessment of phylogenetic sequence conservation by multiple alignment, the search for regions with unusual RNA folding properties, and the analysis of RNA secondary structure by suboptimal free-energy calculations highlighted the significance of this hairpin pair motif. The hairpin pair motif encompasses 70 to 80 nucleotides between the splice donor site and the gag translational initiation codon of these viruses. The motif is composed of two adjacent hairpins both with a perfectly conserved GACG tetraloop. We propose that the novel GACG-hairpin pair motif described here constitutes an essential component of the regulatory machinery in these type C retroviruses.

Base Sequence

Multidomain organization of eukaryotic guanine nucleotide exchange translation initiation factor eIF-2B subunits revealed by analysis of conserved sequence motifs.

Computer-assisted analysis of amino acid sequences using methods for database screening with individual sequences and with multiple alignment blocks reveals a complex multidomain organization of yeast proteins GCD6 and GCD1, and mammalian homolog of GCD6-subunits of the eukaryotic translation initiation factor eIF-2B involved in GDP/GTP exchange on eIF-2. It is shown that these proteins contain a putative nucleotide-binding domain related to a variety of nucleotidyltransferases, most of which are involved in nucleoside diphosphate-sugar formation in bacteria. Three conserved motifs, one of which appears to be a variant of the phosphate-binding site (P-loop) and another that may be considered a specific version of the Mg(2+)-binding site of NTP-utilizing enzymes, were identified in the nucleotidyltransferase-related domain. Together with the third unique motif adjacent to the the P-loop, these motifs comprise the signature of a new superfamily of nucleotide-binding domains. A domain consisting of hexapeptide amino acid repeats with a periodic distribution of bulky hydrophobic residues (isoleucine patch), which previously have been identified in bacterial acetyltransferases, is located toward the C-terminus from the nucleotidyltransferase-related domain. Finally, at the very C-termini of GCD6, eIF-2B epsilon, and two other eukaryotic translation initiation factors, eIF-4 gamma and eIF-5, there is a previously undetected, conserved domain. It is hypothesized that the nucleotidyltransferase-related domain is directly involved in the GDP/GTP exchange, whereas the C-terminal conserved domain may be involved in the interaction of eIF-2B, eIF-4 gamma, and eIF-5 with eIF-2.

Amino Acid Sequence

Helical fold prediction for the cyclin box.

The smooth progression of the eukaryotic cell cycle relies on the periodic activation of members of a family of cell cycle kinases by regulatory proteins called cyclins. Outside of the cell cycle, cyclin homologs play important roles in regulating the assembly of transcription complexes; distant structural relatives of the conserved cyclin core or "box" can also function as general transcription factors (like TFIIB) or survive embedded in the chain of the tumor suppressor, retinoblastoma protein. The present work attempts the prediction of the canonical secondary, supersecondary, and tertiary fold of the minimal cyclin box domain using a combination of techniques that make use of the evolutionary information captured in a multiple alignment of homolog sequences. A tandem set of closely packed, helical modules are predicted to form the cyclin box domain.

Amino Acid Sequence

A new family of carbon-nitrogen hydrolases.

Using computer methods for database search and multiple alignment, statistically significant sequence similarities were identified between several nitrilases with distinct substrate specificity, cyanide hydratases, aliphatic amidases, beta-alanine synthase, and a few other proteins with unknown molecular function. All these proteins appear to be involved in the reduction of organic nitrogen compounds and ammonia production. Sequence conservation over the entire length, as well as the similarity in the reactions catalyzed by the known enzymes in this family, points to a common catalytic mechanism. The new family of enzymes is characterized by several conserved motifs, one of which contains an invariant cysteine that is part of the catalytic site in nitrilases. Another highly conserved motif includes an invariant glutamic acid that might also be involved in catalysis.

Amidohydrolases

Prediction of structurally conserved regions of D-specific hydroxy acid dehydrogenases by multiple alignment with formate dehydrogenase.

We propose a multiple alignment of the sequence of formate dehydrogenase with the D-specific 2-hydroxy acid dehydrogenases family. Structurally conserved regions are predicted for those sequences corresponding to important regions of the catalytic and the coenzyme binding domains defined from the known three-dimensional structure of the formate dehydrogenase, namely the nicotinamide binding site (beta D to beta F) and the beta A-loop-alpha B region containing the typical glycine pattern of the adenosine binding site, the catalytic histidine/aspartic acid pair and an arginine probably involved in the interaction with the carboxyl group of the substrate.

Alcohol Oxidoreductases

Characterization of the nuclear gene encoding mitochondrial aconitase in the marine red alga Gracilaria verrucosa.

We have cloned a nuclear gene from the marine red alga Gracilaria verrucosa that encodes the complete 779 amino-acid mitochondrial aconitase (m-ACN), the first characterized from a photosynthetic organism. The N-terminal 28 deduced amino acids are predicted to constitute the mitochondrial transit peptide, the first described from a red alga. Putative transcriptional cis-acting elements were identified in the upstream untranslated region. The G. verrucosa m-ACN gene (m-ACN) is present in a single copy and is located ca. 1.5 kb upstream from the single-copy polyubiquitin gene. The single spliceosomal intron is located near the 5' end of the region encoding the mature m-ACN in precisely the same location and phase as intron 2 in Caenorhabditis elegans m-ACN; sequences at its 3' and 5' splice junctions and at the predicted lariat branch point conform well to the eukaryote consensus sequences. Multiple protein-sequence alignment of m-ACN, bacterial aconitase (b-ACN) and iron-responsive element-binding protein (IRE-BP), and phylogenetic analyses, revealed that m-ACN does not share a recent common ancestry with either b-ACN or IRE-BP.

Aconitate Hydratase

Expansion of the mammalian 3 beta-hydroxysteroid dehydrogenase/plant dihydroflavonol reductase superfamily to include a bacterial cholesterol dehydrogenase, a bacterial UDP-galactose-4-epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus.

Mammalian 3 beta-hydroxysteroid dehydrogenase and plant dihydroflavonol reductases are descended from a common ancestor. Here we present evidence that Nocardia cholesterol dehydrogenase, E. coli UDP-galactose-4 epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus are homologous to 3 beta-hydroxysteroid dehydrogenase and dihydroflavonol reductase. Analysis of a multiple alignment of these sequences indicates that viral ORFs are most closely related to the mammalian 3 beta-hydroxysteroid dehydrogenases. The ancestral protein of this superfamily is likely to be one that metabolized sugar nucleotides. The sequence similarity between 3 beta-hydroxysteroid dehydrogenase and the viral ORFs is sufficient to suggest that these ORFs have an activity that is similar to 3 beta-hydroxysteroid dehydrogenase or cholesterol dehydrogenase, although the putative substrates are not yet known.

3-Hydroxysteroid Dehydrogenases

Computer-assisted dissection of rolling circle DNA replication.

A comparative analysis of the proteins involved in initiation and termination of rolling circle replication (RCR) was performed using computer-assisted methods of data based screening, motif search and multiple amino acid sequence alignment. Two vast classes of such proteins were delineated, one of these being associated with RCR proper, and the other with mobilization (conjugal transfer) of plasmid DNA. The common denominator of the two classes was found to be a conserved amino acid motif that consists of the sequence HisUHisUUU (U--bulky hydrophobic residue; hereafter HUH motif). Based on analogies with metalloenzymes, it is hypothesized that the two conserved His residues this motif may be involved in metal ion coordination required for the activity of the RCR and mobilization proteins. The proteins of the replication (Rep) class contained two additional conserved motifs, with the motif around the Tyr residue(s) forming the covalent link with nicked DNA being located C-proximally of the HUH motif. This class further split into two large superfamilies and several smaller families, with the proteins belonging to a single but not to different (super)families demonstrating statistically significant similarity to each other. Superfamily I, prototyped by the gene A proteins of small isometric single-stranded (ss) DNA bacteriophages, included also Rep proteins of P2-related double-stranded (ds) DNA bacteriophages, the small phage-plasmid hybrid phasyl, and several cyanobacterial and archaebacterial plasmids. These proteins contained two invariant Tyr residues separated by three partially conserved amino acids, suggesting that they all may share the cleavage-ligation mechanism proposed for phi X174 A protein and involving alternate covalent binding of both tyrosines to DNA (Van Mansfeld, A.D., Van Teeffelen, H.A., Baas, P.D., Jansz, H.S., 1986. Nucl. Acids Res. 14, 4229-4238). Superfamily II included Rep proteins of a number of ssDNA plasmids replicating mainly in gram-positive bacteria that unexpectedly were shown to be related to the Rep proteins of plant geminiviruses. Conservation of the "HUH" motif and a motif around the putative DNA-linking Tyr residue was observed also in the Rep proteins of animal parvoviruses containing linear ssDNA with a terminal hairpin and replicating via the rolling hairpin mechanism. The class of plasmid mobilization (Mob) proteins was characterized by the opposite orientation of the conserved motifs, with the (putative) DNA-linking Tyr being located N-proximally of the "HUH" motif.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

A model for human cytochrome P450 2D6 based on homology modeling and NMR studies of substrate binding.

The cytochrome P450 responsible for the debrisoquine/sparteine polymorphism (P450 2D6) has been produced in large quantities by expression of a modified cDNA in baculovirus. A polyhistidine extension was incorporated at the C-terminus of the expressed protein, which, after purification of the protein on a nickel-agarose column, could be removed proteolytically by treatment with thrombin. Purified yields of P450 2D6 were 2.4 mg from 700 mL of cell culture. The protein had a greater than 90% heme content and was fully active, having no residual absorbance at 420 nm in the reduced CO complex. The quantities produced allowed direct study of the interaction of the substrate codeine with the enzyme by paramagnetic relaxation effects on the NMR spectrum of the substrate. Distances between the heme iron atom and substrate protons were calculated from these experiments, and the orientation of the substrate in the binding pocket was determined. This showed that codeine was bound with the methoxy group of the molecule closest to the heme iron (iron-methyl proton distance of 3.1 +/- 0.1 A), consistent with the observed O-demethylation to morphine. A model of the complex Of P450 2D6 with codeine was built from a multiple sequence and structure alignment of the known crystal structures for P450s, incorporating the experimental constraints derived from the NMR studies. This showed that the overall fold Of P450 2D6 is more similar to that of P450 BM3 than to either P450 cam or P450 terp. Codeine binds to P450 2D6 so that the methoxy group is directly above the A ring of the heme, while the basic nitrogen interacts with the carboxylate of aspartate 301.

Amino Acid Sequence

Mutagenesis and the molecular modeling of the rat angiotensin II receptor (AT1).

The molecular interaction involved in the ligand binding of the rat angiotensin II receptor (AT1A) was studied by site-directed mutagenesis and receptor model building. The three-dimensional structure of AT1A was constructed on the basis of a multiple amino acid sequence alignment of seven transmembrane domain receptors and angiotensin II receptors and after the beta 2 adrenergic receptor model built on the template of the bacteriorhodopsin structure. These data indicated that there are conserved residues that are actively involved in the receptor-ligand interaction. Eleven conserved residues in AT1, His166, Arg167, Glu173, His183, Glu185, Lys199, Trp253, His256, Phe259, Thr260, and Asp263, were targeted individually for site-directed mutation to Ala. Using COS-7 cells transiently expressing these mutated receptors, we found that the binding of angiotensin II was not affected in three of the mutations in the second extracellular loop, whereas the ligand binding affinity was greatly reduced in mutants Lys199-->Ala, Trp253-->Ala, Phe259-->Ala, Asp263-->Ala, and Arg167-->Ala. These amino acid residues appeared to provide binding sites for Ang II. The molecular modeling provided useful structural information for the peptide hormone receptor AT1A. Binding of EXP985, a nonpeptide angiotensin II antagonist, was found to be involved with Arg167 but not Lys199.

Amino Acid Sequence

RNA sequence of potato virus X strain HB.

The genomic RNA of the potato virus X (PVX) strain HB, isolated in Bolivia and able to overcome all known resistance genes, has been cloned and sequenced. The PVXHB RNA sequence is 6432 nucleotides long and contains, similarly to the RNAs of other PVX strains, five open reading frames encoding proteins of M(r)s 165.1K, 24.5K, 12.4K, 7.6K and 25.1K (coat protein), respectively. Multiple amino acid sequence alignments of the coat proteins of four PVX strains identified eight amino acid residues unique for PVXHB. Structural prediction comparisons of the coat proteins of PVXHB and of the other strains suggest a general structural similarity. However, two of the eight amino acid residues unique for strain HB gave rise to a change in the predicted coat protein structure, suggesting a possible involvement in the resistance-breaking activity of PVXHB.

Amino Acid Sequence

Molecular dissection of the beta subunit of F1-ATPase into peptide fragments.

Partial digestion of the native beta subunit of F1-ATPase from the thermophilic Bacillus strain PS3 by three different proteases produced a limited number of peptide fragments. In most cases, the peptides remained associated, and the gross structure of the beta subunit was not destroyed. Furthermore, most peptides were able to reassociate into the form of the beta subunit after denaturating urea treatment. Therefore, the cleaved sites are most likely located in water-exposed loop regions in the tertiary structure of the protein. Almost all peptides were analyzed, and 17 cleaved sites were determined. From the analysis of the distribution of cleaved sites and deletions or insertions in the multiple amino acid sequence alignment of proteins homologous to the beta subunit, locations of five loops and four candidate loops in the beta subunit are suggested. There are two large loops in the central region of the beta subunit sequence, and dicyclohexylcarbodiimide-reactive Glu190 is located in one of them. Tyr341, involved in putative catalytic ATP binding, is also found in one of the loops. Then, taking cleaved sites as a reference, two kinds of expression plasmids, each of which carried genes of two complementary peptide fragments, 1-193 and 198-473 or 1-284 and 285-473, were constructed and expressed in Escherichia coli. For each plasmid, two peptides were coexpressed, associated into a stable beta subunit form in E. coli cells, and purified without dissociation. When these beta subunits were denatured by urea and applied to polyacrylamide gel without denaturant, a protein band with the same mobility as that of the beta subunit appeared, indicating that reassociation of peptide fragments into the form of the beta subunit occurred upon removal of urea. These beta subunits retained the ability to reconstitute the alpha 3 beta 3 gamma complexes even though the efficiency of reconstitution and the recovered ATPase activities were decreased. These complexes were stable at high or low temperature, and ATPase activities were sensitive to inhibition by N3-.

Amino Acid Sequence

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence

Flexible protein sequence patterns. A sensitive method to detect weak structural similarities.

The concept of a flexible protein sequence pattern is defined. In contrast to conventional pattern matching, template or sequence alignment methods, flexible patterns allow residue patterns typical of a complete protein fold to be developed in terms of residue positions (elements), separated by gaps of defined range. An efficient dynamic programming algorithm is presented to enable the best alignment(s) of a pattern with a sequence to be identified. The flexible pattern method is evaluated in detail by reference to the globin protein family, and by comparison to alignment techniques that exploit single sequence, multiple sequence and secondary structural information. A flexible pattern derived from seven globins aligned on structural criteria successfully discriminates all 345 globins from non-globins in the Protein Identification Resource database. Furthermore, a pattern that uses helical regions from just human alpha-haemoglobin identified 337 globins compared to 318 for the best non-pattern global alignment method. Patterns derived from successively fewer, yet more highly conserved positions in a structural alignment of seven globins show that as few as 38 residue positions (25 buried hydrophobic, 4 exposed and 9 others) may be used to uniquely identify the globin fold. The study suggests that flexible patterns gain discriminating power both by discarding regions known to vary within the protein family, and by defining gaps within specific ranges. Flexible patterns therefore provide a convenient and powerful bridge between regular expression pattern matching techniques and more conventional local and global sequence comparison algorithms.

Amino Acid Sequence