Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Eukaryotic DNA polymerase amino acid sequence required for 3'----5' exonuclease activity.

We have identified an amino-proximal sequence motif, Phe-Asp-Ile-Glu-Thr, in Saccharomyces cerevisiae DNA polymerase II that is almost identical to a sequence comprising part of the 3'----5' exonuclease active site of Escherichia coli DNA polymerase I. Similar motifs were identified by amino acid sequence alignment in related, aphidicolin-sensitive DNA polymerases possessing 3'----5' proofreading exonuclease activity. Substitution of Ala for the Asp and Glu residues in the motif reduced the exonuclease activity of partially purified DNA polymerase II at least 100-fold while preserving the polymerase activity. Yeast strains expressing the exonuclease-deficient DNA polymerase II had on average about a 22-fold increase in spontaneous mutation rate, consistent with a presumed proofreading role in vivo. In multiple amino acid sequence alignments of this and two other conserved motifs described previously, five residues of the 3'----5' exonuclease active site of E. coli DNA polymerase I appeared to be invariant in aphidicolin-sensitive DNA polymerases known to possess 3'----5' proofreading exonuclease activity. None of these residues, however, appeared to be identifiable in the catalytic subunits of human, yeast, or Drosophila alpha DNA polymerases.

Amino Acid Sequence

Selection of circularization sites in a group I IVS RNA requires multiple alignments of an internal template-like sequence.

Circularization and reverse circularization of the Tetrahymena thermophila rRNA intervening sequence resemble the first and second steps in splicing, respectively. However, site-specific base substitutions show that different nucleotides are involved in selection of the 5' splice site and the circularization sites. Furthermore, a substitution at the major circularization site that prevents circularization can be suppressed by second substitutions at two different nucleotide positions. A model is proposed in which adjacent and overlapping sequences can function as a binding site, forming a short duplex with the sequence at the circularization site and thus directing circularization and reverse circularization. Because the 5' exon-binding site and three potential circularization binding sites fall within a contiguous eight nucleotide region, this sequence may translocate relative to the catalytic core of the ribozyme in a template-like manner.

Animals

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software

Immunoglobulin-binding FcrA and Enn proteins and M proteins of group A streptococci evolved independently from a common ancestral protein.

Significant sequence homology between M proteins and immunoglobulin (Ig)-binding proteins of group A streptococci suggests that these proteins arose by gene duplication followed by the development of functional diversity due to mutations and intragenic recombinations. The deduced sequence of multiple Ig-binding proteins and M proteins were compared to distinguish between two evolutionary models. Did these functionally distinct genes originate in the distant past from duplication of a common ancestral gene and then functionally evolve independently or did they evolve more recently, one from the other by duplication of a fixed gene? Multiple alignments of conserved sequences of these proteins are consistent with the former hypothesis. Comparison of N termini of Ig-binding proteins revealed less diversity than that of the M proteins' N termini, suggesting that these proteins are under less selective pressure to change.

Amino Acid Sequence

Herpesviral deoxythymidine kinases contain a site analogous to the phosphoryl-binding arginine-rich region of porcine adenylate kinase; comparison of secondary structure predictions and conservation.

Twelve herpesviral deoxythymidine kinases were examined for regions of sequence similarity by multiple alignment. Six highly conserved sites were observed. Site 1 corresponded to a glycine-rich loop that forms part of the ATP-binding pocket in porcine adenylate kinase (PAK), and site 5 corresponded to a region in PAK, located on one lobe of the cleft, that contains arginine residues that bind substrate phosphoryl groups. Site 3, consisting of the motif -DRH-, is thought to be involved in thymine/deoxythymidine recognition; site 4, which is nearby, probably participates in this function as well. The functions of sites 2 and 6 have not been identified. Secondary structure predictions were made by the Garnier method and averaged for each position in the multiple alignment. The structure predicted for all six sites was typically a short flexible region (turn or coil) at or adjacent to the site, flanked by rigid structures (helix or sheet) on either side.

Adenylate Kinase

Prediction of an rRNA methyltransferase domain in human tumor-specific nucleolar protein P120.

Using computer methods for identification of amino acid motifs in sequence databases and multiple alignment, it is shown that human proliferation-associated nucleolar protein P120 contains a putative methyltransferase domain that is conserved in a group of bacterial proteins. It is hypothesized that P120 and the related prokaryotic proteins are rRNA methylases required for division of all types of cells.

Amino Acid Sequence

Novel GACG-hairpin pair motif in the 5' untranslated region of type C retroviruses related to murine leukemia virus.

We searched for the presence of common RNA structural motifs in mammalian type C retroviruses related to murine leukemia viruses and the closely related avian spleen necrosis virus. A novel motif consisting of a pair of hairpins, called hairpin pair motif, was detected in the 5' untranslated regions of the genomes of these retroviruses. A combination of computational analyses that included the assessment of phylogenetic sequence conservation by multiple alignment, the search for regions with unusual RNA folding properties, and the analysis of RNA secondary structure by suboptimal free-energy calculations highlighted the significance of this hairpin pair motif. The hairpin pair motif encompasses 70 to 80 nucleotides between the splice donor site and the gag translational initiation codon of these viruses. The motif is composed of two adjacent hairpins both with a perfectly conserved GACG tetraloop. We propose that the novel GACG-hairpin pair motif described here constitutes an essential component of the regulatory machinery in these type C retroviruses.

Base Sequence

Multidomain organization of eukaryotic guanine nucleotide exchange translation initiation factor eIF-2B subunits revealed by analysis of conserved sequence motifs.

Computer-assisted analysis of amino acid sequences using methods for database screening with individual sequences and with multiple alignment blocks reveals a complex multidomain organization of yeast proteins GCD6 and GCD1, and mammalian homolog of GCD6-subunits of the eukaryotic translation initiation factor eIF-2B involved in GDP/GTP exchange on eIF-2. It is shown that these proteins contain a putative nucleotide-binding domain related to a variety of nucleotidyltransferases, most of which are involved in nucleoside diphosphate-sugar formation in bacteria. Three conserved motifs, one of which appears to be a variant of the phosphate-binding site (P-loop) and another that may be considered a specific version of the Mg(2+)-binding site of NTP-utilizing enzymes, were identified in the nucleotidyltransferase-related domain. Together with the third unique motif adjacent to the the P-loop, these motifs comprise the signature of a new superfamily of nucleotide-binding domains. A domain consisting of hexapeptide amino acid repeats with a periodic distribution of bulky hydrophobic residues (isoleucine patch), which previously have been identified in bacterial acetyltransferases, is located toward the C-terminus from the nucleotidyltransferase-related domain. Finally, at the very C-termini of GCD6, eIF-2B epsilon, and two other eukaryotic translation initiation factors, eIF-4 gamma and eIF-5, there is a previously undetected, conserved domain. It is hypothesized that the nucleotidyltransferase-related domain is directly involved in the GDP/GTP exchange, whereas the C-terminal conserved domain may be involved in the interaction of eIF-2B, eIF-4 gamma, and eIF-5 with eIF-2.

Amino Acid Sequence

A new family of carbon-nitrogen hydrolases.

Using computer methods for database search and multiple alignment, statistically significant sequence similarities were identified between several nitrilases with distinct substrate specificity, cyanide hydratases, aliphatic amidases, beta-alanine synthase, and a few other proteins with unknown molecular function. All these proteins appear to be involved in the reduction of organic nitrogen compounds and ammonia production. Sequence conservation over the entire length, as well as the similarity in the reactions catalyzed by the known enzymes in this family, points to a common catalytic mechanism. The new family of enzymes is characterized by several conserved motifs, one of which contains an invariant cysteine that is part of the catalytic site in nitrilases. Another highly conserved motif includes an invariant glutamic acid that might also be involved in catalysis.

Amidohydrolases

Prediction of structurally conserved regions of D-specific hydroxy acid dehydrogenases by multiple alignment with formate dehydrogenase.

We propose a multiple alignment of the sequence of formate dehydrogenase with the D-specific 2-hydroxy acid dehydrogenases family. Structurally conserved regions are predicted for those sequences corresponding to important regions of the catalytic and the coenzyme binding domains defined from the known three-dimensional structure of the formate dehydrogenase, namely the nicotinamide binding site (beta D to beta F) and the beta A-loop-alpha B region containing the typical glycine pattern of the adenosine binding site, the catalytic histidine/aspartic acid pair and an arginine probably involved in the interaction with the carboxyl group of the substrate.

Alcohol Oxidoreductases

Characterization of the nuclear gene encoding mitochondrial aconitase in the marine red alga Gracilaria verrucosa.

We have cloned a nuclear gene from the marine red alga Gracilaria verrucosa that encodes the complete 779 amino-acid mitochondrial aconitase (m-ACN), the first characterized from a photosynthetic organism. The N-terminal 28 deduced amino acids are predicted to constitute the mitochondrial transit peptide, the first described from a red alga. Putative transcriptional cis-acting elements were identified in the upstream untranslated region. The G. verrucosa m-ACN gene (m-ACN) is present in a single copy and is located ca. 1.5 kb upstream from the single-copy polyubiquitin gene. The single spliceosomal intron is located near the 5' end of the region encoding the mature m-ACN in precisely the same location and phase as intron 2 in Caenorhabditis elegans m-ACN; sequences at its 3' and 5' splice junctions and at the predicted lariat branch point conform well to the eukaryote consensus sequences. Multiple protein-sequence alignment of m-ACN, bacterial aconitase (b-ACN) and iron-responsive element-binding protein (IRE-BP), and phylogenetic analyses, revealed that m-ACN does not share a recent common ancestry with either b-ACN or IRE-BP.

Aconitate Hydratase

Expansion of the mammalian 3 beta-hydroxysteroid dehydrogenase/plant dihydroflavonol reductase superfamily to include a bacterial cholesterol dehydrogenase, a bacterial UDP-galactose-4-epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus.

Mammalian 3 beta-hydroxysteroid dehydrogenase and plant dihydroflavonol reductases are descended from a common ancestor. Here we present evidence that Nocardia cholesterol dehydrogenase, E. coli UDP-galactose-4 epimerase, and open reading frames in vaccinia virus and fish lymphocystis disease virus are homologous to 3 beta-hydroxysteroid dehydrogenase and dihydroflavonol reductase. Analysis of a multiple alignment of these sequences indicates that viral ORFs are most closely related to the mammalian 3 beta-hydroxysteroid dehydrogenases. The ancestral protein of this superfamily is likely to be one that metabolized sugar nucleotides. The sequence similarity between 3 beta-hydroxysteroid dehydrogenase and the viral ORFs is sufficient to suggest that these ORFs have an activity that is similar to 3 beta-hydroxysteroid dehydrogenase or cholesterol dehydrogenase, although the putative substrates are not yet known.

3-Hydroxysteroid Dehydrogenases

Computer-assisted dissection of rolling circle DNA replication.

A comparative analysis of the proteins involved in initiation and termination of rolling circle replication (RCR) was performed using computer-assisted methods of data based screening, motif search and multiple amino acid sequence alignment. Two vast classes of such proteins were delineated, one of these being associated with RCR proper, and the other with mobilization (conjugal transfer) of plasmid DNA. The common denominator of the two classes was found to be a conserved amino acid motif that consists of the sequence HisUHisUUU (U--bulky hydrophobic residue; hereafter HUH motif). Based on analogies with metalloenzymes, it is hypothesized that the two conserved His residues this motif may be involved in metal ion coordination required for the activity of the RCR and mobilization proteins. The proteins of the replication (Rep) class contained two additional conserved motifs, with the motif around the Tyr residue(s) forming the covalent link with nicked DNA being located C-proximally of the HUH motif. This class further split into two large superfamilies and several smaller families, with the proteins belonging to a single but not to different (super)families demonstrating statistically significant similarity to each other. Superfamily I, prototyped by the gene A proteins of small isometric single-stranded (ss) DNA bacteriophages, included also Rep proteins of P2-related double-stranded (ds) DNA bacteriophages, the small phage-plasmid hybrid phasyl, and several cyanobacterial and archaebacterial plasmids. These proteins contained two invariant Tyr residues separated by three partially conserved amino acids, suggesting that they all may share the cleavage-ligation mechanism proposed for phi X174 A protein and involving alternate covalent binding of both tyrosines to DNA (Van Mansfeld, A.D., Van Teeffelen, H.A., Baas, P.D., Jansz, H.S., 1986. Nucl. Acids Res. 14, 4229-4238). Superfamily II included Rep proteins of a number of ssDNA plasmids replicating mainly in gram-positive bacteria that unexpectedly were shown to be related to the Rep proteins of plant geminiviruses. Conservation of the "HUH" motif and a motif around the putative DNA-linking Tyr residue was observed also in the Rep proteins of animal parvoviruses containing linear ssDNA with a terminal hairpin and replicating via the rolling hairpin mechanism. The class of plasmid mobilization (Mob) proteins was characterized by the opposite orientation of the conserved motifs, with the (putative) DNA-linking Tyr being located N-proximally of the "HUH" motif.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

Mutagenesis and the molecular modeling of the rat angiotensin II receptor (AT1).

The molecular interaction involved in the ligand binding of the rat angiotensin II receptor (AT1A) was studied by site-directed mutagenesis and receptor model building. The three-dimensional structure of AT1A was constructed on the basis of a multiple amino acid sequence alignment of seven transmembrane domain receptors and angiotensin II receptors and after the beta 2 adrenergic receptor model built on the template of the bacteriorhodopsin structure. These data indicated that there are conserved residues that are actively involved in the receptor-ligand interaction. Eleven conserved residues in AT1, His166, Arg167, Glu173, His183, Glu185, Lys199, Trp253, His256, Phe259, Thr260, and Asp263, were targeted individually for site-directed mutation to Ala. Using COS-7 cells transiently expressing these mutated receptors, we found that the binding of angiotensin II was not affected in three of the mutations in the second extracellular loop, whereas the ligand binding affinity was greatly reduced in mutants Lys199-->Ala, Trp253-->Ala, Phe259-->Ala, Asp263-->Ala, and Arg167-->Ala. These amino acid residues appeared to provide binding sites for Ang II. The molecular modeling provided useful structural information for the peptide hormone receptor AT1A. Binding of EXP985, a nonpeptide angiotensin II antagonist, was found to be involved with Arg167 but not Lys199.

Amino Acid Sequence

RNA sequence of potato virus X strain HB.

The genomic RNA of the potato virus X (PVX) strain HB, isolated in Bolivia and able to overcome all known resistance genes, has been cloned and sequenced. The PVXHB RNA sequence is 6432 nucleotides long and contains, similarly to the RNAs of other PVX strains, five open reading frames encoding proteins of M(r)s 165.1K, 24.5K, 12.4K, 7.6K and 25.1K (coat protein), respectively. Multiple amino acid sequence alignments of the coat proteins of four PVX strains identified eight amino acid residues unique for PVXHB. Structural prediction comparisons of the coat proteins of PVXHB and of the other strains suggest a general structural similarity. However, two of the eight amino acid residues unique for strain HB gave rise to a change in the predicted coat protein structure, suggesting a possible involvement in the resistance-breaking activity of PVXHB.

Amino Acid Sequence

Molecular dissection of the beta subunit of F1-ATPase into peptide fragments.

Partial digestion of the native beta subunit of F1-ATPase from the thermophilic Bacillus strain PS3 by three different proteases produced a limited number of peptide fragments. In most cases, the peptides remained associated, and the gross structure of the beta subunit was not destroyed. Furthermore, most peptides were able to reassociate into the form of the beta subunit after denaturating urea treatment. Therefore, the cleaved sites are most likely located in water-exposed loop regions in the tertiary structure of the protein. Almost all peptides were analyzed, and 17 cleaved sites were determined. From the analysis of the distribution of cleaved sites and deletions or insertions in the multiple amino acid sequence alignment of proteins homologous to the beta subunit, locations of five loops and four candidate loops in the beta subunit are suggested. There are two large loops in the central region of the beta subunit sequence, and dicyclohexylcarbodiimide-reactive Glu190 is located in one of them. Tyr341, involved in putative catalytic ATP binding, is also found in one of the loops. Then, taking cleaved sites as a reference, two kinds of expression plasmids, each of which carried genes of two complementary peptide fragments, 1-193 and 198-473 or 1-284 and 285-473, were constructed and expressed in Escherichia coli. For each plasmid, two peptides were coexpressed, associated into a stable beta subunit form in E. coli cells, and purified without dissociation. When these beta subunits were denatured by urea and applied to polyacrylamide gel without denaturant, a protein band with the same mobility as that of the beta subunit appeared, indicating that reassociation of peptide fragments into the form of the beta subunit occurred upon removal of urea. These beta subunits retained the ability to reconstitute the alpha 3 beta 3 gamma complexes even though the efficiency of reconstitution and the recovered ATPase activities were decreased. These complexes were stable at high or low temperature, and ATPase activities were sensitive to inhibition by N3-.

Amino Acid Sequence

fRagmentomics: an R package for integrating cell-free DNA fragment features with mutational status to support liquid biopsy interpretation.

SUMMARY: Liquid biopsy offers a non-invasive approach to study tumor-derived genetic material circulating in plasma. Beyond genetic alterations, the fragmentomic features of cell-free DNA-such as fragment size, genomic position, and end-motifs-provide valuable insights into the biological and clinical context of DNA release. fRagmentomics is a user-friendly R package designed to characterize cfDNA fragments overlapping one or multiple small mutations of any type, starting from an aligned sequencing file (BAM). It supports multiple mutation input formats, accommodates one-based and zero-based genomic conventions, resolves mutation representation ambiguities, and accepts any reference file in FASTA format. For each fragment overlapping a mutation of interest, fRagmentomics outputs fragment-level features including its fragment size, end-motifs, and mutational status, along with additional fragment-level or read-level information. The package implements an indel-aware and optionally soft-clip-preserving fragment size computation that improves accuracy over conventional size estimates based solely on aligned positions. AVAILABILITY AND IMPLEMENTATION: fRagmentomics is licensed under GNU General Public License v3.0 and available at https://github.com/ElsaB-Lab/fRagmentomics, https://anaconda.org/elsab-lab/r-fragmentomics and https://bioconductor.org/packages/fRagmentomics, with documentation and a tutorial. CONTACT: yoann.pradat@gustaveroussy.fr, elsa.bernard@gustaveroussy.fr. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Software

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence