Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

OrthoGUI: graphical presentation of Orthostrapper results.

SUMMARY: Orthostrapper is a program that calculates orthology support values for pairs of sequences in a multiple alignment (Storm and Sonnhammer, Bioinformatics, 18, 92-99, 2002). Here we present OrthoGUI, a web interface and display tool for Orthostrapper analysis. OrthoGUI visualizes the Orthostrapper output in both tabular and tree representations, and can also apply a clustering algorithm to identify groups of multiple orthologs, which are indicated by colour coding. AVAILABILITY: http://www.cgb.ki.se/OrthoGUI CONTACT: erik.sonnhammer@cgb.ki.se

ATP-Binding Cassette Transporters↗

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software↗

Immunoglobulin-binding FcrA and Enn proteins and M proteins of group A streptococci evolved independently from a common ancestral protein.

Significant sequence homology between M proteins and immunoglobulin (Ig)-binding proteins of group A streptococci suggests that these proteins arose by gene duplication followed by the development of functional diversity due to mutations and intragenic recombinations. The deduced sequence of multiple Ig-binding proteins and M proteins were compared to distinguish between two evolutionary models. Did these functionally distinct genes originate in the distant past from duplication of a common ancestral gene and then functionally evolve independently or did they evolve more recently, one from the other by duplication of a fixed gene? Multiple alignments of conserved sequences of these proteins are consistent with the former hypothesis. Comparison of N termini of Ig-binding proteins revealed less diversity than that of the M proteins' N termini, suggesting that these proteins are under less selective pressure to change.

Amino Acid Sequence↗

Herpesviral deoxythymidine kinases contain a site analogous to the phosphoryl-binding arginine-rich region of porcine adenylate kinase; comparison of secondary structure predictions and conservation.

Twelve herpesviral deoxythymidine kinases were examined for regions of sequence similarity by multiple alignment. Six highly conserved sites were observed. Site 1 corresponded to a glycine-rich loop that forms part of the ATP-binding pocket in porcine adenylate kinase (PAK), and site 5 corresponded to a region in PAK, located on one lobe of the cleft, that contains arginine residues that bind substrate phosphoryl groups. Site 3, consisting of the motif -DRH-, is thought to be involved in thymine/deoxythymidine recognition; site 4, which is nearby, probably participates in this function as well. The functions of sites 2 and 6 have not been identified. Secondary structure predictions were made by the Garnier method and averaged for each position in the multiple alignment. The structure predicted for all six sites was typically a short flexible region (turn or coil) at or adjacent to the site, flanked by rigid structures (helix or sheet) on either side.

Adenylate Kinase↗

The alpha/beta fold uracil DNA glycosylases: a common origin with diverse fates.

BACKGROUND: Uracil DNA glycosylases (UDGs) are major repair enzymes that protect DNA from mutational damage caused by uracil incorporated as a result of a polymerase error or deamination of cytosine. Four distinct families of UDGs have been identified, which show very limited sequence similarity to each other, although two of them have been shown to possess the same structural fold. The structural and evolutionary relationships between the rest of the UDGs remain uncertain. RESULTS: Using sequence profile searches, multiple alignment analysis and protein structure comparisons, we show here that all known UDGs possess the same fold and must have evolved from a common ancestor. Although all UDGs catalyze essentially the same reaction, significant changes in the configuration of the catalytic residues were detected within their common fold, which probably results in differences in the biochemistry of these enzymes. The extreme sequence divergence of the UDGs, which is unusual for enzymes with the same principal activity, is probably due to the major role of the uracil-flipping caused by the conformational strain enacted by the enzyme on uracil-containing DNA, as compared with the catalytic action of individual polar residues. We predict two previously undetected families of UDGs and delineate a hypothetical scenario for their evolution. CONCLUSIONS: UDGs form a single protein superfamily with a distinct structural fold and a common evolutionary origin. Differences in the catalytic mechanism of the different families combined with the construction of the catalytic pocket have, however, resulted in extreme sequence divergence of these enzymes.

Amino Acid Sequence↗

Hidden Markov multiple event sequence models: A paradigm for the spatio-temporal analysis of fMRI data.

This paper presents a novel, completely unsupervised fMRI brain mapping method that addresses the three problems of hemodynamic response function (HRF) variability, hemodynamic event timing, and fMRI response non-linearity. Spatial and temporal information are directly taken into account into the core of the activation detection process. In practice, activation detection at voxel v is formulated in terms of temporal alignment between sequences of hemodynamic response onsets (HROs) detected in the fMRI signal at v and in the spatial neighborhood of v, and the input sequence of stimuli or stimulus onsets. Event-related and epoch paradigms are considered. The multiple event sequence alignment problem is solved within the probabilistic framework of hidden Markov multiple event sequence models (HMMESMs), a new class of hidden Markov models. Results obtained on real and synthetic data significantly outperform those obtained with the popular statistical parametric mapping (SPM2) method without requiring any prior definition of the expected activation patterns, the HMMESM mapping approach being completely unsupervised.

Adolescent↗

Prediction of an rRNA methyltransferase domain in human tumor-specific nucleolar protein P120.

Using computer methods for identification of amino acid motifs in sequence databases and multiple alignment, it is shown that human proliferation-associated nucleolar protein P120 contains a putative methyltransferase domain that is conserved in a group of bacterial proteins. It is hypothesized that P120 and the related prokaryotic proteins are rRNA methylases required for division of all types of cells.

Amino Acid Sequence↗

Novel GACG-hairpin pair motif in the 5' untranslated region of type C retroviruses related to murine leukemia virus.

We searched for the presence of common RNA structural motifs in mammalian type C retroviruses related to murine leukemia viruses and the closely related avian spleen necrosis virus. A novel motif consisting of a pair of hairpins, called hairpin pair motif, was detected in the 5' untranslated regions of the genomes of these retroviruses. A combination of computational analyses that included the assessment of phylogenetic sequence conservation by multiple alignment, the search for regions with unusual RNA folding properties, and the analysis of RNA secondary structure by suboptimal free-energy calculations highlighted the significance of this hairpin pair motif. The hairpin pair motif encompasses 70 to 80 nucleotides between the splice donor site and the gag translational initiation codon of these viruses. The motif is composed of two adjacent hairpins both with a perfectly conserved GACG tetraloop. We propose that the novel GACG-hairpin pair motif described here constitutes an essential component of the regulatory machinery in these type C retroviruses.

Base Sequence↗

Quod erat demonstrandum? The mystery of experimental validation of apparently erroneous computational analyses of protein sequences.

BACKGROUND: Computational predictions are critical for directing the experimental study of protein functions. Therefore it is paradoxical when an apparently erroneous computational prediction seems to be supported by experiment. RESULTS: We analyzed six cases where application of novel or conventional computational methods for protein sequence and structure analysis led to non-trivial predictions that were subsequently supported by direct experiments. We show that, on all six occasions, the original prediction was unjustified, and in at least three cases, an alternative, well-supported computational prediction, incompatible with the original one, could be derived. The most unusual cases involved the identification of an archaeal cysteinyl-tRNA synthetase, a dihydropteroate synthase and a thymidylate synthase, for which experimental verifications of apparently erroneous computational predictions were reported. Using sequence-profile analysis, multiple alignment and secondary-structure prediction, we have identified the unique archaeal 'cysteinyl-tRNA synthetase' as a homolog of extracellular polygalactosaminidases, and the 'dihydropteroate synthase' as a member of the beta-lactamase-like superfamily of metal-dependent hydrolases. CONCLUSIONS: In each of the analyzed cases, the original computational predictions could be refuted and, in some instances, alternative strongly supported predictions were obtained. The nature of the experimental evidence that appears to support these predictions remains an open question. Some of these experiments might signify discovery of extremely unusual forms of the respective enzymes, whereas the results of others could be due to artifacts.

Acetyltransferases↗

Characterization of an early gene encoding for dUTPase in Rana grylio virus.

dUTPase (DUT) is a ubiquitous and important enzyme responsible for regulating levels of dUTP. Here, an iridovirus DUT was identified and characterized from Rana grylio virus (RGV) which is a pathogen agent in pig frog. The DUT encodes a protein of 164aa with a predicted molecular mass of 17.4 kDa, and its transcriptional initiation site was determined by 5'RACE to start from the nucleotide A at 15 nt upstream of the initiation codon ATG. Sequence comparisons and multiple alignments suggested that RGV DUT was quite similar to other identified DUTs that function as homotrimers. Phylogenetic analysis implied that DUT horizontal transfers might have occurred between the vertebrate hosts and iridoviruses. Furthermore, its temporal expression pattern during RGV infection course was characterized by RT-PCR and Western blot analysis. It begins to transcribe and translate as early as 4h postinfection (p.i.), and remains detectable at 48 h p.i. DUT-EGFP fusion protein was observed in the cytoplasm of pEGFP-N3-Dut transfected EPC cells. Immunofluorescence also confirmed DUT cytoplasm localization in RGV-infected cells. Using drug inhibition analysis by a de novo protein synthesis inhibitor (cycloheximide) and a viral DNA replication inhibitor (cytosine arabinofuranoside), RGV DUT was classified as an early (E) viral gene during the in vitro infection. Moreover, RGV DUT overexpression was shown that there was no effect on RGV replication by viral replication kinetics assay.

Amino Acid Sequence↗

Multidomain organization of eukaryotic guanine nucleotide exchange translation initiation factor eIF-2B subunits revealed by analysis of conserved sequence motifs.

Computer-assisted analysis of amino acid sequences using methods for database screening with individual sequences and with multiple alignment blocks reveals a complex multidomain organization of yeast proteins GCD6 and GCD1, and mammalian homolog of GCD6-subunits of the eukaryotic translation initiation factor eIF-2B involved in GDP/GTP exchange on eIF-2. It is shown that these proteins contain a putative nucleotide-binding domain related to a variety of nucleotidyltransferases, most of which are involved in nucleoside diphosphate-sugar formation in bacteria. Three conserved motifs, one of which appears to be a variant of the phosphate-binding site (P-loop) and another that may be considered a specific version of the Mg(2+)-binding site of NTP-utilizing enzymes, were identified in the nucleotidyltransferase-related domain. Together with the third unique motif adjacent to the the P-loop, these motifs comprise the signature of a new superfamily of nucleotide-binding domains. A domain consisting of hexapeptide amino acid repeats with a periodic distribution of bulky hydrophobic residues (isoleucine patch), which previously have been identified in bacterial acetyltransferases, is located toward the C-terminus from the nucleotidyltransferase-related domain. Finally, at the very C-termini of GCD6, eIF-2B epsilon, and two other eukaryotic translation initiation factors, eIF-4 gamma and eIF-5, there is a previously undetected, conserved domain. It is hypothesized that the nucleotidyltransferase-related domain is directly involved in the GDP/GTP exchange, whereas the C-terminal conserved domain may be involved in the interaction of eIF-2B, eIF-4 gamma, and eIF-5 with eIF-2.

Amino Acid Sequence↗

Helical fold prediction for the cyclin box.

The smooth progression of the eukaryotic cell cycle relies on the periodic activation of members of a family of cell cycle kinases by regulatory proteins called cyclins. Outside of the cell cycle, cyclin homologs play important roles in regulating the assembly of transcription complexes; distant structural relatives of the conserved cyclin core or "box" can also function as general transcription factors (like TFIIB) or survive embedded in the chain of the tumor suppressor, retinoblastoma protein. The present work attempts the prediction of the canonical secondary, supersecondary, and tertiary fold of the minimal cyclin box domain using a combination of techniques that make use of the evolutionary information captured in a multiple alignment of homolog sequences. A tandem set of closely packed, helical modules are predicted to form the cyclin box domain.

Amino Acid Sequence↗

A new family of carbon-nitrogen hydrolases.

Using computer methods for database search and multiple alignment, statistically significant sequence similarities were identified between several nitrilases with distinct substrate specificity, cyanide hydratases, aliphatic amidases, beta-alanine synthase, and a few other proteins with unknown molecular function. All these proteins appear to be involved in the reduction of organic nitrogen compounds and ammonia production. Sequence conservation over the entire length, as well as the similarity in the reactions catalyzed by the known enzymes in this family, points to a common catalytic mechanism. The new family of enzymes is characterized by several conserved motifs, one of which contains an invariant cysteine that is part of the catalytic site in nitrilases. Another highly conserved motif includes an invariant glutamic acid that might also be involved in catalysis.

Amidohydrolases↗

Structural modeling of ataxin-3 reveals distant homology to adaptins.

Spinocerebellar ataxia type 3 (SCA3) is a polyglutamine disorder caused by a CAG repeat expansion in the coding region of a gene encoding ataxin-3, a protein of yet unknown function. Based on a comprehensive computational analysis, we propose a structural model and structure-based functions for ataxin-3. Our predictive strategy comprises the compilation of multiple sequence and structure alignments of carefully selected proteins related to ataxin-3. These alignments are consistent with additional information on sequence motifs, secondary structure, and domain architectures. The application of complementary methods revealed the homology of ataxin-3 to ENTH and VHS domain proteins involved in membrane trafficking and regulatory adaptor functions. We modeled the structure of ataxin-3 using the adaptin AP180 as a template and assessed the reliability of the model by comparison with known sequence and structural features. We could further infer potential functions of ataxin-3 in agreement with known experimental data. Our database searches also identified an as yet uncharacterized family of proteins, which we named josephins because of their pronounced homology to the Josephin domain of ataxin-3.

Adaptor Protein Complex gamma Subunits↗

Classification of common functional loops of kinase super-families.

A structural classification of loops has been obtained from a set of 141 protein structures classified as kinases. A total of 1813 loops was classified into 133 subclasses (9 betabeta(links), 15 betabeta(hairpins), 31 alpha-alpha, 46 alpha-beta and 32 beta-alpha). Functional information and specific features relating subclasses and function were included in the classification. Functional loops such as the P-loop (shared by different folds) or the Gly-rich-loop, among others, were classified into structural motifs. As a result, a common mechanism of catalysis and substrate binding was proved for most kinases. Additionally, the multiple-alignment of loop sequences made within each subclass was shown to be useful for comparative modeling of kinase loops. The classification is summarized in a kinase loop database located at http://sbi.imim.es/archki.

Amino Acid Motifs↗

Identification of proteolipid from an extremely halophilic archaeon Halobacterium salinarum as an N,N'-dicyclohexyl-carbodiimide binding subunit of ATP synthase.

ATP synthesis in an extremely halophilic archaeon, Halobacterium salinarum, was inhibited by N-cyclohexyl-N'-[4-(dimethylamino)-alpha-naphthyl]carbodiimide (NCD-4), a fluorescent analog of N,N'-dicyclohexylcarbodiimide (DCCD). By tracing the fluorescent signal, a hydrophobic 8-kDa protein (proteolipid) was purified from the halobacterial membrane as one of the most DCCD-reactive proteins and its N-terminal amino acid sequence was determined. The gene encoding the proteolipid was found in the region upstream of the genes encoding the two major subunits of halobacterial A-type ATPase [K. Ihara and Y.Mukohata (1991) Arch. Biochem. Biophys. 286, 111-116]. Halobacterial proteolipid was more similar in size to the proteolipid of F-type ATPase than that of V-type ATPase. However, multiple amino acid sequence alignment of proteolipids showed a higher degree of relatedness between V-type and A-type ATPase proteolipids. Together with the recent finding of a triplicate proteolipid encoding gene from the methanogenic archaeon Methanococcus jannaschii [C. J. Bult et al. (1996) Science 273, 1058-1073], proteolipids from archaea seem to have diverse characteristics in comparison with those from eubacteria or from eukaryotes.

Adenosine Triphosphatases↗

Prediction of structurally conserved regions of D-specific hydroxy acid dehydrogenases by multiple alignment with formate dehydrogenase.

We propose a multiple alignment of the sequence of formate dehydrogenase with the D-specific 2-hydroxy acid dehydrogenases family. Structurally conserved regions are predicted for those sequences corresponding to important regions of the catalytic and the coenzyme binding domains defined from the known three-dimensional structure of the formate dehydrogenase, namely the nicotinamide binding site (beta D to beta F) and the beta A-loop-alpha B region containing the typical glycine pattern of the adenosine binding site, the catalytic histidine/aspartic acid pair and an arginine probably involved in the interaction with the carboxyl group of the substrate.

Alcohol Oxidoreductases↗

Functional and structural characterization of the human gene BHLHB5, encoding a basic helix-loop-helix transcription factor.

The genes encoding basic helix-loop-helix (bHLH) transcription factors have been implicated in many aspects of neural development, including cell growth, differentiation, and cell migration. Using both genomic and cDNA mouse and human clones encoding a neural-specific bHLH protein, human BHLHB5 was cloned and mapped to a region on chromosome 8q13 that segregates with Duane syndrome. Genomic sequence analysis of human BHLHB5 and mouse Bhlhb5 revealed that they contain a single exon encoding 381- and 355-amino-acid bHLH proteins, respectively. Multiple amino acid sequence alignments of the Bhlhb5 family members revealed several conserved motifs and an identical 147-amino-acid carboxy-terminal region that contains a 60-amino-acid bHLH domain. A 27-bp trinucleotide repeat (CAG)(9) encoding polyserine was found in human BHLHB5, but only one CAG was found at the corresponding position in the mouse Bhlhb5 and hamster BETA3 genes. Northern blot analysis of human BHLHB5 revealed brain-specific expression with the highest abundance in the cerebellum. Mouse Bhlhb5 can strongly repress a human PAX6 promoter.

Amino Acid Sequence↗