Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Yeast Rio1p is the founding member of a novel subfamily of protein serine kinases involved in the control of cell cycle progression.

Rio1p was identified as a protein serine kinase founding a novel subfamily. It is highly conserved from Archaea to man and only distantly related to previously established protein kinase families. Nevertheless, analysis of multiple protein sequence alignments shows that those amino acid residues that are important for either structure or catalytic activity in conventional protein kinases are also conserved in members of the Rio1p family at the respective positions (corresponding to domains I-XI of protein kinases). Recombinant Rio1p from Escherichia coli and tagged Rio1p from yeast has kinase activity in vitro, and mutation of amino acid residues that are conserved and indispensable for catalytic activity (i.e. ATP-binding motif, catalytic centre) abrogates activity. RIO1 is essential in yeast and plays a role in cell cycle progression. After sporulation of RIO1/rio1 diploids, RIO1-disrupted progeny cease growth after one to three cell divisions and arrest as either large unbudded or large-budded cells. Cells deprived of Rio1p are enlarged and arrest either in G1 or in mitosis mainly with the DNA at the bud neck and short spindles (a phenotype also seen in cells carrying a weak allele), suggesting that Rio1p activity is required for at least at two steps during the cell division cycle: for entrance into S phase and for exit from mitosis. The weak RIO1 allele leads to increased plasmid loss.

Cell Cycle↗

Evolution of p53 in hypoxia-stressed Spalax mimics human tumor mutation.

The tumor suppressor gene p53 controls cellular response to a variety of stress conditions, including DNA damage and hypoxia, leading to growth arrest and/or apoptosis. Inactivation of p53, found in 40-50% of human cancers, confers selective advantage under hypoxic microenvironment during tumor progression. The mole rat, Spalax, spends its entire life cycle underground at decidedly lower oxygen tensions than any other mammal studied. Because a wide range of respiratory adaptations to hypoxic stress evolved in Spalax, we speculated that it might also have developed hypoxia adaptation mechanisms analogous to the genetic/epigenetic alterations acquired during tumor progression. Comparing Spalax with human and mouse p53 revealed an arginine (R) to lysine (K) substitution in Spalax (Arg-174 in human) in the DNA-binding domain, identical to known tumor associated mutations. Multiple p53 sequence alignments with 41 additional species confirmed that Arg-174 is highly conserved. Reporter assays uncovered that Spalax p53 protein is unable to induce apoptosis-regulating target genes, resulting in no expression of apaf1 and partial expression of puma, pten, and noxa. However, cell cycle arrest and p53 stabilization/homeostasis genes were overactivated by Spalax p53. Lys-174 was found critical for apaf1 expression inactivation. A DNA-free p53 structure model predicts that Arg-174 is important for dimerization, whereas Spalax Lys-174 prevents such interactions. Similar neighboring mutations found in human tumors favor growth arrest rather than apoptosis. We hypothesize that, in an analogy with human tumor progression, Spalax underwent remarkable adaptive p53 evolution during 40 million years of underground hypoxic life.

Adaptation, Physiological↗

Conversion of cucumber linoleate 13-lipoxygenase to a 9-lipoxygenating species by site-directed mutagenesis.

Multiple lipoxygenase sequence alignments and structural modeling of the enzyme/substrate interaction of the cucumber lipid body lipoxygenase suggested histidine 608 as the primary determinant of positional specificity. Replacement of this amino acid by a less-space-filling valine altered the positional specificity of this linoleate 13-lipoxygenase in favor of 9-lipoxygenation. These alterations may be explained by the fact that H608V mutation may demask the positively charged guanidino group of R758, which, in turn, may force an inverse head-to-tail orientation of the fatty acid substrate. The R758L+H608V double mutant exhibited a strongly reduced reaction rate and a random positional specificity. Trilinolein, which lacks free carboxylic groups, was oxygenated to the corresponding (13S)-hydro(pero)xy derivatives by both the wild-type enzyme and the linoleate 9-lipoxygenating H608V mutant. These data indicate the complete conversion of a linoleate 13-lipoxygenase to a 9-lipoxygenating species by a single point mutation. It is hypothesized that H608V exchange may alter the orientation of the substrate at the active site and/or its steric configuration in such a way that a stereospecific dioxygen insertion at C-9 may exclusively take place.

Amino Acid Substitution↗

Folding pathway mediated by an intramolecular chaperone. A functional peptide chaperone designed using sequence databases.

Catalytic domains of several prokaryotic and eukaryotic protease families require dedicated N-terminal propeptide domains or "intramolecular chaperones" to facilitate correct folding. Amino acid sequence analysis of these families establishes three important characteristics: (i) propeptides are almost always less conserved than their cognate catalytic domains, (ii) they contain a large number of charged amino acids, and (iii) propeptides within different protease families display insignificant sequence similarity. The implications of these findings are, however, unclear. In this study, we have used subtilisin as our model to redesign a peptide chaperone using information databases. Our goal was to establish the minimum sequence requirements for a functional subtilisin propeptide, because such information could facilitate subsequent design of tailor-made chaperones. A decision-based computer algorithm that maintained conserved residues but varied all non-conserved residues from a multiple protein sequence alignment was developed and utilized to design a novel peptide sequence (ProD). Interestingly, despite a difference of 5 pH units between their isoelectric points and despite displaying only 16% sequence identity with the wild-type propeptide (ProWT), ProD chaperones folding and functions as a potent subtilisin inhibitor. The computed secondary structures and hydrophobic patterns within these two propeptides are similar. However, unlike ProWT, ProD adopts a well defined alpha-beta conformation as an isolated peptide and forms a stoichiometric complex with mature subtilisin. The CD spectra of this complex is similar to ProWT.subtilisin. Our results establish that despite low sequence identity and dramatically different charge distribution, both propeptides adopt similar structural scaffolds. Hence, conserved scaffolds and hydrophobic patterns, but not absolute charge, dictate propeptide function.

Algorithms↗

POLINA: detection and evaluation of single amino acid substitutions in protein superfamilies.

MOTIVATION AND RESULTS: An algorithm is described for the quick identification and evaluation of amino acid substitutions in multiple protein sequence alignment. The strategy is based on the calculation of a relative conservation index for an amino acid at each position of the alignment. The algorithm is implemented in the computer program POLINA (Protein Oriented LINear Analysis) which provides a summary of analysis in a format suitable for import into a graphing program. AVAILABILITY: A copy of source code is available upon request from the authors or can be downloaded via the WWW at http://www.geocities.com/Athens/4654/POLINA.h tml. CONTACT: slevin@aecom.yu.edu; bsatir@aecom.yu.edu

Algorithms↗

A novel RNA-binding motif in omnipotent suppressors of translation termination, ribosomal proteins and a ribosome modification enzyme?

Using computer methods for database search, multiple alignment, protein sequence motif analysis and secondary structure prediction, a putative new RNA-binding motif was identified. The novel motif is conserved in yeast omnipotent translation termination suppressor SUP1, the related DOM34 protein and its pseudogene homologue; three groups of eukaryotic and archaeal ribosomal proteins, namely L30e, L7Ae/S6e and S12e; an uncharacterized Bacillus subtilis protein related to the L7A/S6e group; and Escherichia coli ribosomal protein modification enzyme RimK. We hypothesize that a new type of RNA-binding domain may be utilized to deliver additional activities to the ribosome.

Amino Acid Sequence↗

Concerted evolution of duplicate fla genes in Campylobacter.

Campylobacters have two similar copies (flaA and flaB) of their flagellin gene. It has been hypothesized that the two copies can serve for antigenic phase variation. Analysis of polymorphisms within aligned multiple DNA sequences of the Campylobacter flagellin genes revealed high pairwise homoplasy indexes between flaB/flaB pairs that were not observed between any flaA/flaA pairings or flaA/flaB pairings. Thus it seems there are constraints on the sequence of flaB that distinguish it from flaA. Nevertheless, segments of the two genes that are highly variable between strains are conserved between the flaA and flaB copies of the genes within a strain. The patterns of synonymous and non-synonymous differences suggest that one segment of the flagellin sequence is under selective pressure at the amino acid sequence level. Another segment of the protein is maintained within a strain by conversion or recombination. Comparisons of strict consensus amino acid sequences did not reveal any motifs that are uniquely FlaA or FlaB, but there are differences between FlaA and FlaB in those amino acids available for post-translational modification. The observed pattern of concerted evolution of portions of a structural gene is an unusual finding in bacteria and should be searched for with other duplicated genes. Concerted evolution was unexpected for genes involved in phase variation since it minimizes the antigenic repertoire that can be expressed by a single clone in the face of the host immune response.

Amino Acid Sequence↗

FindTarget: software for subtractive genome analysis.

In silico subtractive/differential genome analysis is a powerful approach for identifying genus- or species-specific genes, or groups of genes that are responsible for a unique phenotype. By this method, one searches for genes present in one group of bacteria and absent in another group. A software package has been developed, named FindTarget, that has a user-friendly web interface to facilitate differential genome analysis. The user chooses the genomes to compare, the similarity criteria and the thresholds to decide if a gene has a counterpart in another genome. The searches are based on BLASTP comparisons of proteomes. FindTarget also includes access to sequences, coloured multiple alignments, phylogenetic trees of conserved proteins and links to public annotated databases which provide a means for validation of the results. To illustrate this approach, a FindTarget search for genes putatively involved in the specificity of cell envelope synthesis of Gram-negative bacteria is presented. The results show that most of the identified genes are clearly involved in cell wall processes, underlining the power of such an approach in general and that of FindTarget in particular.

Bacterial Proteins↗

Multi-virulence-locus sequence typing of Listeria monocytogenes.

A multi-virulence-locus sequence typing (MVLST) scheme was developed for subtyping Listeria monocytogenes, and the results obtained using this scheme were compared to those of pulsed-field gel electrophoresis (PFGE) and the published results of other typing methods, including ribotyping (RT) and multilocus sequence typing (MLST). A set of 28 strains (eight different serotypes and three known genetic lineages) of L. monocytogenes was selected from a strain collection (n > 1,000 strains) to represent the genetic diversity of this species. Internal fragments (ca. 418 to 469 bp) of three virulence genes (prfA, inlB, and inlC) and three virulence-associated genes (dal, lisR, and clpP) were sequenced and analyzed. Multiple DNA sequence alignment identified 10 (prfA), 19 (inlB), 13 (dal), 10 (lisR), 17 (inlC), and 16 (clpP) allelic types and a total of 28 unique sequence types. Comparison of MVLST with automated EcoRI-RT and PFGE with ApaI enzymatic digestion showed that MVLST was able to differentiate strains that were indistinguishable by RT (13 ribotypes; discrimination index = 0.921) or PFGE (22 profiles; discrimination index = 0.970). Comparison of MVLST with housekeeping-gene-based MLST analysis showed that MVLST provided higher discriminatory power for serotype 1/2a and 4b strains than MLST. Cluster analysis based on the intragenic sequences of the selected virulence genes indicated a strain phylogeny closely related to serotypes and genetic lineages. In conclusion, MVLST may improve the discriminatory power of MLST and provide a convenient tool for studying the local epidemiology of L. monocytogenes.

Amino Acid Sequence↗

Culture-independent analysis of fecal enterobacteria in environmental samples by single-cell mRNA profiling.

A culture-independent method called mRNA profiling has been developed for the analysis of fecal enterobacteria and their physiological status in environmental samples. This taxon-specific approach determines the single-cell content of selected gene transcripts whose abundance is either directly or inversely proportional to growth state. Fluorescence in situ hybridization using fluorochrome-labeled oligonucleotide probes was used to measure the cellular concentration of fis and dps mRNA. Relative levels of these transcripts provided a measure of cell growth state and the ability to enumerate fecal enterobacterial cell number. Orthologs were cloned by inverse PCR from several major enterobacterial genera, and probes specific for fecal enterobacteria were designed using multiple DNA sequence alignments. Probe specificity was determined experimentally using pure and mixed cultures of the major enterobacterial genera as well as secondary treated wastewater samples seeded with pure culture inocula. Analysis of the fecal enterobacterial community resident in unseeded secondary treated wastewater detected fluctuations in transcript abundance that were commensurate with incubation time and nutrient availability and demonstrated the utility of the method using environmental samples. mRNA profiling provides a new strategy to improve wastewater disinfection efficiency by accelerating water quality analysis.

Bacterial Outer Membrane Proteins↗

SplitTester: software to identify domains responsible for functional divergence in protein family.

BACKGROUND: Many protein families have undergone functional divergence after gene duplications such that current subgroups of the family carry out overlapping but distinct biological roles. For the protein families with known functional subtypes (a functional split), we developed the software, SplitTester, to identify potential regions that are responsible for the observed distinct functional subtypes within the same protein family. RESULTS: Our software, SplitTester, takes a multiple protein sequences alignment as input, generated from protein members of two subgroups with known functional divergence. SplitTester was designed to construct the neighbor joining tree (a split cluster) from variable-sized sliding windows across the alignment in a process called split-clustering. SplitTester identifies the regions, whose split cluster is consistent with the functional split, but may be inconsistent with the phylogeny of the protein family. We hypothesize that at least some number of these identified regions, which are not following a random mutation process, are responsible for the observed functional split. To test our method, we used reverse transcriptase from a group of Pseudoviridae retrotransposons: to identify residues specific for diverged primer recognition. Candidate regions were then mapped onto the three dimensional structures of reverse transcriptase. The locations of these amino acids within the enzyme are consistent with their biological roles. CONCLUSION: SplitTester aims to identify specific domain sequences responsible for functional divergence of subgroups within a protein family. From the analysis of retroelements reverse transcriptase family, we successfully identified the regions splitting this family according to the primer specificity, implying their functions in the specific primer selection.

Algorithms↗

The nematode leucine-rich repeat-containing, G protein-coupled receptor (LGR) protein homologous to vertebrate gonadotropin and thyrotropin receptors is constitutively active in mammalian cells.

The receptors for LH, FSH, and TSH belong to the large G protein-coupled, seven-transmembrane protein family and are unique in having a large N-terminal extracellular (ecto-) domain containing leucine-rich repeats important for interactions with the large glycoprotein hormone ligands. Recent studies indicated the evolution of an expanding family of homologous leucine-rich repeat-containing, G protein-coupled receptors (LGRs), including the three known glycoprotein hormone receptors; mammalian LGR4 and LGR5; and LGRs in sea anemone, fly, and snail. We isolated nematode LGR cDNA and characterized its gene from the Caenorhabditis elegans genome. This receptor cDNA encodes 929 amino acids consisting of a signal peptide for membrane insertion, an ectodomain with nine leucine-rich repeats, a seven-TM region, and a long C-terminal tail. The nematode LGR has five potential N-linked glycosylation sites in its ectodomain and multiple consensus phosphorylation sites for protein kinase A and C in the cytoplasmic loop and C tail. The nematode receptor gene has 13 exons; its TM region and C tail, unlike mammalian glycoprotein hormone receptors, are encoded by multiple exons. Sequence alignments showed that the TM region of the nematode receptor has 30% identity and 50% similarity to the same region in mammalian glycoprotein hormone receptors. Although human 293T cells expressing the nematode LGR protein do not respond to human glycoprotein hormones, these cells exhibited major increases in basal cAMP production in the absence of ligand stimulation, reaching levels comparable to those in cells expressing a constitutively activated mutant human LH receptor found in patients with familial male-limited precocious puberty. Analysis of cAMP production mediated by chimeric receptors further indicated that the ectodomain and TM region of the nematode LGR and human LH receptor are interchangeable and the TM region of the nematode LGR is responsible for constitutive receptor activation. Thus, the identification and characterization of the nematode receptor provides the basis for understanding the evolutionary relationship of diverse LGRs and for future analysis of mechanisms underlying the activation of glycoprotein hormone receptors and related LGRs.

Amino Acid Motifs↗

Creating lipoxygenases with new positional specificities by site-directed mutagenesis.

In order to analyse the amino acid determinants which alter the positional specificity of plant lipoxygenases (LOXs), multiple LOX sequence alignments and structural modelling of the enzyme-substrate interactions were carried out. These alignments suggested three amino acid residues as the primary determinants of positional specificity. Here we show the generation of two plant LOXs with new positional specificities, a gamma-linoleneate 6-LOX and an arachidonate 11-LOX, by altering only one of these determinants within the active site of two plant LOXs. In the past, site-directed-mutagenesis studies have mainly been carried out with mammalian lipoxygenases (LOXs) [1]. In these experiments two regions have been identified in the primary structure containing sequence determinants for positional specificity. Amino acids aligning with the Sloane determinants [2] are highly conserved among plant LOXs. In contrast, there is amino acid heterogeneity among plant LOXs at the position that aligns with P353 of the rabbit reticulocyte 15-LOX (Borngräber determinants) [3].

Amino Acid Substitution↗

Comparative analysis of the secondary structural motifs of P450BM-3 and the regions located upstream of the calmodulin-binding domain in the nitric oxide synthases.

The proposed method of multiple alignment of secondary structures makes it possible to estimate the multidomain enzymes' structural similarity in those cases where, following alignment, the identity appears to be inadequate or too insignificant to estimate protein relationship in spite of their functional analogy. Multiple alignment of sequences, representing the secondary structural elements of the nitric oxide synthase (NOS) N-terminal "tails", localized upstream of the calmodulin (CAM)-binding domain, and the P450 superfamily representative P450BM-3 has revealed the existence of a common secondary structural motif, standing of 60% identity between the polypeptide chain packing of NOS and the full sequence of P450BM3. This fact may point to the existence of their common ancestor. Presence of the linker olygopeptides within NOS permitted us to determine the boundary of the cytochrome-like part of NOS.

Amino Acid Sequence↗

Differences between pair-wise and multi-sequence alignment methods affect vertebrate genome comparisons.

Producing complete and accurate alignments of multiple genomic sequences is complex and prone to errors, especially with sequences generated from highly diverged species. In this article, we show that multi-sequence (as opposed to pair-wise) alignment methods are substantially better at aligning (or 'capturing') all of the available orthologous sequence from phylogenetically diverse vertebrates (i.e. those separated by relatively long branch lengths). Maximum gains are obtained only when sequences from many species are aligned. Such multi-sequence alignments contain significant amounts of exonic and highly conserved non-exonic sequences that are not captured in pair-wise alignments, thus illustrating the importance of the alignment method used for performing comparative genome analyses.

Animals↗