Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

WAViS server for handling, visualization and presentation of multiple alignments of nucleotide or amino acids sequences.

Web Alignment Visualization Server contains a set of web-tools designed for quick generation of publication-quality color figures of multiple alignments of nucleotide or amino acids sequences. It can be used for identification of conserved regions and gaps within many sequences using only common web browsers. The server is accessible at http://wavis.img.cas.cz.

Computer Graphics↗

Minding the gap: frequency of indels in mtDNA control region sequence data and influence on population genetic analyses.

Insertions and deletions (indels) result in sequences of various lengths when homologous gene regions are compared among individuals or species. Although indels are typically phylogenetically informative, occurrence and incorporation of these characters as gaps in intraspecific population genetic data sets are rarely discussed. Moreover, the impact of gaps on estimates of fixation indices, such as F(ST), has not been reviewed. Here, I summarize the occurrence and population genetic signal of indels among 60 published studies that involved alignments of multiple sequences from the mitochondrial DNA (mtDNA) control region of vertebrate taxa. Among 30 studies observing indels, an average of 12% of both variable and parsimony-informative sites were composed of these sites. There was no consistent trend between levels of population differentiation and the number of gap characters in a data block. Across all studies, the average influence on estimates of PhiST was small, explaining only an additional 1.8% of among population variance (range 0.0-8.0%). Studies most likely to observe an increase in PhiST with the inclusion of gap characters were those with < 20 variable sites, but a near equal number of studies with few variable sites did not show an increase. In contrast to studies at interspecific levels, the influence of indels for intraspecific population genetic analyses of control region DNA appears small, dependent upon total number of variable sites in the data block, and related to species-specific characteristics and the spatial distribution of mtDNA lineages that contain indels.

Animals↗

DiffTool: building, visualizing and querying protein clusters.

UNLABELLED: DiffTool is a resource to build and visualize protein clusters computed from a sequence database. The package provides a clustering tool to construct protein families according to sequence similarities and a web interface to query the corresponding clusters. A subtractive genome analysis tool selects protein families specific for a genome or a group of genomes. For each protein cluster, DiffTool includes access to sequences, coloured multiple alignments and phylogenetic trees. AVAILABILITY: A cluster database built from yeast and complete prokaryotic genomes is queryable at http://bioweb.pasteur.fr/seqanal/difftool. All the Perl sources are freely available to non-profit organizations upon request.

Cluster Analysis↗

The MPI Bioinformatics Toolkit for protein sequence analysis.

The MPI Bioinformatics Toolkit is an interactive web service which offers access to a great variety of public and in-house bioinformatics tools. They are grouped into different sections that support sequence searches, multiple alignment, secondary and tertiary structure prediction and classification. Several public tools are offered in customized versions that extend their functionality. For example, PSI-BLAST can be run against regularly updated standard databases, customized user databases or selectable sets of genomes. Another tool, Quick2D, integrates the results of various secondary structure, transmembrane and disorder prediction programs into one view. The Toolkit provides a friendly and intuitive user interface with an online help facility. As a key feature, various tools are interconnected so that the results of one tool can be forwarded to other tools. One could run PSI-BLAST, parse out a multiple alignment of selected hits and send the results to a cluster analysis tool. The Toolkit framework and the tools developed in-house will be packaged and freely available under the GNU Lesser General Public Licence (LGPL). The Toolkit can be accessed at http://toolkit.tuebingen.mpg.de.

Computational Biology↗

Inter- and intraspecies variations of the 16S-23S rDNA intergenic spacer region of various streptococcal species.

The 16S-23S rDNA intergenic spacer regions (ISR) of different streptococcal species and subspecies were amplified with primers derived from the highly conserved flanking regions of the 16S rRNA and 23S rRNA genes. The single sized amplicons showed a uniform pattern for S. agalactiae, S. dysgalactiae subsp. dysgalactiae (serogroup C), S. dysgalactiae subsp. equisimilis (serogroup G), S. dysgalactiae subsp. dysgalactiae (serogroup L), S. canis, S. phocae, S. uberis, S. parauberis, S. pyogenes and S. equi subsp. equi, respectively. The amplicons of S. equi subsp. zooepidemicus, S. porcinus and S. suis appeared with 3, 5 and 3 different sizes, respectively. ISR of selected strains of each species or subspecies investigated were sequenced and multiple aligned. This allowed a separation of ISR into regions, with 7 regions for S. agalactiae, S. dysgalactiae subsp. dysgalactiae (serogroup C), S. dysgalactiae subsp. equisimilis (serogroup G), S. dysgalactiae subsp. dysgalactiae (serogroup L), S. canis, S. phocae, S. pyogenes and S. suis, 8 regions for S. uberis and S. parauberis and mostly 9 regions for S. equi subsp. equi, S. equi subsp. zooepidemicus and S. porcinus. Region 4, encoding the transfer RNA for alanine (tRNA(Ala)), was present and identical for all isolates investigated. The size and sequence of ISR appears to be a unique marker for streptococci of various species and subspecies and could be used for bacterial identification. In addition the size and sequence variations of ISR of S. equi subsp. zooepidemicus, S. porcinus and S. suis allows a molecular typing of isolates of these species possibly useful in epidemiological aspects.

DNA, Bacterial↗

Homology-extended sequence alignment.

We present a profile-profile multiple alignment strategy that uses database searching to collect homologues for each sequence in a given set, in order to enrich their available evolutionary information for the alignment. For each of the alignment sequences, the putative homologous sequences that score above a pre-defined threshold are incorporated into a position-specific pre-alignment profile. The enriched position-specific profile is used for standard progressive alignment, thereby more accurately describing the characteristic features of the given sequence set. We show that owing to the incorporation of the pre-alignment information into a standard progressive multiple alignment routine, the alignment quality between distant sequences increases significantly and outperforms state-of-the-art methods, such as T-COFFEE and MUSCLE. We also show that although entirely sequence-based, our novel strategy is better at aligning distant sequences when compared with a recent contact-based alignment method. Therefore, our pre-alignment profile strategy should be advantageous for applications that rely on high alignment accuracy such as local structure prediction, comparative modelling and threading.

Algorithms↗

Genetic diversity on 16S rDNA sequence and phylogenic tree analysis in Pasteurella pneumotropica strains isolated from laboratory animals.

To reveal the genetic diversity of Pasteurella pneumotropica, the 16S rDNA sequence and multiple alignments were performed for 35 strains (from 17 mice, 13 rats, 3 hamsters, 1 rabbit, and 1 guinea pig) identified as P. pneumotropica using a commercial biochemical test kit or PCR test and two reference strains (ATCC 35149 and CNP160). Each strain showed a close similarity with one of the following organisms: P. pneumotropica (M75083), Bisgaard taxon22 (AY172726), Pasteurella sp. MCCM00235 (AF224300), Pasteurellaceae gen. sp. Forsyth A3 (AF224301), and Actinobacillus muris (AF024526) on GenBank, and were divided into six clusters on a phylogenic tree. Two reference strains, P. pneumotropica biotype Jawetz and Heyl, were classified at both ends of the clusters. Our conclusion is that P. pneumotropica should be reclassified because of the very wide genetic diversity that exists.

Animals↗

Simplifying amino acid alphabets by means of a branch and bound algorithm and substitution matrices.

MOTIVATION: Protein and DNA are generally represented by sequences of letters. In a number of circumstances simplified alphabets (where one or more letters would be represented by the same symbol) have proved their potential utility in several fields of bioinformatics including searching for patterns occurring at an unexpected rate, studying protein folding and finding consensus sequences in multiple alignments. The main issue addressed in this paper is the possibility of finding a general approach that would allow an exhaustive analysis of all the possible simplified alphabets, using substitution matrices like PAM and BLOSUM as a measure for scoring. RESULTS: The computational approach presented in this paper has led to a computer program called AlphaSimp (Alphabet Simplifier) that can perform an exhaustive analysis of the possible simplified amino acid alphabets, using a branch and bound algorithm together with standard or user-defined substitution matrices. The program returns a ranked list of the highest-scoring simplified alphabets. When the extent of the simplification is limited and the simplified alphabets are maintained above ten symbols the program is able to complete the analysis in minutes or even seconds on a personal computer. However, the performance becomes worse, taking up to several hours, for highly simplified alphabets. AVAILABILITY: AlphaSimp and other accessory programs are available at http://bioinformatics.cribi.unipd.it/alphasimp

Algorithms↗

Identification and characterization of surrogate peptide ligand for orphan G protein-coupled receptor mas using phage-displayed peptide library.

In the present study, a phage-displayed random peptide library was used to identify surrogate peptide ligands for orphan GPCR mas. Sequence analysis of the isolated phage clones indicated a selective enrichment of some peptide sequences. Moreover, multiple alignments of the isolated phage clones gave two conserved peptide motifs from which we synthesized peptide MBP7 for further evaluation. Characterization of the representative phage clones and the synthetic peptide MBP7 by immunocytochemistry revealed a strong punctate cell surface staining in CHO cells expressing mas-GFP fusion protein. The isolated phage clones and synthetic peptide MBP7 induced mas internalization in a stable CHO cell clone (MC0M80) over-expressing mas. In addition, MBP7-stimulated phospholipase C activity and intracellular calcium mobilization in these same cells. In summary, we have demonstrated a systematic approach to derive surrogate peptide ligands for orphan GPCRs. With this technique, we have identified two conserved peptide motifs which allow us to identify potential protein partners for mas, and have generated a peptide agonist MBP7 which will be invaluable for functional characterization of the mas oncogene.

Amino Acid Sequence↗

Improved prediction for N-termini of alpha-helices using empirical information.

The prediction of the secondary structure of proteins from their amino acid sequences remains a key component of many approaches to the protein folding problem. The most abundant form of regular secondary structure in proteins is the alpha-helix, in which specific residue preferences exist at the N-terminal locations. Propensities derived from these observed amino acid frequencies in the Protein Data Bank (PDB) database correlate well with experimental free energies measured for residues at different N-terminal positions in alanine-based peptides. We report a novel method to exploit this data to improve protein secondary structure prediction through identification of the correct N-terminal sequences in alpha-helices, based on existing popular methods for secondary structure prediction. With this algorithm, the number of correctly predicted alpha-helix start positions was improved from 30% to 38%, while the overall prediction accuracy (Q3) remained the same, using cross-validated testing. Although the algorithm was developed and tested on multiple sequence alignment-based secondary structure predictions, it was also able to improve the predictions of start locations by methods that use single sequences to make their predictions. Furthermore, the residue frequencies at N-terminal positions of the improved predictions better reflect those seen at the N-terminal positions of alpha-helices in proteins. This has implications for areas such as comparative modeling, where a more accurate prediction of the N-terminal regions of alpha-helices should benefit attempts to model adjacent loop regions. The algorithm is available as a Web tool, located at http://rocky.bms.umist.ac.uk/elephant.

Databases, Protein↗

Identification of substrate orienting and phosphorylation sites within tryptophan hydroxylase using homology-based molecular modeling.

Tryptophan hydroxylase (TPH) is the initial and rate-limiting enzyme in the biosynthesis of serotonin. The inherent instability of TPH has prevented a crystallographic structure from being resolved. For this reason, multiple sequence alignment-based molecular modeling was utilized to generate a full-length model of human TPH. Previously determined crystal coordinates of two highly homologous proteins, phenylalanine hydroxylase and tyrosine hydroxylase, were used as templates. Analysis of the model aided rational mutagenesis studies to further dissect the regulation and catalysis of TPH. Using rational site-directed mutagenesis, it was determined that Tyr235 (Y235), within the active site of TPH, appears to be involved as a tryptophan substrate orienting residue. The mutants Y235A and Y235L displayed reduced specific activity compared to wild-type TPH ( approximately 5 % residual activity). The K(m) of tryptophan for the Y235A (564 microM) and Y235L (96 microM) mutant was significantly increased compared to wild-type TPH (42 microM). In addition, kinetic analyses were performed on wild-type TPH and a deletion construct that lacks the amino terminal autoregulatory sequence (TPH NDelta15). This sequence in phenylalanine hydroxylase (residues 19 to 33) has previously been proposed to act as a steric regulator of substrate accessibility to the active site. Changes in the steady-state kinetics for tetrahydrobiopterin (BH(4)) and tryptophan for TPH NDelta15 were not observed. Finally, it was demonstrated that both Ser58 and Ser260 are substrates for Ca(2+)/calmodulin-dependent protein kinase II. Additional analysis of this model will aid in deciphering the regulation and substrate specificity of TPH, as well as providing a basis to understand as yet to be identified polymorphisms.

Amino Acid Sequence↗

Genetic relatedness of six North-Indian butterfly species (Lepidoptera :Pieridae) based on 16S rRNA sequence analysis.

The present work involves the assessment of level of genetic relatedness or divergence amongst the six North-Indian species of Lepidoptera belonging to family Pieridae and sub family Pierinae on the basis of sequence variation of 16S ribosomal RNA. The PCR amplified products of these species were directly sequenced using ABI Prism BigDye Terminator Sequencing Kits (Applied Biosystems). The multiple nucleotide sequence alignment analysis has revealed several differences across these species. Significantly high percentage of A + T base composition content ranging between 73.13% (Ixias pyrene ) and 79.20 % (Pieris brassica) was observed in studied species. The percentage divergence in the investigated species of Pieridae family varied from 5.5% to 21.7%. The two species of Catopsilia revealed minimum sequence divergence of only 5.5%, whereas the other two groups of Ixias and Pieris revealed 15.5% and 8.6% sequence divergence, respectively. Pieris canidia and Ixias pyrene are genetically most divergent (21.7%) amongst the studied lepidopteran species. Phylogenetic analysis based on 16S rRNA nucleotide sequence revealed grouping of six species of Lepidoptera in the form of two different clusters, each cluster being represented by two species from the same genera. The separate taxonomic grouping of these Indian species has been observed when compared with several species of Piernae and Coliadinae subfamilies from other country isolates.

Animals↗

The inference of evolutionary trees from molecular data.

1. Procedures for multiple alignment of sequence data, subsequent phylogenetic inference, and testing of the trees derived are presented. 2. The assumptions underlying different approaches and the extent to which they are valid are discussed.

Amino Acid Sequence↗

Sequence and transcriptional analysis of the nourseothricin acetyltransferase-encoding gene nat1 from Streptomyces noursei.

We have determined the nucleotide (nt) sequence of nat1, a gene encoding nourseothricin (Nc) acetyltransferase (AT) from Streptomyces noursei, and its transcriptional start point (tsp). The nt sequence upstream from the coding region is completely different from that of the stat gene (encoding streptothricin AT) from Streptomyces lavendulae [S. Horinouchi, K. Furuya, M. Nishiyama, H. Suzuki and T. Beppu, J. Bacteriol. 169 (1987) 1929-1937], even though the nt sequences of the two genes and the deduced amino acid (aa) sequences of the two enzymes show a high degree of similarity. Another stat gene, derived from a Gram-negative plasmid, showed only deduced aa similarity, but not nt sequence similarity, to the above two. A database search for related aa sequences did not reveal any clear-cut homologies to other types of protein. A multiple aa sequence alignment of several ATs is presented.

Acetyltransferases↗

Evidence for an evolutionary relationship among type-II restriction endonucleases.

Type-II restriction-modification (R-M) systems comprise two enzymes, a DNA methyltransferase (MTase) and a restriction endonuclease (ENase), each of which specifically interact with the same 4-8 bp sequence. All type-II MTases share several amino acid (aa) sequence motifs, which makes an evolutionary relatedness among these enzymes probable. The type-II ENases, in contrast, except for some homologous isoschizomers, do not share significant aa sequence similarity. Therefore, ENases in general have been considered unrelated. Here we show that in addition to the analysis of the genotype (aa sequence), a comparison of the phenotype (recognition sequence) of these enzymes can provide independent information regarding evolutionary relationships, and thereby, help to analyze the significance of weak aa sequence similarities. Multistep Monte-Carlo analyses were employed to demonstrate that the recognition sequences of those ENases, which were found to be related by a progressive multiple aa sequence alignment, are more similar to each other than would be expected by chance. This analysis supports the notion that not only type-II MTases, but also type-II ENases did not arise independently in evolution, but rather evolved from one or a few primordial DNA-modifying and DNA-cleaving enzymes, respectively.

Amino Acid Sequence↗

Designer antibacterial peptides kill fluoroquinolone-resistant clinical isolates.

A significant number of Escherichia coli and Klebsiella pneumoniae bacterial strains in urinary tract infections are resistant to fluoroquinolones. Peptide antibiotics are viable alternatives although these are usually either toxic or insufficiently active. By applying multiple alignment and sequence optimization steps, we designed multifunctional proline-rich antibacterial peptides that maintained their DnaK-binding ability in bacteria and low toxicity in eukaryotes, but entered bacterial cells much more avidly than earlier peptide derivatives. The resulting chimeric and statistical analogues exhibited 8-32 microg/mL minimal inhibitory concentration efficacies in Muller-Hinton broth against a series of clinical pathogens. Significantly, the best peptide, compound 5, A3-APO, retained full antibacterial activity in the presence of mouse serum. Across a set of eight fluoroquinolone-resistant clinical isolates, peptide 5 was 4 times more potent than ciprofloxacin. On the basis of the in vitro efficacy, toxicity, and pharmacokinetics data, we estimate that peptide 5 will be suitable for treating infections in the 3-5 mg/kg dose range.

Amino Acid Sequence↗

Yeast Rio1p is the founding member of a novel subfamily of protein serine kinases involved in the control of cell cycle progression.

Rio1p was identified as a protein serine kinase founding a novel subfamily. It is highly conserved from Archaea to man and only distantly related to previously established protein kinase families. Nevertheless, analysis of multiple protein sequence alignments shows that those amino acid residues that are important for either structure or catalytic activity in conventional protein kinases are also conserved in members of the Rio1p family at the respective positions (corresponding to domains I-XI of protein kinases). Recombinant Rio1p from Escherichia coli and tagged Rio1p from yeast has kinase activity in vitro, and mutation of amino acid residues that are conserved and indispensable for catalytic activity (i.e. ATP-binding motif, catalytic centre) abrogates activity. RIO1 is essential in yeast and plays a role in cell cycle progression. After sporulation of RIO1/rio1 diploids, RIO1-disrupted progeny cease growth after one to three cell divisions and arrest as either large unbudded or large-budded cells. Cells deprived of Rio1p are enlarged and arrest either in G1 or in mitosis mainly with the DNA at the bud neck and short spindles (a phenotype also seen in cells carrying a weak allele), suggesting that Rio1p activity is required for at least at two steps during the cell division cycle: for entrance into S phase and for exit from mitosis. The weak RIO1 allele leads to increased plasmid loss.

Cell Cycle↗

Evolution of p53 in hypoxia-stressed Spalax mimics human tumor mutation.

The tumor suppressor gene p53 controls cellular response to a variety of stress conditions, including DNA damage and hypoxia, leading to growth arrest and/or apoptosis. Inactivation of p53, found in 40-50% of human cancers, confers selective advantage under hypoxic microenvironment during tumor progression. The mole rat, Spalax, spends its entire life cycle underground at decidedly lower oxygen tensions than any other mammal studied. Because a wide range of respiratory adaptations to hypoxic stress evolved in Spalax, we speculated that it might also have developed hypoxia adaptation mechanisms analogous to the genetic/epigenetic alterations acquired during tumor progression. Comparing Spalax with human and mouse p53 revealed an arginine (R) to lysine (K) substitution in Spalax (Arg-174 in human) in the DNA-binding domain, identical to known tumor associated mutations. Multiple p53 sequence alignments with 41 additional species confirmed that Arg-174 is highly conserved. Reporter assays uncovered that Spalax p53 protein is unable to induce apoptosis-regulating target genes, resulting in no expression of apaf1 and partial expression of puma, pten, and noxa. However, cell cycle arrest and p53 stabilization/homeostasis genes were overactivated by Spalax p53. Lys-174 was found critical for apaf1 expression inactivation. A DNA-free p53 structure model predicts that Arg-174 is important for dimerization, whereas Spalax Lys-174 prevents such interactions. Similar neighboring mutations found in human tumors favor growth arrest rather than apoptosis. We hypothesize that, in an analogy with human tumor progression, Spalax underwent remarkable adaptive p53 evolution during 40 million years of underground hypoxic life.

Adaptation, Physiological↗