Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

The histone database: a comprehensive WWW resource for histones and histone fold-containing proteins.

The Histone Database (HDB) is an annotated and searchable collection of all full-length sequences and structures of histone and non-histone proteins containing the histone fold motif. These sequences are both eukaryotic and archaeal in origin. Several new histone fold-containing proteins have been identified, including Spt7p, and a few false positives have been removed from the earlier version of HDB. Database contents include compilations of post-translational modifications for each of the core and linker histones, as well as genomic information in the form of map loci for the human histone gene complement, with the genetic loci linked to Online Mendelian Inheritance in Man (OMIM). Conflicts between similar sequence entries from a number of source databases are also documented. Newly added to the HDB are multiple sequence alignments in which predicted functions of histone fold amino acid residues are annotated. The database is freely accessible through the WWW at http://genome.nhgri.nih.gov/histones/

Amino Acid Sequence↗

LGICdb: the ligand-gated ion channel database.

Ligand-Gated Ion Channels (LGIC) are polymeric transmembrane proteins involved in the fast response to numerous neurotransmitters. All these receptors are formed by homologous subunits and the last two decades revealed an unexpected wealth of genes coding for these subunits. The Ligand-Gated Ion Channel database (LGICdb) has been developed to handle this increasing amount of data. The database aims to provide only one entry for each gene, containing annotated nucleic acid and protein sequences. The repository is carefully structured and the entries can be retrieved by various criteria. In addition to the sequences, the LGICdb provides multiple sequence alignments, phylogenetic analyses and atomic coordinates when available. The database is accessible via the World Wide Web (http://www.pasteur.fr/recherche/banques/LGIC /LGIC.html), where it is continuously updated. The version 16 (September 2000) available for download contained 333 entries covering 34 species.

Databases, Factual↗

TIGRFAMs: a protein family resource for the functional identification of proteins.

TIGRFAMs is a collection of protein families featuring curated multiple sequence alignments, hidden Markov models and associated information designed to support the automated functional identification of proteins by sequence homology. We introduce the term 'equivalog' to describe members of a set of homologous proteins that are conserved with respect to function since their last common ancestor. Related proteins are grouped into equivalog families where possible, and otherwise into protein families with other hierarchically defined homology types. TIGRFAMs currently contains over 800 protein families, available for searching or downloading at www.tigr.org/TIGRFAMs. Classification by equivalog family, where achievable, complements classification by orthology, superfamily, domain or motif. It provides the information best suited for automatic assignment of specific functions to proteins from large-scale genome sequencing projects.

Databases, Factual↗

The Celera Discovery System.

The Celera Discovery System (CDS) is a web-accessible research workbench for mining genomic and related biological information. Users have access to the human and mouse genome sequences with annotation presented in summary form in BioMolecule Reports for genes, transcripts and proteins. Over 40 additional databases are available, including sequence, mapping, mutation, genetic variation, mRNA expression, protein structure, motif and classification data. Data are accessible by browsing reports, through a variety of interactive graphical viewers, and by advanced query capability provided by the LION SRS search engine. A growing number of sequence analysis tools are available, including sequence similarity, pattern searching, multiple sequence alignment and Hidden Markov Model search. A user workspace keeps track of queries and analyses. CDS is widely used by the academic research community and requires a subscription for access. The system and academic pricing information are available at http://cds.celera.com.

Animals↗

A comparison of profile hidden Markov model procedures for remote homology detection.

Profile hidden Markov models (HMMs) are amongst the most successful procedures for detecting remote homology between proteins. There are two popular profile HMM programs, HMMER and SAM. Little is known about their performance relative to each other and to the recently improved version of PSI-BLAST. Here we compare the two programs to each other and to non-HMM methods, to determine their relative performance and the features that are important for their success. The quality of the multiple sequence alignments used to build models was the most important factor affecting the overall performance of profile HMMs. The SAM T99 procedure is needed to produce high quality alignments automatically, and the lack of an equivalent component in HMMER makes it less complete as a package. Using the default options and parameters as would be expected of an inexpert user, it was found that from identical alignments SAM consistently produces better models than HMMER and that the relative performance of the model-scoring components varies. On average, HMMER was found to be between one and three times faster than SAM when searching databases larger than 2000 sequences, SAM being faster on smaller ones. Both methods were shown to have effective low complexity and repeat sequence masking using their null models, and the accuracy of their E-values was comparable. It was found that the SAM T99 iterative database search procedure performs better than the most recent version of PSI-BLAST, but that scoring of PSI-BLAST profiles is more than 30 times faster than scoring of SAM models.

Amino Acid Sequence↗

The TIGRFAMs database of protein families.

TIGRFAMs is a collection of manually curated protein families consisting of hidden Markov models (HMMs), multiple sequence alignments, commentary, Gene Ontology (GO) assignments, literature references and pointers to related TIGRFAMs, Pfam and InterPro models. These models are designed to support both automated and manually curated annotation of genomes. TIGRFAMs contains models of full-length proteins and shorter regions at the levels of superfamilies, subfamilies and equivalogs, where equivalogs are sets of homologous proteins conserved with respect to function since their last common ancestor. The scope of each model is set by raising or lowering cutoff scores and choosing members of the seed alignment to group proteins sharing specific function (equivalog) or more general properties. The overall goal is to provide information with maximum utility for the annotation process. TIGRFAMs is thus complementary to Pfam, whose models typically achieve broad coverage across distant homologs but end at the boundaries of conserved structural domains. The database currently contains over 1600 protein families. TIGRFAMs is available for searching or downloading at www.tigr.org/TIGRFAMs.

Animals↗

Conservation of structure and function among tyrosine recombinases: homology-based modeling of the lambda integrase core-binding domain.

Tyrosine recombinases participate in diverse biological processes by catalyzing recombination between specific DNA sites. Although a conserved protein fold has been described for the catalytic (CAT) domains of five recombinases, structural relationships between their core-binding (CB) domains remain unclear. Despite differences in the specificity and affinity of core-type DNA recognition, a conserved binding mechanism is suggested by the shared two-domain motif in crystal structure models of the recombinases Cre, XerD and Flp. We have found additional evidence for conservation of the CB domain fold. Comparison of XerD and Cre crystal structures showed that their CB domains are closely related; the three central alpha-helices of these domains are superposable to within 1.44 A. A structure-based multiple sequence alignment containing 25 diverse CB domain sequences provided evidence for widespread conservation of both structural and functional elements in this fold. Based upon the Cre and XerD crystal structures, we employed homology modeling to construct a three-dimensional structure for the lambda integrase CB domain. The model provides a conceptual framework within which many previously identified, functionally important amino acid residues were investigated. In addition, the model predicts new residues that may participate in core-type DNA binding or dimerization, thereby providing hypotheses for future genetic and biochemical experiments.

Amino Acid Sequence↗

Assessing functional divergence in EF-1alpha and its paralogs in eukaryotes and archaebacteria.

A number of methods have recently been published that use phylogenetic information extracted from large multiple sequence alignments to detect sites that have changed properties in related protein families. In this study we use such methods to assess functional divergence between eukaryotic EF-1alpha (eEF-1alpha), archaebacterial EF-1alpha (aEF-1alpha) and two eukaryote-specific EF-1alpha paralogs-eukaryotic release factor 3 (eRF3) and Hsp70 subfamily B suppressor 1 (HBS1). Overall, the evolutionary modes of aEF-1alpha, HBS1 and eRF3 appear to significantly differ from that of eEF-1alpha. However, functionally divergent (FD) sites detected between aEF-1alpha and eEF-1alpha only weakly overlap with sites implicated as putative EF-1beta or aminoacyl-tRNA (aa-tRNA) binding residues in EF-1alpha, as expected based on the shared ancestral primary translational functions of these two orthologs. In contrast, FD sites detected between eEF-1alpha and its paralogs significantly overlap with the putative EF-1beta and/or aa-tRNA binding sites in EF-1alpha. In eRF3 and HBS1, these sites appear to be released from functional constraints, indicating that they bind neither eEF-1beta nor aa-tRNA. These results are consistent with experimental observations that eRF3 does not bind to aa-tRNA, but do not support the 'EF-1alpha-like' function recently proposed for HBS1. We re-assess the available genetic data for HBS1 in light of our analyses, and propose that this protein may function in stop codon-independent peptide release.

Amino Acid Sequence↗

STING Millennium: A web-based suite of programs for comprehensive and simultaneous analysis of protein structure and sequence.

STING Millennium Suite (SMS) is a new web-based suite of programs and databases providing visualization and a complex analysis of molecular sequence and structure for the data deposited at the Protein Data Bank (PDB). SMS operates with a collection of both publicly available data (PDB, HSSP, Prosite) and its own data (contacts, interface contacts, surface accessibility). Biologists find SMS useful because it provides a variety of algorithms and validated data, wrapped-up in a user friendly web interface. Using SMS it is now possible to analyze sequence to structure relationships, the quality of the structure, nature and volume of atomic contacts of intra and inter chain type, relative conservation of amino acids at the specific sequence position based on multiple sequence alignment, indications of folding essential residue (FER) based on the relationship of the residue conservation to the intra-chain contacts and Calpha-Calpha and Cbeta-Cbeta distance geometry. Specific emphasis in SMS is given to interface forming residues (IFR)-amino acids that define the interactive portion of the protein surfaces. SMS may simultaneously display and analyze previously superimposed structures. PDB updates trigger SMS updates in a synchronized fashion. SMS is freely accessible for public data at http://www.cbi.cnptia.embrapa.br, http://mirrors.rcsb.org/SMS and http://trantor.bioc.columbia.edu/SMS.

Chymotrypsin↗

The web server of IBM's Bioinformatics and Pattern Discovery group.

We herein present and discuss the services and content which are available on the web server of IBM's Bioinformatics and Pattern Discovery group. The server is operational around the clock and provides access to a variety of methods that have been published by the group's members and collaborators. The available tools correspond to applications ranging from the discovery of patterns in streams of events and the computation of multiple sequence alignments, to the discovery of genes in nucleic acid sequences and the interactive annotation of amino acid sequences. Additionally, annotations for more than 70 archaeal, bacterial, eukaryotic and viral genomes are available on-line and can be searched interactively. The tools and code bundles can be accessed beginning at http://cbcsrv.watson.ibm.com/Tspd.html whereas the genomics annotations are available at http://cbcsrv.watson.ibm.com/Annotations/.

Computational Biology↗

Hairpin-duplex equilibrium reflected in the A-->B transition in an undecamer quasi-palindrome present in the locus control region of the human beta-globin gene cluster.

Our recent work on an A-->G single nucleotide polymorphism (SNP) at the quasi-palindromic sequence d(TGGGG[A/G]CCCCA) of HS4 of the human beta-globin locus control region in an Indian population showed a significant association between the G allele and the occurrence of beta-thalassemia. Using UV-thermal denaturation, gel assay, circular dichroism (CD) and nuclease digestion experiments we have demonstrated that the undecamer quasi- palindromic sequence d(TGGGGACCCCA) (HPA11) and its reported polymorphic (SNP) version d(TGG GGGCCCCA) (HPG11) exist in hairpin-duplex equilibria. The biphasic nature of the melting profiles for both the oligonucleotides persisted at low as well as high salt concentrations. The HPG11 hairpin showed a higher T(m) than HPA11. The presence of unimolecular and bimolecular species was also shown by non-denaturating gel electrophoresis experiments. The CD spectra of both oligonucleotides showed features of the A- as well as B-type conformations and, moreover, exhibited a concentration dependence. The disappearance of the 265 nm positive CD signal in an oligomer concentration-dependent manner is indicative of an A-->B transition. The results give unprecedented insight into the in vitro structure of the quasi-palindromic sequence and provide the first report in which a hairpin-duplex equilibrium has been correlated with an A-->B interconversion of DNA. The nuclease-dependent degradation suggests that HPG11 is more resistant to nuclease than HPA11. Multiple sequence alignment of the HS4 region of the beta-globin gene cluster from different organisms revealed that this quasi-palindromic stretch is unique to Homo sapiens. We propose that quasi-palindromic sequences may form stable mini- hairpins or cruciforms in the HS4 region and might play a role in regulating beta-globin gene expression by affecting the binding of transcription factors.

Base Sequence↗

An archaebacteria-derived glutamyl-tRNA synthetase and tRNA pair for unnatural amino acid mutagenesis of proteins in Escherichia coli.

The addition of novel amino acids to the genetic code of Escherichia coli involves the generation of an aminoacyl-tRNA synthetase and tRNA pair that is 'orthogonal', meaning that it functions independently of the synthetases and tRNAs endogenous to E.coli. The amino acid specificity of the orthogonal synthetase is then modified to charge the corresponding orthogonal tRNA with an unnatural amino acid that is subsequently incorporated into a polypeptide in response to a nonsense or missense codon. Here we report the development of an orthogonal glutamic acid synthetase and tRNA pair. The tRNA is derived from the consensus sequence obtained from a multiple sequence alignment of archaeal tRNA(Glu) sequences. The glutamyl-tRNA synthetase is from the achaebacterium Pyrococcus horikoshii. The new orthogonal pair suppresses amber nonsense codons with an efficiency roughly comparable to that of the orthogonal tyrosine pair derived from Methanococcus jannaschii, which has been used to selectively incorporate a variety of unnatural amino acids into proteins in E.coli. Development of the glutamic acid orthogonal pair increases the potential diversity of unnatural amino acid structures that may be incorporated into proteins in E.coli.

Acylation↗

EyeSite: a semi-automated database of protein families in the eye.

The EyeSite is a web-based database of protein families for proteins that function in the eye and their homologous sequences. The resource clusters proteins at different levels of homology in order to facilitate functional annotation of sequences and modelling of proteins from structural homologues. Eye proteins are organized into the tissue types in which they function and are clustered into homologous families using a novel protocol employing the TribeMCL algorithm. Homologous families are further subdivided into sequence clusters for which multiple sequence alignments are generated. Structural annotations from the CATH domain database are provided for nearly 90% of the sequences, and protein family annotations from the Pfam database for approximately 86%. Homology models have also been generated where appropriate. The EyeSite is stored in a relational database and is extensively linked to other online bioinformatics resources to help relate allelic variants, annotations and clinical details to the derived data in the database. The EyeSite is available for online search, sequence information and model retrieval at http://eyesite.cryst.bbk.ac.uk/.

Amino Acid Sequence↗

PIRSF: family classification system at the Protein Information Resource.

The Protein Information Resource (PIR) is an integrated public resource of protein informatics. To facilitate the sensible propagation and standardization of protein annotation and the systematic detection of annotation errors, PIR has extended its superfamily concept and developed the SuperFamily (PIRSF) classification system. Based on the evolutionary relationships of whole proteins, this classification system allows annotation of both specific biological and generic biochemical functions. The system adopts a network structure for protein classification from superfamily to subfamily levels. Protein family members are homologous (sharing common ancestry) and homeomorphic (sharing full-length sequence similarity with common domain architecture). The PIRSF database consists of two data sets, preliminary clusters and curated families. The curated families include family name, protein membership, parent-child relationship, domain architecture, and optional description and bibliography. PIRSF is accessible from the website at http://pir.georgetown.edu/pirsf/ for report retrieval and sequence classification. The report presents family annotation, membership statistics, cross-references to other databases, graphical display of domain architecture, and links to multiple sequence alignments and phylogenetic trees for curated families. PIRSF can be utilized to analyze phylogenetic profiles, to reveal functional convergence and divergence, and to identify interesting relationships between homeomorphic families, domains and structural classes.

Amino Acid Motifs↗

PredictRegulon: a web server for the prediction of the regulatory protein binding sites and operons in prokaryote genomes.

An interactive web server is developed for predicting the potential binding sites and its target operons for a given regulatory protein in prokaryotic genomes. The program allows users to submit known or experimentally determined binding sites of a regulatory protein as ungapped multiple sequence alignments. It analyses the upstream regions of all genes in a user-selected prokaryote genome and returns the potential binding sites along with the downstream co-regulated genes (operons). The known binding sites of a regulatory protein can also be used to identify its orthologue binding sites in phylogeneticaly related genomes where the trans-acting regulator protein and cognate cis-acting DNA sequences could be conserved. PredictRegulon can be freely accessed from a link on our world wide web server: http://www.cdfd.org.in/predictregulon/.

5' Flanking Region↗

The web server of IBM's Bioinformatics and Pattern Discovery group: 2004 update.

In this report, we provide an update on the services and content which are available on the web server of IBM's Bioinformatics and Pattern Discovery group. The server, which is operational around the clock, provides access to a large number of methods that have been developed and published by the group's members. There is an increasing number of problems that these tools can help tackle; these problems range from the discovery of patterns in streams of events and the computation of multiple sequence alignments, to the discovery of genes in nucleic acid sequences, the identification--directly from sequence--of structural deviations from alpha-helicity and the annotation of amino acid sequences for antimicrobial activity. Additionally, annotations for more than 130 archaeal, bacterial, eukaryotic and viral genomes are now available on-line and can be searched interactively. The tools and code bundles continue to be accessible from http://cbcsrv.watson.ibm.com/Tspd.html whereas the genomics annotations are available at http://cbcsrv.watson.ibm.com/Annotations/.

Anti-Infective Agents↗

CARNAC: folding families of related RNAs.

We present a tool for the prediction of conserved secondary structure elements of a family of homologous non-coding RNAs. Our method does not require any prior multiple sequence alignment. Thus, it successfully applies to datasets with low primary structure similarity. The functionality is demonstrated using three example datasets: sequences of RNase P RNAs, ciliate telomerases and enterovirus messenger RNAs. CARNAC has a web server that can be accessed at the URL http://bioinfo.lifl.fr/carnac.

Internet↗

COLORADO3D, a web server for the visual analysis of protein structures.

COLORADO3D is a World Wide Web server for the visual presentation of three-dimensional (3D) protein structures. COLORADO3D indicates the presence of potential errors (detected by ANOLEA, PROSAII, PROVE or VERIFY3D), identifies buried residues and depicts sequence conservations. As input, the server takes a file of Protein Data Bank (PDB) coordinates and, optionally, a multiple sequence alignment. As output, the server returns a PDB-formatted file, replacing the B-factor column with values of the chosen parameter (structure quality, residue burial or conservation). Thus, the coordinates of the analyzed protein 'colored' by COLORADO3D can be conveniently displayed with structure viewers such as RASMOL in order to visualize the 3D clusters of regions with common features, which may not necessarily be adjacent to each other at the amino acid sequence level. In particular, COLORADO3D may serve as a tool to judge a structure's quality at various stages of the modeling and refinement (during both experimental structure determination and homology modeling). The GeneSilico group used COLORADO3D in the fifth Critical Assessment of Techniques for Protein Structure Prediction (CASP5) to successfully identify well-folded parts of preliminary homology models and to guide the refinement of misthreaded protein sequences. COLORADO3D is freely available for academic use at http://asia.genesilico.pl/colorado3d/.

Algorithms↗