Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

MyHits: a new interactive resource for protein annotation and domain identification.

The MyHits web server (http://myhits.isb-sib.ch) is a new integrated service dedicated to the annotation of protein sequences and to the analysis of their domains and signatures. Guest users can use the system anonymously, with full access to (i) standard bioinformatics programs (e.g. PSI-BLAST, ClustalW, T-Coffee, Jalview); (ii) a large number of protein sequence databases, including standard (Swiss-Prot, TrEMBL) and locally developed databases (splice variants); (iii) databases of protein motifs (Prosite, Interpro); (iv) a precomputed list of matches ('hits') between the sequence and motif databases. All databases are updated on a weekly basis and the hit list is kept up to date incrementally. The MyHits server also includes a new collection of tools to generate graphical representations of pairwise and multiple sequence alignments including their annotated features. Free registration enables users to upload their own sequences and motifs to private databases. These are then made available through the same web interface and the same set of analytical tools. Registered users can manage their own sequences and annotations using only web tools and freeze their data in their private database for publication purposes.

Computer Graphics↗

Comparative analysis of dioxin response elements in human, mouse and rat genomic sequences.

Comparative approaches were used to identify human, mouse and rat dioxin response elements (DREs) in genomic sequences unambiguously assigned to a nucleotide RefSeq accession number. A total of 13 bona fide DREs, all including the substitution intolerant core sequence (GCGTG) and adjacent variable sequences, were used to establish a position weight matrix and a matrix similarity (MS) score threshold to rank identified DREs. DREs with MS scores above the threshold were disproportionately distributed in close proximity to the transcription start site in all three species. Gene expression assays in hepatic mouse tissue confirmed the responsiveness of 192 genes possessing a putative DRE. Previously identified functional DREs in well-characterized AhR-regulated genes including Cyp1a1 and Cyp1b1 were corroborated. Putative DREs were identified in 48 out of 2437 human-mouse-rat orthologous genes between -1500 and the transcriptional start site, of which 19 of these genes possessed positionally conserved DREs as determined by multiple sequence alignment. Seven of these nineteen genes exhibited 2,3,7,8-tetrachlorodibenzo-p-dioxin-mediated regulation, although there were significant discrepancies between in vivo and in vitro results. Interestingly, of the mouse-rat orthologous genes with a DRE between -1500 and +1500, only 37% had an equivalent human ortholog. These results suggest that AhR-mediated gene expression may not be well conserved across species, which could have significant implications in human risk assessment.

Animals↗

EFICAz: a comprehensive approach for accurate genome-scale enzyme function inference.

EFICAz (Enzyme Function Inference by Combined Approach) is an automatic engine for large-scale enzyme function inference that combines predictions from four different methods developed and optimized to achieve high prediction accuracy: (i) recognition of functionally discriminating residues (FDRs) in enzyme families obtained by a Conservation-controlled HMM Iterative procedure for Enzyme Family classification (CHIEFc), (ii) pairwise sequence comparison using a family specific Sequence Identity Threshold, (iii) recognition of FDRs in Multiple Pfam enzyme families, and (iv) recognition of multiple Prosite patterns of high specificity. For FDR (i.e. conserved positions in an enzyme family that discriminate between true and false members of the family) identification, we have developed an Evolutionary Footprinting method that uses evolutionary information from homofunctional and heterofunctional multiple sequence alignments associated with an enzyme family. The FDRs show a significant correlation with annotated active site residues. In a jackknife test, EFICAz shows high accuracy (92%) and sensitivity (82%) for predicting four EC digits in testing sequences that are <40% identical to any member of the corresponding training set. Applied to Escherichia coli genome, EFICAz assigns more detailed enzymatic function than KEGG, and generates numerous novel predictions.

Amino Acid Sequence↗

eBLOCKs: enumerating conserved protein blocks to achieve maximal sensitivity and specificity.

Classifying proteins into families and superfamilies allows identification of functionally important conserved domains. The motifs and scoring matrices derived from such conserved regions provide computational tools that recognize similar patterns in novel sequences, and thus enable the prediction of protein function for genomes. The eBLOCKs database enumerates a cascade of protein blocks with varied conservation levels for each functional domain. A biologically important region is most stringently conserved among a smaller family of highly similar proteins. The same region is often found in a larger group of more remotely related proteins with a reduced stringency. Through enumeration, highly specific signatures can be generated from blocks with more columns and fewer family members, while highly sensitive signatures can be derived from blocks with fewer columns and more members as in a superfamily. By applying PSI-BLAST and a modified K-means clustering algorithm, eBLOCKs automatically groups protein sequences according to different levels of similarity. Multiple sequence alignments are made and trimmed into a series of ungapped blocks. Motifs and position-specific scoring matrices were derived from eBLOCKs and made available for sequence search and annotation. The eBLOCKs database provides a tool for high-throughput genome annotation with maximal specificity and sensitivity. The eBLOCKs database is freely available on the World Wide Web at http://motif.stanford.edu/eblocks/ to all users for online usage. Academic and not-for-profit institutions wishing copies of the program may contact Douglas L. Brutlag (brutlag@stanford.edu). Commercial firms wishing copies of the program for internal installation may contact Jacqueline Tay at the Stanford Office of Technology Licensing (jacqueline.tay@stanford.edu; http://otl.stanford.edu/).

Algorithms↗

The Adaptive Evolution Database (TAED): a phylogeny based tool for comparative genomics.

From 138,662 embryophyte (higher plant) and 348,142 chordate genes, 4216 embryophyte and 15,452 chordate gene families were generated. For each of these gene families, multiple sequence alignments, phylogenetic trees, ratios of non-synonymous to synonymous nucleotide substitution rates (K(a)/K(s)), mappings from gene trees to the NCBI taxonomy and structural links to solved three-dimensional protein structures in the Protein Data Bank (PDB) with Grantham-weighted mutational factors were all calculated. Of the 'gene family trees', 173 embryophyte and 505 chordate branches show K(a)/K(s) >> 1 and are candidates for functional adaptation. The calculated information is available both as a gene family database and as a phylogenetically indexed resource, called 'The Adaptive Evolution Database' (TAED), available at http://www.bioinfo.no/tools/TAED.

Animals↗

QuasiMotiFinder: protein annotation by searching for evolutionarily conserved motif-like patterns.

Sequence signature databases such as PROSITE, which include amino acid segments that are indicative of a protein's function, are useful for protein annotation. Lamentably, the annotation is not always accurate. A signature may be falsely detected in a protein that does not carry out the associated function (false positive prediction, FP) or may be overlooked in a protein that does carry out the function (false negative prediction, FN). A new approach has emerged in which a signature is replaced with a sequence profile, calculated based on multiple sequence alignment (MSA) of homologous proteins that share the same function. This approach, which is superior to the simple pattern search, essentially searches with the sequence of the query protein against an MSA library. We suggest here an alternative approach, implemented in the QuasiMotiFinder web server (http://quasimotifinder.tau.ac.il/), which is based on a search with an MSA of homologous query proteins against the original PROSITE signatures. The explicit use of the average evolutionary conservation of the signature in the query proteins significantly reduces the rate of FP prediction compared with the simple pattern search. QuasiMotiFinder also has a reduced rate of FN prediction compared with simple pattern searches, since the traditional search for precise signatures has been replaced by a permissive search for signature-like patterns that are physicochemically similar to known signatures. Overall, QuasiMotiFinder and the profile search are comparable to each other in terms of performance. They are also complementary to each other in that signatures that are falsely detected in (or overlooked by) one may be correctly detected by the other.

Amino Acid Motifs↗

Phytome: a platform for plant comparative genomics.

Phytome is an online comparative genomics resource that can be applied to functional plant genomics, molecular breeding and evolutionary studies. It contains predicted protein sequences, protein family assignments, multiple sequence alignments, phylogenies and functional annotations for proteins from a large, phylogenetically diverse set of plant taxa. Phytome serves as a glue between disparate plant gene databases both by identifying the evolutionary relationships among orthologous and paralogous protein sequences from different species and by enabling cross-references between different versions of the same gene curated independently by different database groups. The web interface enables sophisticated queries on lineage-specific patterns of gene/protein family proliferation and loss. This rich dataset is serving as a platform for the unification of sequence-anchored comparative maps across taxonomic families of plants. The Phytome web interface can be accessed at the following URL: http://www.phytome.org. Batch homology searches and bulk downloads are available upon free registration.

Databases, Genetic↗

LGICdb: a manually curated sequence database after the genomes.

Ligand-gated ion channels form transmembrane ionic pores controlled by the binding of chemicals. The LGICdb aims to be a non-redundant, manually curated resource offering access to the large number of subunits composing extracellularly activated ligand-gated ion channels, such as nicotinic, ATP, GABA and glutamate ionotropic receptors. Composed of more than 500 human curated entries, the XML native database has been relocated in 2004 to the EBI. Its facilities have been enhanced with a new search system, customized multiple sequence alignments and manipulation of protein structures (http://www.ebi.ac.uk/compneur-srv/LGICdb/). Despite the vast improvement of general sequence resources, the LGICdb still provide sequences unavailable elsewhere.

Databases, Protein↗

Phytophthora functional genomics database (PFGD): functional genomics of phytophthora-plant interactions.

The Phytophthora Functional Genomics Database (PFGD; http://www.pfgd.org), developed by the National Center for Genome Resources in collaboration with The Ohio State University-Ohio Agricultural Research and Development Center (OSU-OARDC), is a publicly accessible information resource for Phytophthora-plant interaction research. PFGD contains transcript, genomic, gene expression and functional assay data for Phytophthora infestans, which causes late blight of potato, and Phytophthora sojae, which affects soybeans. Automated analyses are performed on all sequence data, including consensus sequences derived from clustered and assembled expressed sequence tags. The PFGD search filter interface allows intuitive navigation of transcript and genomic data organized by library and derived queries using modifiers, annotation keywords or sequence names. BLAST services are provided for libraries built from the transcript and genomic sequences. Transcript data visualization tools include Quality Screening, Multiple Sequence Alignment and Features and Annotations viewers. A genomic browser that supports comparative analysis via novel dynamic functional annotation comparisons is also provided. PFGD is integrated with the Solanaceae Genomics Database (SolGD; http://www.solgd.org) to help provide insight into the mechanisms of infection and resistance, specifically as they relate to the genus Phytophthora pathogens and their plant hosts.

Algal Proteins↗

Computational approaches for predicting the biological effect of p53 missense mutations: a comparison of three sequence analysis based methods.

Prediction of the biological effect of missense substitutions has become important because they are often observed in known or candidate disease susceptibility genes. In this paper, we carried out a 3-step analysis of 1514 missense substitutions in the DNA-binding domain (DBD) of TP53, the most frequently mutated gene in human cancers. First, we calculated two types of conservation scores based on a TP53 multiple sequence alignment (MSA) for each substitution: (i) Grantham Variation (GV), which measures the degree of biochemical variation among amino acids found at a given position in the MSA; (ii) Grantham Deviation (GD), which reflects the 'biochemical distance' of the mutant amino acid from the observed amino acid at a particular position (given by GV). Second, we used a method that combines GV and GD scores, Align-GVGD, to predict the transactivation activity of each missense substitution. We compared our predictions against experimentally measured transactivation activity (yeast assays) to evaluate their accuracy. Finally, the prediction results were compared with those obtained by the program Sorting Intolerant from Tolerant (SIFT) and Dayhoff's classification. Our predictions yielded high prediction accuracy for mutants showing a loss of transactivation ( approximately 88% specificity) with lower prediction accuracy for mutants with transactivation similar to that of the wild-type (67.9 to 71.2% sensitivity). Align-GVGD results were comparable to SIFT (88.3 to 90.6% and 67.4 to 70.3% specificity and sensitivity, respectively) and outperformed Dayhoff's classification (80 and 40.9% specificity and sensitivity, respectively). These results further demonstrate the utility of the Align-GVGD method, which was previously applied to BRCA1. Align-GVGD is available online at http://agvgd.iarc.fr.

Amino Acid Sequence↗

The ENCODE Project at UC Santa Cruz.

The goal of the Encyclopedia Of DNA Elements (ENCODE) Project is to identify all functional elements in the human genome. The pilot phase is for comparison of existing methods and for the development of new methods to rigorously analyze a defined 1% of the human genome sequence. Experimental datasets are focused on the origin of replication, DNase I hypersensitivity, chromatin immunoprecipitation, promoter function, gene structure, pseudogenes, non-protein-coding RNAs, transcribed RNAs, multiple sequence alignment and evolutionarily constrained elements. The ENCODE project at UCSC website (http://genome.ucsc.edu/ENCODE) is the primary portal for the sequence-based data produced as part of the ENCODE project. In the pilot phase of the project, over 30 labs provided experimental results for a total of 56 browser tracks supported by 385 database tables. The site provides researchers with a number of tools that allow them to visualize and analyze the data as well as download data for local analyses. This paper describes the portal to the data, highlights the data that has been made available, and presents the tools that have been developed within the ENCODE project. Access to the data and types of interactive analysis that are possible are illustrated through supplemental examples.

Base Sequence↗

GeMprospector--online design of cross-species genetic marker candidates in legumes and grasses.

The web program GeMprospector (URL: http://cgi-www.daimi.au.dk/cgi-chili/GeMprospector/main) allows users to automatically design large sets of cross-species genetic marker candidates targeting either legumes or grasses. The user uploads a collection of ESTs from one or more legume or grass species, and they are compared with a database of clusters of homologous EST and genomic sequences from other legumes or grasses, respectively. Multiple sequence alignments between submitted ESTs and their homologues in the appropriate database form the basis of automated PCR primer design in conserved exons such that each primer set amplifies an intron. The only user input is a collection of ESTs, not necessarily from more than one species, and GeMprospector can boost the potential of such an EST collection by combining it with a large database to produce cross-species genetic marker candidates for legumes or grasses.

Base Sequence↗

iCR: a web tool to identify conserved targets of a regulatory protein across the multiple related prokaryotic species.

Gene regulatory circuits are often commonly shared between two closely related organisms. Our web tool iCR (identify Conserved target of a Regulon) makes use of this fact and identify conserved targets of a regulatory protein. iCR is a special refined extension of our previous tool PredictRegulon- that predicts genome wide, the potential binding sites and target operons of a regulatory protein in a single user selected genome. Like PredictRegulon, the iCR accepts known binding sites of a regulatory protein as ungapped multiple sequence alignment and provides the potential binding sites. However important differences are that the user can select more than one genome at a time and the output reports the genes that are common in two or more species. In order to achieve this, iCR makes use of Cluster of Orthologous Group (COG) indices for the genes. This tool analyses the upstream region of all user-selected prokaryote genome and gives the output based on conservation target orthologs. iCR also reports the Functional class codes based on COG classification for the encoded proteins of downstream genes which helps user understand the nature of the co-regulated genes at the result page itself. iCR is freely accessible at http://www.cdfd.org.in/icr/.

Bacterial Proteins↗

TreeDet: a web server to explore sequence space.

The TreeDet (Tree Determinant) Server is the first release of a system designed to integrate results from methods that predict functional sites in protein families. These methods take into account the relation between sequence conservation and evolutionary importance. TreeDet fully analyses the space of protein sequences in either user-uploaded or automatically generated multiple sequence alignments. The methods implemented in the server represent three main classes of methods for the detection of family-dependent conserved positions, a tree-based method, a correlation based method and a method that employs a principal component analyses coupled to a cluster algorithm. An additional method is provided to highlight the reliability of the position in the alignments. The server is available at http://www.pdg.cnb.uam.es/servers/treedet.

Amino Acid Sequence↗

The MIGenAS integrated bioinformatics toolkit for web-based sequence analysis.

We describe a versatile and extensible integrated bioinformatics toolkit for the analysis of biological sequences over the Internet. The web portal offers convenient interactive access to a growing pool of chainable bioinformatics software tools and databases that are centrally installed and maintained by the RZG. Currently, supported tasks comprise sequence similarity searches in public or user-supplied databases, computation and validation of multiple sequence alignments, phylogenetic analysis and protein-structure prediction. Individual tools can be seamlessly chained into pipelines allowing the user to conveniently process complex workflows without the necessity to take care of any format conversions or tedious parsing of intermediate results. The toolkit is part of the Max-Planck Integrated Gene Analysis System (MIGenAS) of the Max Planck Society available at www.migenas.org (click 'Start Toolkit').

Animals↗

FISH--family identification of sequence homologues using structure anchored hidden Markov models.

The FISH server is highly accurate in identifying the family membership of domains in a query protein sequence, even in the case of very low sequence identities to known homologues. A performance test using SCOP sequences and an E-value cut-off of 0.1 showed that 99.3% of the top hits are to the correct family saHMM. Matches to a query sequence provide the user not only with an annotation of the identified domains and hence a hint to their function, but also with probable 2D and 3D structures, as well as with pairwise and multiple sequence alignments to homologues with low sequence identity. In addition, the FISH server allows users to upload and search their own protein sequence collection or to quarry public protein sequence data bases with individual saHMMs. The FISH server can be accessed at http://babel.ucmp.umu.se/fish/.

Databases, Protein↗

Influenza Virus Database (IVDB): an integrated information resource and analysis platform for influenza virus research.

Frequent outbreaks of highly pathogenic avian influenza and the increasing data available for comparative analysis require a central database specialized in influenza viruses (IVs). We have established the Influenza Virus Database (IVDB) to integrate information and create an analysis platform for genetic, genomic, and phylogenetic studies of the virus. IVDB hosts complete genome sequences of influenza A virus generated by Beijing Institute of Genomics (BIG) and curates all other published IV sequences after expert annotation. Our Q-Filter system classifies and ranks all nucleotide sequences into seven categories according to sequence content and integrity. IVDB provides a series of tools and viewers for comparative analysis of the viral genomes, genes, genetic polymorphisms and phylogenetic relationships. A search system has been developed for users to retrieve a combination of different data types by setting search options. To facilitate analysis of global viral transmission and evolution, the IV Sequence Distribution Tool (IVDT) has been developed to display the worldwide geographic distribution of chosen viral genotypes and to couple genomic data with epidemiological data. The BLAST, multiple sequence alignment and phylogenetic analysis tools were integrated for online data analysis. Furthermore, IVDB offers instant access to pre-computed alignments and polymorphisms of IV genes and proteins, and presents the results as SNP distribution plots and minor allele distributions. IVDB is publicly available at http://influenza.genomics.org.cn.

Databases, Genetic↗

PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.

Peroxisomes are essential organelles of eukaryotic origin, ubiquitously distributed in cells and organisms, playing key roles in lipid and antioxidant metabolism. Loss or malfunction of peroxisomes causes more than 20 fatal inherited conditions. We have created a peroxisomal database (http://www.peroxisomeDB.org) that includes the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae, by gathering, updating and integrating the available genetic and functional information on peroxisomal genes. PeroxisomeDB is structured in interrelated sections 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases', that include hyperlinks to selected features of NCBI, ENSEMBL and UCSC databases. We have designed graphical depictions of the main peroxisomal metabolic routes and have included updated flow charts for diagnosis. Precomputed BLAST, PSI-BLAST, multiple sequence alignment (MUSCLE) and phylogenetic trees are provided to assist in direct multispecies comparison to study evolutionary conserved functions and pathways. Highlights of the PeroxisomeDB include new tools developed for facilitating (i) identification of novel peroxisomal proteins, by means of identifying proteins carrying peroxisome targeting signal (PTS) motifs, (ii) detection of peroxisomes in silico, particularly useful for screening the deluge of newly sequenced genomes. PeroxisomeDB should contribute to the systematic characterization of the peroxisomal proteome and facilitate system biology approaches on the organelle.

Animals↗