Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data Storage And Retrieval”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

EPPS: mining the COG database by an extended phylogenetic patterns search.

SUMMARY: EPPS runs under Microsoft Windows. It is an extended version of the phylogenetic patterns search (PPS). The output condition of PPS is the exact match of a user defined phylogenetic pattern with the pattern represented by the respective cluster of orthologous groups (COG). In contrast, the software described here is less restrictive. The user may define the accuracy of the search by the number of genomes that are allowed not to match the predefined phylogenetic pattern. Thus, EPPS has the advantage to detect COGs even if organisms defined to be included are not or organisms defined to be excluded are present in the output COGs.

Amino Acid Sequence↗

SCide: identification of stabilization centers in proteins.

SUMMARY: SCide is a program to identify stabilization centers from known protein structures. These are residues involved in cooperative long-range contacts, which can be formed between various regions of a single polypeptide chain, or they can belong to different peptides or polypeptides in a complex. The server takes a PDB file as an input, and the result is presented in graphical or text format. AVAILABILITY: SCide is available on the web at http://www.enzim.hu/scide. The source code can be obtained from the authors on request.

Amino Acid Sequence↗

Increased coverage obtained by combination of methods for protein sequence database searching.

MOTIVATION: Sequence alignment methods that compare two sequences (pairwise methods) are important tools for the detection of biological sequence relationships. In genome annotation, multiple methods are often run and agreement between methods taken as confirmation. In this paper, we assess the advantages of combining search methods by comparing seven pairwise alignment methods, including three local dynamic programming algorithms (PRSS, SSEARCH and SCANPS), two global dynamic programming algorithms (GSRCH and AMPS) and two heuristic approximations (BLAST and FASTA), individually and by pairwise intersection and union of their result lists at equal p-value cut-offs. RESULTS: When applied singly, the dynamic programming methods SCANPS and SSEARCH gave significantly better coverage (p=0.01) compared to AMPS, GSRCH, PRSS, BLAST and FASTA. Results ranked by BLAST p-values gave significantly better coverage compared to ranking by BLAST e-values. Of 56 combinations of eight methods considered, 19 gave significant increases in coverage at low error compared to the parent methods at an equal p-value cutoff. The union of results by BLAST (p-value) and FASTA at an equal p-value cutoff gave significantly better coverage than either method individually. The best overall performance was obtained from the intersection of the results from SSEARCH and the GSRCH62 global alignment method. At an error level of five false positives, this combination found 444 true positives, a significant 12.4% increase over SSEARCH applied alone.

Algorithms↗

Serial BLAST searching.

MOTIVATION: The translating BLAST algorithms are powerful tools for finding protein-coding genes because they identify amino acid similarities in nucleotide sequences. Unfortunately, these kinds of searches are computationally intensive and often represent bottlenecks in sequence analysis pipelines. Tuning parameters for speed can make the searches much faster, but one risks losing low-scoring alignments. However, high scoring alignments are relatively resistant to such changes in parameters, and this fact makes it possible to use a serial strategy where a fast, insensitive search is used to pre-screen a database for similar sequences, and a slow, sensitive search is used to produce the sequence alignments. RESULTS: Serial BLAST searches improve both the speed and sensitivity.

Algorithms↗

WILMA-automated annotation of protein sequences.

Large-scale annotation of sets of proteins is a frequently occurring task in association with genome sequencing projects. Here, we present an automated platform for the functional annotation of large sets of protein sequences. Various bioinformatics tools are used to achieve a comprehensive description of protein sequences and to link these results to standard Gene Ontology descriptors for molecular function, biological processes and cellular components. Access to the annotation is provided via a web-interface and database queries. These interfaces allow to formulate proteome wide queries as well as the investigation of details of individual results. WILMA annotations of the proteomes of Homo sapiens, Mus musculus, Arabidopsis thaliana and Caenorhabditis elegans are accessible at http://www.came.sbg.ac.at/wilma/

Amino Acid Sequence↗

GENIUS II: a high-throughput database system for linking ORFs in complete genomes to known protein three-dimensional structures.

GENIUS II is an automated database system in which open reading frames (ORFs) in complete genomes are assigned to known protein three-dimensional (3D) structures. The system uses the multiple intermediate sequence search method in which query and target sequences are linked by intermediate sequences gathered by PSI-BLAST search. By applying the system to 129 complete genomes, 43.8% on average of the ORFs in the genomes were assigned to known 3D structures and the results are available for free at GENIUS II web site.

Algorithms↗

DisProt: a database of protein disorder.

UNLABELLED: The Database of Protein Disorder (DisProt) is a curated database that provides structure and function information about proteins that lack a fixed three-dimensional (3D) structure under putatively native conditions, either in their entirety or in part. Starting from the central premise that intrinsic disorder is an important structural class of protein and in order to meet the increasing interest thereof, DisProt is aimed at becoming a central repository of disorder-related information. For each disordered protein, the database includes the name of the protein, various aliases, accession codes, amino acid sequence, location of the disordered region(s), and methods used for structural (disorder) characterization. If applicable, most entries also list the biological function(s) of each disordered region, how each region of disorder is used for function, as well as provide links to PubMed abstracts and major protein databases. AVAILABILITY: www.disprot.org

Amino Acid Sequence↗

The carbohydrate sequence markup language (CabosML): an XML description of carbohydrate structures.

UNLABELLED: Bioinformatics resources for glycomics are very poor as compared with those for genomics and proteomics. The complexity of carbohydrate sequences makes it difficult to define a common language to represent them, and the development of bioinformatics tools for glycomics has not progressed. In this study, we developed a carbohydrate sequence markup language (CabosML), an XML description of carbohydrate structures. AVAILABILITY: The language definition (XML Schema) and an experimental database of carbohydrate structures using an XML database management system are available at http://www.phoenix.hydra.mki.co.jp/CabosDemo.html CONTACT: kikuchi@hydra.mki.co.jp.

Carbohydrate Sequence↗

MagicMatch--cross-referencing sequence identifiers across databases.

MOTIVATION: At present, mapping of sequence identifiers across databases is a daunting, time-consuming and computationally expensive process, usually achieved by sequence similarity searches with strict threshold values. SUMMARY: We present a rapid and efficient method to map sequence identifiers across databases. The method uses the MD5 checksum algorithm for message integrity to generate sequence fingerprints and uses these fingerprints as hash strings to map sequences across databases. The program, called MagicMatch, is able to cross-link any of the major sequence databases within a few seconds on a modest desktop computer.

Algorithms↗

METIS: multiple extraction techniques for informative sentences.

SUMMARY: METIS is a web-based integrated annotation tool. From single query sequences, the PRECIS component allows users to generate structured protein family reports from sets of related Swiss-Prot entries. These reports may then be augmented with pertinent sentences extracted from online biomedical literature via support vector machine and rule-based sentence classification systems. AVAILABILITY: http://umber.sbs.man.ac.uk/dbbrowser/metis/

Algorithms↗

Mapping PDB chains to UniProtKB entries.

MOTIVATION: UniProtKB/SwissProt is the main resource for detailed annotations of protein sequences. This database provides a jumping-off point to many other resources through the links it provides. Among others, these include other primary databases, secondary databases, the Gene Ontology and OMIM. While a large number of links are provided to Protein Data Bank (PDB) files, obtaining a regularly updated mapping between UniProtKB entries and PDB entries at the chain or residue level is not straightforward. In particular, there is no regularly updated resource which allows a UniProtKB/SwissProt entry to be identified for a given residue of a PDB file. RESULTS: We have created a completely automatically maintained database which maps PDB residues to residues in UniProtKB/SwissProt and UniProtKB/trEMBL entries. The protocol uses links from PDB to UniProtKB, from UniProtKB to PDB and a brute-force sequence scan to resolve PDB chains for which no annotated link is available. Finally the sequences from PDB and UniProtKB are aligned to obtain a residue-level mapping. AVAILABILITY: The resource may be queried interactively or downloaded from http://www.bioinf.org.uk/pdbsws/.

Amino Acid Sequence↗

ZooDDD: a cross-species database for digital differential display analysis.

UNLABELLED: In this article, we combined EST information from the UniGene database and orthologous relationships from the Ensembl database to construct a ZooDDD database. The primary function of ZooDDD is to mine evolutionary conserved, highly expressed, tissue-specific orthologues in model animals. The candidate genes of interest derived from the ZooDDD database will provide biologists with a good step for comparing the expression, functions and evolution of animal genomes. AVAILABILITY: http://bio301.iis.sinica.edu.tw/~ZooDDDNew/main.php.

Base Sequence↗

OSIRIS: a tool for retrieving literature about sequence variants.

UNLABELLED: Sequence variants, in particular single nucleotide polymorphisms (SNPs), are key elements for the identification of genes associated with complex diseases and with particular drug responses. The search for literature about sequence variation is hampered by the large number of allelic variants reported for many genes and by the variability in both gene and sequence variants nomenclatures. We describe OSIRIS, a search tool that integrates different sources of information with the aim to retrieve literature about sequence variation of a gene. In addition, it provides a method to link a dbSNP entry with the articles referring to it. AVAILABILITY: OSIRIS is available for public use at http://ibi.imim.es/

Abstracting and Indexing↗

ARFA: a program for annotating bacterial release factor genes, including prediction of programmed ribosomal frameshifting.

UNLABELLED: Correct annotation of genes encoding release factors in bacterial genomes is often complicated by utilization of +1 programmed ribosomal frameshifting during synthesis of release factor 2, RF2. In the absence of robust computational approaches for predicting ribosomal frameshifting, the success of proper annotation depends on annotators' familiarity with this phenomenon. Here we describe a novel computer tool that allows automatic discrimination of genes encoding class-I bacterial release factors, RF1, RF2 and RFH. Most usefully, this program identifies and automatically annotates +1 frameshifting in RF2 encoding genes. Comparison of ARFA performance with existing annotations of bacterial genomes revealed that only 20% of RF2 genes utilizing ribosomal frameshifting during their expression are annotated correctly. AVAILABILITY: The PHP based web interface of ARFA and the source code are located at http://recode.genetics.utah.edu/arfa

Base Sequence↗

Mining frequent stem patterns from unaligned RNA sequences.

MOTIVATION: In detection of non-coding RNAs, it is often necessary to identify the secondary structure motifs from a set of putative RNA sequences. Most of the existing algorithms aim to provide the best motif or few good motifs, but biologists often need to inspect all the possible motifs thoroughly. RESULTS: Our method RNAmine employs a graph theoretic representation of RNA sequences and detects all the possible motifs exhaustively using a graph mining algorithm. The motif detection problem boils down to finding frequently appearing patterns in a set of directed and labeled graphs. In the tasks of common secondary structure prediction and local motif detection from long sequences, our method performed favorably both in accuracy and in efficiency with the state-of-the-art methods such as CMFinder. AVAILABILITY: The software is available upon request.

Algorithms↗

Identification of new eukaryotic tRNA genes in genomic DNA databases by a multistep weight matrix analysis of transcriptional control regions.

A linear method for the search of eukaryotic nuclear tRNA genes in DNA databases is described. Based on a modified version of the general weight matrix procedure, our algorithm relies on the recognition of two intragenic control regions known as A and B boxes, a transcription termination signal, and on the evaluation of the spacing between these elements. The scanning of the eukaryotic nuclear DNA database using this search algorithm correctly identified 933 of the 940 known tRNA genes (0.74% of false negatives). Thirty new potential tRNA genes were identified, and the transcriptional activity of two of them was directly verified by in vitro transcription. The total false positive rate of the algorithm was 0.014%. Structurally unusual tRNA genes, like those coding for selenocysteine tRNAs, could also be recognized using a set of rules concerning their specific properties, and one human gene coding for such tRNA was identified. Some of the newly identified tRNA genes were found in rather uncommon genomic positions: 2 in centromeric regions and 3 within introns. Furthermore, the presence of extragenically located B boxes in tRNA genes from various organisms could be detected through a specific subroutine of the standard search program.

Algorithms↗

The Blocks database--a system for protein classification.

The Blocks Database contains multiple alignments of conserved regions in protein families. The database can be searched by e-mail and World Wide Web(WWW) servers (http://blocks.fhcrc.org/help) to classify protein and nucleotide sequences.

Amino Acid Sequence↗