Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Genome annotation of Anopheles gambiae using mass spectrometry-derived data.

BACKGROUND: A large number of animal and plant genomes have been completely sequenced over the last decade and are now publicly available. Although genomes can be rapidly sequenced, identifying protein-coding genes still remains a problematic task. Availability of protein sequence data allows direct confirmation of protein-coding genes. Mass spectrometry has recently emerged as a powerful tool for proteomic studies. Protein identification using mass spectrometry is usually carried out by searching against databases of known proteins or transcripts. This approach generally does not allow identification of proteins that have not yet been predicted or whose transcripts have not been identified. RESULTS: We searched 3,967 mass spectra from 16 LC-MS/MS runs of Anopheles gambiae salivary gland homogenates against the Anopheles gambiae genome database. This allowed us to validate 23 known transcripts and 50 novel transcripts. In addition, a novel gene was identified on the basis of peptides that matched a genomic region where no gene was known and no transcript had been predicted. The amino termini of proteins encoded by two predicted transcripts were confirmed based on N-terminally acetylated peptides sequenced by tandem mass spectrometry. Finally, six sequence polymorphisms could be annotated based on experimentally obtained peptide sequences. CONCLUSION: The peptide sequences from this study were mapped onto the genomic sequence using the distributed annotation system available at Ensembl and can be visualized in the context of all other existing annotations. The strategy described in this paper can be used to correct and confirm genome annotations and permit discovery of novel proteins in a high-throughput manner by mass spectrometry.

Animals↗

Mayday--a microarray data analysis workbench.

UNLABELLED: Mayday is a workbench for visualization, analysis and storage of microarray data. It features a graphical user interface and supports the development and integration of existing and new analysis methods. Besides the infrastructural core functionality, Mayday offers a variety of plug-ins, such as various interactive viewers, a connection to the R statistical environment, a connection to SQL-based databases and different data mining methods, including WEKA-library based methods for classification and various clustering methods. In addition, so-called meta information objects are provided for annotation of the microarray data allowing integration of data from different sources, which is a feature that, for instance, is employed in the enhanced heatmap visualization. SUPPLEMENTARY INFORMATION: The software and more detailed information including screenshots and a user guide as well as test data can be found on the Mayday home page http://www.zbit.uni-tuebingen.de/pas/mayday. The core is published under the GPL (GNU Public License) and the associated plug-ins under the LGPL (Lesser GNU Public License).

Computer Graphics↗

sc-PDB: an annotated database of druggable binding sites from the Protein Data Bank.

The sc-PDB is a collection of 6 415 three-dimensional structures of binding sites found in the Protein Data Bank (PDB). Binding sites were extracted from all high-resolution crystal structures in which a complex between a protein cavity and a small-molecular-weight ligand could be identified. Importantly, ligands are considered from a pharmacological and not a structural point of view. Therefore, solvents, detergents, and most metal ions are not stored in the sc-PDB. Ligands are classified into four main categories: nucleotides (< 4-mer), peptides (< 9-mer), cofactors, and organic compounds. The corresponding binding site is formed by all protein residues (including amino acids, cofactors, and important metal ions) with at least one atom within 6.5 angstroms of any ligand atom. The database was carefully annotated by browsing several protein databases (PDB, UniProt, and GO) and storing, for every sc-PDB entry, the following features: protein name, function, source, domain and mutations, ligand name, and structure. The repository of ligands has also been archived by diversity analysis of molecular scaffolds, and several chemoinformatics descriptors were computed to better understand the chemical space covered by stored ligands. The sc-PDB may be used for several purposes: (i) screening a collection of binding sites for predicting the most likely target(s) of any ligand, (ii) analyzing the molecular similarity between different cavities, and (iii) deriving rules that describe the relationship between ligand pharmacophoric points and active-site properties. The database is periodically updated and accessible on the web at http://bioinfo-pharma.u-strasbg.fr/scPDB/.

Algorithms↗

Application of eVOC: controlled vocabularies for unifying gene expression data.

To provide standardised description of gene expression and cross platform querying of databases, we have developed eVOC (http://www.sanbi.ac.za/evoc/), consisting of four orthogonal ontologies which describe Anatomical System, Cell Type, Pathology and Developmental Stage. We have annotated 47 microarray expression data sets and all publicly available human cDNA and SAGE tag libraries. eVOC has been integrated with the public resource EnsMart, which provides linking of transcripts and libraries with expression terms and the human genome sequence (http://www.ensembl.org/Homo_sapiens/martview).

Databases, Genetic↗

Of genomes and proteomes.

The era of complete genome sequences has arrived and with it vast amounts of data which must be annotated, cross referenced, and placed within the regulatory networks which define the physiology of an organism. One eucaryotic and three procaryotic genomes have been completed and the data made available and another 50 sequences are expected to be completed by the end of the decade. One of the first steps in the new post genome era will be to decipher the functions of the huge numbers of new open reading frames. Various approaches to investigate what unknown genes do and how genes interact together within an organism are being undertaken, including (1) the simultaneous measurement of the expression levels of all genes in a cell and (2) the mapping and quantitation of all proteins expressed within a cell. The idea of systematically mapping and identifying the total protein complement of the genome (the "proteome") arose over 20 years ago when the separation of proteins from total cell extracts by two dimensional (2D) gel electrophoresis was developed. This review will focus on the use of 2D gel electrophoresis as the basis for constructing proteome maps and on the rapid advances in mass spectrometry which will allow the large-scale, automated identification of proteins which is necessary for the creation of such databases.

Animals↗

3D-GENOMICS: a database to compare structural and functional annotations of proteins between sequenced genomes.

The 3D-GENOMICS database (http://www.sbg.bio. ic.ac.uk/3dgenomics/) provides structural annotations for proteins from sequenced genomes. In August 2003 the database included data for 93 proteomes. The annotations stored in the database include homologous sequences from various sequence databases, domains from SCOP and Pfam, patterns from Prosite and other predicted sequence features such as transmembrane regions and coiled coils. In addition to annotations at the sequence level, several precomputed cross- proteome comparative analyses are available based on SCOP domain superfamily composition. Annotations are available to the user via a web interface to the database. Multiple points of entry are available so that a user is able to: (i) directly access annotations for a single protein sequence via keywords or accession codes, (ii) examine a sequence of interest chosen from a summary of annotations for a particular proteome, or (iii) access precomputed frequency-based cross-proteome comparative analyses.

Amino Acid Sequence↗

GARSA: genomic analysis resources for sequence annotation.

SUMMARY: Growth of genome data and analysis possibilities have brought new levels of difficulty for scientists to understand, integrate and deal with all this ever-increasing information. In this scenario, GARSA has been conceived aiming to facilitate the tasks of integrating, analyzing and presenting genomic information from several bioinformatics tools and genomic databases, in a flexible way. GARSA is a user-friendly web-based system designed to analyze genomic data in the context of a pipeline. EST and GGS data can be analyzed using the system since it accepts (1) chromatograms, (2) download of sequences from GenBank, (3) Fasta files stored locally or (4) a combination of all three. Quality evaluation of chromatograms, vector removing and clusterization are easily performed as part of the pipeline. A number of local and customizable Blast and CDD analyses can be performed as well as Interpro, complemented with phylogeny analyses. GARSA is being used for the analyses of Trypanosoma vivax (GSS and EST), Trypanosoma rangeli (GSS, EST and ORESTES), Bothrops jararaca (EST), Piaractus mesopotamicus (EST) and Lutzomyia longipalpis (EST). AVAILABILITY: The GARSA system is freely available under GPL license (http://www.biowebdb.org/garsa/). For download requests visit http://www.biowebdb.org/garsa/ or contact Dr Alberto Dávila.

Animals↗

ASAP, a systematic annotation package for community analysis of genomes.

ASAP (a systematic annotation package for community analysis of genomes) is a relational database and web interface developed to store, update and distribute genome sequence data and functional characterization (https://asap.ahabs.wisc.edu/annotation/php/ASAP1.htm). ASAP facilitates ongoing community annotation of genomes and tracking of information as genome projects move from preliminary data collection through post-sequencing functional analysis. The ASAP database includes multiple genome sequences at various stages of analysis, corresponding experimental data and access to collections of related genome resources. ASAP supports three levels of users: public viewers, annotators and curators. Public viewers can currently browse updated annotation information for Escherichia coli K-12 strain MG1655, genome-wide transcript profiles from more than 50 microarray experiments and an extensive collection of mutant strains and associated phenotypic data. Annotators worldwide are currently using ASAP to participate in a community annotation project for the Erwinia chrysanthemi strain 3937 genome. Curation of the E. chrysanthemi genome annotation as well as those of additional published enterobacterial genomes is underway and will be publicly accessible in the near future.

Databases, Genetic↗

The predictive power of the CluSTr database.

SUMMARY: The CluSTr database employs a fully automatic single-linkage hierarchical clustering method based on a similarity matrix. In order to compute the matrix, first all-against-all pair-wise comparisons between protein sequences are computed using the Smith-Waterman algorithm. The statistical significance of the similarity scores is then assessed using a Monte Carlo analysis, yielding Z-values, which are used to populate the matrix. This paper describes automated annotation experiments that quantify the predictive power and hence the biological relevance of the CluSTr data. The experiments utilized the UniProt data-mining framework to derive annotation predictions using combinations of InterPro and CluSTr. We show that this combination of data sources greatly increases the precision of predictions made by the data-mining framework, compared with the use of InterPro data alone. We conclude that the CluSTr approach to clustering proteins makes a valuable contribution to traditional protein classifications. AVAILABILITY: http://www.ebi.ac.uk/clustr/.

Algorithms↗

ASAP: the Alternative Splicing Annotation Project.

Recently, genomics analyses have demonstrated that alternative splicing is widespread in mammalian genomes (30-60% of genes reported to have multiple isoforms), and may be one of their most important mechanisms of functional regulation. However, by comparison with other genomics data such as genome annotation, SNPs, or gene expression, there exists relatively little database infrastructure for the study of alternative splicing. We have constructed an online database ASAP (the Alternative Splicing Annotation Project) for biologists to access and mine the enormous wealth of alternative splicing information coming from genomics and proteomics. ASAP is based on genome-wide analyses of alternative splicing in human (30 793 alternative splice relationships found) from detailed alignment of expressed sequences onto the genomic sequence. ASAP provides precise gene exon-intron structure, alternative splicing, tissue specificity of alternative splice forms, and protein isoform sequences resulting from alternative splicing. Moreover, it can help biologists design probe sequences for distinguishing specific mRNA isoforms. ASAP is intended to be a community resource for collaborative annotation of alternative splice forms, their regulation, and biological functions. The URL for ASAP is http://www.bioinformatics.ucla.edu/ASAP.

Alternative Splicing↗

Navigating the Brookhaven Protein Data Bank.

The Protein Data Bank maintained at Brookhaven National Laboratories has expanded to the point where even experienced users have difficulty understanding, exploring, and exploiting it. This paper describes a text file, an annotation of the Protein Data Bank, which helps users find information on related files and structures. The most recent version of this file includes information on homologous structures, including both sequence homology and structural homology. This file is in ASCII format and is available electronically. It is easy to search locally on any type of computer, using an editor or a pattern-matching program, such as grep.

Base Sequence↗

The Stanford Microarray Database: a user's guide.

The Stanford Microarray Database (SMD) is a DNA microarray research database that provides a large amount of data for public use. This chapter describes the use of the primary tools for searching, browsing, retrieving, and analyzing data available for SMD. With this introduction, researchers and students will be able to examine and analyze a large body of gene expression and other experiments. Additional tools for depositing, annotating, sharing, and analyzing data, available only to registered users, are also described. SMD is available for installation as a local database.

Cluster Analysis↗

Analysis of expressed sequence tags from Musa acuminata ssp. burmannicoides, var. Calcutta 4 (AA) leaves submitted to temperature stresses.

In order to discover genes expressed in leaves of Musa acuminata ssp. burmannicoides var. Calcutta 4 (AA), from plants submitted to temperature stress, we produced and characterized two full-length enriched cDNA libraries. Total RNA from plants subjected to temperatures ranging from 5 degrees C to 25 degrees C and from 25 degrees C to 45 degrees C was used to produce a COLD and a HOT cDNA library, respectively. We sequenced 1,440 clones from each library. Following quality analysis and vector trimming, we assembled 2,286 sequences from both libraries into 1,019 putative transcripts, consisting of 217 clusters and 802 singletons, which we denoted Musa acuminata assembled expressed sequence tagged (EST) sequences (MaAES). Of these MaAES, 22.87% showed no matches with existing sequences in public databases. A global analysis of the MaAES data set indicated that 10% of the sequenced cDNAs are present in both cDNA libraries, while 42% and 48% are present only in the COLD or in the HOT libraries, respectively. Annotation of the MaAES data set categorized them into 22 functional classes. Of the 2,286 high-quality sequences, 715 (31.28%) originated from full-length cDNA clones and resulted in a set of 149 genes.

Base Sequence↗

Proteome analysis of an aerobic hyperthermophilic crenarchaeon, Aeropyrum pernix K1.

We analyzed the proteome of a crenararchaeon, Aeropyrum pernix K1, by using the following four methods: (i) two-dimensional PAGE followed by MALDI-TOF MS, (ii) one-dimensional SDS-PAGE in combination with two-dimensional LC-MS/MS, (iii) multidimensional LC-MS/MS, and (iv) two-dimensional PAGE followed by amino-terminal amino acid sequencing. These methods were found to be complementary to each other, and biases in the data obtained in one method could largely be compensated by the data obtained in the other methods. Consequently a total of 704 proteins were successfully identified, 134 of which were unique to A. pernix K1, and 19 were not described previously in the genomic annotation. We found that the original annotation of the genomic data of this archaeon was not adequate in particular with respect to proteins of 10-20 kDa in size, many of which were described as hypothetical. Furthermore the amino-terminal amino acid sequence analysis indicated that surprisingly the translation of 52% of their genes starts with TTG in contrast to ATG (28%) and GTG (20%). Thus, A. pernix K1 is the first example of an organism in which TTG is the most predominant translational initiation codon.

Aerobiosis↗

GenBank.

The GenBank sequence database continues to expand its data coverage, quality control, annotation content and retrieval services. GenBank is comprised of DNA sequences submitted directly by authors as well as sequences from the other major public databases. An integrated retrieval system, known as Entrez, contains data from GenBank and from the major protein sequence and structural databases, as well as related MEDLINE abstracts. Users may access GenBank over the Internet through the World Wide Web and through special client-server programs for text and sequence similarity searching. FTP, CD-ROM and e-mail servers are alternate means of access.

Base Sequence↗

angaGEDUCI: Anopheles gambiae gene expression database with integrated comparative algorithms for identifying conserved DNA motifs in promoter sequences.

BACKGROUND: The completed sequence of the Anopheles gambiae genome has enabled genome-wide analyses of gene expression and regulation in this principal vector of human malaria. These investigations have created a demand for efficient methods of cataloguing and analyzing the large quantities of data that have been produced. The organization of genome-wide data into one unified database makes possible the efficient identification of spatial and temporal patterns of gene expression, and by pairing these findings with comparative algorithms, may offer a tool to gain insight into the molecular mechanisms that regulate these expression patterns. DESCRIPTION: We provide a publicly-accessible database and integrated data-mining tool, angaGEDUCI, that unifies 1) stage- and tissue-specific microarray analyses of gene expression in An. gambiae at different developmental stages and temporal separations following a bloodmeal, 2) functional gene annotation, 3) genomic sequence data, and 4) promoter sequence comparison algorithms. The database can be used to study genes expressed in particular stages, tissues, and patterns of interest, and to identify conserved promoter sequence motifs that may play a role in the regulation of such expression. The database is accessible from the address http://www.angaged.bio.uci.edu. CONCLUSION: By combining gene expression, function, and sequence data with integrated sequence comparison algorithms, angaGEDUCI streamlines spatial and temporal pattern-finding and produces a straightforward means of developing predictions and designing experiments to assess how gene expression may be controlled at the molecular level.

Algorithms↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl/), maintained at the European Bioinformatics Institute (EBI), incorporates, organizes and distributes nucleotide sequences from public sources. The database is a part of an international collaboration with DDBJ (Japan) and GenBank (USA). Data are exchanged between the collaborating databases on a daily basis to achieve optimal synchrony. The web-based tool, Webin, is the preferred system for individual submission of nucleotide sequences, including Third Party Annotation (TPA) and alignment data. Automatic submission procedures are used for submission of data from large-scale genome sequencing centres and from the European Patent Office. Database releases are produced quarterly. The latest data collection can be accessed via FTP, email and WWW interfaces. The EBI's Sequence Retrieval System (SRS) integrates and links the main nucleotide and protein databases as well as many other specialist molecular biology databases. For sequence similarity searching, a variety of tools (e.g. FASTA and BLAST) are available that allow external users to compare their own sequences against the data in the EMBL Nucleotide Sequence Database, the complete genomic component subsection of the database, the WGS data sets and other databases. All available resources can be accessed via the EBI home page at http://www.ebi.ac.uk.

Animals↗

Fugu ESTs: new resources for transcription analysis and genome annotation.

The draft Fugu rubripes genome was released in 2002, at which time relatively few cDNAs were available to aid in the annotation of genes. The data presented here describe the sequencing and analysis of 24,398 expressed sequence tags (ESTs) generated from 15 different adult and juvenile Fugu tissues, 74% of which matched protein database entries. Analysis of the EST data compared with the Fugu genome data predicts that approximately 10,116 gene tags have been generated, covering almost one-third of Fugu predicted genes. This represents a remarkable economy of effort. Comparison with the Washington University zebrafish EST assemblies indicates strong conservation within fish species, but significant differences remain. This potentially represents divergence of sequence in the 5' terminal exons and UTRs between these two fish species, although clearly, complete EST data sets are not available for either species. This project provides new Fugu resources, and the analysis adds significant weight to the argument that EST programs remain an essential resource for genome exploitation and annotation. This is particularly timely with the increasing availability of draft genome sequence from different organisms and the mounting emphasis on gene function and regulation.

Animals↗