Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

GermOnline, a cross-species community knowledgebase on germ cell differentiation.

GermOnline provides information and microarray expression data for genes involved in mitosis and meiosis, gamete formation and germ line development across species. The database has been developed, and is being curated and updated, by life scientists in cooperation with bioinformaticists. Information is contributed through an online form using free text, images and the controlled vocabulary developed by the GeneOntology Consortium. Authors provide up to three references in support of their contribution. The database is governed by an international board of scientists to ensure a standardized data format and the highest quality of GermOnline's information content. Release 2.0 provides exclusive access to microarray expression data from Saccharomyces cerevisiae and Rattus norvegicus, as well as curated information on approximately 700 genes from various organisms. The locus report pages include links to external databases that contain relevant annotation, microarray expression and proteome data. Conversely, the Saccharomyces Genome Database (SGD), S.cerevisiae GeneDB and Swiss-Prot link to the budding yeast section of GermOnline from their respective locus pages. GermOnline, a fully operational prototype subject-oriented knowledgebase designed for community annotation and array data visualization, is accessible at http://www.germonline.org. The target audience includes researchers who work on mitotic cell division, meiosis, gametogenesis, germ line development, human reproductive health and comparative genomics.

Animals↗

Development of a data management tool for investigating multivariate space and free will experiences in virtual reality.

Virtual reality (VR) has become mature enough to be successfully used in clinical applications such as exposure therapy, pain distraction, and neuropsychological assessment. However, we now need to go beyond the outcome data from this research and conduct the detailed scientific investigations required to better understand what factors influence why VR works (or doesn't) in these types of clinical applications. This knowledge is required to guide the development of VR applications in the key areas of education, training, and rehabilitation and to further evolve existing VR approaches. One of the primary assets obtained with the use of VR is the ability to simulate the complexity of real world environments, within which human performance can be tested and trained. But this asset comes with a price in terms of the capture, quantification and analysis of large, multivariate and concurrent data sources that reflect the naturalistic behavioral interaction that is afforded in a virtual world. As well, while achieving realism has been a main goal in making convincing VR environments, just what constitutes realism and how much is needed is still an open question situated firmly in the research domain. Just as in real "reality," such factors in virtual reality are complex and multivariate, and the understanding of this complexity presents exceptional challenges to the VR researcher. For certain research questions, good behavioral science often requires consistent delivery of stimuli within tightly controlled lab-based experimental conditions. However, for other important research questions we do not want to constrain naturalistic behavior and limit VR's ability to replicate real world conditions, simply because it is easier to study human performance with traditional lab-based methodologies. By doing so we may compromise the very qualities that comprise VR's unique capacity to mimic the experiences and challenges that exist in everyday life. What is really needed to address scientific questions that require natural exploration of a simulated environment are more usable and robust tools to instrument, organize, and visualize the complex data generated by measurements of participant behaviors within a virtual world. This paper briefly describes the rationale and methodology of an initial study in an ongoing research program that aims to investigate human performance within a virtual environment where unconstrained "free will" exploratory behavior is essential to research questions that involve the relationships between physiology, emotion, and memory. After a discussion of the research protocol and the types of data that were collected, we describe a novel tool that was borne from our need to more efficiently capture, manage, and explore the complex data that was generated in this research. An example of a research participant's annotated display from this data management and visualization tool is then presented. It is our view that this tool provides the capacity to better visualize and understand the complex data relationships that may arise in VR research that investigates naturalistic free will behavior and its impact on other human performance variables.

Activities of Daily Living↗

Semantic search among heterogeneous biological databases based on gene ontology.

Semantic search is a key issue in integration of heterogeneous biological databases. In this paper, we present a methodology for implementing semantic search in BioDW, an integrated biological data warehouse. Two tables are presented: the DB2GO table to correlate Gene Ontology (GO) annotated entries from BioDW data sources with GO, and the semantic similarity table to record similarity scores derived from any pair of GO terms. Based on the two tables, multifarious ways for semantic search are provided and the corresponding entries in heterogeneous biological databases in semantic terms can be expediently searched.

Database Management Systems↗

DWARF--a data warehouse system for analyzing protein families.

BACKGROUND: The emerging field of integrative bioinformatics provides the tools to organize and systematically analyze vast amounts of highly diverse biological data and thus allows to gain a novel understanding of complex biological systems. The data warehouse DWARF applies integrative bioinformatics approaches to the analysis of large protein families. DESCRIPTION: The data warehouse system DWARF integrates data on sequence, structure, and functional annotation for protein fold families. The underlying relational data model consists of three major sections representing entities related to the protein (biochemical function, source organism, classification to homologous families and superfamilies), the protein sequence (position-specific annotation, mutant information), and the protein structure (secondary structure information, superimposed tertiary structure). Tools for extracting, transforming and loading data from public available resources (ExPDB, GenBank, DSSP) are provided to populate the database. The data can be accessed by an interface for searching and browsing, and by analysis tools that operate on annotation, sequence, or structure. We applied DWARF to the family of alpha/beta-hydrolases to host the Lipase Engineering database. Release 2.3 contains 6138 sequences and 167 experimentally determined protein structures, which are assigned to 37 superfamilies 103 homologous families. CONCLUSION: DWARF has been designed for constructing databases of large structurally related protein families and for evaluating their sequence-structure-function relationships by a systematic analysis of sequence, structure and functional annotation. It has been applied to predict biochemical properties from sequence, and serves as a valuable tool for protein engineering.

Amino Acid Sequence↗

PipeOnline 2.0: automated EST processing and functional data sorting.

Expressed sequence tags (ESTs) are generated and deposited in the public domain, as redundant, unannotated, single-pass reactions, with virtually no biological content. PipeOnline automatically analyses and transforms large collections of raw DNA-sequence data from chromatograms or FASTA files by calling the quality of bases, screening and removing vector sequences, assembling and rewriting consensus sequences of redundant input files into a unigene EST data set and finally through translation, amino acid sequence similarity searches, annotation of public databases and functional data. PipeOnline generates an annotated database, retaining the processed unigene sequence, clone/file history, alignments with similar sequences, and proposed functional classification, if available. Functional annotation is automatic and based on a novel method that relies on homology of amino acid sequence multiplicity within GenBank records. Records are examined through a function ordered browser or keyword queries with automated export of results. PipeOnline offers customization for individual projects (MyPipeOnline), automated updating and alert service. PipeOnline is available at http://stress-genomics.org.

Automation↗

Organelle DB: a cross-species database of protein localization and function.

To efficiently utilize the growing body of available protein localization data, we have developed Organelle DB, a web-accessible database cataloging more than 25,000 proteins from nearly 60 organelles, subcellular structures and protein complexes in 154 organisms spanning the eukaryotic kingdom. Organelle DB is the first on-line resource devoted to the identification and presentation of eukaryotic proteins localized to organelles and subcellular structures. As such, Organelle DB is a strong resource of data from the human proteome as well as from the major model organisms Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster, Caenorhabditis elegans and Mus musculus. In particular, Organelle DB is a central repository of yeast data, incorporating results--and actual fluorescent imagesfrom ongoing large-scale studies of protein localization in S.cerevisiae. Each protein in Organelle DB is presented with its sequence and, as available, a detailed description of its function; functions were extracted from relevant model organism databases, and links to these databases are provided within Organelle DB. To facilitate data interoperability, we have annotated all protein localizations using vocabulary from the Gene Ontology consortium. We also welcome new data for inclusion in Organelle DB, which may be freely accessed at http://organelledb.lsi.umich.edu.

Animals↗

The tissue microarray data exchange specification: implementation by the Cooperative Prostate Cancer Tissue Resource.

BACKGROUND: Tissue Microarrays (TMAs) have emerged as a powerful tool for examining the distribution of marker molecules in hundreds of different tissues displayed on a single slide. TMAs have been used successfully to validate candidate molecules discovered in gene array experiments. Like gene expression studies, TMA experiments are data intensive, requiring substantial information to interpret, replicate or validate. Recently, an open access Tissue Microarray Data Exchange Specification has been released that allows TMA data to be organized in a self-describing XML document annotated with well-defined common data elements. While this specification provides sufficient information for the reproduction of the experiment by outside research groups, its initial description did not contain instructions or examples of actual implementations, and no implementation studies have been published. The purpose of this paper is to demonstrate how the TMA Data Exchange Specification is implemented in a prostate cancer TMA. RESULTS: The Cooperative Prostate Cancer Tissue Resource (CPCTR) is funded by the National Cancer Institute to provide researchers with samples of prostate cancer annotated with demographic and clinical data. The CPCTR now offers prostate cancer TMAs and has implemented a TMA database conforming to the new open access Tissue Microarray Data Exchange Specification. The bulk of the TMA database consists of clinical and demographic data elements for 299 patient samples. These data elements were extracted from an Excel database using a transformative Perl script. The Perl script and the TMA database are open access documents distributed with this manuscript. CONCLUSIONS: TMA databases conforming to the Tissue Microarray Data Exchange Specification can be merged with other TMA files, expanded through the addition of data elements, or linked to data contained in external biological databases. This article describes an open access implementation of the TMA Data Exchange Specification and provides detailed guidance to researchers who wish to use the Specification.

Confidentiality↗

A MATLAB toolbox for the analysis of articulatory data in the production of speech.

The goal of this paper is to present EMATOOLS, a set of scripts for displaying and annotating acoustic and articulatory data simultaneously in studies on speech production. These scripts were developed with the use of MATLAB, a multiplatform computing environment for numeric computation and visualization. The system is equipped with a mouse-driven graphical interface made up of a number of displays. This interface can be easily customized to speed up routine tasks. The scripts can also be used in a noninteractive way, as stand-alone MATLAB commands. Output data can be imported into any standard spreadsheet. EMATOOLS is freely available from www.lpl.univ-aix.fr/nguyen/ematools.html.

Computer Graphics↗

The International Gene Trap Consortium Website: a portal to all publicly available gene trap cell lines in mouse.

Gene trapping is a method of generating murine embryonic stem (ES) cell lines containing insertional mutations in known and novel genes. A number of international groups have used this approach to create sizeable public cell line repositories available to the scientific community for the generation of mutant mouse strains. The major gene trapping groups worldwide have recently joined together to centralize access to all publicly available gene trap lines by developing a user-oriented Website for the International Gene Trap Consortium (IGTC). This collaboration provides an impressive public informatics resource comprising approximately 45 000 well-characterized ES cell lines which currently represent approximately 40% of known mouse genes, all freely available for the creation of knockout mice on a non-collaborative basis. To standardize annotation and provide high confidence data for gene trap lines, a rigorous identification and annotation pipeline has been developed combining genomic localization and transcript alignment of gene trap sequence tags to identify trapped loci. This information is stored in a new bioinformatics database accessible through the IGTC Website interface. The IGTC Website (www.genetrap.org) allows users to browse and search the database for trapped genes, BLAST sequences against gene trap sequence tags, and view trapped genes within biological pathways. In addition, IGTC data have been integrated into major genome browsers and bioinformatics sites to provide users with outside portals for viewing this data. The development of the IGTC Website marks a major advance by providing the research community with the data and tools necessary to effectively use public gene trap resources for the large-scale characterization of mammalian gene function.

Animals↗

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics↗

G-InforBIO: integrated system for microbial genomics.

BACKGROUND: Genome databases contain diverse kinds of information, including gene annotations and nucleotide and amino acid sequences. It is not easy to integrate such information for genomic study. There are few tools for integrated analyses of genomic data, therefore, we developed software that enables users to handle, manipulate, and analyze genome data with a variety of sequence analysis programs. RESULTS: The G-InforBIO system is a novel tool for genome data management and sequence analysis. The system can import genome data encoded as eXtensible Markup Language documents as formatted text documents, including annotations and sequences, from DNA Data Bank of Japan and GenBank encoded as flat files. The genome database is constructed automatically after importing, and the database can be exported as documents formatted with eXtensible Markup Language or tab-deliminated text. Users can retrieve data from the database by keyword searches, edit annotation data of genes, and process data with G-InforBIO. In addition, information in the G-InforBIO database can be analyzed seamlessly with nine different software programs, including programs for clustering and homology analyses. CONCLUSION: The G-InforBIO system simplifies genome analyses by integrating several available software programs to allow efficient handling and manipulation of genome data. G-InforBIO is freely available from the download site.

Algorithms↗

Automated methods of predicting the function of biological sequences using GO and BLAST.

BACKGROUND: With the exponential increase in genomic sequence data there is a need to develop automated approaches to deducing the biological functions of novel sequences with high accuracy. Our aim is to demonstrate how accuracy benchmarking can be used in a decision-making process evaluating competing designs of biological function predictors. We utilise the Gene Ontology, GO, a directed acyclic graph of functional terms, to annotate sequences with functional information describing their biological context. Initially we examine the effect on accuracy scores of increasing the allowed distance between predicted and a test set of curator assigned terms. Next we evaluate several annotator methods using accuracy benchmarking. Given an unannotated sequence we use the Basic Local Alignment Search Tool, BLAST, to find similar sequences that have already been assigned GO terms by curators. A number of methods were developed that utilise terms associated with the best five matching sequences. These methods were compared against a benchmark method of simply using terms associated with the best BLAST-matched sequence (best BLAST approach). RESULTS: The precision and recall of estimates increases rapidly as the amount of distance permitted between a predicted term and a correct term assignment increases. Accuracy benchmarking allows a comparison of annotation methods. A covering graph approach performs poorly, except where the term assignment rate is high. A term distance concordance approach has a similar accuracy to the best BLAST approach, demonstrating lower precision but higher recall. However, a discriminant function method has higher precision and recall than the best BLAST approach and other methods shown here. CONCLUSION: Allowing term predictions to be counted correct if closely related to a correct term decreases the reliability of the accuracy score. As such we recommend using accuracy measures that require exact matching of predicted terms with curator assigned terms. Furthermore, we conclude that competing designs of BLAST-based GO term annotators can be effectively compared using an accuracy benchmarking approach. The most accurate annotation method was developed using data mining techniques. As such we recommend that designers of term annotators utilise accuracy benchmarking and data mining to ensure newly developed annotators are of high quality.

Benchmarking↗

ArrayExpress--a public repository for microarray gene expression data at the EBI.

ArrayExpress is a public repository for microarray data that supports the MIAME (Minimum Information About a Microarray Experiment) requirements and stores well-annotated raw and normalized data. As of November 2004, ArrayExpress contains data from approximately 12,000 hybridizations covering 35 species. Data can be submitted online or directly from local databases or LIMS in a standard format, and password-protected access to prepublication data is provided for reviewers and authors. The data can be retrieved by accession number or queried by various parameters such as species, author and array platform. A facility to query experiments by gene and sample properties is provided for a growing subset of curated data that is loaded in to the ArrayExpress data warehouse. Data can be visualized and analysed using Expression Profiler, the integrated data analysis tool. ArrayExpress is available at http://www.ebi.ac.uk/arrayexpress.

Animals↗

Bioinformatics support for high-throughput proteomics.

In the "post-genome" era, mass spectrometry (MS) has become an important method for the analysis of proteome data. The rapid advancement of this technique in combination with other methods used in proteomics results in an increasing number of high-throughput projects. This leads to an increasing amount of data that needs to be archived and analyzed. To cope with the need for automated data conversion, storage, and analysis in the field of proteomics, the open source system ProDB was developed. The system handles data conversion from different mass spectrometer software, automates data analysis, and allows the annotation of MS spectra (e.g. assign gene names, store data on protein modifications). The system is based on an extensible relational database to store the mass spectra together with the experimental setup. It also provides a graphical user interface (GUI) for managing the experimental steps which led to the MS data. Furthermore, it allows the integration of genome and proteome data. Data from an ongoing experiment was used to compare manual and automated analysis. First tests showed that the automation resulted in a significant saving of time. Furthermore, the quality and interpretability of the results was improved in all cases.

Algorithms↗

The TRIPLES database: a community resource for yeast molecular biology.

TRIPLES is a web-accessible database of TRansposon-Insertion Phenotypes, Localization and Expression in Saccharomyces cerevisiae-a relational database housing nearly half a million data points generated from an ongoing study using large-scale transposon mutagenesis to characterize gene function in yeast. At present, TRIPLES contains three principal data sets (i.e. phenotypic data, protein localization data and expression data) for over 3500 annotated yeast genes as well as several hundred non-annotated open reading frames. In addition, the TRIPLES web site provides online order forms linked to each data set so that users may request any strain or reagent generated from this project free of charge. In response to user requests, the TRIPLES web site has undergone several recent modifications. Our localization data have been supplemented with approximately 500 fluorescent micrographs depicting actual staining patterns observed upon indirect immunofluorescence analysis of indicated epitope-tagged proteins. These localization data, as well as all other data sets within TRIPLES, are now available in full as tab-delimited text. To accommodate increased reagent requests, all orders are now cataloged in a separate database, and users are notified immediately of order receipt and shipment. Also, TRIPLES is one of five sites incorporated into the new functional analysis tool Function Junction provided by the Saccharomyces Genome Database. TRIPLES may be accessed from the Yale Genome Analysis Center (YGAC) homepage at http://ygac.med.yale.edu.

Computer Graphics↗

Automatic discovery of cross-family sequence features associated with protein function.

BACKGROUND: Methods for predicting protein function directly from amino acid sequences are useful tools in the study of uncharacterized protein families and in comparative genomics. Until now, this problem has been approached using machine learning techniques that attempt to predict membership, or otherwise, to predefined functional categories or subcellular locations. A potential drawback of this approach is that the human-designated functional classes may not accurately reflect the underlying biology, and consequently important sequence-to-function relationships may be missed. RESULTS: We show that a self-supervised data mining approach is able to find relationships between sequence features and functional annotations. No preconceived ideas about functional categories are required, and the training data is simply a set of protein sequences and their UniProt/Swiss-Prot annotations. The main technical aspect of the approach is the co-evolution of amino acid-based regular expressions and keyword-based logical expressions with genetic programming. Our experiments on a strictly non-redundant set of eukaryotic proteins reveal that the strongest and most easily detected sequence-to-function relationships are concerned with targeting to various cellular compartments, which is an area already well studied both experimentally and computationally. Of more interest are a number of broad functional roles which can also be correlated with sequence features. These include inhibition, biosynthesis, transcription and defence against bacteria. Despite substantial overlaps between these functions and their corresponding cellular compartments, we find clear differences in the sequence motifs used to predict some of these functions. For example, the presence of polyglutamine repeats appears to be linked more strongly to the "transcription" function than to the general "nuclear" function/location. CONCLUSION: We have developed a novel and useful approach for knowledge discovery in annotated sequence data. The technique is able to identify functionally important sequence features and does not require expert knowledge. By viewing protein function from a sequence perspective, the approach is also suitable for discovering unexpected links between biological processes, such as the recently discovered role of ubiquitination in transcription.

Algorithms↗

IMGT/LIGM-DB: a systematized approach for ImMunoGeneTics database coherence and data distribution improvement.

IMGT, the international ImMunoGeneTics database (http:(/)/imgt.cnusc.fr:8104), created by Marie-Paule Lefranc, Montpellier, France, is an integrated database specializing in antigen receptors and MHC of all vertebrate species. IMGT includes LIGM-DB, developed for Immunoglobulins and T-cell-receptors. LIGM-DB distributes high quality data with an important increment value added by the LIGM expert annotations. LIGM-DB accurate immunogenetics data is based on the standardization of biological knowledge related to keywords, annotation labels and gene identification. The management of such data resulting from biological research requires an high flexible implementation to quickly reflect up-to-date results, and to integrate new knowledge. We developed a systematized approach and defined LIGM-DB systems which manage and realize the major tasks for the database survey. In this paper, we will focus on the coherence system, which became absolutely crucial to maintain data quality as the database is growing up and as the biological knowledge continues to improve, and on the distribution system which makes LIGM-DB data easy to access, download and reuse. Efforts have been done to improve the data distribution procedures and adapt them to the current bioinformatics needs. Recently, we have developed an API which allows Java programmers to remotely access and integrate LIGM-DB data in other computer environments.

Animals↗

Identification of candidate Purkinje cell-specific markers by gene expression profiling in wild-type and pcd(3J) mice.

The identification of mRNAs that have restricted expression patterns in the brain represents powerful tools with which to characterize and manipulate the nervous system. Here, we describe a strategy using microarray technology (Affymetrix Mouse Genome 430 2.0 Arrays) to identify mRNA transcripts that are candidate markers of cerebellar Purkinje neurons. Initially, gene expression profiles were compared between cerebella of 4-month-old Purkinje cell degeneration (pcd(3J)) mice, in which most Purkinje cells had already degenerated and wild-type littermates with a normal complement of Purkinje neurons. Of 14,563 probe sets expressed in wild-type cerebellum, 797 showed a significant (p<0.0001) reduction in pcd(3J) mice. These probes could represent transcripts with varying levels of specificity for Purkinje cells as well as transcripts in other cell types that decline as a secondary consequence of Purkinje cell loss. Ranking of the probe signals revealed that well-known Purkinje cell-specific transcripts such as calbindin and L7/pcp2 clustered in a group that was <33% of wild-type levels. Therefore, to identify potentially new Purkinje cell-specific transcripts that cluster with the known markers, more stringent selection criteria were applied (<33% of wild-type signal and p<0.0001). With these criteria, 55 independent transcripts were identified of which 33 were annotated genes and 22 were ESTs and RIKEN cDNAs. A literature search revealed that 25 of the 33 annotated genes were expressed in Purkinje cells, with no data being available on the other 8. Thus, the additional 8 annotated and 22 un-annotated genes are clustered with many genes expressed in Purkinje cells making them candidate markers. To confirm the microarray data, eight representative annotated genes were selected including five reported to be in Purkinje neurons and three for which no data was available. Semi-quantitative RT-PCR demonstrated reduced expression of all eight transcripts in cerebella from pcd(3J) mice. The promoters of genes expressed selectively in subsets of neurons can be used to direct heterologous gene expression in transgenic mice and the more restricted the expression pattern the greater their utility. Therefore, microarray analysis was used to assess expression levels of all 55 transcripts in cerebral cortex, striatum, substantia nigra and ventral tegmental area. This permitted the identification of a set of genes whose promoters might have utility for selectively targeting gene expression to cerebellar Purkinje cells.

Animals↗