Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

The Mouse Genome Database (MGD): integrating biology with the genome.

The Mouse Genome Database (MGD) is one component of the Mouse Genome Informatics (MGI) system (http://www.informatics.jax.org), a community database resource for the laboratory mouse. MGD strives to provide a comprehensive knowledgebase about the mouse with experiments and data annotated from both literature and online sources. MGD curates and presents consensus and experimental data representations of genetic, genotype (sequence) and phenotype information including highly detailed reports about genes and gene products. Primary foci of integration are through representations of relationships between genes, sequences and phenotypes. MGD collaborates with other bioinformatics groups to curate a definitive set of information about the laboratory mouse and to build and implement the data and semantic standards that are essential for comparative genome analysis. Recent developments in MGD discussed here include an extensive integration of the mouse sequence data and substantial revisions in the presentation, query and visualization of sequence data.

Animals↗

DynGO: a tool for visualizing and mining of Gene Ontology and its associations.

BACKGROUND: A large volume of data and information about genes and gene products has been stored in various molecular biology databases. A major challenge for knowledge discovery using these databases is to identify related genes and gene products in disparate databases. The development of Gene Ontology (GO) as a common vocabulary for annotation allows integrated queries across multiple databases and identification of semantically related genes and gene products (i.e., genes and gene products that have similar GO annotations). Meanwhile, dozens of tools have been developed for browsing, mining or editing GO terms, their hierarchical relationships, or their "associated" genes and gene products (i.e., genes and gene products annotated with GO terms). Tools that allow users to directly search and inspect relations among all GO terms and their associated genes and gene products from multiple databases are needed. RESULTS: We present a standalone package called DynGO, which provides several advanced functionalities in addition to the standard browsing capability of the official GO browsing tool (AmiGO). DynGO allows users to conduct batch retrieval of GO annotations for a list of genes and gene products, and semantic retrieval of genes and gene products sharing similar GO annotations. The result are shown in an association tree organized according to GO hierarchies and supported with many dynamic display options such as sorting tree nodes or changing orientation of the tree. For GO curators and frequent GO users, DynGO provides fast and convenient access to GO annotation data. DynGO is generally applicable to any data set where the records are annotated with GO terms, as illustrated by two examples. CONCLUSION: We have presented a standalone package DynGO that provides functionalities to search and browse GO and its association databases as well as several additional functions such as batch retrieval and semantic retrieval. The complete documentation and software are freely available for download from the website http://biocreative.ifsm.umbc.edu/dyngo.

Computer Graphics↗

PathFinder: reconstruction and dynamic visualization of metabolic pathways.

MOTIVATION: Beyond methods for a gene-wise annotation and analysis of sequenced genomes new automated methods for functional analysis on a higher level are needed. The identification of realized metabolic pathways provides valuable information on gene expression and regulation. Detection of incomplete pathways helps to improve a constantly evolving genome annotation or discover alternative biochemical pathways. To utilize automated genome analysis on the level of metabolic pathways new methods for the dynamic representation and visualization of pathways are needed. RESULTS: PathFinder is a tool for the dynamic visualization of metabolic pathways based on annotation data. Pathways are represented as directed acyclic graphs, graph layout algorithms accomplish the dynamic drawing and visualization of the metabolic maps. A more detailed analysis of the input data on the level of biochemical pathways helps to identify genes and detect improper parts of annotations. As an Relational Database Management System (RDBMS) based internet application PathFinder reads a list of EC-numbers or a given annotation in EMBL- or Genbank-format and dynamically generates pathway graphs.

Bacillus subtilis↗

GeneCruiser: a web service for the annotation of microarray data.

SUMMARY: GeneCruiser is a web service allowing users to annotate their genomic data by mapping microarray feature identifiers to gene identifiers from databases, such as UniGene, while providing links to web resources, such as the UCSC Genome Browser. It relies on a regularly updated database that retrieves and indexes the mappings between microarray probes and genomic databases. Genes are identified using the Life Sciences Identifier standard. AVAILABILITY: GeneCruiser is freely available in the following forms: Web service and Web application, http://www.genecruiser.org; GenePattern, GeneCruiser access has been integrated into our microarray analysis platform, GenePattern. http://www.genepattern.org.

Animals↗

Managing research data with self-documenting files.

Processing biomedical research data is frequently complex due to the evolutionary nature of experiments and the requisite modification of analysis software. For the past several years we have been evolving a set of software tools designed to improve our ability to respond to evolving experimental designs. These tools allow the investigator to easily manipulate the research data and specify desired data transformations at run time. Coupling of analysis software to research data files is dependent on data files that are commented in a manner similar to that used in programming languages. The resulting annotated data files are self-documenting, and their use facilitates visual interpretation of displayed data as well as automatic processing of subsets of data. Here we present a formal description of a self-documenting file and describe several software tools that facilitate processing of biomedical research data.

Electronic Data Processing↗

The UCSC Table Browser data retrieval tool.

The University of California Santa Cruz (UCSC) Table Browser (http://genome.ucsc.edu/cgi-bin/hgText) provides text-based access to a large collection of genome assemblies and annotation data stored in the Genome Browser Database. A flexible alternative to the graphical-based Genome Browser, this tool offers an enhanced level of query support that includes restrictions based on field values, free-form SQL queries and combined queries on multiple tables. Output can be filtered to restrict the fields and lines returned, and may be organized into one of several formats, including a simple tab- delimited file that can be loaded into a spreadsheet or database as well as advanced formats that may be uploaded into the Genome Browser as custom annotation tracks. The Table Browser User's Guide located on the UCSC website provides instructions and detailed examples for constructing queries and configuring output.

Animals↗

CERTOMICS: trusted single-cell multiomics pipeline for high-resolution profiling of adoptive cellular immunotherapies.

SUMMARY: Adoptive cellular immunontherapies, such as chimeric antigen receptor (CAR) T cell therapy, have transformed cancer treatment, yet challenges such as resistance, relapse, and high costs limit their efficacy and accessibility. A comprehensive understanding of cellular heterogeneity and molecular profiles is essential to improve these therapies. Advanced single-cell multiomics technologies have the power to analyze the complex interactions between CAR-engineered cells, immune cells, and tumor cells. However, standardized single-cell multiomics computational pipelines specifically tailored to CAR-engineered cell products are lacking. Due to the synthetic nature of CAR transgenes, additional steps for reliable identification and characterization of CAR-positive cells are required but not included in existing data-processing workflows. To address this, we present CERTOMICS, a Nextflow-based, CAR-aware pipeline offering enhanced CERTainty in immunophenotyping and data interpretation, tailored for single-cell multiOMICSprofiling of adoptive cellular immunotherapies. The pipeline standardizes processing 10x Genomics single-cell multiomics data and integrates CAR-specific identification and quality control. Additionally, a curated repository of CAR construct sequences and annotation data is provided, serving as an extensible resource to support the analysis and development of CAR T cell therapies. AVAILABILITY AND IMPLEMENTATION: Detailed documentation of this pipeline, along with a resource on latest FDA-approved CAR therapies is available on our website: https://fraunhofer-izi.github.io/Living-Drugs-Wiki/. The data underlying this article are available on GitHub at https://github.com/fraunhofer-izi/CERTOMICS. The code is also published on Zenodo at https://doi.org/10.5281/zenodo.18709693.

Multiomics↗

Definition and detection of alarms in critical care.

Critical care medicine has developed enormously in complexity and even more so in cost over the past twenty years. There has been evidence of remarkable progress in improved outcomes from some conditions, particularly when severely ill patients are treated in well equipped and well managed intensive care units (ICU) which have clear directorship and comprehensive management guidelines and protocols (Zimmermann et al., Crit Care Med 1993; 21:1443-1451). Nevertheless, for some conditions such as severe acute respiratory failure and multiple organ failure, there is considerable debate as to whether there has been any improvement at all (Lee et al., Thorax 1994; 49:596-597. Artigas et al., Adult respiratory distress syndrome, Churchill Livingstone, Edinbugh, London, Madrid, Melbourne, New York, Tokyo, pp. 509-525). Developments in signal processing and monitoring and recording technology have resulted in a vast increase in the quantity of data that is available to clinicians trying to manage critically ill patients (Price, Bailliere's Clin Anaesthesiol 1987; 1:533-556) but there is little evidence that this apparent gain has lead to better clinical decisions or earlier warning of significant instability. One of the tasks of the European Union sponsored IMPROVE group was to attempt to identify significant downward trends in vital parameters sufficiently early to allow clinical intervention to be potent and effective and ultimately improve patient outcome from a wide range of life threatening conditions. The first stage of this task was to define examples of such life threatening deterioration and conduct a survey in representative intensive care units of the incidence of these conditions and the subsequent patient outcomes. This is a preliminary task, the next stage being the gathering of "real time' data from critically ill patients for 24-h sample periods to probe for deteriorating trends and to compile a comprehensive annotated data library of physiological data as a rich resource for future adaptations in signal processing technology and clinical decision support.

Bacterial Infections↗

Visualization for genomics: the Microbial Genome Viewer.

SUMMARY: A Web-based visualization tool, the Microbial Genome Viewer, is presented that allows the user to combine complex genomic data in a highly interactive way. This Web tool enables the interactive generation of chromosome wheels and linear genome maps from genome annotation data stored in a MySQL database. The generated images are in scalable vector graphics (SVG) format, which is suitable for creating high-quality scalable images and dynamic Web representations. Gene-related data such as transcriptome and time-course microarray experiments can be superimposed on the maps for visual inspection. AVAILABILITY: The Microbial Genome Viewer 1.0 is freely available at http://www.cmbi.kun.nl/MGV

Chromosome Mapping↗

Ontology annotation treebrowser : an interactive tool where the complementarity of medical subject headings and gene ontology improves the interpretation of gene lists.

Gene expression and proteomics analysis allow the investigation of thousands of biomolecules in parallel. This results in a long list of interesting genes or proteins and a list of annotation terms in the order of thousands. It is not a trivial task to understand such a gene list and it would require extensive efforts to bring together the overwhelming amounts of associated information from the literature and databases. Thus, it is evident that we need ways of condensing and filtering this information. An excellent way to represent knowledge is to use ontologies, where it is possible to group genes or terms with overlapping context, rather than studying one-dimensional lists of keywords. Therefore, we have built the ontology annotation treebrowser (OAT) to represent, condense, filter and summarise the knowledge associated with a list of genes or proteins. The OAT system consists of two disjointed parts; a MySQL database named OATdb, and a treebrowser engine that is implemented as a web interface. The OAT system is implemented using Perl scripts on an Apache web server and the gene, ontology and annotation data is stored in a relational MySQL database. In OAT, we have harmonized the two ontologies of medical subject headings (MeSH) and gene ontology (GO), to enable us to use knowledge both from the literature and the annotation projects in the same tool. OAT includes multiple gene identifier sets, which are merged internally in the OAT database. We have also generated novel MeSH annotations by mapping accession numbers to MEDLINE entries. The ontology browser OAT was created to facilitate the analysis of gene lists. It can be browsed dynamically, so that a scientist can interact with the data and govern the outcome. Test statistics show which branches are enriched. We also show that the two ontologies complement each other, with surprisingly low overlap, by mapping annotations to the Unified Medical Language System. We have developed a novel interactive annotation browser that is the first to incorporate both MeSH and GO for improved interpretation of gene lists. With OAT, we illustrate the benefits of combining MeSH and GO for understanding gene lists. OAT is available as a public web service at: http://www.ifm.liu.se/bioinfo/oat.

Algorithms↗

dbRIP: a highly integrated database of retrotransposon insertion polymorphisms in humans.

Retrotransposons constitute over 40% of the human genome and play important roles in the evolution of the genome. Since certain types of retrotransposons, particularly members of the Alu, L1, and SVA families, are still active, their recent and ongoing propagation generates a unique and important class of human genomic diversity/polymorphism (for the presence and absence of an insertion) with some elements known to cause genetic diseases. So far, over 2,300, 500, and 80 Alu, L1, and SVA insertions, respectively, have been reported to be polymorphic and many more are yet to be discovered. We present here the Database of Retrotransposon Insertion Polymorphisms (dbRIP; http://falcon.roswellpark.org:9090), a highly integrated and interactive database of human retrotransposon insertion polymorphisms (RIPs). dbRIP currently contains a nonredundant list of 1,625, 407, and 63 polymorphic Alu, L1, and SVA elements, respectively, or a total of 2,095 RIPs. In dbRIP, we deploy the utilities and annotated data of the genome browser developed at the University of California at Santa Cruz (UCSC) for user-friendly queries and integrative browsing of RIPs along with all other genome annotation information. Users can query the database by a variety of means and have access to the detailed information related to a RIP, including detailed insertion sequences and genotype data. dbRIP represents the first database providing comprehensive, integrative, and interactive compilation of RIP data, and it will be a useful resource for researchers working in the area of human genetics.

Databases, Genetic↗

Laboratory Information Management Software for genotyping workflows: applications in high throughput crop genotyping.

BACKGROUND: With the advances in DNA sequencer-based technologies, it has become possible to automate several steps of the genotyping process leading to increased throughput. To efficiently handle the large amounts of genotypic data generated and help with quality control, there is a strong need for a software system that can help with the tracking of samples and capture and management of data at different steps of the process. Such systems, while serving to manage the workflow precisely, also encourage good laboratory practice by standardizing protocols, recording and annotating data from every step of the workflow. RESULTS: A laboratory information management system (LIMS) has been designed and implemented at the International Crops Research Institute for the Semi-Arid Tropics (ICRISAT) that meets the requirements of a moderately high throughput molecular genotyping facility. The application is designed as modules and is simple to learn and use. The application leads the user through each step of the process from starting an experiment to the storing of output data from the genotype detection step with auto-binning of alleles; thus ensuring that every DNA sample is handled in an identical manner and all the necessary data are captured. The application keeps track of DNA samples and generated data. Data entry into the system is through the use of forms for file uploads. The LIMS provides functions to trace back to the electrophoresis gel files or sample source for any genotypic data and for repeating experiments. The LIMS is being presently used for the capture of high throughput SSR (simple-sequence repeat) genotyping data from the legume (chickpea, groundnut and pigeonpea) and cereal (sorghum and millets) crops of importance in the semi-arid tropics. CONCLUSION: A laboratory information management system is available that has been found useful in the management of microsatellite genotype data in a moderately high throughput genotyping laboratory. The application with source code is freely available for academic users and can be downloaded from http://www.icrisat.org/gt-bt/lims/lims.asp.

Algorithms↗

The IBIS project: data collection in London. Improved Monitoring for Brain Dysfunction during Intensive Care and Surgery.

The primary aim of the Improved Monitoring for Brain Dysfunction during Intensive Care and Surgery (IBIS) project was to create a unique and comprehensively annotated data library (DL) of multiple physiological, including neurophysiological, signals. Data collection was undertaken in Kuopio, Finland and London, UK, and comparable protocols were used at all the sites. In London, 43 patients were recruited at the Royal Brompton Hospital, followed by nine at St. Bartholomew's Hospital, all of whom underwent cardiac or combined cardiac and carotid artery surgery. Thirty-seven patients underwent a single operation, while 15 underwent two procedures. The protocols and equipment used, problems specific to the electrically hostile environment and preliminary results are described, including those of clinical interest. The DL is being used for the development of clinically applicable neurophysiological monitoring tools.

Adult↗

MitBASE: a comprehensive and integrated mitochondrial DNA database.

MitBASE is an integrated and comprehensive database of mitochondrial DNA data which collects all available information from different organisms and from intraspecie variants and mutants. Research institutions from different countries are involved, each in charge of developing, collecting and annotating data for the organisms they are specialised in. The design of the actual structure of the database and its implementation in a user-friendly format are the care of the European Bioinformatics Institute. The database can be accessed on the Web at the following address: http://www.ebi.ac. uk/htbin/Mitbase/mitbase.pl. The impact of this project is intended for both basic and applied research. The study of mitochondrial genetic diseases and mitochondrial DNA intraspecie diversity are key topics in several biotechnological fields. The database has been funded within the EU Biotechnology programme.

Animals↗

Developing a corpus of clinical notes manually annotated for part-of-speech.

PURPOSE: This paper presents a project whose main goal is to construct a corpus of clinical text manually annotated for part-of-speech (POS) information. We describe and discuss the process of training three domain experts to perform linguistic annotation. METHODS: Three domain experts were trained to perform manual annotation of a corpus of clinical notes. A part of this corpus was combined with the Penn Treebank corpus of general purpose English text and another part was set aside for testing. The corpora were then used for training and testing statistical part-of-speech taggers. We list some of the challenges as well as encouraging results pertaining to inter-rater agreement and consistency of annotation. RESULTS: We used the Trigrams'n'Tags (TnT) [T. Brants, TnT-a statistical part-of-speech tagger, In: Proceedings of NAACL/ANLP-2000 Symposium, 2000] tagger trained on general English data to achieve 89.79% correctness. The same tagger trained on a portion of the medical data annotated for this project improved the performance to 94.69%. Furthermore, we find that discriminating between different types of discourse represented by different sections of clinical text may be very beneficial to improve correctness of POS tagging. CONCLUSION: Our preliminary experimental results indicate the necessity for adapting state-of-the-art POS taggers to the sublanguage domain of clinical text.

Abstracting and Indexing↗

Microarray phenotyping in Dictyostelium reveals a regulon of chemotaxis genes.

MOTIVATION: Coordinate regulation of gene expression can provide information on gene function. To begin a large-scale analysis of Dictyostelium gene function, we clustered genes based on their expression in wild-type and mutant strains and analyzed their functions. RESULTS: We found 17 modes of wild-type gene expression and refined them into 57 submodes considering mutant data. Annotation analyses revealed correlations between co-expression and function and an unexpected correlation between expression and function of genes involved in various aspects of chemotaxis. Co-regulation of chemotaxis genes was also found in published data from neutrophils. To test the predictive power of the analysis, we examined the phenotypes of mutations in seven co-regulated genes that had no published role in chemotaxis. Six mutants exhibited chemotaxis defects, supporting the idea that function can be inferred from co-expression. The clustering and annotation analyses provide a public resource for Dictyostelium functional genomics.

Animals↗

In silico studies of energy metabolism of normal and diseased heart.

Biotechnology research is developing into genomic analyses that involve the simultaneous monitoring of thousands of genes. The development of various bioinformatics resources that provide efficient access to information is necessary. We have used single-pass sequencing of randomly selected cDNA clones to generate expressed sequence tags (ESTs). These ESTs data has been widely used to study gene expression in a variety of heart libraries [1, 21]. Data annotation on our recent finding allows us to construct the profiles of genes in the energy metabolizing pathways (glycolysis and glycogen metabolism) that are expressed in heart cDNA libraries. In silico studies of genes of energy metabolism yields data that are consistent with results derived from conventional metabolic experiments. The change in gene profiles describing the metabolism of diseased hearts is also presented here.

Cardiomegaly↗

From gene networks to brain networks.

The brain's structural organization is so complex that 2,500 years of analysis leaves pervasive uncertainty about (i) the identity of its basic parts (regions with their neuronal cell types and pathways interconnecting them), (ii) nomenclature, (iii) systematic classification of the parts with respect to topographic relationships and functional systems and (iv) the reliability of the connectional data itself. Here we present a prototype knowledge management system (http://brancusi.usc.edu/bkms/) for analyzing the architecture of brain networks in a systematic, interactive and extendable way. It supports alternative interpretations and models, is based on fully referenced and annotated data and can interact with genomic and functional knowledge management systems through web services protocols.

Animals↗