Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

The Ribosomal Database Project (RDP-II): previewing a new autoaligner that allows regular updates and the new prokaryotic taxonomy.

The Ribosomal Database Project-II (RDP-II) pro-vides data, tools and services related to ribosomal RNA sequences to the research community. Through its website (http://rdp.cme.msu.edu), RDP-II offers aligned and annotated rRNA sequence data, analysis services, and phylogenetic inferences (trees) derived from these data. RDP-II release 8.1 contains 16 277 prokaryotic, 5201 eukaryotic, and 1503 mitochondrial small subunit rRNA sequences in aligned and annotated format. The current public beta release of 9.0 debuts a new regularly updated alignment of over 50 000 annotated (eu)bacterial sequences. New analysis services include a sequence search and selection tool (Hierarchy Browser) and a phylogenetic tree building and visualization tool (Phylip Interface). A new interactive tutorial guides users through the basics of rRNA sequence analysis. Other services include probe checking, phylogenetic placement of user sequences, screening of users' sequences for chimeric rRNA sequences, automated alignment, production of similarity matrices, and services to plan and analyze terminal restriction fragment polymorphism (T-RFLP) experiments. The RDP-II email address for questions or comments is rdpstaff@msu.edu.

Animals↗

A theory of optimal differential gene expression.

We investigate a model of optimal regulation, intended to describe large-scale differential gene expression. Relations between the optimal expression patterns and the function of genes are deduced from an optimality principle: the regulators have to maximise a fitness function which they influence directly via a cost term, and indirectly via their control on important cell variables, such as metabolic fluxes. According to the model, the optimal linear response to small perturbations reflects the regulators' functions, namely their linear influences on the cell variables. The optimal behaviour can be realised by a linear feedback mechanism. Known or assumed properties of response coefficients lead to predictions about regulation patterns. A symmetry relation predicted for deletion experiments is verified with gene expression data. Where the optimality assumption is valid, our results justify the use of expression data for functional annotation and for pathway reconstruction and suggest the use of linear factor models for the analysis of gene expression data.

Adaptation, Physiological↗

Data input module for Birth Defects Systems Manager.

The need for a computational bioinformatics infrastructure to manage the vast digital information from functional genomics and proteomics motivated us to develop Birth Defects Systems Manager (BDSM) as an open resource to facilitate analysis and discovery in developmental biology and developmental toxicity. This report describes the design, development and implementation of the data loading module of BDSM, referred to as LoadBDSM. It includes a shared data directory resource that can be granted various levels of security for different research groups or investigators to manage experimental datasets individually or in groups. LoadBDSM allows the upload of data and experiment details using controlled semantics for developmental exposure (toxicant, dosing scenario, intervention), biological sample (species, tissue, stage) and disease outcome (time, risk, phenotype). It adheres to existing controlled vocabulary plus rules of inference (ontologies) for experiment, data and metadata annotations. LoadBDSM extends the capabilities of BDSM to support the emergence of "embryo-formatics" defined here as the data, information and knowledge from genomic sciences applied to, or derived from, an embryological context. This includes, but is not limited to, delineating pathways and biological regulatory networks for specific chemicals or classes of developmental toxicants, developing novel biomarkers indicative of exposure and/or predictive of adverse effects, and integrating modern computing and information technology with data from molecular biology.

Abnormalities, Drug-Induced↗

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics↗

Perspectives: sequence data base searching in the era of large-scale genomic sequencing.

Large-scale sequencing of human and model organism genomes will have a profound impact on our ability to use sequence data base searching to predict the biochemical functions of sequences of interest. Despite the great value of more sequences in the data bases, a huge increase in data base size will also have adverse effects on data base searches. Upcoming problems will include (1) greatly increased search times, (2) an increase in background noise of high-scoring but biologically irrelevant matches, (3) inaccurate coding region prediction, leading to problems in protein data base searching, and (4) limited first-pass sequence annotation, making it difficult to determine the biological relevance of data base hits. Improved data base annotation tools and construction of smaller data bases of representative and highly-annotated sequences for first-pass analyses will be essential to deal with the impending flood of new genomic sequence.

Animals↗

Defining a new annotation object for DICOM image: a practical approach.

In this article, we present a new way of creating annotation objects for DICOM images, using the redundant data channel. Various types of annotations, including types containing color information, are possible and annotation objects can overlap the original DICOM image on a screen. Annotation objects can be created easily using a digital pen. Scanned images used in an electronic patient record can be added to objects. Although there are various ways of manipulating annotation objects, such as insertion, addition and modification of annotation objects in the DICOM image, the original clinical image is not affected because a redundant data channel is used for the annotation. The proposed method is expected to be very useful to medium and small clinics that cannot afford picture archiving and communication systems, as the DICOM standard makes provision for the annotation of clinical images in various ways.

Humans↗

Managing clinical research data: software tools for hypothesis exploration.

Data representation, data file specification, and the communication of data between software systems are playing increasingly important roles in clinical data management. This paper describes the concept of a self-documenting file that contains annotations or comments that aid visual inspection of the data file. We describe access of data from annotated files and illustrate data analysis with a few examples derived from the UNIX operating environment. Use of annotated files provides the investigator with both a useful representation of the primary data and a repository of comments that describe some of the context surrounding data capture.

Data Interpretation, Statistical↗

The NeuARt II system: a viewing tool for neuroanatomical data based on published neuroanatomical atlases.

BACKGROUND: Anatomical studies of neural circuitry describing the basic wiring diagram of the brain produce intrinsically spatial, highly complex data of great value to the neuroscience community. Published neuroanatomical atlases provide a spatial framework for these studies. We have built an informatics framework based on these atlases for the representation of neuroanatomical knowledge. This framework not only captures current methods of anatomical data acquisition and analysis, it allows these studies to be collated, compared and synthesized within a single system. RESULTS: We have developed an atlas-viewing application ('NeuARt II') in the Java language with unique functional properties. These include the ability to use copyrighted atlases as templates within which users may view, save and retrieve data-maps and annotate them with volumetric delineations. NeuARt II also permits users to view multiple levels on multiple atlases at once. Each data-map in this system is simply a stack of vector images with one image per atlas level, so any set of accurate drawings made onto a supported atlas (in vector graphics format) could be uploaded into NeuARt II. Presently the database is populated with a corpus of high-quality neuroanatomical data from the laboratory of Dr Larry Swanson (consisting 64 highly-detailed maps of PHAL tract-tracing experiments, made up of 1039 separate drawings that were published in 27 primary research publications over 17 years). Herein we take selective examples from these data to demonstrate the features of NeuArt II. Our informatics tool permits users to browse, query and compare these maps. The NeuARt II tool operates within a bioinformatics knowledge management platform (called 'NeuroScholar') either as a standalone or a plug-in application. CONCLUSION: Anatomical localization is fundamental to neuroscientific work and atlases provide an easily-understood framework that is widely used by neuroanatomists and non-neuroanatomists alike. NeuARt II, the neuroinformatics tool presented here, provides an accurate and powerful way of representing neuroanatomical data in the context of commonly-used brain atlases for visualization, comparison and analysis. Furthermore, it provides a framework that supports the delivery and manipulation of mapped data either as a standalone system or as a component in a larger knowledge management system.

Anatomy, Artistic↗

Apollo: a sequence annotation editor.

The well-established inaccuracy of purely computational methods for annotating genome sequences necessitates an interactive tool to allow biological experts to refine these approximations by viewing and independently evaluating the data supporting each annotation. Apollo was developed to meet this need, enabling curators to inspect genome annotations closely and edit them. FlyBase biologists successfully used Apollo to annotate the Drosophila melanogaster genome and it is increasingly being used as a starting point for the development of customized annotation editing tools for other genome projects.

Animals↗

Data-poor categorization and passage retrieval for gene ontology annotation in Swiss-Prot.

BACKGROUND: In the context of the BioCreative competition, where training data were very sparse, we investigated two complementary tasks: 1) given a Swiss-Prot triplet, containing a protein, a GO (Gene Ontology) term and a relevant article, extraction of a short passage that justifies the GO category assignment; 2) given a Swiss-Prot pair, containing a protein and a relevant article, automatic assignment of a set of categories. METHODS: Sentence is the basic retrieval unit. Our classifier computes a distance between each sentence and the GO category provided with the Swiss-Prot entry. The Text Categorizer computes a distance between each GO term and the text of the article. Evaluations are reported both based on annotator judgements as established by the competition and based on mean average precision measures computed using a curated sample of Swiss-Prot. RESULTS: Our system achieved the best recall and precision combination both for passage retrieval and text categorization as evaluated by official evaluators. However, text categorization results were far below those in other data-poor text categorization experiments The top proposed term is relevant in less that 20% of cases, while categorization with other biomedical controlled vocabulary, such as the Medical Subject Headings, we achieved more than 90% precision. We also observe that the scoring methods used in our experiments, based on the retrieval status value of our engines, exhibits effective confidence estimation capabilities. CONCLUSION: From a comparative perspective, the combination of retrieval and natural language processing methods we designed, achieved very competitive performances. Largely data-independent, our systems were no less effective that data-intensive approaches. These results suggests that the overall strategy could benefit a large class of information extraction tasks, especially when training data are missing. However, from a user perspective, results were disappointing. Further investigations are needed to design applicable end-user text mining tools for biologists.

Computational Biology↗

GenBank.

The GenBank nucleotide sequence database now contains sequence data and associated annotation corresponding to 85,000,000 nucleotides in 67,000 entries from a total of 3,000 organisms. The input stream of data coming into the database is primarily as direct submissions from the scientific community on electronic media, with little or no data being keyboarded from the printed page by the databank staff. The data are maintained in a relational database management system and are made available in flatfile form through on-line access, and through various network and off-line computer-readable media. The data are also distributed in relational form through satellite copies at a number of institutions in the U.S. and elsewhere. In addition, GenBank provides the U.S. distribution center for the BIOSCI electronic bulletin board service.

Animals↗

FunnyBase: a systems level functional annotation of Fundulus ESTs for the analysis of gene expression.

BACKGROUND: While studies of non-model organisms are critical for many research areas, such as evolution, development, and environmental biology, they present particular challenges for both experimental and computational genomic level research. Resources such as mass-produced microarrays and the computational tools linking these data to functional annotation at the system and pathway level are rarely available for non-model species. This type of "systems-level" analysis is critical to the understanding of patterns of gene expression that underlie biological processes. RESULTS: We describe a bioinformatics pipeline known as FunnyBase that has been used to store, annotate, and analyze 40,363 expressed sequence tags (ESTs) from the heart and liver of the fish, Fundulus heteroclitus. Primary annotations based on sequence similarity are linked to networks of systematic annotation in Gene Ontology (GO) and the Kyoto Encyclopedia of Genes and Genomes (KEGG) and can be queried and computationally utilized in downstream analyses. Steps are taken to ensure that the annotation is self-consistent and that the structure of GO is used to identify higher level functions that may not be annotated directly. An integrated framework for cDNA library production, sequencing, quality control, expression data generation, and systems-level analysis is presented and utilized. In a case study, a set of genes, that had statistically significant regression between gene expression levels and environmental temperature along the Atlantic Coast, shows a statistically significant (P < 0.001) enrichment in genes associated with amine metabolism. CONCLUSION: The methods described have application for functional genomics studies, particularly among non-model organisms. The web interface for FunnyBase can be accessed at http://genomics.rsmas.miami.edu/funnybase/super_craw4/. Data and source code are available by request at jpaschall@bioinfobase.umkc.edu.

Animals↗

Specialized microbial databases for inductive exploration of microbial genome sequences.

BACKGROUND: The enormous amount of genome sequence data asks for user-oriented databases to manage sequences and annotations. Queries must include search tools permitting function identification through exploration of related objects. METHODS: The GenoList package for collecting and mining microbial genome databases has been rewritten using MySQL as the database management system. Functions that were not available in MySQL, such as nested subquery, have been implemented. RESULTS: Inductive reasoning in the study of genomes starts from "islands of knowledge", centered around genes with some known background. With this concept of "neighborhood" in mind, a modified version of the GenoList structure has been used for organizing sequence data from prokaryotic genomes of particular interest in China. GenoChore http://bioinfo.hku.hk/genochore.html, a set of 17 specialized end-user-oriented microbial databases (including one instance of Microsporidia, Encephalitozoon cuniculi, a member of Eukarya) has been made publicly available. These databases allow the user to browse genome sequence and annotation data using standard queries. In addition they provide a weekly update of searches against the world-wide protein sequences data libraries, allowing one to monitor annotation updates on genes of interest. Finally, they allow users to search for patterns in DNA or protein sequences, taking into account a clustering of genes into formal operons, as well as providing extra facilities to query sequences using predefined sequence patterns. CONCLUSION: This growing set of specialized microbial databases organize data created by the first Chinese bacterial genome programs (ThermaList, Thermoanaerobacter tencongensis, LeptoList, with two different genomes of Leptospira interrogans and SepiList, Staphylococcus epidermidis) associated to related organisms for comparison.

Algorithms↗

TEAM: a tool for the integration of expression, and linkage and association maps.

The identification of genes primarily responsible for complex genetic disorders is a daunting task. Despite the assignment of many susceptibility loci, there has only been limited success in identifying disease genes based solely on positional information from genome-wide screens. The incorporation of several complementary strategies in a single integrated approach should facilitate and further enhance the efficacy of this search for genes. To permit the integration of linkage, association and expression data, together with functional annotations, we have developed a Java-based software tool: TEAM (tool for the integration of expression, and linkage and association maps). TEAM includes a genome viewer, capable of overlaying karyobands, genes, markers, linkage graphs, association data, gene expression levels and functional annotations in one composite view. Data management, analysis and filtering functionality was implemented and extended with links to the Ensembl, Unigene and Gene Ontology databases to facilitate gene annotation. Filtering functionality can help prevent the exclusion of poorly annotated, but differentially expressed, genes that reside in candidate regions that show linkage or association. Here we demonstrate the program's functionality in our study on coeliac disease (OMIM 212750), a multifactorial gluten-sensitive enteropathy. We performed a combined data analysis of a genome-wide linkage screen in 82 Dutch families with affected siblings and the microarray expression profiles of 18,110 cDNAs in 22 intestinal biopsies.

Celiac Disease↗

A functional annotation of subproteomes in human plasma.

The data collected by Human Proteome Organization's Plasma Proteome Pilot project phase was analyzed by members of our working group. Accordingly, a functional annotation of the human plasma proteome was carried out. Here, we report the findings of our analyses. First, bioinformatic analyses were undertaken to determine the likely sources of plasma proteins and to develop a protein interaction network of proteins identified in this project. Second, annotation of these proteins was performed in the context of functional subproteomes involved in the coagulation pathway, the mononuclear phagocytic system, the inflammation pathway, the cardiovascular system, and the liver; as well as the subset of proteins associated with DNA binding activities. Our analyses contributed to the Plasma Proteome Database (http://www.plasmaproteomedatabase.org), an annotated database of plasma proteins identified by HPPP as well as from other published studies. In addition, we address several methodological considerations including the selective enrichment of post-translationally modified proteins by the use of multi-lectin chromatography as well as the use of peptidomic techniques to characterize the low molecular weight proteins in plasma. Furthermore, we have performed additional analyses of peptide identification data to annotate cleavage of signal peptides, sites of intra-membrane proteolysis and post-translational modifications. The HPPP-organized, multi-laboratory effort, as described herein, resulted in much synergy and was essential to the success of this project.

Blood Coagulation↗

The role SWISS-PROT and TrEMBL play in the genome research environment.

SWISS-PROT, a curated protein sequence data bank, contains not only sequence data but also annotation relevant to a particular sequence. The annotation added to each entry is done by a team of biologists and comes, primarily, from articles in journals reporting the actual sequencing and sometimes characterisation. Review articles and collaboration with external experts also play a role along with the use of secondary databases like PROSITE and Pfam in addition to a variety of feature prediction methods. Annotation added by these methods is checked for relevance and likelihood to a particular sequence. The onset of genome sequencing has led to a dramatic increase in sequence data to be included in SWISS-PROT. This has led to the production of TrEMBL (Translation of the EMBL database). TrEMBL consists of entries in a SWISS-PROT format that are derived from the translation of all coding sequences in the EMBL nucleotide sequence database, that are not in SWISS-PROT. Unlike SWISS-PROT entries those in TrEMBL are awaiting manual annotation. However, rather than just representing basic sequence and source information, steps have been taken to add features and annotation automatically. In taking these steps it is hoped that TrEMBL entries are enhanced with some indication as to what a protein is, could or may be.

Amino Acid Sequence↗

CysView: protein classification based on cysteine pairing patterns.

CysView is a web-based application tool that identifies and classifies proteins according to their disulfide connectivity patterns. It accepts a dataset of annotated protein sequences in various formats and returns a graphical representation of cysteine pairing patterns. CysView displays cysteine patterns for those records in the data with disulfide annotations. It allows the viewing of records grouped by connectivity patterns. CysView's utility as an analysis tool was demonstrated by the rapid and correct classification of scorpion toxin entries from GenPept on the basis of their disulfide pairing patterns. It has proved useful for rapid detection of irrelevant and partial records, or those with incomplete annotations. CysView can be used to support distant homology between proteins. CysView is publicly available at http://research.i2r.a-star.edu.sg/CysView/.

Computer Graphics↗

Toward the development of a gene index to the human genome: an assessment of the nature of high-throughput EST sequence data.

A rigorous analysis of the Merck-sponsored EST data with respect to known gene sequences increases the utility of the data set and helps refine methods for building a gene index. A highly curated human transcript data base was used as a reference data set of known genes. A detailed analysis of EST sequences derived from known genes was performed to assess the accuracy of EST sequence annotation. The EST data was screened to remove low-quality and low-complexity sequences. A set of high-quality ESTs similar to the transcript data base was identified using BLAST; this subset of ESTs was compared with the set of known genes using the Smith-Waterman algorithm. Error rates of several types were assessed based on a flexible match criterion defining sequence identity. The rate of lane-tracking errors is very low, approximately 0.5%. Insert size data is accurate within approximately 20%. Reversed clone and internal priming error rates are approximately 5% and 2.5%, respectively, contributing to the incorrect identification of reads as 3' ends of genes. Follow-up investigation reveals that a significant number of clones, miscategorized as reversed, represent overlapping genes on the opposite strand of entries in the transcript data base. Relevance of these results to the creation of a high-quality index to the human genome capable of supporting diverse genomic investigations is discussed.

Algorithms↗