Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

AgBase: a functional genomics resource for agriculture.

BACKGROUND: Many agricultural species and their pathogens have sequenced genomes and more are in progress. Agricultural species provide food, fiber, xenotransplant tissues, biopharmaceuticals and biomedical models. Moreover, many agricultural microorganisms are human zoonoses. However, systems biology from functional genomics data is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation and agricultural research communities are smaller with limited funding compared to many model organism communities. DESCRIPTION: To facilitate systems biology in these traditionally agricultural species we have established "AgBase", a curated, web-accessible, public resource http://www.agbase.msstate.edu for structural and functional annotation of agricultural genomes. The AgBase database includes a suite of computational tools to use GO annotations. We use standardized nomenclature following the Human Genome Organization Gene Nomenclature guidelines and are currently functionally annotating chicken, cow and sheep gene products using the Gene Ontology (GO). The computational tools we have developed accept and batch process data derived from different public databases (with different accession codes), return all existing GO annotations, provide a list of products without GO annotation, identify potential orthologs, model functional genomics data using GO and assist proteomics analysis of ESTs and EST assemblies. Our journal database helps prevent redundant manual GO curation. We encourage and publicly acknowledge GO annotations from researchers and provide a service for researchers interested in GO and analysis of functional genomics data. CONCLUSION: The AgBase database is the first database dedicated to functional genomics and systems biology analysis for agriculturally important species and their pathogens. We use experimental data to improve structural annotation of genomes and to functionally characterize gene products. AgBase is also directly relevant for researchers in fields as diverse as agricultural production, cancer biology, biopharmaceuticals, human health and evolutionary biology. Moreover, the experimental methods and bioinformatics tools we provide are widely applicable to many other species including model organisms.

Agriculture↗

Bioinformatics for venom and toxin sciences.

Venomous animals produce a myriad of important pharmacological components. The individual components, or venoms (toxins), are used in ion channel and receptor studies, drug discovery, and formulation of insecticides. The toxin data are scattered across public databases which provide sequence and structural descriptions, but very limited functional annotation. The exponential growth of newly identified toxin data has created a need for better data management. Venominformatics is a systematic bioinformatics approach in which classified, consolidated and cleaned venom data are stored into repositories and integrated with advanced bioinformatics tools for the analysis of structure and function of toxins. Venominformatics complements experimental studies and helps reduce the number of essential experiments.

Animals↗

PATRIC: the VBI PathoSystems Resource Integration Center.

The PathoSystems Resource Integration Center (PATRIC) is one of eight Bioinformatics Resource Centers (BRCs) funded by the National Institute of Allergy and Infection Diseases (NIAID) to create a data and analysis resource for selected NIAID priority pathogens, specifically proteobacteria of the genera Brucella, Rickettsia and Coxiella, and corona-, calici- and lyssaviruses and viruses associated with hepatitis A and E. The goal of the project is to provide a comprehensive bioinformatics resource for these pathogens, including consistently annotated genome, proteome and metabolic pathway data to facilitate research into counter-measures, including drugs, vaccines and diagnostics. The project's curation strategy has three prongs: 'breadth first' beginning with whole-genome and proteome curation using standardized protocols, a 'targeted' approach addressing the specific needs of researchers and an integrative strategy to leverage high-throughput experimental data (e.g. microarrays, proteomics) and literature. The PATRIC infrastructure consists of a relational database, analytical pipelines and a website which supports browsing, querying, data visualization and the ability to download raw and curated data in standard formats. At present, the site warehouses complete sequences for 17 bacterial and 332 viral genomes. The PATRIC website (https://patric.vbi.vt.edu) will continually grow with the addition of data, analysis and functionality over the course of the project.

Bioterrorism↗

Imagene: an integrated computer environment for sequence annotation and analysis.

MOTIVATION: To be fully and efficiently exploited, data coming from sequencing projects together with specific sequence analysis tools need to be integrated within reliable data management systems. Systems designed to manage genome data and analysis tend to give a greater importance either to the data storage or to the methodological aspect, but lack a complete integration of both components. RESULTS: This paper presents a co-operative computer environment (called Imagenetrade mark) dedicated to genomic sequence analysis and annotation. Imagene has been developed by using an object-based model. Thanks to this representation, the user can directly manipulate familiar data objects through icons or lists. Imagene also incorporates a solving engine in order to manage analysis tasks. A global task is solved by successive divisions into smaller sub-tasks. During program execution, these sub-tasks are graphically displayed to the user and may be further re-started at any point after task completion. In this sense, Imagene is more transparent to the user than a traditional menu-driven package. Imagene also provides a user interface to display, on the same screen, the results produced by several tasks, together with the capability to annotate these results easily. In its current form, Imagene has been designed particularly for use in microbial sequencing projects. AVAILABILITY: Imagene best runs on SGI (Irix 6.3 or higher) workstations. It is distributed free of charge on a CD-ROM, but requires some Ilog licensed software to run. Some modules also require separate license agreements. Please contact the authors for specific academic conditions and other Unix platforms. CONTACT: imagene home page: http://wwwabi.snv.jussieu.fr/imagene

Bacillus subtilis↗

Microarray analysis of gene expression: considerations in data mining and statistical treatment.

DNA microarray represents a powerful tool in biomedical discoveries. Harnessing the potential of this technology depends on the development and appropriate use of data mining and statistical tools. Significant current advances have made microarray data mining more versatile. Researchers are no longer limited to default choices that generate suboptimal results. Conflicting results in repeated experiments can be resolved through attention to the statistical details. In the current dynamic environment, there are many choices and potential pitfalls for researchers who intend to incorporate microarrays as a research tool. This review is intended to provide a simple framework to understand the choices and identify the pitfalls. Specifically, this review article discusses the choice of microarray platform, preprocessing raw data, differential expression and validation, clustering, annotation and functional characterization of genes, and pathway construction in light of emergent concepts and tools.

Cluster Analysis↗

HIVbase: a PC/Windows-based software offering storage and querying power for locally held HIV-1 genetic, experimental and clinical data.

BACKGROUND: Human immunodeficiency virus (HIV) research involves ongoing, repetitious sequencing of the HIV genome and the massive accumulation of associated investigational data. As a result, the storage of annotated DNA and/or protein sequences, as well as information retrieval, have become increasingly difficult tasks, with scientists extracting less information from their collected data than they should. OBJECTIVES: Our objective was to design and develop a software package to aid researchers in the storage, analysis and exploration of their HIV-associated data. RESULTS: HIVbase contains familiar, easy-to-use interfaces and functionality for integrating many types of disparate data. The software contains tools that allow for the mass import of raw genetic data, eliminate repetitious sequence translations, have the ability to identify automatically and store HIV regions of interest from nucleic acid or protein sequences, allow for the export of data in commonly used analysis-ready formats, and for unique querying approaches.

Algorithms↗

PhyloGibbs: a Gibbs sampling motif finder that incorporates phylogeny.

A central problem in the bioinformatics of gene regulation is to find the binding sites for regulatory proteins. One of the most promising approaches toward identifying these short and fuzzy sequence patterns is the comparative analysis of orthologous intergenic regions of related species. This analysis is complicated by various factors. First, one needs to take the phylogenetic relationship between the species into account in order to distinguish conservation that is due to the occurrence of functional sites from spurious conservation that is due to evolutionary proximity. Second, one has to deal with the complexities of multiple alignments of orthologous intergenic regions, and one has to consider the possibility that functional sites may occur outside of conserved segments. Here we present a new motif sampling algorithm, PhyloGibbs, that runs on arbitrary collections of multiple local sequence alignments of orthologous sequences. The algorithm searches over all ways in which an arbitrary number of binding sites for an arbitrary number of transcription factors (TFs) can be assigned to the multiple sequence alignments. These binding site configurations are scored by a Bayesian probabilistic model that treats aligned sequences by a model for the evolution of binding sites and "background" intergenic DNA. This model takes the phylogenetic relationship between the species in the alignment explicitly into account. The algorithm uses simulated annealing and Monte Carlo Markov-chain sampling to rigorously assign posterior probabilities to all the binding sites that it reports. In tests on synthetic data and real data from five Saccharomyces species our algorithm performs significantly better than four other motif-finding algorithms, including algorithms that also take phylogeny into account. Our results also show that, in contrast to the other algorithms, PhyloGibbs can make realistic estimates of the reliability of its predictions. Our tests suggest that, running on the five-species multiple alignment of a single gene's upstream region, PhyloGibbs on average recovers over 50% of all binding sites in S. cerevisiae at a specificity of about 50%, and 33% of all binding sites at a specificity of about 85%. We also tested PhyloGibbs on collections of multiple alignments of intergenic regions that were recently annotated, based on ChIP-on-chip data, to contain binding sites for the same TF. We compared PhyloGibbs's results with the previous analysis of these data using six other motif-finding algorithms. For 16 of 21 TFs for which all other motif-finding methods failed to find a significant motif, PhyloGibbs did recover a motif that matches the literature consensus. In 11 cases where there was disagreement in the results we compiled lists of known target genes from the literature, and found that running PhyloGibbs on their regulatory regions yielded a binding motif matching the literature consensus in all but one of the cases. Interestingly, these literature gene lists had little overlap with the targets annotated based on the ChIP-on-chip data. The PhyloGibbs code can be downloaded from http://www.biozentrum.unibas.ch/~nimwegen/cgi-bin/phylogibbs.cgi or http://www.imsc.res.in/~rsidd/phylogibbs. The full set of predicted sites from our tests on yeast are available at http://www.swissregulon.unibas.ch.

Algorithms↗

MIPS bacterial genomes functional annotation benchmark dataset.

MOTIVATION: Any development of new methods for automatic functional annotation of proteins according to their sequences requires high-quality data (as benchmark) as well as tedious preparatory work to generate sequence parameters required as input data for the machine learning methods. Different program settings and incompatible protocols make a comparison of the analyzed methods difficult. RESULTS: The MIPS Bacterial Functional Annotation Benchmark dataset (MIPS-BFAB) is a new, high-quality resource comprising four bacterial genomes manually annotated according to the MIPS functional catalogue (FunCat). These resources include precalculated sequence parameters, such as sequence similarity scores, InterPro domain composition and other parameters that could be used to develop and benchmark methods for functional annotation of bacterial protein sequences. These data are provided in XML format and can be used by scientists who are not necessarily experts in genome annotation. AVAILABILITY: BFAB is available at http://mips.gsf.de/proj/bfab

Bacterial Proteins↗

MADCAP: isolation of novel nAb-naïve AAV capsids from metagenomic data.

UNLABELLED: Gene therapy using adeno-associated virus (AAV) vectors offers promising treatment for genetic disorders, but significant limitations restrict clinical application. Current AAV serotypes exhibit strong liver tropism and require high doses for extra-hepatic targeting, and pre-existing antibodies (NAbs) exclude up to 50% of potential patients. Evolutionarily distant isolates can evade neutralization but typically transduce human tissues poorly and require extensive engineering. We developed MADCAP (Metagenomic AAV Discovery and Capsid Annotation Pipeline) to systematically mine metagenomic data for functional, clinically relevant AAV capsids. We hypothesized that these sources might contain capsids that do not circulate widely in humans, can transduce human cells, and avoid neutralization. We screened 4.2 million metagenomic samples and identified 139 novel AAV capsid isolates which were tested for viral capsid assembly, viability, neutralization evasion, and tissue transduction in non-human primates. While natural serotypes (AAV1, AAV2, AAV9) were neutralized at low dilutions of pooled human immunoglobulin (IVIG), 68% of tested MADCAP capsids exhibited minimal to undetectable neutralization even at supra-physiological IVIG concentrations. Systemically delivered MADCAP capsids effectively transduced multiple clinically relevant tissues in non-human primates. Two capsids, MC46 and MC55, demonstrated improved CNS tropism compared to AAV9 while maintaining comparable production yields. In passive transfer studies, MC46 retained full transduction efficiency in the presence of human antibodies, while AAV9 transduction was completely lost. This work establishes metagenomic mining as a powerful tool for accelerating AAV capsid discovery, identifying isolates with favorable tissue tropisms and resistance to broadly neutralizing antibodies. IMPORTANCE: This work provides proof of concept that potentially clinically relevant AAVs can be isolated from metagenomic data. Our findings lay the groundwork for accelerated discovery of AAV capsids which could potentially increase the accessibility and effectiveness of AAV gene therapy.

AAV↗

Conditional Diffusion Model-Based Method for Annotation of Antibiotic Resistance Gene Properties.

The crisis of bacterial antibiotic resistance, which has led to a decline in the effectiveness of antibiotics originally used to combat bacterial infections, has emerged as an urgent challenge for public health. Antibiotic resistance genes (ARGs) are one of the key reasons for bacteria to develop resistance to antibiotics. Therefore, accurately identifying and annotating the critical properties of ARGs is of great importance for addressing the antibiotic resistance emergency. Although existing deep learning models demonstrate remarkable effectiveness in extracting local features from sequence data, they still face limitations in the capacity to further gain the enriched latent representations within the data. To address the critical challenge of extracting higher-quality representations from ARGs sequence data, we propose a novel ARGs properties annotation method based on the conditional diffusion model which is used to learn latent representations through domain-specific knowledge injection. Specifically, during the conditional information integration phase, we systematically incorporate ARGs' domain knowledge to guide the diffusion process in generating high-quality latent representations. To overcome information redundancy caused by direct concatenation of conditional information and intermediate features, we design a cross-attention mechanism that enables feature fusion between heterogeneous information sources, thereby enhancing further the quality of obtained representations. Experimental results on widely used data sets demonstrate the framework's effectiveness in achieving superior prediction performance compared to existing methods.

Anti-Bacterial Agents↗

Phytophthora functional genomics database (PFGD): functional genomics of phytophthora-plant interactions.

The Phytophthora Functional Genomics Database (PFGD; http://www.pfgd.org), developed by the National Center for Genome Resources in collaboration with The Ohio State University-Ohio Agricultural Research and Development Center (OSU-OARDC), is a publicly accessible information resource for Phytophthora-plant interaction research. PFGD contains transcript, genomic, gene expression and functional assay data for Phytophthora infestans, which causes late blight of potato, and Phytophthora sojae, which affects soybeans. Automated analyses are performed on all sequence data, including consensus sequences derived from clustered and assembled expressed sequence tags. The PFGD search filter interface allows intuitive navigation of transcript and genomic data organized by library and derived queries using modifiers, annotation keywords or sequence names. BLAST services are provided for libraries built from the transcript and genomic sequences. Transcript data visualization tools include Quality Screening, Multiple Sequence Alignment and Features and Annotations viewers. A genomic browser that supports comparative analysis via novel dynamic functional annotation comparisons is also provided. PFGD is integrated with the Solanaceae Genomics Database (SolGD; http://www.solgd.org) to help provide insight into the mechanisms of infection and resistance, specifically as they relate to the genus Phytophthora pathogens and their plant hosts.

Algal Proteins↗

The mouse gene expression database GXD

The gene expression database (GXD) is being developed to store and integrate expression information for mouse development. GXD addresses many issues that apply to gene expression databases in general, and its data structures and supporting software tools are generalized in design and thus readily adaptable to other life stages and species. Integration of GXD with the mouse genome database (MGD) and interconnections with other relevant databases will place the gene expression data into the larger biological and analytical context. Here, we describe the design and implementation of GXD and illustrate, in particular, the gene expression annotator, an electronic system for submitting expression data to the database.Copyright 1997 Academic Press Limited Copyright 1997Academic Press Limited

Journal Article↗

BioWarehouse: a bioinformatics database warehouse toolkit.

BACKGROUND: This article addresses the problem of interoperation of heterogeneous bioinformatics databases. RESULTS: We introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research. CONCLUSION: BioWarehouse embodies significant progress on the database integration problem for bioinformatics.

Computational Biology↗

Automated complete slide digitization: a medium for simultaneous viewing by multiple pathologists.

Developments in telepathology robotic systems have evolved the concept of a 'virtual microscope' handling 'digital slides'. Slide digitization is a method of archiving salient histological features in numerical (digital) form. The value and potential of this have begun to be recognized by several international centres. Automated complete slide digitization has application at all levels of clinical practice and will benefit undergraduate, postgraduate, and continuing education. Unfortunately, as the volume of potential data on a histological slide represents a significant problem in terms of digitization, storage, and subsequent manipulation, the reality of virtual microscopy to date has comprised limited views at inadequate resolution. This paper outlines a system refined in the authors' laboratory, which employs a combination of enhanced hardware, image capture, and processing techniques designed for telepathology. The system is able to scan an entire slide at high magnification and create a library of such slides that may exist on an internet server or be distributed on removable media (such as CD-ROM or DVD). A digital slide allows image data manipulation at a level not possible with conventional light microscopy. Combinations of multiple users, multiple magnifications, annotations, and addition of ancillary textual and visual data are now possible. This demonstrates that with increased sophistication, the applications of telepathology technology need not be confined to second opinion, but can be extended on a wider front.

Analog-Digital Conversion↗

Multiple signal integration by decision tree induction to detect artifacts in the neonatal intensive care unit.

The high incidence of false alarms in the intensive care unit (ICU) necessitates the development of improved alarming techniques. This study aimed to detect artifact patterns across multiple physiologic data signals from a neonatal ICU using decision tree induction. Approximately 200 h of bedside data were analyzed. Artifacts in the data streams were visually located and annotated retrospectively by an experienced clinician. Derived values were calculated for successively overlapping time intervals of raw values, and then used as feature attributes for the induction of models trying to classify 'artifact' versus 'not artifact' cases. The results are very promising, indicating that integration of multiple signals by applying a classification system to sets of values derived from physiologic data streams may be a viable approach to detecting artifacts in neonatal ICU data.

Artifacts↗

Facilitating narrative medical discussions of type 1 diabetes with computer visualizations and photography.

OBJECTIVE: Patient-centered approaches to medicine suggest the need for physicians to become more aware of concerns and needs expressed in patient narratives. However, patients and physicians have different goals and discourse styles during consultations. We attempt to bridge these differences by providing patients with ways to collect, visualize, and describe behavioral and biomedical data. METHODS: We describe an intervention where individuals with type 1 diabetes photograph health-related behaviors. These images and blood glucose records are displayed in computer visualizations and used during patient-physician interviews. RESULTS: Qualitative analyses of interview data with patients who photographed their lives suggest the range of difficulties associated with diabetes self-management. The visualizations helped them articulate concerns about stress, peer relations, and unhealthy routines. CONCLUSION: Interventions that combine biomedical and biopsychosocial data during patient-physician consultations may be beneficial for patients, helping them reflect on correlations between behaviors and health. Physicians are provided with contextual evidence to better understand patient issues around diabetes management. PRACTICE IMPLICATIONS: We suggest that this and similar interventions could be used as an occasional diagnostic to help patients articulate stories of their health-related practices. Annotated archives of photographs and glucose data may also be useful tools for sharing diabetes experiences with newly diagnosed patients.

Adaptation, Psychological↗

Technologies for integrating biological data.

The process of building a new database relevant to some field of study in biomedicine involves transforming, integrating and cleansing multiple data sources, as well as adding new material and annotations. This paper reviews some of the requirements of a general solution to this data integration problem. Several representative technologies and approaches to data integration in biomedicine are surveyed. Then some interesting features that separate the more general data integration technologies from the more specialised ones are highlighted.

Database Management Systems↗

GeneXplorer: an interactive web application for microarray data visualization and analysis.

BACKGROUND: When publishing large-scale microarray datasets, it is of great value to create supplemental websites where either the full data, or selected subsets corresponding to figures within the paper, can be browsed. We set out to create a CGI application containing many of the features of some of the existing standalone software for the visualization of clustered microarray data. RESULTS: We present GeneXplorer, a web application for interactive microarray data visualization and analysis in a web environment. GeneXplorer allows users to browse a microarray dataset in an intuitive fashion. It provides simple access to microarray data over the Internet and uses only HTML and JavaScript to display graphic and annotation information. It provides radar and zoom views of the data, allows display of the nearest neighbors to a gene expression vector based on their Pearson correlations and provides the ability to search gene annotation fields. CONCLUSIONS: The software is released under the permissive MIT Open Source license, and the complete documentation and the entire source code are freely available for download from CPAN http://search.cpan.org/dist/Microarray-GeneXplorer/.

Computer Graphics↗