Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Systematic analysis of snake neurotoxins' functional classification using a data warehousing approach.

MOTIVATION: Sequence annotations, functional and structural data on snake venom neurotoxins (svNTXs) are scattered across multiple databases and literature sources. Sequence annotations and structural data are available in the public molecular databases, while functional data are almost exclusively available in the published articles. There is a need for a specialized svNTXs database that contains NTX entries, which are organized, well annotated and classified in a systematic manner. RESULTS: We have systematically analyzed svNTXs and classified them using structure-function groups based on their structural, functional and phylogenetic properties. Using conserved motifs in each phylogenetic group, we built an intelligent module for the prediction of structural and functional properties of unknown NTXs. We also developed an annotation tool to aid the functional prediction of newly identified NTXs as an additional resource for the venom research community. AVAILABILITY: We created a searchable online database of NTX proteins sequences (http://research.i2r.a-star.edu.sg/Templar/DB/snake_neurotoxin). This database can also be found under Swiss-Prot Toxin Annotation Project website (http://www.expasy.org/sprot/).

Animals↗

GARBAN: genomic analysis and rapid biological annotation of cDNA microarray and proteomic data.

SUMMARY: Genomic Analysis and Rapid Biological ANnotation (GARBAN) is a new tool that provides an integrated framework to analyze simultaneously and compare multiple data sets derived from microarray or proteomic experiments. It carries out automated classifications of genes or proteins according to the criteria of the Gene Ontology Consortium at a level of depth defined by the user. Additionally, it performs clustering analysis of all sets based on functional categories or on differential expression levels. GARBAN also provides graphical representations of the biological pathways in which all the genes/proteins participate. AVAILABILITY: http://garban.tecnun.es.

Algorithms↗

DAtA: database of Arabidopsis thaliana annotation.

The Database of Arabidopsis thaliana Annotation (D At A) was created to enable easy access to and analysis of all the Arabidopsis genome project annotation. The database was constructed using the completed A.thaliana genomic sequence data currently in GenBank. An automated annotation process was used to predict coding sequences for GenBank records that do not include annotation. D At A also contains protein motifs and protein similarities derived from searches of the proteins in D At A with motif databases and the non-redundant protein database. The database is routinely updated to include new GenBank submissions for Arabidopsis genomic sequences and new Blast and protein motif search results. A web interface to D At A allows coding sequences to be searched by name, comment, blast similarity or motif field. In addition, browse options present lists of either all the protein names or identified motifs present in the sequenced A.thaliana genome. The database can be accessed at http://baggage. stanford.edu/group/arabprotein/

Arabidopsis↗

Expression array annotation using the BioMediator biological data integration system and the BioConductor analytic platform.

This paper presents the implementation of a model for expression array annotation (EAA) using the BioMediator biological data integration system along with BioConductor, an analytic tools platform. The model presented addresses the need for annotation sources identified during BioConductor inverted exclamation mark s development. Annotation provides us with well-curated genomic background knowledge for expression array analysis and interpretation. Annotation requests are constructed and posted to the query interface of the EAA package (the EAA model implemented as a component of BioConductor). The software enumerates all possible annotation paths for queries. These are then transformed to PQL queries and processed by BioMediator. Annotation entities returned from the EAA package answer the annotation request.

Computational Biology↗

ArrayExpress: a public database of gene expression data at EBI.

ArrayExpress is a public repository for microarray-based gene expression data, resulting from the implementation of the MAGE object model to ensure accurate data structuring and the MIAME standard, which defines the annotation requirements. ArrayExpress accepts data as MAGE-ML files for direct submissions or data from MIAMExpress, the MIAME compliant web-based annotation and submission tool of EBI. A team of curators supports the submission process, providing assistance in data annotation. Data retrieval is performed through a dedicated web interface. Relevant results may be exported to ExpressionProfiler, the EBI based expression analysis tool available online (http://www.ebi.ac.uk/arrayexpress).

Computational Biology↗

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans↗

Annotation of environmental OMICS data: application to the transcriptomics domain.

Researchers working on environmentally relevant organisms, populations, and communities are increasingly turning to the application of OMICS technologies to answer fundamental questions about the natural world, how it changes over time, and how it is influenced by anthropogenic factors. In doing so, the need to capture meta-data that accurately describes the biological "source" material used in such experiments is growing in importance. Here, we provide an overview of the formation of the "Env" community of environmental OMICS researchers and its efforts at considering the meta-data capture needs of those working in environmental OMICS. Specifically, we discuss the development to date of the Env specification, an informal specification including descriptors related to geographic location, environment, organism relationship, and phenotype. We then describe its application to the description of environmental transcriptomic experiments and how we have used it to extend the Minimum Information About a Microarray Experiment (MIAME) data standard to create a domain-specific extension that we have termed MIAME/Env. Finally, we make an open call to the community for participation in the Env Community and its future activities.

Ecology↗

Unsupervised pattern recognition: an introduction to the whys and wherefores of clustering microarray data.

Clustering has become an integral part of microarray data analysis and interpretation. The algorithmic basis of clustering -- the application of unsupervised machine-learning techniques to identify the patterns inherent in a data set -- is well established. This review discusses the biological motivations for and applications of these techniques to integrating gene expression data with other biological information, such as functional annotation, promoter data and proteomic data.

Algorithms↗

Clustering biological annotations and gene expression data to identify putatively co-regulated biological processes.

MOTIVATION: Functional profiling is a key step of microarray gene expression data analysis. Identifying co-regulated biological processes could help for better understanding of underlying biological interactions within the studied biological frame. RESULTS: We present herein an original approach designed to search for putatively co-regulated biological processes sharing a significant number of co-expressed genes. An R language implementation named "FunCluster" was built and tested on two gene expression data sets. A discriminatory functional analysis of the first data set, related to experiments performed on separated adipocytes and stroma vascular fraction cells of human white adipose tissue, highlighted the prevalent role of nonadipose cells in the synthesis of inflammatory and immunity molecules in human adiposity. On the second data set, resulting from a model investigating insulin coordinated regulation of gene expression in human skeletal muscle, FunCluster analysis spotlighted novel functional classes of putatively co-regulated biological processes related to protein metabolism and the regulation of muscular contraction. AVAILABILITY: Supplementary information about the FunCluster tool is available on-line at http://corneliu.henegar.info/FunCluster.htm.

Algorithms↗

The NITE XML Toolkit: flexible annotation for multimodal language data.

Multimodal corpora that show humans interacting via language are now relatively easy to collect. Current tools allow one either to apply sets of time-stamped codes to the data and consider their timing and sequencing or to describe some specific linguistic structure that is present in the data, built over the top of some form of transcription. To further our understanding of human communication, the research community needs code sets with both timings and structure, designed flexibly to address the research questions at hand. The NITE XML Toolkit offers library support that software developers can call upon when writing tools for such code sets and, thus, enables richer analyses than have previously been possible. It includes data handling, a query language containing both structural and temporal constructs, components that can be used to build graphical interfaces, sample programs that demonstrate how to use the libraries, a tool for running queries, and an experimental engine that builds interfaces on the basis of declarative specifications.

Communication↗

An annotated bibliography of methods for analysing correlated categorical data.

This paper provides an annotated bibliography of over 100 articles concerning methods for analysing correlated categorical response data. Most of the papers listed here concern categorical regression models and estimation, with particular emphasis on binary responses. The papers are classified by several characteristics which group them according to common themes. The bibliography serves as a reference of methods for analysts of correlated categorical data, as well as for persons interested in methodologic work in this active area of statistical research.

Clinical Trials as Topic↗

An extensible automated protein annotation tool: standardizing input and output using validated XML.

MOTIVATION: There is a frequent need to apply a large range of local or remote prediction and annotation tools to one or more sequences. We have created a tool able to dispatch one or more sequences to assorted services by defining a consistent XML format for data and annotations. RESULTS: By analyzing annotation tools, we have determined that annotations can be described using one or more of the six forms of data: numeric or textual annotation of residues, domains (residue ranges) or whole sequences. With this in mind, XML DTDs have been designed to store the input and output of any server. Plug-in wrappers to a number of services have been written which are called from a master script. The resulting APATML is then formatted for display in HTML. Alternatively further tools may be written to perform post-analysis.

Amino Acid Sequence↗

ASD: the Alternative Splicing Database.

Alternative splicing is widespread in mammalian gene expression, and variant splice patterns are often specific to different stages of development, particular tissues or a disease state. There is a need to systematically collect data on alternatively spliced exons, introns and splice isoforms, and to annotate this data. The Alternative Splicing Database consortium has been addressing this need, and is committed to maintaining and developing a value-added database of alternative splice events, and of experimentally verified regulatory mechanisms that mediate splice variants. In this paper we present two of the products from this project: namely, a database of computationally delineated alternative splice events as seen in alignments of EST/cDNA sequences with genome sequences, and a database of alternatively spliced exons collected from literature. The reported splice events are from nine different organisms and are annotated for various biological features including expression states and cross-species conservation. The data are presented on our ASD web pages (http://www.ebi.ac.uk/asd).

Alternative Splicing↗

Database tools for integrating and searching membrane property data correlated with neuronal morphology.

A critical problem in neuroscience is the lack of database tools for integrating neuronal property data. We report here the development of a combined object oriented-relational database (NeuronDB, http://senselab.med.yale.edu/neurondb) that meets these needs by providing tools for integrating data within neurons and comparing data across neurons. It focuses on three types of neuronal properties voltage-gated channels, neurotransmitter receptors, and neurotransmitters. The data are organized in relation to different regions of neurons as represented in canonical forms; using simple canonical models of complex cells as a vehicle for indexing information permits the database to be searchable across different neurons. Using these multidimensional search tools, users can locate specific properties in specific regions of a neuron; obtain integrated summaries of all properties within a region; and carry out searches to compare properties across equivalent compartments in different neurons. These tools thus permit searches of the multidimensional neuron property space equivalent to homology searches of sequence databases. NeuronDB is accessible over the Internet; it provides immediate links to citation indexes and abstracts supporting the deposited data, and annotations that indicate the state of acceptance of the data. Users are encouraged to contribute data. The ability to input the data from NeuronDB directly to NEURON and GENESIS is being developed. As a shared Web resource, NeuronDB should enhance the efforts of neuroscientists and neuronal modellers to analyze and compare the functional operations of different types of neurons.

Automation↗

NASCArrays: a repository for microarray data generated by NASC's transcriptomics service.

NASC operates an Affymetrix 'GeneChip' (microarray) service for the Arabidopsis thaliana community. All data produced by the service are publicly available through our microarray data base 'NASCArrays' published at http://affymetrix. arabidopsis.info. The data are accessible through text searching and a series of data mining tools. All data are annotated with sample preparation details, and the original Affymetrix data are available for download. The database aims to be MIAME supportive and provide a coordinated resource for re searchers interested in the transcriptome of Arabidopsis. Using this database, data produced will be shared with other databases worldwide.

Arabidopsis↗

SOURCE: a unified genomic resource of functional annotations, ontologies, and gene expression data.

The explosion in the number of functional genomic datasets generated with tools such as DNA microarrays has created a critical need for resources that facilitate the interpretation of large-scale biological data. SOURCE is a web-based database that brings together information from a broad range of resources, and provides it in manner particularly useful for genome-scale analyses. SOURCE's GeneReports include aliases, chromosomal location, functional descriptions, GeneOntology annotations, gene expression data, and links to external databases. We curate published microarray gene expression datasets and allow users to rapidly identify sets of co-regulated genes across a variety of tissues and a large number of conditions using a simple and intuitive interface. SOURCE provides content both in gene and cDNA clone-centric pages, and thus simplifies analysis of datasets generated using cDNA microarrays. SOURCE is continuously updated and contains the most recent and accurate information available for human, mouse, and rat genes. By allowing dynamic linking to individual gene or clone reports, SOURCE facilitates browsing of large genomic datasets. Finally, SOURCEs batch interface allows rapid extraction of data for thousands of genes or clones at once and thus facilitates statistical analyses such as assessing the enrichment of functional attributes within clusters of genes. SOURCE is available at http://source.stanford.edu.

Animals↗

Integration of GO annotations in Correspondence Analysis: facilitating the interpretation of microarray data.

MOTIVATION: The functional interpretation of microarray datasets still represents a time-consuming and challenging task. Up to now functional categories that are relevant for one or more experimental context(s) have been commonly extracted from a set of regulated genes and presented in long lists. RESULTS: To facilitate interpretation, we integrated Gene Ontology (GO) annotations into Correspondence Analysis to display genes, experimental conditions and gene-annotations in a single plot. The position of the annotations in these plots can be directly used for the functional interpretation of clusters of genes or experimental conditions without the need for comparing long lists of annotations. Correspondence Analysis is not limited in the number of experimental conditions that can be compared simultaneously, allowing an easy identification of characterizing annotations even in complex experimental settings. Due to the rapidly increasing amount of annotation data available, we apply an annotation filter. Hereby the number of displayed annotations can be significantly reduced to a set of descriptive ones, further enhancing the interpretability of the plot. We validated the method on transcription data from Saccharomyces cerevisiae and human pancreatic adenocarcinomas. AVAILABILITY: The M-CHiPS software is accessible for collaborators at http://www.mchips.org

Algorithms↗

CDACHIE: chromatin domain annotation by integrating chromatin interaction and epigenomic data with contrastive learning.

MOTIVATION: Chromatin domain annotation identifies functional genomic regions, such as active and inactive zones, based on epigenomic features like histone modifications, DNA methylation, and chromatin accessibility. While recent methods have utilized both chromatin interaction data (e.g. Hi-C) and epigenomic data, they often overlook the direct relationship between these data types. RESULTS: In this study, we introduce Chromatin Domain Annotation using Contrastive Learning for Hi-C and Epigenomic Data (CDACHIE), a method for identifying chromatin domains from Hi-C and epigenomic data. Our approach leverages contrastive learning to generate aligned representative vectors for both data types at each genomic bin. The concatenated vectors are then clustered using K-means to classify distinct chromatin domain types. CDACHIE achieves superior performance in Variance Explained, evaluated across gene expression, replication timing, and ChIA-PET data. This highlights its robust ability to integrate semantic associations between Hi-C and epigenomic features within the embedding space. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/maruyama-lab-design/CDACHIE. An archival snapshot of the code used in this study is available on Zenodo: https://doi.org/10.5281/zenodo.15751780.

Chromatin↗