Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

A virtual laboratory notebook for simulation models.

In this paper we describe how we have adopted the laboratory notebook as a metaphor for interacting with computer simulation models. This 'virtual' notebook stores the simulation output and meta-data (which is used to record the scientist's interactions with the simulation). The meta-data stored consists of annotations (equivalent to marginal notes in a laboratory notebook), a history tree and a log of user interactions. The history tree structure records when in 'simulation' time, and from what starting point in the tree changes are made to the parameters by the user. Typically these changes define a new run of the simulation model (which is represented as a new branch of the history tree). The tree shows the structure of the changes made to the simulation and the log is required to keep the order in which the changes occurred. Together they form a record which you would normally find in a laboratory notebook. The history tree is plotted in simulation parameter space. This shows the scientist's interactions with the simulation visually and allows direct manipulation of the parameter information presented, which in turn is used to control directly the state of the simulation. The interactions with the system are graphical and usually involve directly selecting or dragging data markers and other graphical control devices around in parameter space. If the graphical manipulators do not provide precise enough control then textual manipulation is still available which allows numerical values to be entered by hand. The Virtual Laboratory Notebook, by providing interesting interactions with the visual view of the history tree, provides a mechanism for giving the user complex and novel ways of interacting with biological computer simulation models.

Computational Biology↗

Genomes as geography: using GIS technology to build interactive genome feature maps.

BACKGROUND: Many commonly used genome browsers display sequence annotations and related attributes as horizontal data tracks that can be toggled on and off according to user preferences. Most genome browsers use only simple keyword searches and limit the display of detailed annotations to one chromosomal region of the genome at a time. We have employed concepts, methodologies, and tools that were developed for the display of geographic data to develop a Genome Spatial Information System (GenoSIS) for displaying genomes spatially, and interacting with genome annotations and related attribute data. In contrast to the paradigm of horizontally stacked data tracks used by most genome browsers, GenoSIS uses the concept of registered spatial layers composed of spatial objects for integrated display of diverse data. In addition to basic keyword searches, GenoSIS supports complex queries, including spatial queries, and dynamically generates genome maps. Our adaptation of the geographic information system (GIS) model in a genome context supports spatial representation of genome features at multiple scales with a versatile and expressive query capability beyond that supported by existing genome browsers. RESULTS: We implemented an interactive genome sequence feature map for the mouse genome in GenoSIS, an application that uses ArcGIS, a commercially available GIS software system. The genome features and their attributes are represented as spatial objects and data layers that can be toggled on and off according to user preferences or displayed selectively in response to user queries. GenoSIS supports the generation of custom genome maps in response to complex queries about genome features based on both their attributes and locations. Our example application of GenoSIS to the mouse genome demonstrates the powerful visualization and query capability of mature GIS technology applied in a novel domain. CONCLUSION: Mapping tools developed specifically for geographic data can be exploited to display, explore and interact with genome data. The approach we describe here is organism independent and is equally useful for linear and circular chromosomes. One of the unique capabilities of GenoSIS compared to existing genome browsers is the capacity to generate genome feature maps dynamically in response to complex attribute and spatial queries.

Animals↗

Computational Proteomics Analysis System (CPAS): an extensible, open-source analytic system for evaluating and publishing proteomic data and high throughput biological experiments.

The open-source Computational Proteomics Analysis System (CPAS) contains an entire data analysis and management pipeline for Liquid Chromatography Tandem Mass Spectrometry (LC-MS/MS) proteomics, including experiment annotation, protein database searching and sequence management, and mining LC-MS/MS peptide and protein identifications. CPAS architecture and features, such as a general experiment annotation component, installation software, and data security management, make it useful for collaborative projects across geographical locations and for proteomics laboratories without substantial computational support.

Computational Biology↗

Putting microarrays in a context: integrated analysis of diverse biological data.

In recent years, multiple types of high-throughput functional genomic data that facilitate rapid functional annotation of sequenced genomes have become available. Gene expression microarrays are the most commonly available source of such data. However, genomic data often sacrifice specificity for scale, yielding very large quantities of relatively lower-quality data than traditional experimental methods. Thus sophisticated analysis methods are necessary to make accurate functional interpretation of these large-scale data sets. This review presents an overview of recently developed methods that integrate the analysis of microarray data with sequence, interaction, localisation and literature data, and further outlines current challenges in the field. The focus of this review is on the use of such methods for gene function prediction, understanding of protein regulation and modelling of biological networks.

Algorithms↗

The Eukaryotic Promoter Database, EPD: new entry types and links to gene expression data.

The Eukaryotic Promoter Database (EPD) is an annotated, non-redundant collection of eukaryotic Pol II promoters, for which the transcription start site has been determined experimentally. Access to promoter sequences is provided by pointers to positions in nucleotide sequence entries. The annotation part of an entry includes a description of the initiation site mapping data, exhaustive cross-references to the EMBL nucleotide sequence database, SWISS-PROT, TRANSFAC and other databases, as well as bibliographic references. EPD is structured in a way that facilitates dynamic extraction of biologically meaningful promoter subsets for comparative sequence analysis. World Wide Web-based interfaces have been developed which enable the user to view EPD entries in different formats, to select and extract promoter sequences according to a variety of criteria, and to navigate to related databases exploiting different cross-references. The EPD web site also features yearly updated base frequency matrices for major eukaryotic promoter elements. EPD can be accessed at http://www.epd.isb-sib.ch.

Animals↗

Challenges in data management for functional genomics.

Biological databases face challenges in four main areas: (1). integration, interoperation and federation; (2). ontologies and definitions of semantics; (3). community annotation; and (4). integration of data analysis tools with databases. Each of these areas provides interesting targets for research and development.

Computational Biology↗

The IclR family of transcriptional activators and repressors can be defined by a single profile.

In the last decade enormous advances in life sciences have been possible due to the information obtained from DNA sequencing projects. The optimal interpretation and analysis of genome sequence data requires the precise annotation and classification of proteins deduced from open reading frames, which is usually done with the help of family-specific signatures. Here we report a novel profile for the IclR type of transcriptional activators and repressors. In contrast to profiles for other families of transcriptional regulators, the new IclR profile is located outside the helix-turn-helix DNA-binding motif. We provide evidence that the new profile is more specific than any of the existing signatures for this family of regulators. More than 500 representatives of this family were identified with this profile. A database on bacterial regulators (http://www.bactregulators.org) was built to compile and regroup the sequences with the aid of the new profile.

Amino Acid Sequence↗

Complex functionality of gene groups identified from high-throughput data.

Relating experimental data to biological knowledge is necessary to cope with the avalanches of new data emerging from recent developments in high-throughput technologies. Automatic functional profiling becomes the de facto standard approach for the secondary analysis of high-throughput data. A number of tools employing available gene functional annotations have been developed for this purpose. However, current annotations are derived mostly from traditional analysis of the individual gene function. The complex biological phenomena carried out by the concerted activity of many genes often requires the definition of new complex functionality (related to a group of genes), which is, in many cases, not available in current annotation vocabularies. Functional profiling with annotation terms related to the description of individual biological functions of a gene may fail to provide reasonable interpretation of biological relationships in a set of genes involved in complex biological phenomena. We introduce a novel procedure to profile a complex functionality of a gene set. Complex functionality is constructed as a combination of available annotation terms. By profiling ChIP-chip data from Saccharomyces cerevisiae we demonstrate that this technique produces deeper insights into the results of high-throughput experiments that are beyond the known facts described in the functional classifications.

Animals↗

Recent additions and improvements to the Onto-Tools.

The Onto-Tools suite is composed of an annotation database and six seamlessly integrated, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner and Pathway-Express. The Onto-Tools database has been expanded to include various types of data from 12 new databases. Our database now integrates different types of genomic data from 19 sequence, gene, protein and annotation databases. Additionally, our database is also expanded to include complete Gene Ontology (GO) annotations. Using the enhanced database and GO annotations, Onto-Express now allows functional profiling for 24 organisms and supports 17 different types of input IDs. Onto-Translate is also enhanced to fully utilize the capabilities of the new Onto-Tools database with an ultimate goal of providing the users with a non-redundant and complete mapping from any type of identification system to any other type. Currently, Onto-Translate allows arbitrary mappings between 29 types of IDs. Pathway-Express is a new tool that helps the users find the most interesting pathways for their input list of genes. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Animals↗

A High-Resolution Stereo-Seq Spatial Transcriptomic Resource for Adult Holstein Cattle Liver.

The bovine liver is a highly compartmentalized organ that plays essential roles in continuous gluconeogenesis and nitrogen recycling; however, its spatial molecular architecture has remained largely uncharacterized due to the limitations of traditional bulk and single-cell approaches. To address this gap, Spatial Enhanced Resolution Omics-sequencing (Stereo-seq) was utilized to generate a subcellular-resolution (500 nm) transcriptomic map of an adult Holstein cattle liver, and a refined reference-guided workflow was implemented to overcome standard annotation limitations in livestock. Raw sequencing data were processed using the Stereo-seq Analysis Workflow and analyzed with Stereopy, Seurat, SingleR, and reference-guided workflows. Spatial aggregation was evaluated at Bin20, Bin50, Bin100, Bin150, and Bin200. Increasing bin size increased molecular identifier counts and detected-gene complexity while progressively reducing spatial granularity. Bin50, corresponding to 50 × 50 DNA nanoballs and an approximate nominal footprint of 25 × 25 µm, was therefore selected as a practical intermediate aggregation level for the primary analyses. Quality-control assessment, Leiden clustering, UMAP visualization, reference-based cell-type annotation, cluster-marker analysis, and spatial mapping of canonical hepatic genes demonstrated preservation of biologically interpretable liver transcriptional organization. Raw sequencing data processed spatial matrices, annotated objects, and analysis code are publicly available to support reanalysis and computational benchmarking. In summary, we present a Stereo-seq spatial transcriptomic resource generated from liver tissue of an adult Holstein cow. This initial resource provides a valuable foundation for future studies of bovine liver biology, comparative genomics, and the spatial basis of livestock health and production traits.

Animals↗

Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research.

SUMMARY: We present here Blast2GO (B2G), a research tool designed with the main purpose of enabling Gene Ontology (GO) based data mining on sequence data for which no GO annotation is yet available. B2G joints in one application GO annotation based on similarity searches with statistical analysis and highlighted visualization on directed acyclic graphs. This tool offers a suitable platform for functional genomics research in non-model species. B2G is an intuitive and interactive desktop application that allows monitoring and comprehension of the whole annotation and analysis process. AVAILABILITY: Blast2GO is freely available via Java Web Start at http://www.blast2go.de. SUPPLEMENTARY MATERIAL: http://www.blast2go.de -> Evaluation.

Algorithms↗

Image analysis for automatic segmentation of cytoplasms and classification of Rac1 activation.

BACKGROUND: Rac1 is a GTP-binding molecule involved in a wide range of cellular processes. Using digital image analysis, agonist-induced translocation of green fluorescent protein (GFP) Rac1 to the cellular membrane can be estimated quantitatively for individual cells. METHODS: A fully automatic image analysis method for cell segmentation, feature extraction, and classification of cells according to their activation, i.e., GFP-Rac1 translocation and ruffle formation at stimuli, is described. Based on training data produced by visual annotation of four image series, a statistical classifier was created. RESULTS: The results of the automatic classification were compared with results from visual inspection of the same time sequences. The automatic classification differed from the visual classification at about the same level as visual classifications performed by two different skilled professionals differed from each other. Classification of a second image set, consisting of seven image series with different concentrations of agonist, showed that the classifier could detect an increased proportion of activated cells at increased agonist concentration. CONCLUSIONS: Intracellular activities, such as ruffle formation, can be quantified by fully automatic image analysis, with an accuracy comparable to that achieved by visual inspection. This analysis can be done at a speed of hundreds of cells per second and without the subjectivity introduced by manual judgments.

Animals↗

[The new indicator set for health reporting activities in the German States].

In May 2003, the third revised version of the indicator set for health reporting activities was confirmed by the health ministries of all German States (Bundesländer). Modeled on the restructured indicator set which has been annotated with meta-data descriptions, most Bundesländer have now started to collect data for their specific health reporting activities. Thanks to the support provided by national data holders and the Federal Statistical Office, it has been possible to further enlarge the database and for the first time also ensure access via the Federal Statistical Office. In this contribution the authors describe the methodological and statistical principles of the indicator set. Another aspect is the benefit of the indicator set for the health reporting activities in the German States.

Aged↗

TET2 promotes monocyte inflammatory activation in asthma via ALKBH5-m6A regulation and PI3K signaling: evidence from m6A-SNP and single-cell analyses.

Asthma is a complex inflammatory airway disease with strong genetic determinants, yet the functional relevance of most asthma-associated non-coding variants remains unclear. Emerging evidence suggests that N6-methyladenosine (m6A) modification may serve as a critical epitranscriptomic link between genetic variation and immune regulation. In this study, we aimed to systematically identify functionally relevant m6A-regulated genes in asthma by integrating large-scale GWAS data, m6A-SNP annotations, and single-cell transcriptomic analyses, and to investigate their roles in monocyte-driven airway inflammation. We identified TET2 as a key m6A-regulated gene associated with both asthma and lung function, which was selectively upregulated in monocytes during asthma and accompanied by activation of inflammatory and PI3K signaling pathways. Mechanistic experiments further demonstrated that inflammatory stimulation induced ALKBH5 expression, reduced m6A modification of TET2 mRNA, and increased TET2 protein levels, thereby promoting PI3K/AKT signaling and pro-inflammatory cytokine production, whereas inhibition of TET2 or ALKBH5 attenuated these effects. Collectively, these findings demonstrate that ALKBH5-mediated m6A regulation of TET2 enhances PI3K/AKT signaling in monocytes, thereby promoting inflammatory responses in asthma. Our study establishes TET2 as a key m6A-regulated gene linking genetic susceptibility to monocyte-driven inflammation, and highlights the ALKBH5-m6A-TET2 axis as a potential therapeutic target for modulating aberrant immune responses in asthma.

Humans↗

Sequence-based protein structure prediction using a reduced state-space hidden Markov model.

This work describes the use of a hidden Markov model (HMM), with a reduced number of states, which simultaneously learns amino acid sequence and secondary structure for proteins of known three-dimensional structure and it is used for two tasks: protein class prediction and fold recognition. The Protein Data Bank and the annotation of the SCOP database are used for training and evaluation of the proposed HMM for a number of protein classes and folds. Results demonstrate that the reduced state-space HMM performs equivalently, or even better in some cases, on classifying proteins than a HMM trained with the amino acid sequence. The major advantage of the proposed approach is that a small number of states is employed and the training algorithm is of low complexity and thus relatively fast.

Algorithms↗

EMMA: a platform for consistent storage and efficient analysis of microarray data.

As a high throughput technique, microarray experiments produce large data sets, consisting of measured data, laboratory protocols, and experimental settings. We have implemented the open source platform EMMA to store and analyze these data. The system provides automated pipelines for data processing and has a modular architecture that can be easily extended. EMMA features detailed reports about spots and their corresponding measurements. In addition to routine data analysis algorithms, the system can be integrated with other components that contain additional data sources (e.g. genome annotation systems).

Algorithms↗

Sec and Tat Mediated Secretion Safeguards Mycobacterium tuberculosis Membrane Homeostasis.

Protein secretion is essential for the growth and virulence of Mycobacterium tuberculosis, yet the organization and function of its secretion pathways remain poorly understood. We reviewed the existing literature, combined it with systematic queries, and finalized annotations based on experimental data and computational predictions to compile a curated list of 92 secretory components and 198 reactions involved in Sec, twin-arginine translocation (Tat), and ESX pathways. Using CRISPRi, targeted depletion of SecA1 or TatAC impaired both in vitro growth and ex vivo survival. Label-free quantitative secretome analysis revealed decreased export of substrates dependent on SecA1 and TatAC, with enrichment of cytosolic proteins in culture filtrates, indicating increased membrane dysbiosis. Membrane proteomics showed elevated levels of proteins engaged in intermediary and lipid metabolism, while proteins associated with the cell wall and cell processes decreased, suggesting weakened membrane integrity. Loss of SecA1 or TatAC increased membrane permeability, with the effect being more pronounced in the case of TatAC, and caused structural abnormalities seen under electron microscopy. Overall, our integrated multi-omics and functional genetics studies demonstrate that the SecA1 and Tat pathways are essential for maintaining membrane homeostasis in Mycobacterium tuberculosis. These results suggest that essential secretory proteins may be promising targets for therapeutic intervention.

Mycobacterium tuberculosis↗

Taking advantage of sophisticated pacemaker diagnostics.

The ever-increasing complexity of pacing systems, combined with functions that vary from one manufacturer to another, can pose challenges during analysis of device function. Standard pacemaker diagnostics are measured data, electrogram telemetry, maker annotations and event counters, albeit with their current limitations. New diagnostic features discussed include time-based diagnostics, histograms of sensed amplitudes, pacing thresholds, or impedance trending. Mode-switching algorithms, combined with diagnostic features, facilitate the use of dual-chamber devices in patients with paroxysmal atrial tachyarrhythmias. The introduction of electrogram storage into pacemakers further improves diagnostic capabilities and allows a permanent validation and optimization of diagnostic and therapeutic algorithms. External diagnostic devices, which provide Holter recordings with continuous marker annotations and patient-triggered diagnostics, are additional features that will become increasingly important.

Arrhythmias, Cardiac↗