Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

A probabilistic learning approach to whole-genome operon prediction.

We present a computational approach to predicting operons in the genomes of prokaryotic organisms. Our approach uses machine learning methods to induce predictive models for this task from a rich variety of data types including sequence data, gene expression data, and functional annotations associated with genes. We use multiple learned models that individually predict promoters, terminators and operons themselves. A key part of our approach is a dynamic programming method that uses our predictions to map every known and putative gene in a given genome into its most probable operon. We evaluate our approach using data from the E. coli K-12 genome.

Gene Expression Profiling↗

PartiGene--constructing partial genomes.

UNLABELLED: Expressed sequence tags (ESTs) offer a low-cost approach to gene discovery and are being used by an increasing number of laboratories to obtain sequence information for a wide variety of organisms. The challenge lies in processing and organizing this data within a genomic context to facilitate large scale analyses. Here we present PartiGene, an integrated sequence analysis suite that uses freely available public domain software to (1) process raw trace chromatograms into sequence objects suitable for submission to dbEST; (2) place these sequences within a genomic context; (3) perform customizable first-pass annotation of the data; and (4) present the data as HTML tables and an SQL database resource. PartiGene has been used to create a number of non-model organism database resources including NEMBASE (http://www.nematodes.org) and LumbriBase (http://www.earthworms.org/). The packages are readily portable, freely available and can be run on simple Linux-based workstations. AVAILABILITY: PartiGene is available from http://www.nematodes.org/PartiGene and also forms part of the EST analysis software, associated with the Natural Environmental Research Council (UK) Bio-Linux project (http://envgen.nox.ac.uk/biolinux.html).

Chromatography↗

Mining biochemical information: lessons taught by the ribosome.

The publication of the crystal structures of the ribosome offers an opportunity to retrospectively evaluate the information content of hundreds of qualitative biochemical and biophysical studies of these structures. We assessed the correspondence between more than 2,500 experimental proximity measurements and the distances observed in the ribosomal crystals. Although detailed experimental procedures and protocols are unique in almost each analyzed paper, the data can be grouped into subsets with similar patterns and analyzed in an integrative fashion. We found that, for crosslinking, footprinting, and cleavage data, the corresponding distances observed in crystal structures generally did not exceed the maximum values expected (from the estimated length of the agent and maximal anticipated deviations from the conformations found in crystals). However, the distribution of distances had heavier tails than those typically assumed when building three-dimensional models, and the fraction of incompatible distances was greater than expected. Some of these incompatibilities can be attributed to the experimental methods used. In addition, the accuracy of these procedures appears to be sensitive to the different reactivities, flexibilities, and interactions among the components. These findings demonstrate the necessity of a very careful analysis of data used for structural modeling and consideration of all possible parameters that could potentially influence the quality of measurements. We conclude that experimental proximity measurements can provide useful distance information for structural modeling, but with a broad distribution of inferred distance ranges. We also conclude that development of automated modeling approaches would benefit from better annotations of experimental data for detection and interpretation of their significance.

Algorithms↗

Assessing the protease and protease inhibitor content of the human genome.

The revealing of the entire complement of protease and protease inhibitor sequences by the Human Genome Project will be of great importance to both academic and pharmaceutical research. Although the finishing phase is not yet complete, a selection of secondary annotation sources and comparisons with completed model organism genomes already allow useful estimates to be made. Conservative extrapolation suggests a total of approximately 1.8% for human proteases. This is close to the figures for yeast (1.7%) and worm (1.8%) but lower than the fly (3.4%) which has a large trypsin-like protease content. Using estimates for the human proteome of between 40,000 and 60,000 genes would extrapolate to 700-1,100 proteases, compared with approximately 360 currently represented as GenBank mRNAs. Preliminary comparisons between domain annotations for predicted human gene products and completed proteins suggest the genomic protease family and mechanistic class distributions will broadly reflect those in the current transcript data. The protease:inhibitor ratio at the mRNA level is currently approximately 9:1, but genome annotation data indicate that inhibitory domains are more widespread than this ratio would indicate.

Databases, Factual↗

springScape: visualisation of microarray and contextual bioinformatic data using spring embedding and an 'information landscape'.

The interpretation of microarray and other high-throughput data is highly dependent on the biological context of experiments. However, standard analysis packages are poor at simultaneously presenting both the array and related bioinformatic data. We have addressed this challenge by developing a system springScape based on 'spring embedding' and an 'information landscape' allowing several related data sources to be dynamically combined while highlighting one particular feature. Each data source is represented as a network of nodes connected by weighted edges. The networks are combined and embedded in the 2-D plane by spring embedding such that nodes with a high similarity are drawn close together. Complex relationships can be discovered by varying the weight of each data source and observing the dynamic response of the spring network. By modifying Procrustes analysis, we find that the visualizations have an acceptable degree of reproducibility. The 'information landscape' highlights one particular data source, displaying it as a smooth surface whose height is proportional to both the information being viewed and the density of nodes. The algorithm is demonstrated using several microarray data sets in combination with protein-protein interaction data and GO annotations. Among the features revealed are the spatio-temporal profile of gene expression and the identification of GO terms correlated with gene expression and protein interactions. The power of this combined display lies in its interactive feedback and exploitation of human visual pattern recognition. Overall, springScape shows promise as a tool for the interpretation of microarray data in the context of relevant bioinformatic information.

Algorithms↗

Large-scale mutational analysis for the annotation of the mouse genome.

After sequencing the human and mouse genomes, the annotation of these sequences with biological functions is an important challenge in genomic research. A major tool to analyse gene function on the organismal level is the analysis of mutant phenotypes. Because of its genetic and physiological similarity to man, the mouse has become the model organism of choice for the study of genetic diseases. In addition, there is at the moment no other vertebrate for which versatile techniques to manipulate the genome are as well developed. Several mouse mutagenesis projects have provided the proof-of-principle that a systematic and comprehensive mutagenesis of every gene in the mammalian genome will be feasible. An exhaustive functional annotation of the mammalian genome can only be achieved in a combination of phenotype- and gene-driven approaches in large- and small-scale academic and private projects. Major challenges will be to develop standardised phenotyping protocols for the clinical and pathological characterisation of mouse mutants, the improvement of mutation detection methods and the dissemination of resources and data. Beyond gene annotation, it will be necessary to understand how gene functions are integrated into the complex network of regulatory interactions in the cell.

Animals↗

PaVESy: Pathway Visualization and Editing System.

UNLABELLED: A data managing system for editing and visualization of biological pathways is presented. The main component of PaVESy (Pathway Visualization and Editing System) is a relational SQL database system. The database design allows storage of biological objects, such as metabolites, proteins, genes and respective relations, which are required to assemble metabolic and regulatory biological interactions. The database model accommodates highly flexible annotation of biological objects by user-defined attributes. In addition, specific roles of objects are derived from these attributes in the context of user-defined interactions, e.g. in the course of pathway generation or during editing of the database content. Furthermore, the user may organize and arrange the database content within a folder structure and is free to group and annotate database objects of interest within customizable subsets. Thus, we allow an individualized view on the database content and facilitate user customization. A JAVA-based class library was developed, which serves as the database programming interface to PaVESy. This API provides classes, which implement the concepts of object persistence in SQL databases, such as entries, interactions, annotations, folders and subsets. We created editing and visualization tools for navigation in and visualization of the database content. User approved pathway assemblies are stored and may be retrieved for continued modification, annotation and export. Data export is interfaced with a range of network visualization programs, such as Pajek or other software allowing import of SBML or GML data format. AVAILABILITY: http://pavsey.mpimp-golm.mpg.de

Database Management Systems↗

cuticleDB: a relational database of Arthropod cuticular proteins.

BACKGROUND: The insect exoskeleton or cuticle is a bi-partite composite of proteins and chitin that provides protective, skeletal and structural functions. Little information is available about the molecular structure of this important complex that exhibits a helicoidal architecture. Scores of sequences of cuticular proteins have been obtained from direct protein sequencing, from cDNAs, and from genomic analyses. Most of these cuticular protein sequences contain motifs found only in arthropod proteins. DESCRIPTION: cuticleDB is a relational database containing all structural proteins of Arthropod cuticle identified to date. Many come from direct sequencing of proteins isolated from cuticle and from sequences from cDNAs that share common features with these authentic cuticular proteins. It also includes proteins from the Drosophila melanogaster and the Anopheles gambiae genomes, that have been predicted to be cuticular proteins, based on a Pfam motif (PF00379) responsible for chitin binding in Arthropod cuticle. The total number of the database entries is 445: 370 derive from insects, 60 from Crustacea and 15 from Chelicerata. The database can be accessed from our web server at http://bioinformatics.biol.uoa.gr/cuticleDB. CONCLUSIONS: CuticleDB was primarily designed to contain correct and full annotation of cuticular protein data. The database will be of help to future genome annotators. Users will be able to test hypotheses for the existence of known and also of yet unknown motifs in cuticular proteins. An analysis of motifs may contribute to understanding how proteins contribute to the physical properties of cuticle as well as to the precise nature of their interaction with chitin.

Amino Acid Motifs↗

Microarray RNA transcriptional profiling: part II. Analytical considerations and annotation.

This review summarizes the various data filtration, transformation and normalization processes for different array platforms (cDNA, oligos, one- and two-color), data analysis methods and their validation, and databases and annotation for RNA transcriptional profiling microarrays. This review is intended to introduce the beginner to the analyses and interpretation of gene expression studies using a nonmathematical approach for easier comprehension. Microarray analysis is not a trivial undertaking as there is no single method that works well for all, and results obtained from these analyses should be considered as a complement to other approaches.

Computational Biology↗

Molecular abnormalities in oocytes from women with polycystic ovary syndrome revealed by microarray analysis.

CONTEXT: Polycystic ovary syndrome (PCOS), the most common cause of anovulatory infertility, is characterized by increased ovarian androgen production and arrested follicle development and is frequently associated with insulin resistance. These PCOS phenotypes are associated with exaggerated ovarian responsiveness to FSH and increased pregnancy loss. OBJECTIVE: The objective of this study was to examine whether the perturbations in follicle growth and the intrafollicular environment affect gene expression and ultimately development of the PCOS oocyte. DESIGN: Oocyte cDNA was subjected to microarray and PCR analysis. SETTING: This study was conducted at a university laboratory. PATIENTS: The study comprised 10 normal ovulatory women and nine women with PCOS. INTERVENTION: The intervention was GnRH analog/recombinant human FSH therapy for in vitro fertilization. MAIN OUTCOME MEASURE: The main outcome measure was mRNA abundance of oocyte-expressed genes. RESULTS: Cluster analysis revealed differences in global gene expression profiles between normal and PCOS oocytes. Of the 8123 transcripts expressed in the oocytes, 374 genes showed significant differences in mRNA abundance in PCOS oocytes. Annotation of the data demonstrated that a subset of these genes was associated with chromosome alignment and segregation during mitosis and/or meiosis. Furthermore, 68 of the differentially expressed genes contained putative androgen receptor and/or peroxisome proliferating receptor gamma binding sites. CONCLUSIONS: These analyses demonstrated that normal and PCOS oocytes that are morphologically indistinguishable and of high quality exhibit different gene expression profiles. Promoter analysis suggests that androgens and other activators of nuclear receptors may play a role in differential gene expression in the PCOS oocyte. Likewise, annotation of the differentially expressed genes suggests that defects in meiosis or early embryonic development may contribute to reduced developmental competency of PCOS oocytes.

Base Sequence↗

The Cerefy Neuroradiology Atlas: a Talairach-Tournoux atlas-based tool for analysis of neuroimages available over the internet.

The article introduces an atlas-assisted method and a tool called the Cerefy Neuroradiology Atlas (CNA), available over the Internet for neuroradiology and human brain mapping. The CNA contains an enhanced, extended, and fully segmented and labeled electronic version of the Talairach-Tournoux brain atlas, including parcelated gyri and Brodmann's areas. To our best knowledge, this is the first online, publicly available application with the Talairach-Tournoux atlas. The process of atlas-assisted neuroimage analysis is done in five steps: image data loading, Talairach landmark setting, atlas normalization, image data exploration and analysis, and result saving. Neuroimage analysis is supported by a near-real-time, atlas-to-data warping based on the Talairach transformation. The CNA runs on multiple platforms; is able to process simultaneously multiple anatomical and functional data sets; and provides functions for a rapid atlas-to-data registration, interactive structure labeling and annotating, and mensuration. It is also empowered with several unique features, including interactive atlas warping facilitating fine tuning of atlas-to-data fit, navigation on the triplanar formed by the image data and the atlas, multiple-images-in-one display with interactive atlas-anatomy-function blending, multiple label display, and saving of labeled and annotated image data. The CNA is useful for fast atlas-assisted analysis of neuroimage data sets. It increases accuracy and reduces time in localization analysis of activation regions; facilitates to communicate the information on the interpreted scans from the neuroradiologist to other clinicians and medical students; increases the neuroradiologist's confidence in terms of anatomy and spatial relationships; and serves as a user-friendly, public domain tool for neuroeducation. At present, more than 700 users from five continents have subscribed to the CNA.

Atlases as Topic↗

Sodium Overload-Related Molecular Subtypes and a Four-Gene Prognostic Signature Predict Survival, Immune Landscape, and Therapeutic Response in Acute Myeloid Leukemia.

Sodium overload has recently emerged as a critical metabolic stressor involved in cancer progression; however, its molecular characteristics and clinical relevance in acute myeloid leukemia (AML) remain unexplored. RNA-seq data sets, clinical annotations, and mutational profiles of AML patients were annotations from The Cancer Genome Atlas and integrated with Genotype-Tissue Expression normal samples. Sodium overload-related genes (SORGs) were obtained from GeneCards. Differentially expressed SORGs (DESORGs) screened by applying the limma statistical model, followed by univariate Cox proportional hazards regression, consensus clustering, functional enrichment, immune infiltration analysis, and pathway evaluation. A prognostic signature was developed through least absolute shrinkage and selection operator regression followed by multivariate Cox modeling. The model's performance was further verified in two external GEO data sets (GSE71014 and GSE37642). Nomogram construction, subgroup analysis, tumor mutational burden (TMB) assessment, drug sensitivity prediction, transcription factor (TF) analysis, and competing endogenous RNA (ceRNA) network analyses were also performed. A total of 57 DESORGs were identified, and 2 sodium overload-related molecular subtypes exhibited distinct survival, immune infiltration, and inflammatory pathway activation. A robust four-gene signature (DOCK1, GABRE, HTR7, ACSM1) stratified patients into high- and low-risk categories with significantly different survival across training and validation cohorts. High-risk patients displayed increased immune infiltration, higher TMB, reduced sensitivity to multiple chemotherapeutic drugs, and inferior predicted response to PD-L1 blockade. TF and ceRNA networks revealed multilayered transcriptional and post-transcriptional regulation of the signature genes. This study identifies sodium overload-related molecular heterogeneity in AML and establishes a validated four-gene prognostic signature that integrates genomic, immunologic, and therapeutic features, offering potential utility for personalized risk assessment and treatment optimization.

Humans↗

DNA Data Bank of Japan (DDBJ) for genome scale research in life science.

The DNA Data Bank of Japan (DDBJ, http://www.ddbj.nig.ac.jp) has made an effort to collect as much data as possible mainly from Japanese researchers. The increase rates of the data we collected, annotated and released to the public in the past year are 43% for the number of entries and 52% for the number of bases. The increase rates are accelerated even after the human genome was sequenced, because sequencing technology has been remarkably advanced and simplified, and research in life science has been shifted from the gene scale to the genome scale. In addition, we have developed the Genome Information Broker (GIB, http://gib.genes.nig.ac.jp) that now includes more than 50 complete microbial genome and Arabidopsis genome data. We have also developed a database of the human genome, the Human Genomics Studio (HGS, http://studio.nig.ac.jp). HGS provides one with a set of sequences being as continuous as possible in any one of the 24 chromosomes. Both GIB and HGS have been updated incorporating newly available data and retrieval tools.

Animals↗

Ab initio prediction of transcription factor targets using structural knowledge.

Current approaches for identification and detection of transcription factor binding sites rely on an extensive set of known target genes. Here we describe a novel structure-based approach applicable to transcription factors with no prior binding data. Our approach combines sequence data and structural information to infer context-specific amino acid-nucleotide recognition preferences. These are used to predict binding sites for novel transcription factors from the same structural family. We demonstrate our approach on the Cys(2)His(2) Zinc Finger protein family, and show that the learned DNA-recognition preferences are compatible with experimental results. We use these preferences to perform a genome-wide scan for direct targets of Drosophila melanogaster Cys(2)His(2) transcription factors. By analyzing the predicted targets along with gene annotation and expression data we infer the function and activity of these proteins.

Journal Article↗

PLET1 (C11orf34), a highly expressed and processed novel gene in pig and mouse placenta, is transcribed but poorly spliced in human.

Sequencing of porcine cDNAs identified a novel EST with high frequency in placenta tissue. Full-length PLET1 (placenta-expressed transcript 1, also called C11orf34) matched a mouse cDNA and many bovine and mouse ESTs but no human transcripts or ESTs. However, the porcine cDNA matched several putative exons within a human genomic DNA fragment on chromosome 11. This human locus is in a region of conserved synteny with pig chromosome 9, to which the porcine gene was subsequently mapped. RNA blot hybridization showed that this gene had high expression in porcine and mouse conceptus and throughout placenta development. In situ hybridization using mouse placenta showed PLET1 expression in trophoblast cells of the labyrinth, as well as in spongiotrophoblast and glycogen trophoblast cells. However, no expression of PLET1 was detected by RNA blot analysis of human placenta, although RT-PCR analysis detected very small amounts of partially spliced RNA that were significantly less abundant than the RNA levels in mouse placenta. Donor and acceptor splicing site sequences in the exons of the human gene are poorly conserved and may be the cause of inefficient splicing found specifically in human tissue. Our data correct GenomeScan annotation of this region of the human genome and describe functional gene discovery in mammals not recognized in human EST projects.

Amino Acid Sequence↗

Ebbie: automated analysis and storage of small RNA cloning data using a dynamic web server.

BACKGROUND: DNA sequencing is used ubiquitously: from deciphering genomes to determining the primary sequence of small RNAs (smRNAs). The cloning of smRNAs is currently the most conventional method to determine the actual sequence of these important regulators of gene expression. Typical smRNA cloning projects involve the sequencing of hundreds to thousands of smRNA clones that are delimited at their 5' and 3' ends by fixed sequence regions. These primers result from the biochemical protocol used to isolate and convert the smRNA into clonable PCR products. Recently we completed a smRNA cloning project involving tobacco plants, where analysis was required for approximately 700 smRNA sequences. Finding no easily accessible research tool to enter and analyze smRNA sequences we developed Ebbie to assist us with our study. RESULTS: Ebbie is a semi-automated smRNA cloning data processing algorithm, which initially searches for any substring within a DNA sequencing text file, which is flanked by two constant strings. The substring, also termed smRNA or insert, is stored in a MySQL and BlastN database. These inserts are then compared using BlastN to locally installed databases allowing the rapid comparison of the insert to both the growing smRNA database and to other static sequence databases. Our laboratory used Ebbie to analyze scores of DNA sequencing data originating from an smRNA cloning project. Through its built-in instant analysis of all inserts using BlastN, we were able to quickly identify 33 groups of smRNAs from approximately 700 database entries. This clustering allowed the easy identification of novel and highly expressed clusters of smRNAs. Ebbie is available under GNU GPL and currently implemented on http://bioinformatics.org/ebbie/. CONCLUSION: Ebbie was designed for medium sized smRNA cloning projects with about 1,000 database entries. Ebbie can be used for any type of sequence analysis where two constant primer regions flank a sequence of interest. The reliable storage of inserts, and their annotation in a MySQL database, BlastN comparison of new inserts to dynamic and static databases make it a powerful new tool in any laboratory using DNA sequencing. Ebbie also prevents manual mistakes during the excision process and speeds up annotation and data-entry. Once the server is installed locally, its access can be restricted to protect sensitive new DNA sequencing data. Ebbie was primarily designed for smRNA cloning projects, but can be applied to a variety of RNA and DNA cloning projects.

Algorithms↗

How well are protein structures annotated in secondary databases?

We investigated to what extent Protein Data Bank (PDB) entries are annotated with second-party information based on existing cross-references between PDB and 15 other databases. We report 2 interesting findings. First, there is a clear "annotation gap" for structures less than 7 years old for secondary databases that are manually curated. Second, the examined databases overlap with each other quite well, dividing the PDB into 2 well-annotated thirds and one poorly annotated third. Both observations should be taken into account in any study depending on the selection of protein structures by their annotation.

Amino Acid Sequence↗

Linking genotype to phenotype: the International Rice Information System (IRIS).

The International Rice Information System (IRIS, http://www.iris.irri.org) is the rice implementation of the International Crop Information System (ICIS, http://www.icis.cgiar.org), a database system for the management and integration of global information on genetic resources and germplasm improvement for any crop. Building upon the germplasm genealogy and field data components of ICIS, IRIS is being extended to handle diverse rice genomics data including: genetic mapping, genome annotation, genotype, mutant, transcripteome, proteome and metabolomic data. Users can access information in the database through stand-alone programs and WWW interfaces offering specialist views to researchers with different interests.

Database Management Systems↗