Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

MicroArray Facility: a laboratory information management system with extended support for Nylon based technologies.

BACKGROUND: High throughput gene expression profiling (GEP) is becoming a routine technique in life science laboratories. With experimental designs that repeatedly span thousands of genes and hundreds of samples, relying on a dedicated database infrastructure is no longer an option.GEP technology is a fast moving target, with new approaches constantly broadening the field diversity. This technology heterogeneity, compounded by the informatics complexity of GEP databases, means that software developments have so far focused on mainstream techniques, leaving less typical yet established techniques such as Nylon microarrays at best partially supported. RESULTS: MAF (MicroArray Facility) is the laboratory database system we have developed for managing the design, production and hybridization of spotted microarrays. Although it can support the widely used glass microarrays and oligo-chips, MAF was designed with the specific idiosyncrasies of Nylon based microarrays in mind. Notably single channel radioactive probes, microarray stripping and reuse, vector control hybridizations and spike-in controls are all natively supported by the software suite. MicroArray Facility is MIAME supportive and dynamically provides feedback on missing annotations to help users estimate effective MIAME compliance. Genomic data such as clone identifiers and gene symbols are also directly annotated by MAF software using standard public resources. The MAGE-ML data format is implemented for full data export. Journalized database operations (audit tracking), data anonymization, material traceability and user/project level confidentiality policies are also managed by MAF. CONCLUSION: MicroArray Facility is a complete data management system for microarray producers and end-users. Particular care has been devoted to adequately model Nylon based microarrays. The MAF system, developed and implemented in both private and academic environments, has proved a robust solution for shared facilities and industry service providers alike.

Cloning, Molecular↗

NetAffx: Affymetrix probesets and annotations.

NetAffx (http://www.affymetrix.com) details and annotates probesets on Affymetrix GeneChip microarrays. These annotations include (i) static information specific to the probeset composition; (ii) sequence annotations extracted from public databases; and (iii) protein sequence-level annotations derived from public domain programs, as well as libraries of hidden Markov models (HMMs) developed at Affymetrix. For each probeset, NetAffx lists the probe sequences, and the consensus sequence interrogated by the probes; for the larger chip sets, interactive maps display this sequence data in genomic context. Sequence annotations include Gene Ontology (GO) terms and depiction of GO graph relationships; predicted protein domains and motifs; orthologous sequences; links to relevant pathways; and links to public databases including UniGene, LocusLink, SWISS-PROT and OMIM.

Animals↗

The early origins and development of the scatterplot.

Of all the graphic forms used today, the scatterplot is arguably the most versatile, polymorphic, and generally useful invention in the history of statistical graphics. Its use by Galton led to the discovery of correlation and regression, and ultimately to much of present multivariate statistics. So, it is perhaps surprising that there is no one widely credited with the invention of this idea. Even more surprising is that there are few contenders for this title, and this question seems not to have been raised before. This article traces some of the developments in the history of this graphical method, the origin of the term scatterplot, the role it has played in the history of science, and some of its modern descendants. We suggest that the origin of this method can be traced to its unique advantage: the possibility to discover regularity in empirical data by smoothing and other graphic annotations to enhance visual perception.

Animals↗

Mining statistically significant associations for exploratory analysis of human sleep data.

We introduce a specialized association rule mining technique that can extract patterns from complex sleep data comprising polysomnographic recordings, clinical summaries, and sleep questionnaire responses. The rules mined can describe associations among temporally annotated events and questionnaire or summary data; e.g., the likelihood that an occurrence of a rapid eye movement (REM) sleep stage during the second 100 sleep epochs of the night is associated with moderate caffeine intake. We use chi2 analysis to ensure statistical significance of the mined rules at the level P < 0.05. Our results, obtained by mining sleep-related data from 242 human subjects, reveal clinically interesting associations among the polysomnographic and summary variables. Our experience suggests that association mining may also be useful for selection of variables prior to using logistic regression.

Algorithms↗

Expression profiling in transformed human B cells: influence of Btk mutations and comparison to B cell lymphomas using filter and oligonucleotide arrays.

We have used both Clontech Atlas Human Hematology/Immunology cDNA microarrays, containing 588 genes, and Affymetrix oligonucleotide U95Av2 human array complementary to more than 12,500 genes to get a global view of genes expressed in Epstein-Barr virus (EBV)-transformed B cells and genes regulated by Bruton's tyrosine kinase (Btk). We compared EBV-transformed wild-type (WT) B cells from a healthy individual, WT1 and an X-linked agammaglobulinemia (XLA) patient cell line, XLA1, using the Clontech filters arrays. Eleven genes were > or =1.9-fold induced in absence of functional Btk. Furthermore, we analyzed a second patient cell line, XLA2, and compared this to two WT cell lines using oligonucleotide arrays. A total of 391 genes were found to be differentially expressed, including kinases and transcriptions factors. Furthermore, one expressed sequence tag and eight complementary DNA clones with unknown function were down-regulated in XLA2, indicating their biological role. Higher-fold inductions, Fyn (39.5), Hck (15.5) and Cyp1B1 (5.8), were observed using oligonucleotide array and were confirmed using real-time PCR for Fyn (20.8), Hck (6.7) and Cyp1B1 (10). Two genes, B cell translocation gene1 (BTG1) and B cell-specific OCT binding factor-1 (OBF-1) were induced > or =1.9-fold in both XLA1 and XLA2 analyzed by Atlas filter arrays andAffymetrix chips, respectively. Data from both filter and oligonucleotide arrays were compared to the gene clusters of a previously published lymphoma expression profile by linking to the UniGene transcript database. Our findings demonstrate for the first time the use of microarray to study the influence of Btk mutations and the use of functional annotation and validation of expression data by comparison of microarray analyses.

Agammaglobulinaemia Tyrosine Kinase↗

inGeno--an integrated genome and ortholog viewer for improved genome to genome comparisons.

BACKGROUND: Systematic genome comparisons are an important tool to reveal gene functions, pathogenic features, metabolic pathways and genome evolution in the era of post-genomics. Furthermore, such comparisons provide important clues for vaccines and drug development. Existing genome comparison software often lacks accurate information on orthologs, the function of similar genes identified and genome-wide reports and lists on specific functions. All these features and further analyses are provided here in the context of a modular software tool "inGeno" written in Java with Biojava subroutines. RESULTS: InGeno provides a user-friendly interactive visualization platform for sequence comparisons (comprehensive reciprocal protein--protein comparisons) between complete genome sequences and all associated annotations and features. The comparison data can be acquired from several different sequence analysis programs in flexible formats. Automatic dot-plot analysis includes output reduction, filtering, ortholog testing and linear regression, followed by smart clustering (local collinear blocks; LCBs) to reveal similar genome regions. Further, the system provides genome alignment and visualization editor, collinear relationships and strain-specific islands. Specific annotations and functions are parsed, recognized, clustered, logically concatenated and visualized and summarized in reports. CONCLUSION: As shown in this study, inGeno can be applied to study and compare in particular prokaryotic genomes against each other (gram positive and negative as well as close and more distantly related species) and has been proven to be sensitive and accurate. This modular software is user-friendly and easily accommodates new routines to meet specific user-defined requirements.

Base Sequence↗

Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies.

The spliced alignment of expressed sequence data to genomic sequence has proven a key tool in the comprehensive annotation of genes in eukaryotic genomes. A novel algorithm was developed to assemble clusters of overlapping transcript alignments (ESTs and full-length cDNAs) into maximal alignment assemblies, thereby comprehensively incorporating all available transcript data and capturing subtle splicing variations. Complete and partial gene structures identified by this method were used to improve The Institute for Genomic Research Arabidopsis genome annotation (TIGR release v.4.0). The alignment assemblies permitted the automated modeling of several novel genes and >1000 alternative splicing variations as well as updates (including UTR annotations) to nearly half of the approximately 27 000 annotated protein coding genes. The algorithm of the Program to Assemble Spliced Alignments (PASA) tool is described, as well as the results of automated updates to Arabidopsis gene annotations.

Algorithms↗

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins↗

A genomic perspective on human proteases.

Over 400 human proteases documented in secondary databases can already be delineated in genomic sequence. A Genome Ontology annotation of 30585 sequences in the provisional human proteome set recognises 498 proteases, i.e. 1.6%. Homology searches against finished sequence and comparisons between mouse and zebrafish are likely to increase this total. However, the data already indicate that the mechanistic class, sequence family and domain distribution of the genomic complement of proteases is unlikely to shift significantly from that already observed in the transcript data. Genomically derived novel sequences will require bioinformatic analysis and biochemical verification. The increasing availability of annotated genomic data will enable studies of splice variants, transcriptional control, polymorphisms, pseudogenes, inactive homologues and evolution. Comparative work on complete human protease families should produce a more integrated picture of their biochemistry and physiology. Genomic data will also lead to the identification of new protease involvement in disease processes and their evaluation as drug targets.

Alternative Splicing↗

Reliability and reproducibility issues in DNA microarray measurements.

DNA microarrays enable researchers to monitor the expression of thousands of genes simultaneously. However, the current technology has several limitations. Here we discuss problems related to the sensitivity, accuracy, specificity and reproducibility of microarray results. The existing data suggest that for relatively abundant transcripts the existence and direction (but not the magnitude) of expression changes can be reliably detected. However, accurate measurements of absolute expression levels and the reliable detection of low abundance genes are difficult to achieve. The main problems seem to be the sub-optimal design or choice of probes and some incorrect probe annotations. Well-designed data-analysis approaches can rectify some of these problems.

Animals↗

Toward metabolic phenomics: analysis of genomic data using flux balances.

Small genome sequencing and annotations are leading to the definition of metabolic genotypes in an increasing number of organisms. Proteomics is beginning to give insights into the use of the metabolic genotype under given growth conditions. These data sets give the basis for systemically studying the genotype-phenotype relationship. Methods of systems science need to be employed to analyze, interpret, and predict this complex relationship. These endeavors will lead to the development of a new field, tentatively named phenomics. This article illustrates how the metabolic characteristics of annotated small genomes can be analyzed using flux balance analysis (FBA). A general algorithm for the formulation of in silico metabolic genotypes is described. Illustrative analyses of the in silico Escherichia coli K-12 metabolic genotypes are used to show how FBA can be used to study the capabilities of this strain.

Biotechnology↗

Comprehensive gene expression analysis by transcript profiling.

After the completion of the genomic sequence of Arabidopsis thaliana, it is now a priority to identify all the genes, their patterns of expression and functions. Transcript profiling is playing a substantial role in annotating and determining gene functions, having advanced from one-gene-at-a-time methods to technologies that provide a holistic view of the genome. In this review, comprehensive transcript profiling methodologies are described, including two that are used extensively by the authors, cDNA-AFLP and cDNA microarraying. Both these technologies illustrate the requirement to integrate molecular biology, automation, LIMS and data analysis. With so much uncharted territory in the Arabidopsis genome, and the desire to tackle complex biological traits, such integrated systems will provide a rich source of data for the correlative, functional annotation of genes.

Gene Expression Profiling↗

An integrated gene annotation and transcriptional profiling approach towards the full gene content of the Drosophila genome.

BACKGROUND: While the genome sequences for a variety of organisms are now available, the precise number of the genes encoded is still a matter of debate. For the human genome several stringent annotation approaches have resulted in the same number of potential genes, but a careful comparison revealed only limited overlap. This indicates that only the combination of different computational prediction methods and experimental evaluation of such in silico data will provide more complete genome annotations. In order to get a more complete gene content of the Drosophila melanogaster genome, we based our new D. melanogaster whole-transcriptome microarray, the Heidelberg FlyArray, on the combination of the Berkeley Drosophila Genome Project (BDGP) annotation and a novel ab initio gene prediction of lower stringency using the Fgenesh software. RESULTS: Here we provide evidence for the transcription of approximately 2,600 additional genes predicted by Fgenesh. Validation of the developmental profiling data by RT-PCR and in situ hybridization indicates a lower limit of 2,000 novel annotations, thus substantially raising the number of genes that make a fly. CONCLUSIONS: The successful design and application of this novel Drosophila microarray on the basis of our integrated in silico/wet biology approach confirms our expectation that in silico approaches alone will always tend to be incomplete. The identification of at least 2,000 novel genes highlights the importance of gathering experimental evidence to discover all genes within a genome. Moreover, as such an approach is independent of homology criteria, it will allow the discovery of novel genes unrelated to known protein families or those that have not been strictly conserved between species.

Animals↗

The proteomics of N-terminal methionine cleavage.

Methionine aminopeptidase (MAP) is a ubiquitous, essential enzyme involved in protein N-terminal methionine excision. According to the generally accepted cleavage rules for MAP, this enzyme cleaves all proteins with small side chains on the residue in the second position (P1'), but many exceptions are known. The substrate specificity of Escherichia coli MAP1 was studied in vitro with a large (>120) coherent array of peptides mimicking the natural substrates and kinetically analyzed in detail. Peptides with Val or Thr at P1' were much less efficiently cleaved than those with Ala, Cys, Gly, Pro, or Ser in this position. Certain residues at P2', P3', and P4' strongly slowed the reaction, and some proteins with Val and Thr at P1' could not undergo Met cleavage. These in vitro data were fully consistent with data for 862 E. coli proteins with known N-terminal sequences in vivo. The specificity sites were found to be identical to those for the other type of MAPs, MAP2s, and a dedicated prediction tool for Met cleavage is now available. Taking into account the rules of MAP cleavage and leader peptide removal, the N termini of all proteins were predicted from the annotated genome and compared with data obtained in vivo. This analysis showed that proteins displaying N-Met cleavage are overrepresented in vivo. We conclude that protein secretion involving leader peptide cleavage is more frequent than generally thought.

Amino Acids↗

AffyMiner: mining differentially expressed genes and biological knowledge in GeneChip microarray data.

BACKGROUND: DNA microarrays are a powerful tool for monitoring the expression of tens of thousands of genes simultaneously. With the advance of microarray technology, the challenge issue becomes how to analyze a large amount of microarray data and make biological sense of them. Affymetrix GeneChips are widely used microarrays, where a variety of statistical algorithms have been explored and used for detecting significant genes in the experiment. These methods rely solely on the quantitative data, i.e., signal intensity; however, qualitative data are also important parameters in detecting differentially expressed genes. RESULTS: AffyMiner is a tool developed for detecting differentially expressed genes in Affymetrix GeneChip microarray data and for associating gene annotation and gene ontology information with the genes detected. AffyMiner consists of the functional modules, GeneFinder for detecting significant genes in a treatment versus control experiment and GOTree for mapping genes of interest onto the Gene Ontology (GO) space; and interfaces to run Cluster, a program for clustering analysis, and GenMAPP, a program for pathway analysis. AffyMiner has been used for analyzing the GeneChip data and the results were presented in several publications. CONCLUSION: AffyMiner fills an important gap in finding differentially expressed genes in Affymetrix GeneChip microarray data. AffyMiner effectively deals with multiple replicates in the experiment and takes into account both quantitative and qualitative data in identifying significant genes. AffyMiner reduces the time and effort needed to compare data from multiple arrays and to interpret the possible biological implications associated with significant changes in a gene's expression.

Algorithms↗

CMAtlas: a comprehensive DNA methylation atlas for exploring epigenetic alterations in 34 human cancer types.

MOTIVATION: Aberrant DNA methylation is a fundamental epigenetic hallmark of cancer. However, existing resources often lack technological diversity and comprehensive cancer coverage. Furthermore, most platforms fail to achieve deep multi-omics integration and tend to ignore cancer-type-specific methylation features, limiting their utility in precision oncology and drug discovery. RESULTS: We developed Cancer Methylation Atlas (CMAtlas), a comprehensive platform integrating 13&#xa0;753 samples across 34 cancer types. By applying technology-tailored pipelines to data from various profiling technologies, we identified 830&#xa0;725 tumor-specific differentially methylated elements (DMEs) and 1&#xa0;480&#xa0;098 differentially methylated regions (DMRs), alongside 1&#xa0;154&#xa0;256 cancer-type-specific DMEs and 329&#xa0;154 DMRs. The platform demonstrates high cross-platform consistency and strong concordance between tumor tissues and cell lines, ensuring the robustness of our findings. All DMEs and DMRs are annotated with multi-omics data (RNA expression, somatic mutations, and chromatin accessibility) and clinical relevance (survival associations and cell-free DNA profiling). We further demonstrate the utility of CMAtlas by identifying prognostic aberrant methylation in colorectal cancer driver genes. AVAILABILITY AND IMPLEMENTATION: CMAtlas is freely accessible at {{https://cmatlas.renlab.cn/}}. The platform offers an intuitive web interface supporting gene-centric and cancer-centric queries, alongside customizable analysis modules designed to facilitate user-specific research needs.

Humans↗

Gene discovery in Plasmodium vivax through sequencing of ESTs from mixed blood stages.

Despite the significance of Plasmodium vivax as the most widespread human malaria parasite and a major public health problem, gene expression in this parasite is poorly understood. To accelerate gene discovery and facilitate the annotation phase of the P. vivax genome project, we have undertaken a transcriptome approach to study gene expression in the mixed blood stages of a P. vivax field isolate. Using a cDNA library constructed from purified blood stages, we have obtained single-pass sequences for approximately 21,500 expressed sequence tags (ESTs), the largest number of transcript tags obtained so far for this species. Cluster analysis revealed that the library is highly redundant, resulting in 5407 clusters. Clustered ESTs were searched against public protein databases for functional annotation, and more than one-third showed a significant match, the majority of these to Plasmodium falciparum proteins. The most abundant clusters were to genes encoding ribosomal proteins and proteins involved in metabolism, consistent with the predominance of trophozoites in the field isolate sample. In spite of the scarcity of other parasite stages in the field isolate, we could identify genes that are expressed in rings, schizonts and gametocytes. This study should facilitate our understanding of the gene expression in P. vivax asexual stages and provide valuable data for gene prediction and annotation of the P. vivax genome sequence.

Animals↗

IMGT gene identification and Colliers de Perles of human immunoglobulins with known 3D structures.

A new database, IMGT/3Dstructure-DB, was developed and implemented in the IMGT (international ImMunoGeneTics database) information system (http://imgt.cines.fr) to provide a unique expertised resource on immunoglobulin and T-cell receptor structural data. Corresponding protein sequences were annotated with IMGT tools, which allow the precise identification of the genes expressed in these proteins, and the description of framework and complementarity determining regions according to the IMGT standardized nomenclature and IMGT unique numbering. Two-dimensional graphical representations of the V-DOMAINs, designated as Colliers de Perles, are automatically produced. A query Web interface allows interactive search of the IMGT/3D structure-DB data. In this article, IMGT gene identification and Colliers de Perles of human immunoglobulins with known 3D structures in the Protein Data Bank are presented.

Alleles↗