Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data annotation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

Predicting gene function from gene expressions and ontologies.

We introduce a methodology for inducing predictive rule models for functional classification of gene expressions from microarray hybridisation experiments. The basic learning method is the rough set framework for rule induction. The methodology is different from the commonly used unsupervised clustering approaches in that it exploits background knowledge of gene function in a supervised manner. Genes are annotated using Ashburner's Gene Ontology and the functional classes used for learning are mined from these annotations. From the original expression data, we extract a set of biologically meaningful features that are used for learning. A rule model is induced from the data described in terms of these features. Its predictive quality is fine-turned via cross-validation on subsets of the known genes prior to classification of unknown genes. The predictive and descriptive quality of such a rule model is demonstrated on the fibroblast serum response data previously analysed by Iyer et. al. Our analysis shows that the rules are capable of representing the complex relationship between gene expressions and function, and that it is possible to put forward high quality hypotheses about the function of unknown genes.

Algorithms↗

PineappleDB: an online pineapple bioinformatics resource.

BACKGROUND: A world first pineapple EST sequencing program has been undertaken to investigate genes expressed during non-climacteric fruit ripening and the nematode-plant interaction during root infection. Very little is known of how non-climacteric fruit ripening is controlled or of the molecular basis of the nematode-plant interaction. PineappleDB was developed to provide the research community with access to a curated bioinformatics resource housing the fruit, root and nematode infected gall expressed sequences. DESCRIPTION: PineappleDB is an online, curated database providing integrated access to annotated expressed sequence tag (EST) data for cDNA clones isolated from pineapple fruit, root, and nematode infected root gall vascular cylinder tissues. The database currently houses over 5600 EST sequences, 3383 contig consensus sequences, and associated bioinformatic data including splice variants, Arabidopsis homologues, both MIPS based and Gene Ontology functional classifications, and clone distributions. The online resource can be searched by text or by BLAST sequence homology. The data outputs provide comprehensive sequence, bioinformatic and functional classification information. CONCLUSION: The online pineapple bioinformatic resource provides the research community with access to pineapple fruit and root/gall sequence and bioinformatic data in a user-friendly format. The search tools enable efficient data mining and present a wide spectrum of bioinformatic and functional classification information. PineappleDB will be of broad appeal to researchers investigating pineapple genetics, non-climacteric fruit ripening, root-knot nematode infection, crassulacean acid metabolism and alternative RNA splicing in plants.

Alternative Splicing↗

Differential gene expression profile in omental adipose tissue in women with polycystic ovary syndrome.

CONTEXT: The polycystic ovary syndrome (PCOS) is frequently associated with visceral obesity, suggesting that omental adipose tissue might play an important role in the pathogenesis of the syndrome. OBJECTIVE: The objective was to study the expression profiles of omental fat biopsy samples obtained from morbidly obese women with or without PCOS at the time of bariatric surgery. DESIGN: This was a case-control study. SETTINGS: We conducted the study in an academic hospital. PATIENTS: Eight PCOS patients and seven nonhyperandrogenic women submitted to bariatric surgery because of morbid obesity. INTERVENTIONS: Biopsy samples of omental fat were obtained during bariatric surgery. MAIN OUTCOME MEASURE: The main outcome measure was high-density oligonucleotide arrays. RESULTS: After statistical analysis, we identified changes in the expression patterns of 63 genes between PCOS and control samples. Gene classification was assessed through data mining of Gene Ontology annotations and cluster analysis of dysregulated genes between both groups. These methods highlighted abnormal expression of genes encoding certain components of several biological pathways related to insulin signaling and Wnt signaling, oxidative stress, inflammation, immune function, and lipid metabolism, as well as other genes previously related to PCOS or to the metabolic syndrome. CONCLUSION: The differences in the gene expression profiles in visceral adipose tissue of PCOS patients compared with nonhyperandrogenic women involve multiple genes related to several biological pathways, suggesting that the involvement of abdominal obesity in the pathogenesis of PCOS is more ample than previously thought and is not restricted to the induction of insulin resistance.

Adipose Tissue↗

Similarities and differences in genome-wide expression data of six organisms.

Comparing genomic properties of different organisms is of fundamental importance in the study of biological and evolutionary principles. Although differences among organisms are often attributed to differential gene expression, genome-wide comparative analysis thus far has been based primarily on genomic sequence information. We present a comparative study of large datasets of expression profiles from six evolutionarily distant organisms: S. cerevisiae, C. elegans, E. coli, A. thaliana, D. melanogaster, and H. sapiens. We use genomic sequence information to connect these data and compare global and modular properties of the transcription programs. Linking genes whose expression profiles are similar, we find that for all organisms the connectivity distribution follows a power-law, highly connected genes tend to be essential and conserved, and the expression program is highly modular. We reveal the modular structure by decomposing each set of expression data into coexpressed modules. Functionally related sets of genes are frequently coexpressed in multiple organisms. Yet their relative importance to the transcription program and their regulatory relationships vary among organisms. Our results demonstrate the potential of combining sequence and expression data for improving functional gene annotation and expanding our understanding of how gene expression and diversity evolved.

Animals↗

Evidence of a large-scale functional organization of mammalian chromosomes.

Evidence from inbred strains of mice indicates that a quarter or more of the mammalian genome consists of chromosome regions containing clusters of functionally related genes. The intense selection pressures during inbreeding favor the coinheritance of optimal sets of alleles among these genetically linked, functionally related genes, resulting in extensive domains of linkage disequilibrium (LD) among a set of 60 genetically diverse inbred strains. Recombination that disrupts the preferred combinations of alleles reduces the ability of offspring to survive further inbreeding. LD is also seen between markers on separate chromosomes, forming networks with scale-free architecture. Combining LD data with pathway and genome annotation databases, we have been able to identify the biological functions underlying several domains and networks. Given the strong conservation of gene order among mammals, the domains and networks we find in mice probably characterize all mammals, including humans.

Animals↗

Gene expression profiles of mouse aorta and cultured vascular smooth muscle cells differ widely, yet show common responses to dioxin exposure.

Exposure to environmental toxicants may play a role in the onset and progression of cardiovascular disease. Many environmental agents, such as dioxin, are risk factors for atherosclerosis because they may exacerbate an underlying disease by altering gene expression patterns. Expression profiling of vascular tissues allows the simultaneous analysis of thousands of genes and may provide predictive information particularly useful in early disease stages. Often, however, in vivo experiments are unfeasible for material or ethical reasons, and data from cultured cells must be used instead, even though it may not be known whether cultured cells and live tissues share common global responses to the same toxicant. In a search for genes responsive to dioxin exposure, we used oligonucleotide microarrays with DNA sequences from 13,433 genes to compare global gene expression profiles of C57BL/6 mice aortas with cultured vascular smooth muscle cells (vSMCs) of the same mice. Aorta segments and vSMCs differed in the expression of more than 4500 genes, many showing expression differences greater than 1000-fold. Integration of microarray data into Gene Ontology Project annotations showed that many of the genes differentially expressed belonged to the same biological process or metabolic pathway. Notwithstanding these results, a subset of 35 genes responded in the same fashion to dioxin exposure in both systems. Genes in this subset encoded phase I and phase II detoxification enzymes, signal transduction kinases and phosphatases, and proteins involved in DNA repair and the cell cycle. We conclude that vSMCS may be useful aorta surrogates to study early gene expression responses to dioxin exposure, provided that analyses focus on this subset of genes.

Animals↗

Integrated multi-omics identification of m6A-SNP-related diagnostic biomarkers in amyotrophic lateral sclerosis.

BACKGROUND: Amyotrophic lateral sclerosis (ALS) lacks reliable and minimally invasive biomarkers for early diagnosis. m6A-associated single-nucleotide polymorphisms (m6A-SNPs) may influence RNA methylation and gene expression, offering opportunities to identify clinically relevant diagnostic markers. METHODS: We integrated eQTLGen cis-eQTL data, RMVar m6A-SNP annotations, and ALS transcriptomic datasets to identify m6A-SNP-related genes. Random Forest and LASSO regression were combined to screen robust diagnostic markers. A nomogram was constructed and validated using independent cohorts. Immune infiltration, predicted m6A modification sites, and potential RBP-SNP interactions were assessed. Peripheral blood samples from ALS patients were used for exploratory validation of gene expression and global m6A levels. RESULTS: We identified 109 ALS-associated m6A-SNP-related genes with cis-eQTL signals and narrowed these to seven candidate diagnostic markers (TMED5, OXR1, BRI3, FEM1C, SUZ12, EIF2AK4, and TJAP1). The seven-gene model outperformed the individual markers in the training cohort and retained moderate discrimination in the independent validation cohort. ALS samples showed differences in inferred immune-cell composition, including monocytes, neutrophils, and T-cell subsets. The selected SNP loci were located near predicted m6A sites and annotated RBP-binding regions. Exploratory clinical validation showed significant upregulation of FEM1C and SUZ12 at both mRNA and protein levels, accompanied by reduced global m6A modification. CONCLUSIONS: Through multi-omics integration and exploratory clinical validation, this study identifies m6A-SNP-related candidate markers associated with ALS. The findings support further evaluation of m6A-related signatures for ALS discrimination and molecular characterization, while larger independent cohorts and additional calibration are required before clinical application.

Humans↗

Novel candidate targets of beta-catenin/T-cell factor signaling identified by gene expression profiling of ovarian endometrioid adenocarcinomas.

The activity of beta-catenin (beta-cat), a key component of the Wnt signaling pathway, is deregulated in about 40% of ovarian endometrioid adenocarcinomas (OEAs), usually as a result of CTNNB1 gene mutations. The function of beta-cat in neoplastic transformation is dependent on T-cell factor (TCF) transcription factors, but specific genes activated by the interaction of beta-cat with TCFs in OEAs and other cancers with Wnt pathway defects are largely unclear. As a strategy to identify beta-cat/TCF transcriptional targets likely to contribute to OEA pathogenesis, we used oligonucleotide microarrays to compare gene expression in primary OEAs with mutational defects in beta-cat regulation (n = 11) to OEAs with intact regulation of beta-cat activity (n = 17). Both hierarchical clustering and principal component analysis based on global gene expression distinguished beta-cat-defective tumors from those with intact beta-cat regulation. We identified 81 potential beta-cat/TCF targets by selecting genes with at least 2-fold increased expression in beta-cat-defective versus beta-cat regulation-intact tumors and significance in a t test (P < 0.05). Seven of the 81 genes have been previously reported as Wnt/beta-cat pathway targets (i.e., BMP4, CCND1, CD44, FGF9, EPHB3, MMP7, and MSX2). Differential expression of several known and candidate target genes in the OEAs was confirmed. For the candidate target genes CST1 and EDN3, reporter and chromatin immunoprecipitation assays directly implicated beta-cat and TCF in their regulation. Analysis of presumptive regulatory elements in 67 of the 81 candidate genes for which complete genomic sequence data were available revealed an apparent difference in the location and abundance of consensus TCF-binding sites compared with the patterns seen in control genes. Our findings imply that analysis of gene expression profiling data from primary tumor samples annotated with detailed molecular information may be a powerful approach to identify key downstream targets of signaling pathways defective in cancer cells.

Binding Sites↗

[System of distributed storage and analysis of genomic information].

A distributed computing system is developed to search and analyze genetic databases using parallel computing technologies. Queries are processed by a local network PC cluster. A universal task and data exchange format is developed for effective query processing. A multilevel hierarchic task batching procedure is elaborated to generate multiple subtasks and distribute them over cluster units under dynamic priority levels and with dynamic distribution of replicated source data subbases. Primary source data preparation and generation of annotation word indices are used to significantly reduce query processing time.

Databases, Genetic↗

BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis.

biomaRt is a new Bioconductor package that integrates BioMart data resources with data analysis software in Bioconductor. It can annotate a wide range of gene or gene product identifiers (e.g. Entrez-Gene and Affymetrix probe identifiers) with information such as gene symbol, chromosomal coordinates, Gene Ontology and OMIM annotation. Furthermore biomaRt enables retrieval of genomic sequences and single nucleotide polymorphism information, which can be used in data analysis. Fast and up-to-date data retrieval is possible as the package executes direct SQL queries to the BioMart databases (e.g. Ensembl). The biomaRt package provides a tight integration of large, public or locally installed BioMart databases with data analysis in Bioconductor creating a powerful environment for biological data mining.

Algorithms↗

Functional information in SWISS-PROT: the basis for large-scale characterisation of protein sequences.

With the rapid growth of sequence databases, there is an increasing need for reliable functional characterisation and annotation of newly predicted proteins. To cope with such large data volumes, faster and more effective means of protein sequence characterisation and annotation are required. One promising approach is automatic large-scale functional characterisation and annotation, which is generated with limited human interaction. However, such an approach is heavily dependent on reliable data sources. The SWISS-PROT protein sequence database plays an essential role here owing to its high level of functional information.

Animals↗

The mouse genome database (MGD): new features facilitating a model system.

The mouse genome database (MGD, http://www.informatics.jax.org/), the international community database for mouse, provides access to extensive integrated data on the genetics, genomics and biology of the laboratory mouse. The mouse is an excellent and unique animal surrogate for studying normal development and disease processes in humans. Thus, MGD's primary goals are to facilitate the use of mouse models for studying human disease and enable the development of translational research hypotheses based on comparative genotype, phenotype and functional analyses. Core MGD data content includes gene characterization and functions, phenotype and disease model descriptions, DNA and protein sequence data, polymorphisms, gene mapping data and genome coordinates, and comparative gene data focused on mammals. Data are integrated from diverse sources, ranging from major resource centers to individual investigator laboratories and the scientific literature, using a combination of automated processes and expert human curation. MGD collaborates with the bioinformatics community on the development of data and semantic standards, and it incorporates key ontologies into the MGD annotation system, including the Gene Ontology (GO), the Mammalian Phenotype Ontology, and the Anatomical Dictionary for Mouse Development and the Adult Anatomy. MGD is the authoritative source for mouse nomenclature for genes, alleles, and mouse strains, and for GO annotations to mouse genes. MGD provides a unique platform for data mining and hypothesis generation where one can express complex queries simultaneously addressing phenotypic effects, biochemical function and process, sub-cellular location, expression, sequence, polymorphism and mapping data. Both web-based querying and computational access to data are provided. Recent improvements in MGD described here include the incorporation of single nucleotide polymorphism data and search tools, the addition of PIR gene superfamily classifications, phenotype data for NIH-acquired knockout mice, images for mouse phenotypic genotypes, new functional graph displays of GO annotations, and new orthology displays including sequence information and graphic displays.

Animals↗

The SWISS-PROT protein sequence data bank: current status.

SWISS-PROT is an annotated protein sequence database established in 1986 and maintained collaboratively, since 1988, by the Department of Medical Biochemistry of the University of Geneva and the EMBL Data Library. The SWISS-PROT protein sequence data bank consist of sequence entries. Sequence entries are composed of different lines types, each with their own format. For standardization purposes the format of SWISS-PROT follows as closely as possible that of the EMBL Nucleotide Sequence Database. A sample SWISS-PROT entry is shown in Figure 1.

Amino Acid Sequence↗

CASTp: computed atlas of surface topography of proteins with structural and topographical mapping of functionally annotated residues.

Cavities on a proteins surface as well as specific amino acid positioning within it create the physicochemical properties needed for a protein to perform its function. CASTp (http://cast.engr.uic.edu) is an online tool that locates and measures pockets and voids on 3D protein structures. This new version of CASTp includes annotated functional information of specific residues on the protein structure. The annotations are derived from the Protein Data Bank (PDB), Swiss-Prot, as well as Online Mendelian Inheritance in Man (OMIM), the latter contains information on the variant single nucleotide polymorphisms (SNPs) that are known to cause disease. These annotated residues are mapped to surface pockets, interior voids or other regions of the PDB structures. We use a semi-global pair-wise sequence alignment method to obtain sequence mapping between entries in Swiss-Prot, OMIM and entries in PDB. The updated CASTp web server can be used to study surface features, functional regions and specific roles of key residues of proteins.

Amino Acids↗

CardioSignal: a database of transcriptional regulation in cardiac development and hypertrophy.

BACKGROUND: Although extensive research has characterized intricate genetic programs in heart system, the information generated is highly fragmented. Here we have developed a new database called CardioSignal, which was designed for integration of regulatory information on the transcriptional regulation involved in heart development and cardiac hypertrophy. METHODS: Data about sequences, positions and functional annotation of transcription binding sites, cis-regulatory modules as well as promoters were collected from scientific literature. Genes involved in both processes were also manually gathered, particularly those preferentially expressed in the heart. Data was stored in MySQL database and Perl was used as the server-side programming language. RESULTS: Currently, CardioSignal contains 677 cardiac genes from twenty species. Among them are 128 cardiac transcription factors. Of the approximately 179 individual promoters from six species, the database also documented 247 experimentally verified binding sites and 64 cis-regulatory modules. CardioSignal may be searched for the promoter of a specific gene by specifying a gene name, Entrez geneID, swissProt accession number and so on. Downstream targets of transcriptional factors and cardiac regulatory modules can also be retrieved through a user-friendly web interface. Also available is experimental supporting evidence. Computational analysis tools were implemented for on-the-fly motif finding and comparative genomic analysis respectively. CONCLUSIONS: CardioSignal offers a unique resource as it contains simultaneously the promoter collected while correlating the information of transcription factor binding sites and cis-regulatory modules from heart system. We are hopeful that its implementation will contribute toward the elucidation of the complex processes in cardiac development and hypertrophy.

Animals↗

Analysis and organization of protein sequence data: a retrospective spanning four decades.

Protein sequence data are as useful and valuable today as was envisioned by pioneering sequencers and by the organizers of the first sequence database. Sequence analysis was first the province of specialists who developed search, comparison, and tree-building methods. Microcomputers, communication satellites, and the Internet have made these methods accessible to any scientist. The rapid increase in the data has driven a succession of changes in how databases are compiled, distributed, and accessed. Large public databases have become international collaborations. Although they need to develop still more efficient ways to accumulate, organize, annotate, and standardize huge amounts of data, inadequate support is available for such efforts. Thus there will be greater reliance on direct input from the scientific community. The World Wide Web is essential but not sufficient for integrated access to related databases.

Amino Acid Sequence↗

iModMix: integrative module analysis for multi-omics data.

SUMMARY: Integrative Module Analysis for Multi-omics Data (iModMix) is a biology-agnostic framework that enables the discovery of novel associations across any type of quantitative abundance data, including but not limited to transcriptomics, proteomics, and metabolomics. Instead of relying on pathway annotations or prior biological knowledge, iModMix constructs data-driven modules using graphical lasso to estimate sparse networks from omics features. These modules are summarized into eigenfeatures and correlated across datasets for horizontal integration, while preserving the distinct feature sets and interpretability of each omics type. iModMix operates directly on matrices containing expression or abundances for a wide range of features, including but not limited to genes, proteins, and metabolites. Because it does not rely on annotations (e.g., KEGG identifiers), it can seamlessly incorporate both identified and unidentified metabolites, addressing a key limitation of many existing metabolomics tools. iModMix is available as a user-friendly R Shiny application requiring no programming expertise (https://imodmix.moffitt.org), and as a Bioconductor R package for advanced users (https://bioconductor.org/packages/release/bioc/html/iModMix.html). The tool includes several public and in-house datasets to illustrate its utility in identifying novel multi-omics relationships in diverse biological contexts. AVAILABILITY AND IMPLEMENTATION: iModMix is freely available from Bioconductor (https://bioconductor.org/packages/release/bioc/html/iModMix.html), and the example dataset package (iModMixData) is also available from Bioconductor (https://bioconductor.org/packages/release/ data/experiment/html/iModMixData.html). The R package source code and Docker are available from GitHub: https://github.com/biodatalab/iModMix. Shiny application can be accessed at: https://imodmix.moffitt.org.

Multiomics↗

Systematic gene function prediction from gene expression data by using a fuzzy nearest-cluster method.

BACKGROUND: Quantitative simultaneous monitoring of the expression levels of thousands of genes under various experimental conditions is now possible using microarray experiments. However, there are still gaps toward whole-genome functional annotation of genes using the gene expression data. RESULTS: In this paper, we propose a novel technique called Fuzzy Nearest Clusters for genome-wide functional annotation of unclassified genes. The technique consists of two steps: an initial hierarchical clustering step to detect homogeneous co-expressed gene subgroups or clusters in each possibly heterogeneous functional class; followed by a classification step to predict the functional roles of the unclassified genes based on their corresponding similarities to the detected functional clusters. CONCLUSION: Our experimental results with yeast gene expression data showed that the proposed method can accurately predict the genes' functions, even those with multiple functional roles, and the prediction performance is most independent of the underlying heterogeneity of the complex functional classes, as compared to the other conventional gene function prediction approaches.

Algorithms↗