Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Cross-species global and subset gene expression profiling identifies genes involved in prostate cancer response to selenium.

BACKGROUND: Gene expression technologies have the ability to generate vast amounts of data, yet there often resides only limited resources for subsequent validation studies. This necessitates the ability to perform sorting and prioritization of the output data. Previously described methodologies have used functional pathways or transcriptional regulatory grouping to sort genes for further study. In this paper we demonstrate a comparative genomics based method to leverage data from animal models to prioritize genes for validation. This approach allows one to develop a disease-based focus for the prioritization of gene data, a process that is essential for systems that lack significant functional pathway data yet have defined animal models. This method is made possible through the use of highly controlled spotted cDNA slide production and the use of comparative bioinformatics databases without the use of cross-species slide hybridizations. RESULTS: Using gene expression profiling we have demonstrated a similar whole transcriptome gene expression patterns in prostate cancer cells from human and rat prostate cancer cell lines both at baseline expression levels and after treatment with physiologic concentrations of the proposed chemopreventive agent Selenium. Using both the human PC3 and rat PAII prostate cancer cell lines have gone on to identify a subset of one hundred and fifty-four genes that demonstrate a similar level of differential expression to Selenium treatment in both species. Further analysis and data mining for two genes, the Insulin like Growth Factor Binding protein 3, and Retinoic X Receptor alpha, demonstrates an association with prostate cancer, functional pathway links, and protein-protein interactions that make these genes prime candidates for explaining the mechanism of Selenium's chemopreventive effect in prostate cancer. These genes are subsequently validated by western blots showing Selenium based induction and using tissue microarrays to demonstrate a significant association between downregulated protein expression and tumorigenesis, a process that is the reverse of what is seen in the presence of Selenium. CONCLUSIONS: Thus the outlined process demonstrates similar baseline and selenium induced gene expression profiles between rat and human prostate cancers, and provides a method for identifying testable functional pathways for the action of Selenium's chemopreventive properties in prostate cancer.

Adenocarcinoma↗

GOLD.db: genomics of lipid-associated disorders database.

BACKGROUND: The GOLD.db (Genomics of Lipid-Associated Disorders Database) was developed to address the need for integrating disparate information on the function and properties of genes and their products that are particularly relevant to the biology, diagnosis management, treatment, and prevention of lipid-associated disorders. DESCRIPTION: The GOLD.db http://gold.tugraz.at provides a reference for pathways and information about the relevant genes and proteins in an efficiently organized way. The main focus was to provide biological pathways with image maps and visual pathway information for lipid metabolism and obesity-related research. This database provides also the possibility to map gene expression data individually to each pathway. Gene expression at different experimental conditions can be viewed sequentially in context of the pathway. Related large scale gene expression data sets were provided and can be searched for specific genes to integrate information regarding their expression levels in different studies and conditions. Analytic and data mining tools, reagents, protocols, references, and links to relevant genomic resources were included in the database. Finally, the usability of the database was demonstrated using an example about the regulation of Pten mRNA during adipocyte differentiation in the context of relevant pathways. CONCLUSIONS: The GOLD.db will be a valuable tool that allow researchers to efficiently analyze patterns of gene expression and to display them in a variety of useful and informative ways, allowing outside researchers to perform queries pertaining to gene expression results in the context of biological processes and pathways.

Adipocytes↗

CMD: a Cotton Microsatellite Database resource for Gossypium genomics.

BACKGROUND: The Cotton Microsatellite Database (CMD) http://www.cottonssr.org is a curated and integrated web-based relational database providing centralized access to publicly available cotton microsatellites, an invaluable resource for basic and applied research in cotton breeding. DESCRIPTION: At present CMD contains publication, sequence, primer, mapping and homology data for nine major cotton microsatellite projects, collectively representing 5,484 microsatellites. In addition, CMD displays data for three of the microsatellite projects that have been screened against a panel of core germplasm. The standardized panel consists of 12 diverse genotypes including genetic standards, mapping parents, BAC donors, subgenome representatives, unique breeding lines, exotic introgression sources, and contemporary Upland cottons with significant acreage. A suite of online microsatellite data mining tools are accessible at CMD. These include an SSR server which identifies microsatellites, primers, open reading frames, and GC-content of uploaded sequences; BLAST and FASTA servers providing sequence similarity searches against the existing cotton SSR sequences and primers, a CAP3 server to assemble EST sequences into longer transcripts prior to mining for SSRs, and CMap, a viewer for comparing cotton SSR maps. CONCLUSION: The collection of publicly available cotton SSR markers in a centralized, readily accessible and curated web-enabled database provides a more efficient utilization of microsatellite resources and will help accelerate basic and applied research in molecular breeding and genetic mapping in Gossypium spp.

Chromosome Mapping↗

Survey and analysis of microsatellites from transcript sequences in Phytophthora species: frequency, distribution, and potential as markers for the genus.

BACKGROUND: Members of the genus Phytophthora are notorious pathogens with world-wide distribution. The most devastating species include P. infestans, P. ramorum and P. sojae. In order to develop molecular methods for routinely characterizing their populations and to gain a better insight into the organization and evolution of their genomes, we used an in silico approach to survey and compare simple sequence repeats (SSRs) in transcript sequences from these three species. We compared the occurrence, relative abundance, relative density and cross-species transferability of the SSRs in these oomycetes. RESULTS: The number of SSRs in oomycetes transcribed sequences is low and long SSRs are rare. The in silico transferability of SSRs among the Phytophthora species was analyzed for all sets generated, and primers were selected on the basis of similarity as possible candidates for transferability to other Phytophthora species. Sequences encoding putative pathogenicity factors from all three Phytophthora species were also surveyed for presence of SSRs. However, no correlation between gene function and SSR abundance was observed. The SSR survey results, and the primer pairs designed for all SSRs from the three species, were deposited in a public database. CONCLUSION: In all cases the most common SSRs were trinucleotide repeat units with low repeat numbers. A proportion (7.5%) of primers could be transferred with 90% similarity between at least two species of Phytophthora. This information represents a valuable source of molecular markers for use in population genetics, genetic mapping and strain fingerprinting studies of oomycetes, and illustrates how genomic databases can be exploited to generate data-mining filters for SSRs before experimental validation.

Codon↗

Computational and experimental analysis identifies Arabidopsis genes specifically expressed during early seed development.

BACKGROUND: Plant seeds are complex organs in which maternal tissues, embryo and endosperm, follow distinct but coordinated developmental programs. Some morphogenetic and metabolic processes are exclusively associated with seed development. The goal of this study was to explore the feasibility of incorporating the available online bioinformatics databases to discover Arabidopsis genes specifically expressed in certain organs, in our case immature seeds. RESULTS: A total of 11,032 EST sequences obtained from isolated immature seeds were used as the initial dataset (178 of them newly described here). A pilot study was performed using EST virtual subtraction followed by microarray data analysis, using the Genevestigator tool. These techniques led to the identification of 49 immature seed-specific genes. The findings were validated by RT-PCR analysis and in situ hybridization. CONCLUSION: We conclude that the combined in silico data analysis is an effective data mining strategy for the identification of tissue-specific gene expression.

Arabidopsis↗

An online database for brain disease research.

BACKGROUND: The Stanley Medical Research Institute online genomics database (SMRIDB) is a comprehensive web-based system for understanding the genetic effects of human brain disease (i.e. bipolar, schizophrenia, and depression). This database contains fully annotated clinical metadata and gene expression patterns generated within 12 controlled studies across 6 different microarray platforms. DESCRIPTION: A thorough collection of gene expression summaries are provided, inclusive of patient demographics, disease subclasses, regulated biological pathways, and functional classifications. CONCLUSION: The combination of database content, structure, and query speed offers researchers an efficient tool for data mining of brain disease complete with information such as: cross-platform comparisons, biomarkers elucidation for target discovery, and lifestyle/demographic associations to brain diseases.

Bipolar Disorder↗

Developmental gene regulation during tomato fruit ripening and in-vitro sepal morphogenesis.

BACKGROUND: Red ripe tomatoes are the result of numerous physiological changes controlled by hormonal and developmental signals, causing maturation or differentiation of various fruit tissues simultaneously. These physiological changes affect visual, textural, flavor, and aroma characteristics, making the fruit more appealing to potential consumers for seed dispersal. Developmental regulation of tomato fruit ripening has, until recently, been lacking in rigorous investigation. We previously indicated the presence of up-regulated transcription factors in ripening tomato fruit by data mining in TIGR Tomato Gene Index. In our in-vitro system, green tomato sepals cultured at 16 to 22 degrees C turn red and swell like ripening tomato fruit while those at 28 degrees C remain green. RESULTS: Here, we have further examined regulation of putative developmental genes possibly involved in tomato fruit ripening and development. Using molecular biological methods, we have determined the relative abundance of various transcripts of genes during in vitro sepal ripening and in tomato fruit pericarp at three stages of development. A number of transcripts show similar expression in fruits to RIN and PSY1, ripening-associated genes, and others show quite different expression. CONCLUSIONS: Our investigation has resulted in confirmation of some of our previous database mining results and has revealed differences in gene expression that may be important for tomato cultivar variation. We present new and intriguing information on genes that should now be studied in a more focused fashion.

Flowers↗

PineappleDB: an online pineapple bioinformatics resource.

BACKGROUND: A world first pineapple EST sequencing program has been undertaken to investigate genes expressed during non-climacteric fruit ripening and the nematode-plant interaction during root infection. Very little is known of how non-climacteric fruit ripening is controlled or of the molecular basis of the nematode-plant interaction. PineappleDB was developed to provide the research community with access to a curated bioinformatics resource housing the fruit, root and nematode infected gall expressed sequences. DESCRIPTION: PineappleDB is an online, curated database providing integrated access to annotated expressed sequence tag (EST) data for cDNA clones isolated from pineapple fruit, root, and nematode infected root gall vascular cylinder tissues. The database currently houses over 5600 EST sequences, 3383 contig consensus sequences, and associated bioinformatic data including splice variants, Arabidopsis homologues, both MIPS based and Gene Ontology functional classifications, and clone distributions. The online resource can be searched by text or by BLAST sequence homology. The data outputs provide comprehensive sequence, bioinformatic and functional classification information. CONCLUSION: The online pineapple bioinformatic resource provides the research community with access to pineapple fruit and root/gall sequence and bioinformatic data in a user-friendly format. The search tools enable efficient data mining and present a wide spectrum of bioinformatic and functional classification information. PineappleDB will be of broad appeal to researchers investigating pineapple genetics, non-climacteric fruit ripening, root-knot nematode infection, crassulacean acid metabolism and alternative RNA splicing in plants.

Alternative Splicing↗

Comparison of transcripts in Phalaenopsis bellina and Phalaenopsis equestris (Orchidaceae) flowers to deduce monoterpene biosynthesis pathway.

BACKGROUND: Floral scent is one of the important strategies for ensuring fertilization and for determining seed or fruit set. Research on plant scents has hampered mainly by the invisibility of this character, its dynamic nature, and complex mixtures of components that are present in very small quantities. Most progress in scent research, as in other areas of plant biology, has come from the use of molecular and biochemical techniques. Although volatile components have been identified in several orchid species, the biosynthetic pathways of orchid flower fragrance are far from understood. We investigated how flower fragrance was generated in certain Phalaenopsis orchids by determining the chemical components of the floral scent, identifying floral expressed-sequence-tags (ESTs), and deducing the pathways of floral scent biosynthesis in Phalaneopsis bellina by bioinformatics analysis. RESULTS: The main chemical components in the P. bellina flower were shown by gas chromatography-mass spectrometry to be monoterpenoids, benzenoids and phenylpropanoids. The set of floral scent producing enzymes in the biosynthetic pathway from glyceraldehyde-3-phosphate (G3P) to geraniol and linalool were recognized through data mining of the P. bellina floral EST database (dbEST). Transcripts preferentially expressed in P. bellina were distinguished by comparing the scent floral dbEST to that of a scentless species, P. equestris, and included those encoding lipoxygenase, epimerase, diacylglycerol kinase and geranyl diphosphate synthase. In addition, EST filtering results showed that transcripts encoding signal transduction and Myb transcription factors and methyltransferase, in addition to those for scent biosynthesis, were detected by in silico hybridization of the P. bellina unigene database against those of the scentless species, rice and Arabidopsis. Altogether, we pinpointed 66% of the biosynthetic steps from G3P to geraniol, linalool and their derivatives. CONCLUSION: This systems biology program combined chemical analysis, genomics and bioinformatics to elucidate the scent biosynthesis pathway and identify the relevant genes. It integrates the forward and reverse genetic approaches to knowledge discovery by which researchers can study non-model plants.

Acyclic Monoterpenes↗

Exploring cancer register data to find risk factors for recurrence of breast cancer--application of Canonical Correlation Analysis.

BACKGROUND: A common approach in exploring register data is to find relationships between outcomes and predictors by using multiple regression analysis (MRA). If there is more than one outcome variable, the analysis must then be repeated, and the results combined in some arbitrary fashion. In contrast, Canonical Correlation Analysis (CCA) has the ability to analyze multiple outcomes at the same time. One essential outcome after breast cancer treatment is recurrence of the disease. It is important to understand the relationship between different predictors and recurrence, including the time interval until recurrence. This study describes the application of CCA to find important predictors for two different outcomes for breast cancer patients, loco-regional recurrence and occurrence of distant metastasis and to decrease the number of variables in the sets of predictors and outcomes without decreasing the predictive strength of the model. METHODS: Data for 637 malignant breast cancer patients admitted in the south-east region of Sweden were analyzed. By using CCA and looking at the structure coefficients (loadings), relationships between tumor specifications and the two outcomes during different time intervals were analyzed and a correlation model was built. RESULTS: The analysis successfully detected known predictors for breast cancer recurrence during the first two years and distant metastasis 2-4 years after diagnosis. Nottingham Histologic Grading (NHG) was the most important predictor, while age of the patient at the time of diagnosis was not an important predictor. CONCLUSION: In cancer registers with high dimensionality, CCA can be used for identifying the importance of risk factors for breast cancer recurrence. This technique can result in a model ready for further processing by data mining methods through reducing the number of variables to important ones.

Adult↗

Surface-antigen expression profiling of B cell chronic lymphocytic leukemia: from the signature of specific disease subsets to the identification of markers with prognostic relevance.

Studies of gene expression profiling have been successfully used for the identification of molecules to be employed as potential prognosticators. In analogy with gene expression profiling, we have recently proposed a novel method to identify the immunophenotypic signature of B-cell chronic lymphocytic leukemia subsets with different prognosis, named surface-antigen expression profiling. According to this approach, surface marker expression data can be analysed by data mining tools identical to those employed in gene expression profiling studies, including unsupervised and supervised algorithms, with the aim of identifying the immunophenotypic signature of B-cell chronic lymphocytic leukemia subsets with different prognosis. Here we provide an overview of the overall strategy employed for the development of such an "outcome class-predictor" based on surface-antigen expression signatures. In addition, we will also discuss how to transfer the obtained information into the routine clinical practice by providing a flow-chart indicating how to select the most relevant antigens and build-up a prognostic scoring system by weighing each antigen according to its predictive power. Although referred to B-cell chronic lymphocytic leukemia, the methodology discussed here can be also useful in the study of diseases other than B-cell chronic lymphocytic leukemia, when the purpose is to identify novel prognostic determinants.

Journal Article↗

The Adaptive Evolution Database (TAED).

BACKGROUND: Developing an understanding of the molecular basis for the divergence of species lies at the heart of biology. The Adaptive Evolution Database (TAED) serves as a starting point to link events that occur at the same time in the evolutionary history (tree of life) of species, based upon coding sequence evolution analyzed with the Master Catalog. The Master Catalog is a collection of evolutionary models, including multiple sequence alignments, phylogenetic trees, and reconstructed ancestral sequences, for all independently evolving protein sequence modules encoded by genes in GenBank [1]. RESULTS: We have estimated from these models the ratio of nonsynonymous to synonymous nucleotide substitution (Ka/Ks), for each branch in their respective evolutionary trees of every subtree containing only chordata or only embryophyta proteins. Branches with high Ka/Ks values represent candidate episodes in the history of the family where the protein may have undergone positive selection, a phenomenon in molecular evolution where the mutant form of a gene must have conferred more fitness than the ancestral form. Such episodes are frequently associated with change in function. We have found that an unexpectedly large number of families (between 10 and 20% of those families examined) have at least one branch with a notably high Ka/Ks value (putative adaptive evolution). As a resource for biologists wishing to understand the interaction between protein sequences and the Darwinian processes that shape these sequences, we have collected these into The Adaptive Evolution Database (TAED). CONCLUSIONS: Placed in a phylogenetic perspective, candidate genes that are undergoing evolution at the same time in the same lineage can be viewed together. This framework based upon coding sequence evolution can be readily expanded to include other types of evolution. In its present form, TAED provides a resource for bioinformaticists interested in data mining and for experimental evolutionists seeking candidate examples of adaptive evolution for further experimental study.

Animals↗

The adaptive evolution database (TAED).

BACKGROUND: The Master Catalog is a collection of evolutionary families, including multiple sequence alignments, phylogenetic trees and reconstructed ancestral sequences, for all protein-sequence modules encoded by genes in GenBank. It can therefore support large-scale genomic surveys, of which we present here The Adaptive Evolution Database (TAED). In TAED, potential examples of positive adaptation are identified by high values for the normalized ratio of nonsynonymous to synonymous nucleotide substitution rates (KA/KS values) on branches of an evolutionary tree between nodes representing reconstructed ancestral sequences. RESULTS: Evolutionary trees and reconstructed ancestral sequences were extracted from the Master Catalog for every subtree containing proteins from the Chordata only or the Embryophyta only. Branches with high KA/KS values were identified. These represent candidate episodes in the history of the protein family when the protein may have undergone positive selection, where the mutant form conferred more fitness than the ancestral form. Such episodes are frequently associated with change in function. An unexpectedly large number of families (between 10% and 20% of those families examined) were found to have at least one branch with high KA/KS values above arbitrarily chosen cut-offs (1 and 0.6). Most of these survived a robustness test and were collected into TAED. CONCLUSIONS: TAED is a raw resource for bioinformaticists interested in data mining and for experimental evolutionists seeking candidate examples of adaptive evolution for further experimental study. It can be expanded to include other evolutionary information (for example changes in gene regulation or splicing) placed in a phylogenetic perspective.

Adaptation, Physiological↗

Long terminal repeat retrotransposons of Oryza sativa.

BACKGROUND: Long terminal repeat (LTR) retrotransposons constitute a major fraction of the genomes of higher plants. For example, retrotransposons comprise more than 50% of the maize genome and more than 90% of the wheat genome. LTR retrotransposons are believed to have contributed significantly to the evolution of genome structure and function. The genome sequencing of selected experimental and agriculturally important species is providing an unprecedented opportunity to view the patterns of variation existing among the entire complement of retrotransposons in complete genomes. RESULTS: Using a new data-mining program, LTR_STRUC, (LTR retrotransposon structure program), we have mined the GenBank rice (Oryza sativa) database as well as the more extensive (259 Mb) Monsanto rice dataset for LTR retrotransposons. Almost two-thirds (37) of the 59 families identified consist of copia-like elements, but gypsy-like elements outnumber copia-like elements by a ratio of approximately 2:1. At least 17% of the rice genome consists of LTR retrotransposons. In addition to the ubiquitous gypsy- and copia-like classes of LTR retrotransposons, the rice genome contains at least two novel families of unusually small, non-coding (non-autonomous) LTR retrotransposons. CONCLUSIONS: Each of the major clades of rice LTR retrotransposons is more closely related to elements present in other species than to the other clades of rice elements, suggesting that horizontal transfer may have occurred over the evolutionary history of rice LTR retrotransposons. Like LTR retrotransposons in other species with relatively small genomes, many rice LTR retrotransposons are relatively young, indicating a high rate of turnover.

Animals↗

Long terminal repeat retrotransposons of Mus musculus.

BACKGROUND: Long terminal repeat (LTR) retrotransposons make up a large fraction of the typical mammalian genome. They comprise about 8% of the human genome and approximately 10% of the mouse genome. On account of their abundance, LTR retrotransposons are believed to hold major significance for genome structure and function. Recent advances in genome sequencing of a variety of model organisms has provided an unprecedented opportunity to evaluate better the diversity of LTR retrotransposons resident in eukaryotic genomes. RESULTS: Using a new data-mining program, LTR_STRUC, in conjunction with conventional techniques, we have mined the GenBank mouse (Mus musculus) database and the more complete Ensembl mouse dataset for LTR retrotransposons. We report here that the M. musculus genome contains at least 21 separate families of LTR retrotransposons; 13 of these families are described here for the first time. CONCLUSIONS: All families of mouse LTR retrotransposons are members of the gypsy-like superfamily of retroviral-like elements. Several different families of unrelated non-autonomous elements were identified, suggesting that the evolution of non-autonomy may be a common event. High sequence similarity between several LTR retrotransposons identified in this study and those found in distantly-related species suggests that horizontal transfer has been a significant factor in the evolution of mouse LTR retrotransposons.

Animals↗

An Ambystoma mexicanum EST sequencing project: analysis of 17,352 expressed sequence tags from embryonic and regenerating blastema cDNA libraries.

BACKGROUND: The ambystomatid salamander, Ambystoma mexicanum (axolotl), is an important model organism in evolutionary and regeneration research but relatively little sequence information has so far been available. This is a major limitation for molecular studies on caudate development, regeneration and evolution. To address this lack of sequence information we have generated an expressed sequence tag (EST) database for A. mexicanum. RESULTS: Two cDNA libraries, one made from stage 18-22 embryos and the other from day-6 regenerating tail blastemas, generated 17,352 sequences. From the sequenced ESTs, 6,377 contigs were assembled that probably represent 25% of the expressed genes in this organism. Sequence comparison revealed significant homology to entries in the NCBI non-redundant database. Further examination of this gene set revealed the presence of genes involved in important cell and developmental processes, including cell proliferation, cell differentiation and cell-cell communication. On the basis of these data, we have performed phylogenetic analysis of key cell-cycle regulators. Interestingly, while cell-cycle proteins such as the cyclin B family display expected evolutionary relationships, the cyclin-dependent kinase inhibitor 1 gene family shows an unusual evolutionary behavior among the amphibians. CONCLUSIONS: Our analysis reveals the importance of a comprehensive sequence set from a representative of the Caudata and illustrates that the EST sequence database is a rich source of molecular, developmental and regeneration studies. To aid in data mining, the ESTs have been organized into an easily searchable database that is freely available online.

Ambystoma↗

Using ontologies to describe mouse phenotypes.

The mouse is an important model of human genetic disease. Describing phenotypes of mutant mice in a standard, structured manner that will facilitate data mining is a major challenge for bioinformatics. Here we describe a novel, compositional approach to this problem which combines core ontologies from a variety of sources. This produces a framework with greater flexibility, power and economy than previous approaches. We discuss some of the issues this approach raises.

Animals↗

Anatomical ontologies: names and places in biology.

Ontology has long been the preserve of philosophers and logicians. Recently, ideas from this field have been picked up by computer scientists as a basis for encoding knowledge and with the hope of achieving interoperability and intelligent system behavior. In bioinformatics, ontologies might allow hitherto impossible query and data-mining activities. We review the use of anatomy ontologies to represent space in biological organisms, specifically mouse and human.

Anatomy↗