Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,171 records · Page 65Linked to original sources

HAGR: the Human Ageing Genomic Resources.

The Human Ageing Genomic Resources (HAGR) is a collection of online resources for studying the biology of human ageing. HAGR features two main databases: GenAge and AnAge. GenAge is a curated database of genes related to human ageing. Entries were primarily selected based on genetic perturbations in animal models and human diseases as well as an extensive literature review. Each entry includes a variety of automated and manually curated information, including, where available, protein-protein interactions, the relevant literature, and a description of the gene and how it relates to human ageing. The goal of GenAge is to provide the most complete and comprehensive database of genes related to human ageing on the Internet as well as render an overview of the genetics of human ageing. AnAge is an integrative database describing the ageing process in several organisms and featuring, if available, maximum life span, taxonomy, developmental schedules and metabolic rate, making AnAge a unique resource for the comparative biology of ageing. Associated with the databases are data-mining tools and software designed to investigate the role of genes and proteins in the human ageing process as well as analyse ageing across different taxa. HAGR is freely available to the academic community at http://genomics.senescence.info.

Aging↗

WormBase: a comprehensive data resource for Caenorhabditis biology and genomics.

WormBase (http://www.wormbase.org), the model organism database for information about Caenorhabditis elegans and related nematodes, continues to expand in breadth and depth. Over the past year, WormBase has added multiple large-scale datasets including SAGE, interactome, 3D protein structure datasets and NCBI KOGs. To accommodate this growth, the International WormBase Consortium has improved the user interface by adding new features to aid in navigation, visualization of large-scale datasets, advanced searching and data mining. Internally, we have restructured the database models to rationalize the representation of genes and to prepare the system to accept the genome sequences of three additional Caenorhabditis species over the coming year.

Animals↗

Nuclear Receptor Signaling Atlas (www.nursa.org): hyperlinking the nuclear receptor signaling community.

The nuclear receptor signaling (NRS) field has generated a substantial body of information on nuclear receptors, their ligands and coregulators, with the ultimate goal of constructing coherent models of the biological and clinical significance of these molecules. As a component of the Nuclear Receptor Signaling Atlas (NURSA)--the development of a functional atlas of nuclear receptor biology--the NURSA Bioinformatics Resource is developing a strategy to organize and integrate legacy and future information on these molecules in a single web-based resource (www.nursa.org). This entails parallel efforts of (i) developing an appropriate software framework for handling datasets from NURSA laboratories and (ii) designing strategies for the curation and presentation of public data relevant to NRS. To illustrate our approach, we have described here in detail the development of a web-based interface for the NURSA quantitative PCR nuclear receptor expression dataset, incorporating bioinformatics analysis which provides novel perspectives on functional relationships between these molecules. We anticipate that the free and open access of the community to a platform for data mining and hypothesis generation strategies will be a significant contribution to the progress of research in this field.

Animals↗

DRASTIC--INSIGHTS: querying information in a plant gene expression database.

DRASTIC--Database Resource for the Analysis of Signal Transduction In Cells (http://www.drastic.org.uk/) has been created as a first step towards a data-based approach for constructing signal transduction pathways. DRASTIC is a relational database of plant expressed sequence tags and genes up- or down-regulated in response to various pathogens, chemical exposure or other treatments such as drought, salt and low temperature. More than 17700 records have been obtained from 306 treatments affecting 73 plant species from 512 peer-reviewed publications with most emphasis being placed on data from Arabidopsis thaliana. DRASTIC has been developed by the Scottish Crop Research Institute and the University of Abertay Dundee and allows rapid identification of plant genes that are up- or down-regulated by multiple treatments and those that are regulated by a very limited (or perhaps a single) treatment. The INSIGHTS (INference of cell SIGnaling HypoTheseS) suite of web-based tools allows intelligent data mining and extraction of information from the DRASTIC database. Potential response pathways can be visualized and comparisons made between gene expression patterns in response to various treatments. The knowledge gained informs plant signalling pathways and systems biology investigations.

Arabidopsis↗

New Onto-Tools: Promoter-Express, nsSNPCounter and Onto-Translate.

The Onto-Tools suite is composed of an annotation database and eight complementary, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner, Pathway-Express, Promoter-Express and nsSNPCounter. Promoter-Express is a new tool added to the Onto-Tools ensemble that facilitates the identification of transcription factor binding sites active in specific conditions. nsSNPCounter is another new tool that allows computation and analysis of synonymous and non-synonymous codon substitutions for studying evolutionary rates of protein coding genes. Onto-Translate has also been enhanced to expand its scope and accuracy by fully utilizing the capabilities of the Onto-Tools database. Currently, Onto-Translate allows arbitrary mappings between 28 types of IDs for 53 organisms. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Binding Sites↗

The ASAP II database: analysis and comparative genomics of alternative splicing in 15 animal species.

We have greatly expanded the Alternative Splicing Annotation Project (ASAP) database: (i) its human alternative splicing data are expanded approximately 3-fold over the previous ASAP database, to nearly 90,000 distinct alternative splicing events; (ii) it now provides genome-wide alternative splicing analyses for 15 vertebrate, insect and other animal species; (iii) it provides comprehensive comparative genomics information for comparing alternative splicing and splice site conservation across 17 aligned genomes, based on UCSC multigenome alignments; (iv) it provides an approximately 2- to 3-fold expansion in detection of tissue-specific alternative splicing events, and of cancer versus normal specific alternative splicing events. We have also constructed a novel database linking orthologous exons and orthologous introns between genomes, based on multigenome alignment of 17 animal species. It can be a valuable resource for studies of gene structure evolution. ASAP II provides a new web interface enabling more detailed exploration of the data, and integrating comparative genomics information with alternative splicing data. We provide a set of tools for advanced data-mining of ASAP II with Pygr (the Python Graph Database Framework for Bioinformatics) including powerful features such as graph query, multigenome alignment query, etc. ASAP II is available at http://www.bioinformatics.ucla.edu/ASAP2.

Alternative Splicing↗

The Rat Genome Database, update 2007--easing the path from disease to data and back again.

The Rat Genome Database (RGD, http://rgd.mcw.edu) is one of the core resources for rat genomics and recent developments have focused on providing support for disease-based research using the rat model. Recognizing the importance of the rat as a disease model we have employed targeted curation strategies to curate genes, QTL and strain data for neurological and cardiovascular disease areas. This work has centered on rat but also includes data for mouse and human to create 'disease portals' that provide a unified view of the genes, QTL and strain models for these diseases across the three species. The disease curation efforts combined with normal curation activities have served to greatly increase the content of the database, particularly for biological information, including gene ontology, disease, pathway and phenotype ontology annotations. In addition to improving the features and database content, community outreach has been expanded to demonstrate how investigators can leverage the resources at RGD to facilitate their research and to elicit suggestions and needs for future developments. We have published a number of papers that provide additional information on the ontology annotations and the tools at RGD for data mining and analysis to better enable researchers to fully utilize the database.

Animals↗

DroSpeGe: rapid access database for new Drosophila species genomes.

The Drosophila species comparative genome database DroSpeGe (http://insects.eugenes.org/DroSpeGe/) provides genome researchers with rapid, usable access to 12 new and old Drosophila genomes, since its inception in 2004. Scientists can use, with minimal computing expertise, the wealth of new genome information for developing new insights into insect evolution. New genome assemblies provided by several sequencing centers have been annotated with known model organism gene homologies and gene predictions to provided basic comparative data. TeraGrid supplies the shared cyberinfrastructure for the primary computations. This genome database includes homologies to Drosophila melanogaster and eight other eukaryote model genomes, and gene predictions from several groups. BLAST searches of the newest assemblies are integrated with genome maps. GBrowse maps provide detailed views of cross-species aligned genomes. BioMart provides for data mining of annotations and sequences. Common chromosome maps identify major synteny among species. Potential gain and loss of genes is suggested by Gene Ontology groupings for genes of the new species. Summaries of essential genome statistics include sizes, genes found and predicted, homology among genomes, phylogenetic trees of species and comparisons of several gene predictions for sensitivity and specificity in finding new and known genes.

Animals↗

Differential selection after duplication in mammalian developmental genes.

Gene duplication provides the opportunity for subsequent refinement of distinct functions of the duplicated copies. Either through changes in coding sequence or changes in regulatory regions, duplicate copies appear to obtain new or tissue-specific functions. If this divergence were driven by natural selection, we would expect duplicated copies to have differentiated patterns of substitutions. We tested this hypothesis using genes that duplicated before the human/mouse split and whose orthologous relations were clear. The null hypothesis is that the number of amino acid changes between humans and mice was distributed similarly across different paralogs. We used a method modified from Tang and Lewontin to detect heterogeneity in the amino acid substitution pattern between those different paralogs. Our results show that many of the paralogous gene pairs appear to be under differential selection in the human/mouse comparison. The properties that led to diversification appear to have arisen before the split of the human and mouse lineages. Further study of the diverged genes revealed insights regarding the patterns of amino acid substitution that resulted in differences in function and/or expression of these genes. This approach has utility in the study of newly identified members of gene families in genomewide data mining and for contrasting the merits of alternative hypotheses for the evolutionary divergence of function of duplicated genes.

Amino Acid Substitution↗

Transcriptome profiling in root nodules and arbuscular mycorrhiza identifies a collection of novel genes induced during Medicago truncatula root endosymbioses.

Transcriptome profiling based on cDNA array hybridizations and in silico screening was used to identify Medicago truncatula genes induced in both root nodules and arbuscular mycorrhiza (AM). By array hybridizations, we detected several hundred genes that were upregulated in the root nodule and the AM symbiosis, respectively, with a total of 75 genes being induced during both interactions. The second approach based on in silico data mining yielded several hundred additional candidate genes with a predicted symbiosis-enhanced expression. A subset of the genes identified by either expression profiling tool was subjected to quantitative real-time reverse-transcription polymerase chain reaction for a verification of their symbiosis-induced expression. That way, induction in root nodules and AM was confirmed for 26 genes, most of them being reported as symbiosis-induced for the first time. In addition to delivering a number of novel symbiosis-induced genes, our approach identified several genes that were induced in only one of the two root endosymbioses. The spatial expression patterns of two symbiosis-induced genes encoding an annexin and a beta-tubulin were characterized in transgenic roots using promoter-reporter gene fusions.

Annexins↗

Comparison of cytochrome P450 (CYP) genes from the mouse and human genomes, including nomenclature recommendations for genes, pseudogenes and alternative-splice variants.

OBJECTIVES: Completion of both the mouse and human genome sequences in the private and public sectors has prompted comparison between the two species at multiple levels. This review summarizes the cytochrome P450 (CYP) gene superfamily. For the first time, we have the ability to compare complete sets of CYP genes from two mammals. Use of the mouse as a model mammal, and as a surrogate for human biology, assumes reasonable similarity between the two. It is therefore of interest to catalog the genetic similarities and differences, and to clarify the limits of extrapolation from mouse to human. METHODS: Data-mining methods have been used to find all the mouse and human CYP sequences; this includes 102 putatively functional genes and 88 pseudogenes in the mouse, and 57 putatively functional genes and 58 pseudogenes in the human. Comparison is made between all these genes, especially the seven main CYP gene clusters. RESULTS AND CONCLUSIONS: The seven CYP clusters are greatly expanded in the mouse with 72 functional genes versus only 27 in the human, while many pseudogenes are present; presumably this phenomenon will be seen in many other gene superfamily clusters. Complete identification of all pseudogene sequences is likely to be clinically important, because some of these highly similar exons can interfere with PCR-based genotyping assays. A naming procedure for each of four categories of CYP pseudogenes is proposed, and we encourage various gene nomenclature committees to consider seriously the adoption and application of this pseudogene nomenclature system.

Alternative Splicing↗

Psychiatric genetics in silico: databases and tools for psychiatric geneticists.

Bioinformatics can significantly impact the laboratory genetics process from the study design phase to conclusive identification of a disease gene. The present review will highlight key databases to enhance psychiatric genetic study design, based on full use of genomics data and the golden path sequence. It will address methods to ensure comprehensive genetic data mining, using the best available genomic and genetic databases such as the University of California Santa Cruz human genome browser, Ensembl, Mapview, dbSNP and GDB, and locus-specific databases such as Online Mendelian Inheritance In Man. Using the golden path sequence as a template, with the necessary quality checks, it is possible to design detailed genetic studies from sequence information alone. Drawing together this diverse information, it is possible to characterize a locus or gene in silico to a very detailed level. This in turn can have real cost and efficiency benefits by assisting in the identification of markers that are most likely to be informative, or by highlighting the best candidate genes for study.

Databases, Bibliographic↗

Computer algorithm for automated work group classification from free text: the DREAM technique.

OBJECTIVE: This study developed and tested a computer method to automatically assign subjects to aggregate work groups based on their free text work descriptions. METHODS: The Double Root Extended Automated Matcher (DREAM) algorithm classifies individuals based on pairs of subjects' free text word roots in common with those of standard classification systems and several explicitly defined linkages between term roots and aggregates. RESULTS: DREAM effectively analyzed free text from 5887 participants in a multisite chronic obstructive pulmonary disease prevention study (Lung Health Study). For a test set of 533 cases, DREAMs classifications compared favorably with those of a four-human panel. The humans rated the accuracy of DREAM as good or better in 80% of the test cases. CONCLUSIONS: Automated text interpretation is a promising tool for analyzing large data sets for applications in data mining, research, and surveillance. Work descriptive information is most useful when it can link an individual to aggregate entities that have occupational health relevance. Determining the appropriate group requires considerable expertise. This article describes a new method for making such assignments using a computer algorithm to reduce dependence on the limited number of occupational health experts. In addition, computer algorithms foster consistency of assignments.

Algorithms↗

Monitoring the serological proteome: the latest modality in prostate cancer detection.

PURPOSE: Various strategies have recently emerged to improve the diagnostic prediction of prostate cancer (CaP). One such strategy includes the mass profiling of serum protein fractions selectively adsorbed onto chemically modified probes. In the current study we further validated this approach, while offering a more versatile, less expensive and yet equally predictive alternative to existing technologies. MATERIALS AND METHODS: A solid core lipophilic C-18 resin was used to extract and enrich the low molecular weight protein fraction from patient serum for further analysis by mass spectrometry. Mass spectra generated from a 48 patient training set were data mined using multivariate analysis to identify diagnostically significant protein peaks. These peaks were then used to test a blinded study set comprising 168 patients with common statistical algorithms and commercially available software packages. RESULTS: A total of 36 peaks generated from the training set were used to test the combined set of 168 serum samples obtained from 98 healthy individuals and 70 patients with CaP. We report a sensitivity of 94.1% and a specificity of 99.0% with 1 false-positive, 4 false-negative and 5 nondiagnosed cases. CONCLUSIONS: Our results further indicate that mass profiling of serological proteins provides a means for the accurate detection of CaP. In addition, our approach was found to be superior to chip based protocols, generating rich, sharp, highly reproducible spectra attainable in a high throughput manner and at minimal cost. This technique is also scaleable for subsequent protein characterization using multidimensional protein identification technologies. Finally, analyses of mass spectra with commercially available statistical applications was found to be highly effective in generating highly discriminatory m/z values for CaP diagnosis.

Biomarkers, Tumor↗

Folate deficiency induced hyperhomocysteinemia changes the expression of thrombosis-related genes.

Hyperhomocysteinemia (HH) is an independent risk factor for thrombosis although the precise pathogenesis is still unresolved. Previous studies have demonstrated that HH changes whole blood coagulation by increasing the velocity, increasing the firmness of the formed clot, and by prolonging the initiation phase of the coagulation. With the aim of elucidating the genetic pathogenesis which might be responsible for the changes in whole blood coagulation, we applied oligo-array technology to RNA from buffycoat-cells comparing animals suffering from hyperhomocysteinemia (42 micromol/l) with controls (6 micromol/l). Data mining identified a number of relevant genes, and the expression pattern was validated by real time reverse transcriptase-polymerase chain reaction. An upregulation of integrin beta-3, Rap 1b, glycoprotein V, platelet-endothelial cell adhesion molecule-1 (PECAM-1) and von Willebrand factor (vWF) led us to deduce increased platelet activation/aggregation. Coagulation factor XIIIa was upregulated and may contribute in increasing the firmness of the formed clot. Impaired fibrinolysis was anticipated, since an upregulation of plasminogen activator inhibitor-1 (PAI-1) and a downregulation of tissue-type plasminogen activator (t-PA) were detected. Reduced spontaneous contact activation was anticipated due to a downregulation of the kallikrein gene. Upregulation of selectins may contribute to increased tethering and rolling of leukocytes. In conclusion, folate deficiency induced hyperhomocysteinemia changes in the gene expression of buffy coat cells which was characterized by increased platelet activation, impaired fibrinolysis and a reduced contact activation of the coagulation. These changes may contribute to explain the increased risk of thrombosis seen in hyperhomocysteinemia individuals. This pattern of the hyperhomocysteinemia-affected genes may represent a reference for further studies at the protein level to define the folate depletion effects in blood cells.

Animals↗

Insufflation techniques in gynecologic laparoscopy.

Our objectives were to assess the safety and efficacy of different insufflation methods in women undergoing laparoscopy and to develop a model for selection of the appropriate insufflation technique based on the patient's characteristics and surgeon's experience. We performed a retrospective analysis of laparoscopic procedures on 3086 women over a 13-year period at the University of Louisville Hospital, Louisville, KY. All laparoscopic procedures were performed on an outpatient basis by residents under faculty supervision. Five different insufflation techniques were evaluated: standard transumbilical insufflation, open laparoscopy, transuterine insufflation, subcostal insufflation, and direct trocar insertion technique. Body mass index and previous abdominal surgeries were identified as the most important factors in the selection of the most successful insufflation method based on the surgeon's experience, using data mining techniques. During the first insufflation attempt, we were successful at achieving a pneumoperitoneum 94.7% of the time. This number increased to 98.1% when we switched to a second alternative insufflation method. In all, there were 5 complications out of 3086 patients (0.16%) after all insufflation techniques.

Adolescent↗

Organic materials for second-harmonic generation: advances in relating structure to function.

The relationships between molecular structure and the nonlinear optical phenomenon second-harmonic generation (SHG) are discussed. New-found relationships built up from basic structural axioms that were deduced in the 1970s and 1980s are the particular focus of this article, using structural results from X-ray and neutron-diffraction studies. The molecular and supramolecular manifestations of the SHG effect are borne out, although ways to optimize the effect on the molecular scale feature predominantly, since control of SHG on the supramolecular scale remains difficult given present limitations. The use of a variety of templates to generate head-to-tail oriented host-guest species thereby bypassing such limitations is described. The paper concludes with a look ahead at next generation 'octupolar' SHG-active compounds, the prediction of new series of SHG-active compounds via data-mining computational procedures, and developments in diffraction technology that may enable structural movies of a molecule to be captured during the SHG process. A practical assessment of the viability of organic SHG materials for industrial application is reviewed with a positive outcome, thus indicating a promising future for organic SHG materials.

Journal Article↗