Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Automatic construction of gene relation networks using text mining and gene expression data.

Microarray gene expression analysis is a powerful high-throughput technique that enables researchers to monitor the expression of thousands of genes simultaneously. Using this methodology huge amounts of data are produced which have to be analysed. Clustering algorithms are used to group genes together based on a predefined distance measure. However, clustering algorithms do not necessarily group the genes in a biological meaningful way. Additional information is needed to improve the identification of disease relevant genes. The primary objective of our project is to support the analysis of microarray gene expression data by construction of gene relation networks (GRNs). Required information can not be found in a structured representation like a database. In contrast, a large number of relations are described in biomedical literature. The main outcome of this project is the implementation of a software system that provides clinicians and researchers with a tool that supports the analysis of microarray gene expression data by mapping known relationships from the biomedical literature to local gene expression experiments.

Abbreviations as Topic↗

Mining HIV protease cleavage data using genetic programming with a sum-product function.

MOTIVATION: In order to design effective HIV inhibitors, studying and understanding the mechanism of HIV protease cleavage specification is critical. Various methods have been developed to explore the specificity of HIV protease cleavage activity. However, success in both extracting discriminant rules and maintaining high prediction accuracy is still challenging. The earlier study had employed genetic programming with a min-max scoring function to extract discriminant rules with success. However, the decision will finally be degenerated to one residue making further improvement of the prediction accuracy difficult. The challenge of revising the min-max scoring function so as to improve the prediction accuracy motivated this study. RESULTS: This paper has designed a new scoring function called a sum-product function for extracting HIV protease cleavage discriminant rules using genetic programming methods. The experiments show that the new scoring function is superior to the min-max scoring function. AVAILABILITY: The software package can be obtained by request to Dr Zheng Rong Yang.

Algorithms↗

Mining ChIP-chip data for transcription factor and cofactor binding sites.

MOTIVATION: Identification of single motifs and motif pairs that can be used to predict transcription factor localization in ChIP-chip data, and gene expression in tissue-specific microarray data. RESULTS: We describe methodology to identify de novo individual and interacting pairs of binding site motifs from ChIP-chip data, using an algorithm that integrates localization data directly into the motif discovery process. We combine matrix-enumeration based motif discovery with multivariate regression to evaluate candidate motifs and identify motif interactions. When applied to the HNF localization data in liver and pancreatic islets, our methods produce motifs that are either novel or improved known motifs. All motif pairs identified to predict localization are further evaluated according to how well they predict expression in liver and islets and according to how conserved are the relative positions of their occurrences. We find that interaction models of HNF1 and CDP motifs provide excellent prediction of both HNF1 localization and gene expression in liver. Our results demonstrate that ChIP-chip data can be used to identify interacting binding site motifs. AVAILABILITY: Motif discovery programs and analysis tools are available on request from the authors.

Algorithms↗

Global protein function annotation through mining genome-scale data in yeast Saccharomyces cerevisiae.

As we are moving into the post genome-sequencing era, various high-throughput experimental techniques have been developed to characterize biological systems on the genomic scale. Discovering new biological knowledge from the high-throughput biological data is a major challenge to bioinformatics today. To address this challenge, we developed a Bayesian statistical method together with Boltzmann machine and simulated annealing for protein functional annotation in the yeast Saccharomyces cerevisiae through integrating various high-throughput biological data, including yeast two-hybrid data, protein complexes and microarray gene expression profiles. In our approach, we quantified the relationship between functional similarity and high-throughput data, and coded the relationship into 'functional linkage graph', where each node represents one protein and the weight of each edge is characterized by the Bayesian probability of function similarity between two proteins. We also integrated the evolution information and protein subcellular localization information into the prediction. Based on our method, 1802 out of 2280 unannotated proteins in yeast were assigned functions systematically.

Bayes Theorem↗

Theoretical models for determining 222Rn and 220Rn progeny levels in Canadian underground U mines--a comparison with experimental data.

Use has been made of several theoretical models to predict radiation levels in underground U mines. The models used are the Evans model, the Thomas-Epps mine model, the isolated mine model, the mine tunnel with no air flow and the mine tunnel with air flow (Beckman and Holub). Calculations based on the above models have been extended to include 220Rn gas and its progeny, a common occurrence in some Canadian U mines. Theoretical predictions include 222Rn and 220Rn progenies working levels and concentrations, as well as some ratios of great practical interest. The main differences between the models are pointed out and comparison is made with experimental data gathered during the last 3 yr in several Canadian underground U mines. In general terms, the mine models reported in the literature and presented here are neither satisfactory nor clearly distinguishable for practical application on the basis of the experimental data so far collected. The reason for such lack of agreement lies mainly in the unrealistic assumptions on which the models are based. Compounded is the gross oversimplification of quite complex dynamic situations encountered in actual practice. Deficiencies inherent to the models are noted and suggestions to improve the applicability of mine models to practical situations are indicated.

Air Pollutants↗

Using 'off the shelf', computer programs to mine additional insights from published data: diurnal variation in potency of ACTH stimulation of cortisol secretion revealed.

We describe the use of available computer programs to mine additional insights from published data. Graphs from four studies of ACTH and cortisol plasma levels collected throughout the 24 h day in humans were scanned into digital form and the data points extracted. To investigate the magnitude of ACTH stimulation of cortisol secretion across the 24 h, Monte Carlo methods were used to fit the parameters of a computer model of the ACTH-adrenal axis to the extracted data. ACTH was found to have a greater effect on cortisol secretion during the peaks of the cycle than at the nadir. This finding could not be explained by previously published dose response curves of ACTH effect. This implies that other modulators influence the effect of ACTH on the adrenal. This study also demonstrates how available computer programs can be used to examine models of physiologic regulation using data already available in the literature.

Adrenocorticotropic Hormone↗

The accuracy of self-reported regulatory data: the case of coal mine dust.

Coal-mine owners are required to measure miner exposures to respirable dust so that compliance with Federal health regulations can be monitored. This study analyzes the problem of possible underreporting of dust exposures. Using two statistical approaches, data for three mining occupations in 54 large underground coal mines during 1976-1978 are examined for evidence of underreporting. First, regression estimates compare dust concentrations reported by coal-mine owners with those reported by government health inspectors. Then, the statistical distribution of concentrations reported by coal-mine owners are examined for the size and nature of their deviation from log-normality. Both approaches suggest widespread underreporting.

Air Pollutants, Occupational↗

Pulmonary inflammation and crystalline silica in respirable coal mine dust: dose-response.

This study describes the quantitative relationships between early pulmonary responses and the estimated lung-burden or cumulative exposure of respirable-quartz or coal mine dust. Data from a previous bronchoalveolar lavage (BAL) study in coal miners (n = 20) and nonminers (n = 16) were used including cell counts of alveolar macrophages (AMs) and polymorphonuclear leukocytes (PMNs), and the antioxidant superoxide dismutase (SOD) levels. Miners' individual working lifetime particulate exposures were estimated from work histories and mine air sampling data, and quartz lung-burdens were estimated using a lung dosimetry model. Results show that quartz, as either cumulative exposure or estimated lung-burden, was a highly statistically significant predictor of PMN response (P < 0.0001); however cumulative coal dust exposure did not significantly add to the prediction of PMNs (P = 0.2) above that predicted by cumulative quartz exposure (P < 0.0001). Despite the small study size, radiographic category was also significantly related to increasing levels of both PMNs and quartz lung burden (P-values < 0.04). SOD in BAL fluid rose linearly with quartz lung burden (P < 0.01), but AM count in BAL fluid did not (P > 0.4). This study demonstrates dose-response relationships between respirable crystalline silica in coal mine dust and pulmonary inflammation, antioxidant production, and radiographic small opacities.

Adult↗

Genetic, cytogenetic, and carcinogenic effects of radon: a review.

Radon exposure has been linked to lung carcinogenesis in both human and animal studies. Studies of smoking and nonsmoking uranium miners indicate that radon alone is a risk factor for lung cancer at the levels encountered by these miners, although the possibility exists that other substances in the mine environment affect the radon-induced response. The relevance of data from mines to the lower-exposure home environment is often questioned; still, a recent study of miners exposed to relatively low radon concentrations demonstrated a statistically significant increase for lung and laryngeal cancer deaths. In two major series of experiments with rats, the primary carcinogenic effect found was respiratory tract tumors, and evidence for an inverse exposure-rate effect was also noted. Although this inverse dose-rate effect also has been described in underground miner studies, it may not similarly apply to radon in the home environment. This observation is due to the fact that, below a certain exposure, cells are hit once or not at all, and one would not expect any dose-rate effect, either normal or inverse. Because some chromosome aberrations persist in cycling cells as stable events, cytogenetic studies with radon are being performed to help complete the understanding of the events leading to radon-induced neoplasia. Radon has been found to induce 13 times as much cytogenetic damage (as measured by the occurrence of micronuclei) than a similar dose of 60Co. A wide variety of mutation systems have demonstrated alpha-particle mutagenesis; recent investigations have focused on the molecular basis of alpha-induced mutagenesis. Gene mutations are induced by radon in a linear and dose-dependent fashion, and with a high biological effect relative to low-LET irradiation. Studies of the hprt locus show that approximately half of the alpha-induced mutations arise by complete deletion of the gene; the remaining mutations are split between partial deletions, rearrangements, and events not detectable by Southern blot or PCR exon analysis. Although other mutation systems do not show the same spectra as observed in the hprt gene (suggesting that the gene environment affects response), DNA deletions or multilocus lesions of various size appear to be predominant after radon exposure. As data emerge regarding radon-induced changes at the chromosomal and molecular level, the mechanisms involved in radon carcinogenesis are being clarified. This information should increase the understanding of risk at the low exposure levels typically found in the home.

Animals↗

Geographic variability in radon exhalation at a rehabilitated uranium mine in the Northern Territory, Australia.

In this study, dry season radon flux densities and radon fluxes have been determined at the rehabilitated Nabarlek uranium mine in northern Australia using conventional charcoal canisters. Environmental background levels amounted to 31+/- 15 milli Becquerel per m(2) per second (mBq m(-2) s(-1)). Radon flux densities within the fenced rehabilitated mine area showed large variations with a maximum of 6500 mBq m(-2) s(-1) at an area south of the former pit characterised by a disequilibrium between (226)Ra and (238)U. Radon flux densities were also high above the areas of the former pit (mean 971 mBq m(-2) s(-1)) and waste rock dump (mean 335 mBq m(-2) s(-1)). The lower limit for the total pre-mining radon flux from the fenced area (140 ha) was estimated to 214 kBq s(-1), post-mining radon flux amounted to 174 kBq s(-1). Our study highlights that the results of radon flux studies are vitally dependant on the selection of individual survey points. We suggest the use of a randomised system for both the selection of survey points and the placement of charcoal canisters at each survey point, to avoid over estimation of radon flux densities. It is also important to emphasize the significance of having reliable pre-mining radiological data available to assess the success of rehabilitation of a uranium mine site.

Air Pollutants, Radioactive↗

Using molecular similarity to construct accurate semiempirical electronic structure theories.

Ab initio electronic structure methods give accurate results for small systems, but do not scale well to large systems. Chemical insight tells us that molecular functional groups will behave approximately the same way in all molecules, large or small. This molecular similarity is exploited in semiempirical methods, which couple simple electronic structure theories with parameters for the transferable characteristics of functional groups. We propose that high-level calculations on small molecules provide a rich source of parametrization data. In principle, we can select a functional group, generate a large amount of ab initio data on the group in various small-molecule environments, and "mine" this data to build a sophisticated model for the group's behavior in large environments. This work details such a model for electron correlation: a semiempirical, subsystem-based correlation functional that predicts a subsystem's two-electron density matrix as a functional of its one-electron density matrix. This model is demonstrated on two small systems: chains of linear, minimal-basis (H-H)(5), treated as a sum of four overlapping (H-H)(2) subsystems; and the aldehyde group of a set of HOC-R molecules. The results provide an initial demonstration of the feasibility of the approach.

Journal Article↗

Comparative analysis of multiple genome-scale data sets.

The ongoing analyses of published genome-scale data sets is evidence that different approaches are required to completely mine this data. We report the use of novel tools for both visualization and data set comparison to analyze yeast gene-expression (cell cycle and exit from stationary phase/G(0)) and protein-interaction studies. This analysis led to new insights about each data set. For example, G(1)-regulated genes are not co-regulated during exit from stationary phase, indicating that the cells are not synchronized. The tight clustering of other genes during exit from stationary-phase data set further indicates the physiological responses during G(0) exit are separable from cell-cycle events. Comparison of the two data sets showed that ribosomal-protein genes cluster tightly during exit from stationary phase, but are found in three significantly different clusters in the cell-cycle data set. Two protein-interaction data sets were also compared with the gene-expression data. Visual analysis of the complete data sets showed no clear correlation between co-expression of genes and protein interactions, in contrast to published reports examining subsets of the protein-interaction data. Neither two-hybrid study identified a large number of interactions between ribosomal proteins, consistent with recent structural data, indicating that for both data sets, the identification of false-positive interactions may be lower than previously thought.

Cell Cycle↗

Metabolic Dysregulation of the Lysophospholipid/Autotaxin Axis in the Chromosome 9p21 Gene SNP rs10757274.

BACKGROUND: Common chromosome 9p21 single nucleotide polymorphisms (SNPs) increase coronary heart disease risk, independent of traditional lipid risk factors. However, lipids comprise large numbers of structurally related molecules not measured in traditional risk measurements, and many have inflammatory bioactivities. Here, we applied lipidomic and genomic approaches to 3 model systems to characterize lipid metabolic changes in common Chr9p21 SNPs, which confer &#x2248;30% elevated coronary heart disease risk associated with altered expression of ANRIL, a long ncRNA. METHODS: Untargeted and targeted lipidomics was applied to plasma from NPHSII (Northwick Park Heart Study II) homozygotes for AA or GG in rs10757274, followed by correlation and network analysis. To identify candidate genes, transcriptomic data from shRNA downregulation of ANRIL in HEK-293 cells was mined. Transcriptional data from vascular smooth muscle cells differentiated from induced pluripotent stem cells of individuals with/without Chr9p21 risk, nonrisk alleles, and corresponding knockout isogenic lines were next examined. Last, an in-silico analysis of miRNAs was conducted to identify how ANRIL might control lysoPL (lysophosphospholipid)/lysoPA (lysophosphatidic acid) genes. RESULTS: Elevated risk GG correlated with reduced lysoPLs, lysoPA, and ATX (autotaxin). Five other risk SNPs did not show this phenotype. LysoPL-lysoPA interconversion was uncoupled from ATX in GG plasma, suggesting metabolic dysregulation. Significantly altered expression of several lysoPL/lysoPA metabolizing enzymes was found in HEK cells lacking ANRIL. In the vascular smooth muscle cells data set, the presence of risk alleles associated with altered expression of several lysoPL/lysoPA enzymes. Deletion of the risk locus reversed the expression of several lysoPL/lysoPA genes to nonrisk haplotype levels. Genes that were altered across both cell data sets were DGKA, MBOAT2, PLPP1, and LPL. The in-silico analysis identified 4 ANRIL-regulated miRNAs that control lysoPL genes as miR-186-3p, miR-34a-3p, miR-122-5p, and miR-34a-5p. CONCLUSIONS: A Chr9p21 risk SNP associates with complex alterations in immune-bioactive phospholipids and their metabolism. Lipid metabolites and genomic pathways associated with coronary heart disease pathogenesis in Chr9p21 and ANRIL-associated disease are demonstrated.

Chromosomes, Human, Pair 9↗

PolyMAPr: programs for polymorphism database mining, annotation, and functional analysis.

Pharmacogenomic and disease-association studies rely on identifying a comprehensive set of polymorphisms within candidate genes. Public SNP databases are a rich source of polymorphism data, but mining them effectively requires overcoming at least four challenges: ensuring accurate annotations for genes and polymorphisms, eliminating both inter- and intra-database redundancy, integrating data from multiple public sources with data generated locally, and prioritizing the variants for further study. PolyMAPr (Polymorphism Mining and Annotation Programs)' was developed to overcome these challenges and to improve the efficiency of database mining and polymorphism annotation. PolyMAPr takes as input a file containing a list of genes to be processed and files containing each annotated gene sequence. Polymorphic sequences obtained from public databases (dbSNP, CGAP, and JSNP) or through local SNP discovery efforts, as well as oligonucleotide sequences (e.g., PCR primers), are mapped to the annotated gene sequences and named according to suggested nomenclature guidelines. The functional effects of nonsynonymous coding-region SNPs (cSNPs) and any variants that might alter exon splicing enhancer (ESE) sites, putative transcription factor binding sites, or intron-exon splice sites are predicted. The output files are accessible though a browser interface. In addition, the results are also provided in Extensible Markup Language (XML) format to facilitate uploading them into a local relational database. PolyMAPr increases the efficiency of mining public databases for genetic variants within candidate genes and provides a mechanism by which data from multiple sources (both public and private) can be uniformly integrated, thereby significantly reducing the effort required to obtain a comprehensive set of polymorphisms for pharmacogenomic and disease-association studies. PolyMAPr can be obtained from http://pharmacogenomics.wustl.edu.

Databases, Nucleic Acid↗

Beyond the data deluge: data integration and bio-ontologies.

Biomedical research is increasingly a data-driven science. New technologies support the generation of genome-scale data sets of sequences, sequence variants, transcripts, and proteins; genetic elements underpinning understanding of biomedicine and disease. Information systems designed to manage these data, and the functional insights (biological knowledge) that come from the analysis of these data, are critical to mining large, heterogeneous data sets for new biologically relevant patterns, to generating hypotheses for experimental validation, and ultimately, to building models of how biological systems work. Bio-ontologies have an essential role in supporting two key approaches to effective interpretation of genome-scale data sets: data integration and comparative genomics. To date, bio-ontologies such as the Gene Ontology have been used primarily in community genome databases as structured controlled terminologies and as data aggregators. In this paper we use the Gene Ontology (GO) and the Mouse Genome Informatics (MGI) database as use cases to illustrate the impact of bio-ontologies on data integration and for comparative genomics. Despite the profound impact ontologies are having on the digital categorization of biological knowledge, new biomedical research and the expanding and changing nature of biological information have limited the development of bio-ontologies to support dynamic reasoning for knowledge discovery.

Animals↗

Seeding information management capacity to support operational management in hospitals.

There are vast amounts of regularly reported data in the information systems of hospitals, state and federal governments. The increase in accessibility offered by platforms such as the Health Information Exchange (HIE) in New South Wales (NSW) creates a new level of opportunity. Administrative data can also speak to clinical and managerial issues. The capacity to mine these data and use the information for improving quality and efficiency has not been well developed at the "coal face" of operational management. Whilst it has been both possible and useful to track utilisation of services to hospitals and patients as cost and volume, it has not been of interest to track these same data to the operational locus of care--the nursing unit, the operating room, the imaging department. With HIE-type systems, the information is now more readily available and operational managers know this. The challenge is to develop the interdisciplinary capacity to query administrative data to facilitate clinical and managerial decision-making. We report here a possible model of a systematic approach to developing this capacity and some of the results of equipping operational and clinical managers to study problems in their own work settings. These efforts have required no additional internal resources, while the payoffs have been considerable.

Data Collection↗

The life sciences Global Image Database (GID).

Although a vast amount of life sciences data is generated in the form of images, most scientists still store images on extremely diverse and often incompatible storage media, without any type of metadata structure, and thus with no standard facility with which to conduct searches or analyses. Here we present a solution to unlock the value of scientific images. The Global Image Database (GID) is a web-based (http://www.gwer.ch/qv/gid/gid.ht m ) structured central repository for scientific annotated images. The GID was designed to manage images from a wide spectrum of imaging domains ranging from microscopy to automated screening. The annotations in the GID define the source experiment of the images by describing who the authors of the experiment are, when the images were created, the biological origin of the experimental sample and how the sample was processed for visualization. A collection of experimental imaging protocols provides details of the sample preparation, and labeling, or visualization procedures. In addition, the entries in the GID reference these imaging protocols with the probe sequences or antibody names used in labeling experiments. The GID annotations are searchable by field or globally. The query results are first shown as image thumbnail previews, enabling quick browsing prior to original-sized annotated image retrieval. The development of the GID continues, aiming at facilitating the management and exchange of image data in the scientific community, and at creating new query tools for mining image data.

Biological Science Disciplines↗