Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

WHO/ISBRA Study on State and Trait Markers of Alcohol Use and Dependence: analysis of demographic, behavioral, physiologic, and drinking variables that contribute to dependence and seeking treatment. International Society on Biomedical Research on Alcoholism.

BACKGROUND: Discussions between the World Health Organization (WHO) and the International Society on Biomedical Research on Alcoholism (ISBRA) identified the need for a multiple-center international study on state and trait markers of alcohol abuse and alcohol dependence. The reasoning behind the generation of such a project included the need to understand the alcohol use characteristics of diverse populations and the performance of biological markers of alcohol use in a variety of settings throughout the world. A second major reason for initiating this study was to collect DNA for well-structured and stratified association studies between genetic markers and/or "candidate" genes and behavioral/physiological phenotypes of importance to predisposition to alcohol dependence. METHODS: An extensive interview instrument was developed with leadership from the U.S. National Institute on Alcohol Abuse and Alcoholism (NIAAA). The instrument was translated from English to Finnish, French, German, Japanese, and Portuguese (Brazilian). One thousand eight hundred sixty-three subjects were recruited at five clinical centers (Montreal, Canada; Helsinki, Finland; Sapporo, Japan; São Paulo, Brazil; and Sydney, Australia). The subjects responded to the structured interview and provided blood and urine samples for biochemical analysis. This article focuses on the demographic characteristics of the study subjects, their drinking habits, alcohol-dependence characteristics, comorbid psychiatric and other drug variables, and predictors for seeking treatment for alcohol dependence. Multiple logistic regression models were constructed and used to explore variables that contribute to various levels of alcohol consumption, to a diagnosis of alcohol dependence, and to seeking treatment for alcohol dependence. ANOVA with post hoc comparisons, chi2, and Pearson moment calculations were used as necessary to assess additional relationships between variables. RESULTS: A number of factors previously noted in disparate studies were confirmed in our analysis. Men consumed more alcohol than women, Asians consumed less alcohol than whites or Blacks, alcohol-dependent subjects consumed more alcohol than nondependent subjects, alcohol consumption increased with age, and an increased level of education (university or postgraduate education) reduced the percentage of such individuals in the category designated as heavy drinkers (>210 g alcohol/week) and in the group who were currently in treatment for dependence. However, our analysis allowed for much more detailed comparisons; for example, although men drank more than women on a g/day basis, the differences were less pronounced on g/kg/day basis, and alcohol-dependent women drank equal amounts of alcohol as alcohol-dependent men on a g/kg/day basis. Antisocial personality characteristics or reports of trouble sleeping when an individual stops drinking were associated with higher alcohol intake. The most important of the tested factors that contributed to a DSM-IV diagnosis of dependence, however, was the report of anxiety if an individual stopped drinking. In terms of the various criteria within the DSM-IV criteria for alcohol dependence, no one criterion seemed to be prominent for individuals who sought alcohol dependence treatment, but the higher the number of criteria met by the individual, the higher was the probability that he or she would be in treatment. CONCLUSIONS: This initial report is the beginning of the "data mining" of this rich data set. The cross-national/cross-cultural aspects of this study allowed for multiple comparisons of variables across several ethnic/racial groups and allowed for assessment of biochemical markers for alcohol intake and predisposition to alcohol dependence in diverse settings.

Adult↗

Bioinformatics and its impact on clinical research methods. Findings from the Section on Bioinformatics.

OBJECTIVES: To summarize current excellent research in the field of bioinformatics. METHOD: Synopsis of the articles selected for the IMIA Yearbook 2006. RESULTS: Current research in the field of bioinformatics clearly shows ongoing unification of experimental findings and clinical outcomes. Microarray data, gene sequences and clinical data are more and more perceived as different but related facets of one entity. Significant work is done in the area of text and data mining in order to bring together patient data and biochemical phenomena by means of ontologies. A strong trend in the clinical field is performance of exhaustive studies on DNA material derived from patients that suffer from diseases that are already known to be inherited. Examination of appropriate methods covers data and text mining, ontologies as well as machine learning and classification. CONCLUSIONS: The best paper selection of articles on bioinformatics shows examples of excellent research on methods used for studying inherited diseases and their underlying genetic dispositions. Clinical studies, inclusion of experimental findings like microarray data, and of knowledge representation formats all lead to a better understanding the linkage between gene sequences, biological functions and clinical findings in the form of healthy state or physiological disorders.

Awards and Prizes↗

GENEVESTIGATOR. Arabidopsis microarray database and analysis toolbox.

High-throughput gene expression analysis has become a frequent and powerful research tool in biology. At present, however, few software applications have been developed for biologists to query large microarray gene expression databases using a Web-browser interface. We present GENEVESTIGATOR, a database and Web-browser data mining interface for Affymetrix GeneChip data. Users can query the database to retrieve the expression patterns of individual genes throughout chosen environmental conditions, growth stages, or organs. Reversely, mining tools allow users to identify genes specifically expressed during selected stresses, growth stages, or in particular organs. Using GENEVESTIGATOR, the gene expression profiles of more than 22,000 Arabidopsis genes can be obtained, including those of 10,600 currently uncharacterized genes. The objective of this software application is to direct gene functional discovery and design of new experiments by providing plant biologists with contextual information on the expression of genes. The database and analysis toolbox is available as a community resource at https://www.genevestigator.ethz.ch.

Arabidopsis↗

Mining proteases in the genome databases.

Protease data mining can take advantage both of the many specialist, Web-available databases that cover the genetic, protein and nucleic acid sequence information that is specific to a variety of organisms, and of a flexible, but defined, classification system. However, precomputed data, such as gene predictions, should be used with care. Unless there is definitive supporting information, ideally sequencing of a cDNA to show that the predictions are accurate, followed by expression and biochemical characterization of the predicted protein, the predicted gene and its product remains a possibility, rather than a certainty.

Animals↗

Computational inference of regulatory pathways in microbes: an application to phosphorus assimilation pathways in Synechococcus sp. WH8102.

We present a computational protocol for inference of regulatory and signaling pathways in a microbial cell, through literature search, mining "high-throughput'' biological data of various types, and computer-assisted human inference. This protocol consists of four key components: (a) construction of template pathways for microbial organisms related to the target genome, which either have been extensively studied and/or have a significant amount of (relevant) experimental data, (b) inference of initial pathway models for the target genome, through combining the template pathway models and target genome-specific information, (c) refinement and expansion of the initial pathway models through applications of various data mining tools, including phylogenetic profile analysis, inference of protein-protein interactions, and prediction of transcription factor binding sites, and (d) validation and refinement of the pathway models using pathway-specific experimental data or other information. To demonstrate the effectiveness of this procedure, we have applied it to the construction of the phosphorus assimilation pathways in cyanobacterium sp. WH8102. We present, in this paper, a model of the core components of this pathway.

Bacterial Proteins↗

Nonlinear mapping networks.

Among the many dimensionality reduction techniques that have appeared in the statistical literature, multidimensional scaling and nonlinear mapping are unique for their conceptual simplicity and ability to reproduce the topology and structure of the data space in a faithful and unbiased manner. However, a major shortcoming of these methods is their quadratic dependence on the number of objects scaled, which imposes severe limitations on the size of data sets that can be effectively manipulated. Here we describe a novel approach that combines conventional nonlinear mapping techniques with feed-forward neural networks, and allows the processing of data sets orders of magnitude larger than those accessible with conventional methodologies. Rooted on the principle of probability sampling, the method employs a classical algorithm to project a small random sample, and then "learns" the underlying nonlinear transform using a multilayer neural network trained with the back-propagation algorithm. Once trained, the neural network can be used in a feed-forward manner to project the remaining members of the population as well as new, unseen samples with minimal distortion. Using examples from the fields of image processing and combinatorial chemistry, we demonstrate that this method can generate projections that are virtually indistinguishable from those derived by conventional approaches. The ability to encode the nonlinear transform in the form of a neural network makes nonlinear mapping applicable to a wide variety of data mining applications involving very large data sets that are otherwise computationally intractable.

Journal Article↗

Identification of genetic pathways activated by the androgen receptor during the induction of proliferation in the ventral prostate gland.

The androgen receptor (AR), when complexed with 5alpha-dihydrotestosterone (DHT), supports the survival and proliferation of prostate cells, a process critical for normal development, benign prostatic hypertrophy, and tumorigenesis. However, the androgen-responsive genetic pathways that control prostate cell division and differentiation are largely unknown. To identify such pathways, we examined gene expression in the ventral prostate 6 and 24 h after DHT administration to androgen-depleted rats. 234 transcripts were expressed significantly differently from controls (p < 0.05) at both time points and were subjected to extensive data mining. Functional clustering of the data reveals that the majority of these genes can be classified as participating in induction of secretory activity, metabolic activation, and intracellular signaling/signal transduction, indicating that AR rapidly modulates the expression of genes involved in proliferation and differentiation in the prostate. Notably AR represses the expression of several key cell cycle inhibitors, while modulating members of the wnt and notch signaling pathways, multiple growth factors, and peptide hormone signaling systems, and genes involved in MAP kinase and calcium signaling. Analysis of these data also suggested that p53 activity is negatively regulated by AR activation even though p53 RNA was unchanged. Experiments in LNCaP prostate cancer cells reveal that AR inhibits p53 protein accumulation in the nucleus, providing a post-transcriptional mechanism by which androgens control prostate cell growth and survival. In summary these data provide a comprehensive view of the earliest events in AR-mediated prostate cell proliferation in vivo, and suggest that nuclear exclusion of p53 is a critical step in prostate growth.

Androgens↗

Nonclinical vehicle use in studies by multiple routes in multiple species.

The laboratory toxicologist is frequently faced with the challenge of selecting appropriate vehicles or developing utilitarian formulations for use in in vivo nonclinical safety assessment studies. Although there are many vehicles available that may meet physical and chemical requirements for chemical or pharmaceutical formulation, there are wide differences in species and route of administration specific to tolerances to these vehicles. In current practice, these differences are largely approached on a basis of individual experience as there is only scattered literature on individual vehicles and no comprehensive treatment or information source. This approach leads to excessive animal use and unplanned delays in testing and development. To address this need, a consulting firm and three contract research organizations conducted a rigorous data mining operation of control (vehicle) data from studies dating from 1991 to present. The results identified 65 single component vehicles used in 368 studies across multiple species (dog, primate, rat, mouse, rabbit, guinea pig, minipig, chick embryo, and cat) by multiple routes. Reported here are the results of this effort, including maximum tolerated use levels by species, route, and duration of study, with accompanying dose limiting toxicity. Also included are basic chemical information and a review of available literature on each vehicle, as well as guidance on volume limits and pH by route and some basic guidance on nonclinical formulation development.

Animals↗

Postural instability and consequent falls and hip fractures associated with use of hypnotics in the elderly: a comparative review.

The aim of this review is to establish the relationship between treatment with hypnotics and the risk of postural instability and as a consequence, falls and hip fractures, in the elderly. A review of the literature was performed through a search of the MEDLINE, Ingenta and PASCAL databases from 1975 to 2005. We considered as hypnotics only those drugs approved for treating insomnia, i.e. some benzodiazepines and the more recently launched 'Z'-compounds, i.e. zopiclone, zolpidem and zaleplon. Large-scale surveys consistently report increases in the frequency of falls and hip fractures when hypnotics are used in the elderly (2-fold risk). Benzodiazepines are the major class of hypnotics involved in this context; falls and fractures in patients taking Z-compounds are less frequently reported, and in this respect, zolpidem is considered as at risk in only one study. It is important to note, however, that drug adverse effect relationships are difficult to establish with this type of epidemiological data-mining. On the other hand, data obtained in laboratory settings, where confounding factors can be eliminated, prove that benzodiazepines are the most deleterious hypnotics at least in terms of their effects on body sway. Z-compounds are considered safer, probably because of their pharmacokinetic properties as well as their selective pharmacological activities at benzodiazepine-1 (BZ(1)) receptors. The effects of hypnotics on balance, gait and equilibrium are the consequence of differential negative impacts on vigilance and cognitive functions, and are highly dose- and time-dependent. Z-compounds have short half-lives and have less cognitive and residual effects than older medications. Some practical rules need to be followed when prescribing hypnotics in order to prevent falls and hip fractures as much as possible in elderly insomniacs, whether institutionalised or not. These are: (i) establish a clear diagnosis of the sleep disorder; (ii) take into account chronic conditions leading to balance and gait difficulties (motor and cognitive status); (iii) search for concomitant prescription of psychotropics and sedatives; (iv) use half the recommended adult dosage; and (v) declare any adverse effect to pharmacovigilance centres. Comparative pharmacovigilance studies focused on the impact of hypnotics on postural stability are very much needed.

Accidental Falls↗

A smooth response surface algorithm for constructing a gene regulatory network.

A smooth response surface (SRS) algorithm is developed as an elaborate data mining technique for analyzing gene expression data and constructing a gene regulatory network. A three-dimensional SRS is generated to capture the biological relationship between the target and activator-repressor. This new technique is applied to functionally describe triplets of activators, repressors, and targets, and their regulations in gene expression data. A diagnostic strategy is built into the algorithm to evaluate the scores of the triplets so that those with low scores are kept and a regulatory network is constructed based on this information and existing biological knowledge. The predictions based on the identified triplets in two yeast gene expression data sets agree with some experimental data in the literature. It provides a novel model with attractive mathematical and statistical features that make the algorithm valuable for mining expression or concentration information, assist in determining the function of uncharacterized proteins, and can lead to a better understanding of coherent pathways.

Algorithms↗

Mining association rules from clinical databases: an intelligent diagnostic process in healthcare.

Data mining is the process of discovering interesting knowledge, such as patterns, associations, changes, anomalies and significant structures, from large amounts of data stored in databases, data warehouses, or other information repositories. Mining Associations is one of the techniques involved in the process mentioned above and used in this paper. Association is the discovery of association relationships or correlations among a set of items. The algorithm that was implemented is a basic algorithm for mining association rules, known as a priori. In Healthcare, association rules are considered to be quite useful as they offer the possibility to conduct intelligent diagnosis and extract invaluable information and build important knowledge bases quickly and automatically. The problem of identifying new, unexpected and interesting patterns in medical databases in general, and diabetic data repositories in specific, is considered in this paper. We have applied the a priori algorithm to a database containing records of diabetic patients and attempted to extract association rules from the stored real parameters. The results indicate that the methodology followed may be of good value to the diagnostic procedure, especially when large data volumes are involved. The followed process and the implemented system offer an efficient and effective tool in the management of diabetes. Their clinical relevance and utility await the results of prospective clinical studies currently under investigation.

Algorithms↗

Profiling gene expression using onto-express.

Gene expression profiles obtained through microarray or data mining analyses often exist as vast data strings. To interpret the biology of these genetic profiles, investigators must analyze this data in the context of other information such as the biological, biochemical, or molecular function of the translated proteins. This is particularly challenging for a human analyst because large quantities of less than relevant data often bury such information. To address this need we implemented an automated routine, called Onto-Express (http://vortex.cs.wayne.edu:8080), to systematically translate genetic fingerprints into functional profiles. Using strings of accession or cluster identification numbers, Onto-Express searches the public databases and returns tables that correlate expression profiles with the cytogenetic locations, biochemical and molecular functions, biological processes, cellular components, and cellular roles of the translated proteins. The profiles created by Onto-Express fundamentally increase the value of gene expression analyses by facilitating the translation of quantitative value sets to records that contain biological implications.

Gene Expression Profiling↗

QSAR analysis of phenolic antioxidants using MOLMAP descriptors of local properties.

Molecular maps of atom-level properties (MOLMAPs) were developed to represent the diversity of chemical bonds existing in a molecule. Chemical reactivity, being related to the ability for bond breaking and bond making, is primarily determined by the properties of bonds available in a molecule. In order to use physicochemical properties of individual bonds for an entire molecule, and at the same time having a fixed-length molecular representation, all the bonds of a molecule are mapped into a fixed-size 2D self-organizing map (MOLMAP). This article illustrates the application of MOLMAP descriptors to QSAR, with a study of the radical scavenging activity of 47 naturally occurring phenolic antioxidants. Counterpropagation neural networks (CPG NNs) were trained with MOLMAP descriptors selected using genetic algorithms to predict antioxidant activity. The model was subsequently validated by the leave-one-out (LOO) procedure obtaining a q(2) of 0.71. Random Forests were grown with the entire set of MOLMAP descriptors giving 70% of correct classifications as potent, active or inactive in a LOO experiment. Interpretations of both models in terms of discriminant variables were concordant and allowed identifying bonds and substructures that are mostly responsible for antioxidant activity. This work shows how MOLMAPs can be used for data mining of structural and biological activity data, leading to the extraction of relationships between local properties and activity.

Algorithms↗

Protein-protein interaction map of the Trypanosoma cruzi ribosomal P protein complex.

The large subunit of the eukaryotic ribosome possesses a long and protruding stalk formed by the ribosomal P proteins. Four out of five ribosomal P proteins of Trypanosoma cruzi, TcP0, TcP1alpha, TcP2alpha, and TcP2beta had been previously characterized. Data mining of the T. cruzi genome data base allowed the identification of the fifth member of this protein group, a novel P1 protein, named P1beta. To gain insight into the assembly of the stalk, a yeast two-hybrid based protein interaction map was generated. A parasite specific profile of interactions amongst the ribosomal P proteins of T. cruzi was evident. The TcP0 protein was able to interact with all both P1 and both P2 proteins. Moreover, the interactions between P2beta with P1alpha as well as with P2alpha were detected, as well as the ability of TcP2beta to homodimerize. A quantitative evaluation of the interactions established that the strongest interacting pair was TcP0-TcP1beta.

Amino Acid Sequence↗

Object oriented database and electronic notebook for transmission electron microscopy.

As high-resolution biological transmission electron microscopy (TEM) has increased in popularity over recent years, the volume of data and number of projects underway has risen dramatically. A robust tool for effective data management is essential to efficiently process large data sets and extract maximum information from the available data. We present the Electron Microscopy Electronic Notebook (EMEN), a portable, object-oriented, web-based tool for TEM data archival and project management. EMEN has several unique features. First, the database is logically organized and annotated so multiple collaborators at different geographical locations can easily access and interpret the data without assistance. Second, the database was designed to provide flexibility to the user, so it can be used much as a lab notebook would be, while maintaining a structure suitable for data mining and direct interaction with data-processing software. Finally, as an object-oriented database, the database structure is dynamic and can be easily extended to incorporate information not defined in the original database specification.

Databases, Factual↗

HGVbase: a curated resource describing human DNA variation and phenotype relationships.

The Human Genome Variation Database (HGVbase; http://hgvbase.cgb.ki.se) has provided a curated summary of human DNA variation for more than 5 years, thus facilitating research into DNA sequence variation and human phenotypes. The database has undergone many changes and improvements to accommodate increasing volumes and new types of data. The focus of HGVbase has recently shifted towards information on haplotypes and phenotypes, relationships between phenotypes and DNA variation, and collaborative efforts to provide a global resource for genome-phenome data. Open sharing and precise phenotype definitions are necessary to advance the current understanding of common diseases that are typified by complex aetiologies, small genetic effect sizes and multiple confounding factors that obscure positive study results. Association data will increasingly be collected as part of this new project thrust. This report describes the evolving features of HGVbase, and covers in detail the technological choices we have made to enable efficient storage and data mining of increasingly large and complex data sets.

Computational Biology↗

Normalization of cDNA microarray data using wavelet regressions.

Normalization is an essential step in microarray data mining and analysis. For cDNA microarray data, the primary purpose of normalization is removing the intensity-dependent bias across different slides within an experimental group or between multiple groups. The locally weighted regression (lowess) procedure has been widely used for this purpose but can be comparatively time consuming when the dataset becomes relatively large. In this study, we applied wavelet regressions, a new smoothing method for recovering a regression function from data that is supposed to outperform other methods in many cases, such as spline or local polynomial fitting, to normalize two cDNA microarray datasets. Relative to the lowess procedure, we found that wavelet regressions not only produced reliable normalization results but also ran much faster. The computing speed represents one of the most important advantages over other algorithms, especially when one is interested in analyzing a large microarray experiment involving hundreds of slides.

Algorithms↗

Data warehouse implementation with clinical pharmacokinetic/pharmacodynamic data.

OBJECTIVES: We have created a data warehouse for human pharmacokinetic (PK) and pharmacodynamic (PD) data generated primarily within the Clinical PK Group of the Drug Metabolism and Pharmacokinetics (DM&PK) Department of DuPont Pharmaceuticals. METHODS: Data which enters an Oracle-based LIMS directly from chromatography systems or through files from contract research organizations are accessed via SAS/PH.Kinetics, GLP-compliant data analysis software residing on individual users' workstations. Upon completion of the final PK or PD analysis, data are pushed to a predefined location. Data analyzed/created with other software (i.e., WinNonlin, NONMEM, Adapt, etc.) are added to this file repository as well. The warehouse creates views to these data and accumulates metadata on all data sources defined in the warehouse. The warehouse is managed via the SAS/Warehouse Administrator product that defines the environment, creates summarized data structures, and schedules data refresh. RESULTS: The clinical PK/PD warehouse encompasses laboratory, biometric, PK and PD data streams. Detailed logical tables for each compound are created/updated as the clinical PK/PD data warehouse is populated. The data model defined to the warehouse is based on a star schema. Summarized data structures such as multidimensional data bases (MDDB), infomarts, and datamarts are created from detail tables. Data mining and querying of highly summarized data as well as drill-down to detail data is possible via the creation of exploitation tools which front-end the warehouse data. Based on periodic refreshing of the warehouse data, these applications are able to access the most current data available and do not require a manual interface to update/populate the data store. Prototype applications have been web-enabled to facilitate their usage to varied data customers across platform and location. The warehouse also contains automated mechanisms for the construction of study data listings and SAS transport files for eventual incorporation into an electronic submission. CONCLUSIONS: This environment permits the management of online analytical processing via a single administrator once the data model and warehouse configuration have been designed. The expansion of the current environment will eventually connect data from all phases of research and development ensuring the return on investment and hopefully efficiencies in data processing unforeseen with earlier legacy systems.

Adult↗