Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Multidimensional protein identification technology: current status and future prospects.

Protein profiling using high-throughput tandem mass spectrometry has become a powerful method for analyzing changes in global protein expression patterns in cells and tissues as a function of developmental, physiologic and disease processes. This review summarizes the utility and practical application of multidimensional protein identification technology as a platform for comprehensive proteomic profiling of complex biologic samples. The strengths and potential problems and limitations associated with this powerful technology are discussed, with an emphasis placed on one of the biggest challenges currently facing large-scale expression profiling projects -- namely, data analysis. Complementary bioinformatic computational data mining strategies, such as clustering, functional annotation and statistical inference, are also discussed as these are increasingly necessary for interpreting the results of global proteomic profiling studies.

Animals↗

Atypical antipsychotics and pituitary tumors: a pharmacovigilance study.

STUDY OBJECTIVE: To analyze the disproportionality of reporting of hyperprolactinemia, galactorrhea, and pituitary tumors with seven widely used antipsychotic drugs. DESIGN: Retrospective pharmacovigilance study. DATA SOURCE: United States Food and Drug Administration's Adverse Event Reporting System (AERS) database. INTERVENTION: We initially identified higher-than-expected postmarketing reports of pituitary tumors associated with risperidone, a potent dopamine D2-receptor antagonist antipsychotic, by analyzing reporting patterns of these tumors in the AERS database. To further examine this association, we analyzed disproportionate reporting patterns of pituitary tumor reports for seven antipsychotics with different affinities for blocking D2 receptors: aripiprazole, clozapine, olanzapine, quetiapine, risperidone, ziprasidone, and haloperidol. MEASUREMENTS AND MAIN RESULTS: To conduct both of these analyses, we used the Multi-item Gamma Poisson Shrinker (MGPS) data mining algorithm applied to the AERS database. The MGPS uses a Bayesian model to calculate adjusted observed:expected ratios of drug-adverse event associations (Empiric Bayes Geometric Mean [EBGM] values) in huge drug safety databases. The higher the adjusted reporting ratio, or EBGM value, the greater the strength of the association between a drug and an adverse event. Risperidone had the highest adjusted reporting ratios for hyperprolactinemia (EBGM 34.9, 90% confidence interval [CI] 32.8-37.1]), galactorrhea (EBGM 19.9, 90% CI 18.6-21.4), and pituitary tumor (EBGM 18.7, 90% CI 14.9-23.3) among the seven antipsychotics, and one of the highest scores for all drugs in the AERS database. Some tumors were associated with visual field defects, hemorrhage, convulsions, surgery, and severe (>10-fold) prolactin elevations. The EBGM values for risperidone for these adverse events were higher in women, but high EBGM values for these events were also seen in men and children. Moreover, the rank order of the EBGM values for pituitary tumors corresponded to the affinities of these seven drugs for D2 receptors. CONCLUSION: Treatment with potent D2-receptor antagonists, such as risperidone, may be associated with pituitary tumors. These findings are consistent with animal (mice) studies and raise the need for clinical awareness and longitudinal studies.

Adolescent↗

Pharmacogenomics and its potential impact on drug and formulation development.

Recent advances in genomic research have provided the basis for new insights into the importance of genetic and genomic markers during the different stages of drug development. A new field of research, pharmacogenomics, which studies the relationship between drug effects and the genome, has emerged. Structural pharmacogenomics maps the complete DNA sequences of whole genomes (genotypes) including individual variations, and functional pharmacogenomics assesses the expression levels of thousands of genes in one single experiment. Together, these two areas of pharmacogenomics have generated massive databases, which have become a challenge for the research field of informatics and have fostered a new branch of research, bioinformatics. If skillfully used, the databases generated by pharmacogenomics together with data mining on the Web promise to improve the drug development process in a variety of areas: identification of drug targets, evaluation of toxicity, classification of diseases, evaluation of formulations, assessment of drug response and treatment, post-marketing applications, and development of personalized medicines.

Animals↗

The secretin G-protein-coupled receptor family: teleost receptors.

Twenty-one members of the secretin family (family 2) of G-protein-coupled receptors (GPCRs) were identified via directed cloning and data-mining of the Fugu Genome Consortium database, representing the most comprehensive description of secretin GPCRs in a teleost fish to date. Duplicated genes were identified for many of the family members, namely the receptors for pituitary adenylate cyclase-activating polypeptide (PACAP)/vasoactive intestinal peptide (VIP), calcitonin, calcitonin gene-related peptide (CGRP), growth hormone releasing hormone (GHRH), glucagon receptor/glucagon-like peptide (GLP) and parathyroid hormone-related peptide (PTHrP)/PTH. Mining of other teleost genomes (zebrafish and Tetraodon) revealed that the duplicated genes identified in the Takifugu genome were also present in these fish. Additional database searching of the Escherichia coli, yeast, Drosophila, Caenorhabditis elegans and Ciona genomes revealed that the family 2 of GPCRs were only present in the multicellular organisms. Orthologues of all the human secretin receptors were identified with the exception of secretin itself. Additional database searches in the Fugu Genome Consortium database also failed to reveal a secretin ligand and so it is hypothesised that both the receptor and the ligand evolved after the divergence of teleost/tetrapod lineages. Phylogenetic analysis at both the protein and the DNA level provided strong support for each of the individual receptor family groupings, but weak support between groups, making evolutionary inferences difficult. A more critical analysis of the PACAP/VIP receptor family confirmed previous hypotheses that the vasoactive intestinal peptide receptor (VPAC(1)R) gene is the ancestral form of the receptor.

Animals↗

Gene expression profiling of bovine endometrium during the oestrous cycle: detection of molecular pathways involved in functional changes.

The endometrium plays a central role among the reproductive tissues in the context of early embryo-maternal communication and pregnancy. It undergoes typical changes during the sexual/oestrous cycle, which are regulated by the ovarian hormones progesterone and oestrogen. To identify the underlying molecular mechanisms we have performed the first holistic screen of transcriptome changes in bovine intercaruncular endometrium at two stages of the cycle--end of day 0 (late oestrus, low progesterone) and day 12 (dioestrus, high progesterone). A combination of subtracted cDNA libraries and cDNA array hybridisation revealed 133 genes showing at least a 2-fold change of their mRNA abundance, 65 with higher levels at oestrus and 68 at dioestrus. Interestingly, genes were identified which showed differential expression between different uterine sections as well. The most prominent example was the UTMP (uterine milk protein) mRNA, which was markedly upregulated in the cranial part of the ipsilateral uterine horn at oestrus. A Gene Ontology classification of the genes with known function characterised the oestrus time by elevated expression of genes, for example related to cell adhesion, cell motility and extracellular matrix and the dioestrus time by higher expression of mRNAs encoding for a variety of enzymes and transport proteins, in particular ion channels. Searching in pathway databases and literature data-mining revealed physiological processes and signalling cascades, e.g. the transforming growth factor-beta signalling pathway and retinoic acid signalling, which are potentially involved in the regulation of changes of the endometrium during the oestrous cycle.

Animals↗

A societal outcomes map for health research and policy.

The linkages between decisions about health research and policy and actual health outcomes may be extraordinarily difficult to specify. We performed a pilot application of a "road mapping" and technology assessment technique to perinatal health to illustrate how this technique can clarify the relations between available options and improved health outcomes. We used a combination of data-mining techniques and qualitative analyses to set up the underlying structure of a societal health outcomes road map. Societal health outcomes road mapping may be a useful tool for enhancing the ability of the public health community, policymakers, and other stakeholders, such as research administrators, to understand health research and policy options.

Health Policy↗

Visible-near infrared reflectance spectroscopy for rapid, nondestructive assessment of wetland soil quality.

Recent evidence supports using visible-near infrared reflectance spectroscopy (VNIRS) for sensing soil quality; advantages include low-cost, nondestructive, rapid analysis that retains high analytical accuracy for numerous soil performance measures. Research has primarily targeted agricultural applications (precision agriculture, performance diagnostics), but implications for assessing ecological systems are equally significant. Our objective was to extend chemometrics for sensing soil quality to wetlands. Hydric soils posed two challenges. First, wetland soils exhibit a wider range of organic matter concentrations, particularly in riparian areas where levels range from <1% in sedimentation zones to >90% in backwater floodplains; this may mute spectral responses from other soil fractions. Second, spectral inference of cation concentrations in terrestrial soils is for oxidized species; under reducing conditions in wetlands, oxidation state variability is observed, which strongly affects chroma. Riparian soils (n = 273) from western Florida exhibiting substantial target parameter variability were compiled. After minimal pre-processing, soils were scanned under artificial illumination using a laboratory spectrometer. A multivariate data mining technique (regression trees) was used to relate post-processed reflectance spectra to laboratory observations (pH, organic content, cation concentrations, total N, C, and P, extracellular enzyme activity). High validation accuracy was generally observed (r2(validation) > 0.8, RPD > 2.0, where RPD is the ratio of the standard deviation of an attribute to the observed standard error of validation); where accuracy was lower, categorical models (classification trees) successfully screened samples based on diagnostic functional thresholds (validation odds ratio > 10). Graphical models verified significant association between predictions and observations for all parameters, conditioning on biogeochemical covariates. Visible-near infrared reflectance spectroscopy offers both cost and statistical power advantages; hydric conditions do not appear to constrain application.

Cost Control↗

Machine learning for detecting gene-gene interactions: a review.

Complex interactions among genes and environmental factors are known to play a role in common human disease aetiology. There is a growing body of evidence to suggest that complex interactions are 'the norm' and, rather than amounting to a small perturbation to classical Mendelian genetics, interactions may be the predominant effect. Traditional statistical methods are not well suited for detecting such interactions, especially when the data are high dimensional (many attributes or independent variables) or when interactions occur between more than two polymorphisms. In this review, we discuss machine-learning models and algorithms for identifying and characterising susceptibility genes in common, complex, multifactorial human diseases. We focus on the following machine-learning methods that have been used to detect gene-gene interactions: neural networks, cellular automata, random forests, and multifactor dimensionality reduction. We conclude with some ideas about how these methods and others can be integrated into a comprehensive and flexible framework for data mining and knowledge discovery in human genetics.

Algorithms↗

Prediction of protein function in the absence of significant sequence similarity.

Tremendous progress in DNA sequencing has yielded the genomes of a host of important organisms. The utilisation of these resources requires understanding of the function of each gene. Standard methods of functional assignment involve sequence alignment to a gene of known function; however such methods often fail to find any significant matches. Here we discuss a number of recent alternative methods that may be of use when sequence alignment fails. Function can be defined in a number of ways including E.C. number and MIPS and KEGG functional classes. Phylogenetic profiles show the pattern of presence or absence of a protein between genomes. Protein-protein interactions can be identified by searching for interacting pairs of proteins that are fused to a single protein chain in another organism. The gene neighbour method uses the observation that if the genes that encode two proteins are close on a chromosome, the proteins tend to be functionally related. More general methods use sequence properties such as amino acid composition, mean hydrophobicity, predicted secondary structure and post-translational modification sites. Data mining methods devise rules in the form of IF... THEN statements that make predictions of function using sequence based attributes, predicted secondary structure and sequence similarity. Finally, structural features can be used, after modelling the structure of a protein from its sequence or solving its structure. Protein fold class can be strongly indicative of function, while other structural features, such as secondary structure content, cleft size and 3D structural motifs are also useful.

Computer Simulation↗

Molecular mapping in the CNS.

Since ancient times the operation of the brain has elicited more than usual interest. Data mining of the human genome is revealing that many CNS abnormalities have a genetic component. As yet this information can not be used directly to cure or ameliorate specific CNS disorders although this is regarded as having great potential for future therapies. Current CNS drug design and 3D QSAR is based on knowing either the structures of key proteins and how smaller molecules interact with them to obtain a pharmacological response, or on hypothesising about key structural features and interactions by a variety of molecular modelling and computational techniques. Methods used include conformational analyses, pharmacophore development and QSAR which are now being actively applied to increase our understanding of how molecules interact with specific sites within the CNS as a basis for the design of new pharmacologically active compounds. In this review we give an overview of the latest strategies used in 3D-QSAR based drug design and survey the most recent applications of these strategies to the CNS. By way of example, accounts are given of computer-based research aimed at drugs targeting GABA, glutamate, dopamine and opioid receptors.

Central Nervous System Agents↗

A role for information collection, management, and integration in structure-function studies of G-protein coupled receptors.

Elucidation of protein function is greatly facilitated by the availability of an atomic resolution structure or a reliable molecular model. The difficulty of obtaining atomic resolution structures of membrane proteins in general, and of G-protein coupled receptors (GPCRs) in particular, has made the information available from sequence analysis, mutagenesis, and the literature on related GPCRs exceptionally important. Here, we review previous studies of GPCR structure-function from the perspectives of sequence analysis, management of mutagenesis and ligand binding data, and literature data mining. The knowledge derived from these information resources not only constitutes the prerequisites for reliable molecular modeling, but also can provide other insights into GPCR functions. Finally, we review approaches for information integration and applying knowledge discovery techniques to structure-function studies of GPCRs, including molecular modeling itself.

Allosteric Regulation↗

Evaluation of multiple models to distinguish closely related forms of disease using DNA microarray data: an application to multiple myeloma.

MOTIVATION: Standard laboratory classification of the plasma cell dyscrasia monoclonal gammopathy of undetermined significance (MGUS) and the overt plasma cell neoplasm multiple myeloma (MM) is quite accurate, yet, for the most part, biologically uninformative. Most, if not all, cancers are caused by inherited or acquired genetic mutations that manifest themselves in altered gene expression patterns in the clonally related cancer cells. Microarray technology allows for qualitative and quantitative measurements of the expression levels of thousands of genes simultaneously, and it has now been used both to classify cancers that are morphologically indistinguishable and to predict response to therapy. It is anticipated that this information can also be used to develop molecular diagnostic models and to provide insight into mechanisms of disease progression, e.g., transition from healthy to benign hyperplasia or conversion of a benign hyperplasia to overt malignancy. However, standard data analysis techniques are not trivial to employ on these large data sets. Methodology designed to handle large data sets (or modified to do so) is needed to access the vital information contained in the genetic samples, which in turn can be used to develop more robust and accurate methods of clinical diagnostics and prognostics. RESULTS: Here we report on the application of a panel of statistical and data mining methodologies to classify groups of samples based on expression of 12,000 genes derived from a high density oligonucleotide microarray analysis of highly purified plasma cells from newly diagnosed MM, MGUS, and normal healthy donors. The three groups of samples are each tested against each other. The methods are found to be similar in their ability to predict group membership; all do quite well at predicting MM vs. normal and MGUS vs. normal. However, no method appears to be able to distinguish explicitly the genetic mechanisms between MM and MGUS. We believe this might be due to the lack of genetic differences between these two conditions, and may not be due to the failure of the models. We report the prediction errors for each of the models and each of the methods. Additionally, we report ROC curves for the results on group prediction. AVAILABILITY: Logistic regression: standard software, available, for example in SAS. Decision trees and boosted trees: C5.0 from www.rulequest.com. SVM: SVM-light is publicly available from svmlight.joachims.org. Naïve Bayes and ensemble of voters are publicly available from www.biostat.wisc.edu/~mwaddell/eov.html. Nearest Shrunken Centroids is publicly available from http://www-stat.stanford.edu/~tibs/PAM.

Journal Article↗

Radiology reporting: returning to our image-centric roots.

OBJECTIVE: Despite extraordinary advances in imaging and information technologies, the form and content of radiology reporting has changed little in the discipline's more than 100-year history. In this commentary, we outline the challenges that have confronted innovations such as speech recognition and structured reporting and call for a radical rethinking of the reporting process. By combining new applications with the expanding power of radiology and hospital information systems, the attention of the radiologist--and his or her referring colleagues--could be more focused on the image and its meaning. CONCLUSION: One promising result of such a change in focus could be improved and more reliable communication, already an area of heightened concern in the imaging community. Moreover, such a shift away from the printed word to image-centered content could lead to benefits in shared image viewing; more streamlined and timely reporting; data mining of aggregate results; and image archives, and, ultimately, enhancement of the consultative value of the radiologist's contribution to patient care and treatment.

Communication↗

[A review of potential signals generated by an automated method on 3324 pharmacovigilance case reports].

Automated signal generation aims to focus the attention of pharmacovigilance experts on drug-ADR associations which are disproportionally present in a spontaneous reporting system. Since 1986, we could find several signals using classic pharmacovigilance techniques with case reports registered in our pharmacovigilance regional centre. From this dataset 3,324 cases were related to spontaneous reporting. Drug-ADR associations were generated by using a Data Mining Algorithm (DMA) proposed by Evans et al. Potential signals were evaluated by reviewing case reports related to the unlabelled associations. The DMA generated 523 associations of which 107 were not described in the SPC. Most potential signals were false positives. Although the DMA generated little additional knowledge compared to signals already detected using classic techniques, the whole process helped us to focus our case review on a very small subset of the whole dataset (9.6%).

Adverse Drug Reaction Reporting Systems↗

Global distribution of Mycobacterium tuberculosis spoligotypes.

We present a short summary of recent observations on the global distribution of the major clades of the Mycobacterium tuberculosis complex, the causative agent of tuberculosis. This global distribution was defined by data-mining of an international spoligotyping database, SpolDB3. This database contains 11708 patterns from as many clinical isolates originating from more than 90 countries. The 11708 spoligotypes were clustered into 813 shared types. A total of 1300 orphan patterns (clinical isolates showing a unique spoligotype) were also detected.

Databases, Factual↗

Cyclin D1 and molecular chaperones: implications for tumorigenesis.

We recently investigated the mechanisms of cyclin D1 action in human cancer using global analyses of gene expression. With an experimentally-determined expression signature for cyclin D1 overexpression, gene expression data from human tumors, and a novel data-mining method, we were able to reveal a previously unappreciated and apparently predominant functional interdependency between cyclin D1 and C/EBPbeta. Many of the genes we found to be affected by cyclin D1 overexpression are recognized as molecular chaperones or their regulators. Might this provide insights to the role of the cyclin D1-C/EBPbeta axis in carcinogenesis?

Animals↗

Emerging research: a view from one research center.

The Health Management Research Center at the University of Michigan has assembled a database on health risks, medical care costs, an in some cases, productivity measures for over 2,000,000 individuals. For employees of its corporate consortium members, the database contains seven to eighteen years of data. Working with this data, the research team has observed a number of emerging trends. These trends have been stable in this data set for a number of years, but some of them are yet to be subjected to rigorous external peer review. The trends are summarized below. 1) Annual participation rates of 20% to 30% in Health Risk Appraisal are typical; over 10 years, 80% participate at least once, 60% at least twice and 40% at least three times. 2) Among the employers in the data base, excess risk factors account for 21% to 31% of medical care costs, with a mean of 25%. 3) Medical care costs increase as the number risk factors and age increase. As risk factors increase, medical costs increase; as risk factors decrease, medical care costs decrease. The mean cost increase per risk factor increased ($350) may be more than double the mean cost decrease per risk factor decreased ($150). 4) Cost savings greatest among those who participate in programs multiple times. 5) Absenteeism seems to be higher and other measures of productivity lower for those with health risk factors. 6) Programs designed to keep healthy people healthy in addition to reducing the risks of those with multiple risks will probably provide the greatest return to the employers. 7) Best results may be achieved by focusing efforts on employees who have clusters of risk factors associated with low perceived health status. 8) A corporate wellness score which combines risk factor levels and participation rates may provide a "corporate wellness score" which can be used to compare health status across employer. 9) Increased use of longitudinal data sets, fuzzy cut points for data categories and data mining techniques may allow breakthroughs in future analysis efforts.

Academies and Institutes↗

Risk assessment prediction from genome sequences: promises and dreams.

The application of bacterial genomics opens new avenues of research on foodborne pathogens. Foodborne pathogens must be able to colonize their hosts and survive transmission from host to host. Different groups of genes are involved in the processes of survival, colonization, and virulence, and such genes are potential targets for risk assessment and intervention strategies. Filtering from genome sequences the genes relevant to these processes is a major challenge, and although many tools are already available for analyses, this type of data mining is just beginning. For the simplest application, gene comparison, it is important to know how gene function, for instance in virulence, is being defined and tested. In other genomic applications, reserachers look for specific properties or characteristics of (virulence) genes to identify novel gene candidates. Each approach has pitfalls, and gene candidates must be tested in the lab to confirm their function. Models for colonization and virulence are available for most although not all pathogens. Models for survival and stress responses are needed to increase the utilization of genomic approaches to risk assessment. Here, I discuss how genome sequences are likely to help in microbial risk assessment of foodborne pathogens and how dreams may become promises.

Animals↗