Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

From data collection to knowledge data discovery: a medical application of data mining.

Prison inmates are exposed to a variety of major risk factors (psychiatric disorders, suicide attempts, illicit drug use). From 1986 to 1996, the USA prison population more than doubled while in France, it increased from 35655 in 1980 to 51623 in 1995. In spite of these findings, very little information concerning the inmates population is available. At the present time, there is a desire to adopt a policy based on the prevention of recidivism, on adequate release planning and referrals to community-based services. The aim of the RAPPEL project was to build an information system for assessing the social and health status of prison inmates. The pilot project was set up at the prison of Loos and allowed the collection and analysis of nearly 15000 records. The aim of this paper is to present the extension of the project consisting in developing a regional network grouping 11 jails. Information locally available will serve as the basis for the information system of regional jails. Data mining techniques will provide solutions for the extraction of new information. Three data mining tools were experimented : association rules, classification trees and clustering. Further extension consists in a distributed approach allowing direct access to the information system by WEB tools.

Classification↗

Using data mining to characterize DNA mutations by patient clinical features.

In most hereditary cancer syndromes, finding a correspondence between various genetic mutations within a gene (genotype) and a patient's clinical cancer history (phenotype) is challenging; to date there are few clinically meaningful correlations between specific DNA intragenic mutations and corresponding cancer types. To define possible genotype and phenotype correlations, we evaluated the application of data mining methodology whereby the clinical cancer histories of gene-mutation-positive patients were used to define valid or "true" patterns for a specific DNA intragenic mutation. The clinical histories of patients with their corresponding detailed attributes without the same oncologic intragenic mutation were labeled incorrect or "false" patterns. The results of data mining technology yielded characterizing rules for the true cases that constituted clinical features which predicted the intragenic mutation. Some of the initial results derived correlations already independently known in the literature, adding to the confidence of using this methodological approach.

Age of Onset↗

Whole-body gene expression by data mining.

To date, a comprehensive survey of the expression of lysyl oxidase (LOX), lysyl oxidase-like 1 (LOXL1), and lysyl oxidase-like 2 (LOXL2) has yet to be performed. The use of in vitro strategies to accomplish this task would prove daunting as it is both time-consuming and costly. We present a new in silico data mining strategy that directly addresses these limitations. Sequences corresponding to the 3' untranslated regions of LOX, LOXL1, and LOXL2 were individually queried against the human expressed sequence tag database (dbEST). In this manner, the entire tissue repertoire available in the dbEST was surveyed. This provided an estimate of the levels of mRNA transcripts in a variety of adult and fetal tissues. We have also employed this strategy to determine the pattern of expression and levels of a newly discovered gene, CGI-15. The veracity of this technique has been independently assessed by semiquantitative PCR analysis. The application of this technology is bounded only by the ever-growing information available in the GenBank, UniGene, and human EST databases. The utility of our data mining strategy to establish relative transcript levels in numerous tissues is presented.

3' Untranslated Regions↗

Knowledge representation forms for data mining methodologies as applied in thoracic surgery.

Typical ways of disseminating and using results of clinical research are scientific journals and reports. Presentation forms are condensed and comprehensible mainly to the experts following the specific topics. A vast amount of information remains unutilized due to the complex form of presenting the knowledge. Subject of this research is to explore possibilities of representation and also visualization of the results obtained using data mining methodologies. The intention is to formulate more than scientific ways to communicate facts that are of interest for the clinicians, medical students and even patients. Internet technologies as already widely established media support knowledge representation forms such as hypertext documents and structured knowledge components. The "Assist Me" decision support system for surgical treatment of cardiac patients integrates several forms of data mining and representation methodologies. We are showing a feasibility study in which scientific outcomes were forwarded to a broad group of potential users.

Artificial Intelligence↗

Reviewing mobile phases used on Chiralcel OD through an application of data mining tools to CHIRBASE database.

During the past decade, thousands of compounds have been resolved on Chiralcel OD (a cellulose-based chiral stationary phase) under diverse eluting conditions. Many researches have documented the effects of mobile phase on enantioselectivity for a given family of samples but today no comprehensive study aimed at identifying the associations between the structural features present on solute and appropriate mobile phase conditions has yet been proposed. In this review of mobile phases used on Chiralcel OD, we try to go far beyond a simple enumeration of eluting conditions and an effort is made to explore the utility of data mining tools for assessing the knowledge contained in CHIRBASE database. We have extracted from CHIRBASE the chemical features of 2363 chiral compounds separated on Chiralcel OD and their corresponding mobile phases. This data set was submitted to data mining programs for molecular pattern recognition and mobile phase predictions for new cases. Some substructural characteristics of solutes were related to the efficient use of some specific mobile phases. For example, the application of CH3CN/salt buffer at pH 6-7 was found convenient for reversed-phase separation of compounds bearing a tertiary amine functional group. Furthermore, a cluster analysis allowed the arrangement of the mobile phases according to similarity found in molecular patterns of solutes. A decision tree, which may lead to a more rational choice of the mobile phase under reversed-phase conditions, is also proposed.

Carbamates↗

Creating structure features by data mining the PDB to use as molecular-replacement models.

Mathematical data-mining techniques to generate a representative set of protein fragments are described. Protein fragments are used as search models within the macromolecular phasing method of molecular replacement to attempt to phase protein data without a homologous model correctly. Preliminary investigations using these fragments indicate that molecular replacement with AMoRe is not sensitive enough to phase myoglobin or insulin data sufficiently for successful refinement. The results suggest that more advanced molecular replacement techniques may be successful, though at present these are not computationally practical.

Amino Acid Motifs↗

Integrated data mining and network pharmacology to explore the prescription patterns from a senior TCM oncologist's clinical practice in treating chemotherapy-induced hand-foot syndrome.

Hand-foot syndrome (HFS) is a common and refractory adverse effect of chemotherapy lacking specific therapeutic strategies currently. Traditional Chinese medicine (TCM) has shown empirical efficacy in clinical HFS management. This study integrated data mining and network pharmacology to systematically elucidate the medication principles and molecular mechanisms underlying Professor Gang Xie's prescriptions for HFS. All medical records from Professor Xie's specialist clinic (January 2020 to March 2025) were retrospectively collected and standardized in Excel. Prescriptions were analyzed through frequency statistics, association and clustering. Active ingredients of core herb pairs and their disease-related targets were identified using TCMSP, HERB, GeneCards, PharmGKB and GEO databases. Protein-protein interaction (PPI) networks, gene ontology (GO), and Kyoto encyclopedia of genes and genomes (KEGG) pathway analyses were performed. Molecular docking validated interactions between key bioactive compounds and targets. This study involved 217 prescriptions containing 150 herbs. Core herb combinations comprised Radix Astragali (Huangqi), Poria (Fuling), and Radix Pseudostellariae (Taizishen), predominantly classified as spleen-tonifying agents with warm properties, targeting lung, spleen, and stomach meridians. Network analysis identified 67 bioactive compounds and 899 disease targets. Quercetin, kaempferol, acacetin and luteolin were identified the key ingredients. The core targets (TP53, STAT3, PIK3CA, HSP90AA1, AKT1, CTNNB1, PI3KR1, MAPK1) were enriched in MAPK and PI3K-Akt signaling pathways. Molecular docking confirmed strong binding affinity between key compounds and targets. Professor Xie's therapeutic strategy for HFS emphasizes "spleen fortification, phlegm elimination, and stasis resolution." The core herb combination likely exerts anti-HFS effects via modulation of MAPK and PI3K-Akt pathways, providing a pharmacological basis for TCM-driven HFS management.

Network Pharmacology↗

A data mining approach to characterizing medical code usage patterns.

This research describes a synthetic data mining approach to identifying diagnostic (ICD-9) and procedure (CPT) code usage patterns in two US. hospitals, with the goal of determining the adequacy and effectiveness of the current coding classification systems. We combine relative frequency measurements with measures of industry concentration borrowed from industrial economics in order to (1) ascertain the extent to which physicians utilize the available codes in classifying patients and (2) discover the factors that impinge on code usage. Our results partition the domain into areas for which the coding systems perform well and those areas for which the systems perform relatively poorly. The goal is to use this approach to understand how coding systems are used and to highlight areas for targeted improvement of the current coding

Data Interpretation, Statistical↗

Sharing medical data for patient path analysis with data mining method.

The Agora Data project started in October 1997 in France. The objective was to share medical data between several medical institutions to analysis medical care pathways for patients that suffer from low back pain. The analysis of the medical records decomposed in three steps allowed us to produce knowledge on medical contacts of patients with the health care system. In order to study the relations between these contacts, we created medical path of patients within the framework of the possible contacts we had isolated. This work relates the implementation and the first results of the pilot study.

Confidentiality↗

Antipsychotic drugs and heart muscle disorder in international pharmacovigilance: data mining study.

OBJECTIVES: To examine the relation between antipsychotic drugs and myocarditis and cardiomyopathy. DESIGN: Data mining using bayesian statistics implemented in a neural network architecture. SETTING: International database on adverse drug reactions run by the World Health Organization programme for international drug monitoring. MAIN OUTCOME MEASURES: Reports mentioning antipsychotic drugs, cardiomyopathy, or myocarditis. RESULTS: A strong signal existed for an association between clozapine and cardiomyopathy and myocarditis. An association was also seen with other antipsychotics as a group. The association was based on sufficient cases with adequate documentation and apparent lack of confounding to constitute a signal. Associations between myocarditis or cardiomyopathy and lithium, chlorpromazine, fluphenazine, haloperidol, and risperidone need further investigation. CONCLUSIONS: Some antipsychotic drugs seem to be linked to cardiomyopathy and myocarditis. The study shows the potential of bayesian neural networks in analysing data on drug safety.

Antipsychotic Agents↗

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease↗

Selected techniques for data mining in medicine.

Widespread use of medical information systems and explosive growth of medical databases require traditional manual data analysis to be coupled with methods for efficient computer-assisted analysis. This paper presents selected data mining techniques that can be applied in medicine, and in particular some machine learning techniques including the mechanisms that make them better suited for the analysis of medical databases (derivation of symbolic rules, use of background knowledge, sensitivity and specificity of induced descriptions). The importance of the interpretability of results of data analysis is discussed and illustrated on selected medical applications.

Adult↗

XML-based visual data mining in medicine.

Medical databases in general are characterized by a high degree of complexity in terms of quantity of items, number of parameter values and data types (free text, categorical, numerical and other). Substantial domain knowledge is required for adequate formalization of medical entities. In this context we developed medical database plot (mdplot), a data mining tool to visualize both structure and quality of data in medical databases to identify items suitable for evaluation. Data models are provided in XML format. Missing data is identified to enable targeted efforts to improve data quality prior to analysis. Database items are classified as 1:1- related to the patient (i.e. variables are collected once per patient) and 1:n related. mdplot provides a list of all classes contained in a database, the number of records each and a condensed bar chart for semi-quantitative description of completeness according to four types of items: categorical, numerical, text and other. All items in a category are grouped from left to right, the height of each bar represents the proportion of non-missing values with respect to the total number of records in the class; thus the amount of content in a specific class is visualized. By selection of a specific class, a detailed description of it is provided including mean completeness in each item category as well as number of values per item. The new methodology was applied to a cardiological research database consisting of 619 items on 88 patients.

Atrial Fibrillation↗

The microarray explorer tool for data mining of cDNA microarrays: application for the mammary gland.

The Microarray Explorer (MAExplorer) is a versatile Java-based data mining bioinformatic tool for analyzing quantitative cDNA expression profiles across multiple microarray platforms and DNA labeling systems. It may be run as either a stand-alone application or as a Web browser applet over the Internet. With this program it is possible to (i) analyze the expression of individual genes, (ii) analyze the expression of gene families and clusters, (iii) compare expression patterns and (iv) directly access other genomic databases for clones of interest. Data may be downloaded as required from a Web server or in the case of the stand-alone version, reside on the user's computer. Analyses are performed in real-time and may be viewed and directly manipulated in images, reports, scatter plots, histograms, expression profile plots and cluster analyses plots. A key feature is the clone data filter for constraining a working set of clones to those passing a variety of user-specified logical and statistical tests. Reports may be generated with hypertext Web access to UniGene, GenBank and other Internet databases for sets of clones found to be of interest. Users may save their explorations on the Web server or local computer and later recall or share them with other scientists in this groupware Web environment. The emphasis on direct manipulation of clones and sets of clones in graphics and tables provides a high level of interaction with the data, making it easier for investigators to test ideas when looking for patterns. We have used the MAExplorer to profile gene expression patterns of 1500 duplicated genes isolated from mouse mammary tissue. We have identified genes that are preferentially expressed during pregnancy and during lactation. One gene we identified, carbonic anhydrase III, is highly expressed in mammary tissue from virgin and pregnant mice and in gene knock-out mice with underdeveloped mammary epithelium. Other genes, which include those encoding milk proteins, are preferentially expressed during lactation.

Animals↗

Data mining goes multidimensional.

The success of a healthcare organization depends on its ability to acquire, store, analyze and compare data across many parts of the enterprise, by many individuals. While relational databases have been around since the 1970s, their two-dimensional structure has limited--or made impossible--the kind of cross-dimensional trend analysis so necessary to healthcare today. Enter online analytical processing (OLAP), in which servers store data in multiple dimensions, opening a world of opportunity for data-mining across the enterprise. In this issue of HEALTHCARE INFORMATICS, we feature our first report from the National Software Testing Laboratories (NSTL) about technologies that will change the way healthcare does business. A division of The McGraw-Hill Companies, NSTL is an independent software and hardware testing lab offering services that include compatibility testing, bug testing, comparison testing, documentation evaluation and usability.

Computer User Training↗

Comparative genomics using data mining tools.

We have analysed the genomes of representatives of three kingdoms of life, namely, archaea, eubacteria and eukaryota using data mining tools based on compositional analyses of the protein sequences. The representatives chosen in this analysis were Methanococcus jannaschii, Haemophilus influenzae and Saccharomyces cerevisiae. We have identified the common and different features between the three genomes in the protein evolution patterns. M. jannaschii has been seen to have a greater number of proteins with more charged amino acids whereas S. cerevisiae has been observed to have a greater number of hydrophilic proteins. Despite the differences in intrinsic compositional characteristics between the proteins from the different genomes we have also identified certain common characteristics. We have carried out exploratory Principal Component Analysis of the multivariate data on the proteins of each organism in an effort to classify the proteins into clusters. Interestingly, we found that most of the proteins in each organism cluster closely together, but there are a few 'outliers'. We focus on the outliers for the functional investigations, which may aid in revealing any unique features of the biology of the respective organisms

Archaeal Proteins↗

Data warehouse and data mining in a surgical clinic.

Hospitals and clinics have taken advantage of information systems to streamline many clinical and administrative processes. However, the potential of health care information technology as a source of data for clinical and administrative decision support has not been fully explored. In response to pressure for timely information, many hospitals are developing clinical data warehouses. This paper attempts to identify problem areas in the process of developing a data warehouse to support data mining in surgery. Based on the experience from a data warehouse in surgery several solutions are discussed.

Databases, Bibliographic↗

yMGV: a database for visualization and data mining of published genome-wide yeast expression data.

The yeast Microarray Global Viewer (yMGV) is an on-line database providing a synthetic view of the transcriptional expression profiles of Saccharomyces cerevisiae genes in most of the published expression datasets. yMGV displays a one-screen graphical representation of gene expression variations for each published genome-wide experiment, allowing quick retrieval of experimental conditions affecting expression of this gene. yMGV also provides tools to isolate groups of genes sharing similar transcription profiles in a defined subset of experiments. Additionally, yMGV furnishes a set of statistical tools for critical assessment of published data. We therefore believe that yMGV is an efficient tool that affords a quick and comprehensive overview of microarray data and generates new gene classifications. As of 20 March 2001 the yMGV database contains 6 000 000 measurements, representing genome-wide expression comparisons of 932 experiments from 39 microarray publications. The yMGV interface is available at http://transcriptome.ens.fr/ymgv/.

Computational Biology↗