Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Creating structure features by data mining the PDB to use as molecular-replacement models.

Mathematical data-mining techniques to generate a representative set of protein fragments are described. Protein fragments are used as search models within the macromolecular phasing method of molecular replacement to attempt to phase protein data without a homologous model correctly. Preliminary investigations using these fragments indicate that molecular replacement with AMoRe is not sensitive enough to phase myoglobin or insulin data sufficiently for successful refinement. The results suggest that more advanced molecular replacement techniques may be successful, though at present these are not computationally practical.

Amino Acid Motifs↗

Bayesian analysis, pattern analysis, and data mining in health care.

PURPOSE OF REVIEW: To discuss the current role of data mining and Bayesian methods in biomedicine and heath care, in particular critical care. RECENT FINDINGS: Bayesian networks and other probabilistic graphical models are beginning to emerge as methods for discovering patterns in biomedical data and also as a basis for the representation of the uncertainties underlying clinical decision-making. At the same time, techniques from machine learning are being used to solve biomedical and health-care problems. SUMMARY: With the increasing availability of biomedical and health-care data with a wide range of characteristics there is an increasing need to use methods which allow modeling the uncertainties that come with the problem, are capable of dealing with missing data, allow integrating data from various sources, explicitly indicate statistical dependence and independence, and allow integrating biomedical and clinical background knowledge. These requirements have given rise to an influx of new methods into the field of data analysis in health care, in particular from the fields of machine learning and probabilistic graphical models.

Bayes Theorem↗

Integrated data mining and network pharmacology to explore the prescription patterns from a senior TCM oncologist's clinical practice in treating chemotherapy-induced hand-foot syndrome.

Hand-foot syndrome (HFS) is a common and refractory adverse effect of chemotherapy lacking specific therapeutic strategies currently. Traditional Chinese medicine (TCM) has shown empirical efficacy in clinical HFS management. This study integrated data mining and network pharmacology to systematically elucidate the medication principles and molecular mechanisms underlying Professor Gang Xie's prescriptions for HFS. All medical records from Professor Xie's specialist clinic (January 2020 to March 2025) were retrospectively collected and standardized in Excel. Prescriptions were analyzed through frequency statistics, association and clustering. Active ingredients of core herb pairs and their disease-related targets were identified using TCMSP, HERB, GeneCards, PharmGKB and GEO databases. Protein-protein interaction (PPI) networks, gene ontology (GO), and Kyoto encyclopedia of genes and genomes (KEGG) pathway analyses were performed. Molecular docking validated interactions between key bioactive compounds and targets. This study involved 217 prescriptions containing 150 herbs. Core herb combinations comprised Radix Astragali (Huangqi), Poria (Fuling), and Radix Pseudostellariae (Taizishen), predominantly classified as spleen-tonifying agents with warm properties, targeting lung, spleen, and stomach meridians. Network analysis identified 67 bioactive compounds and 899 disease targets. Quercetin, kaempferol, acacetin and luteolin were identified the key ingredients. The core targets (TP53, STAT3, PIK3CA, HSP90AA1, AKT1, CTNNB1, PI3KR1, MAPK1) were enriched in MAPK and PI3K-Akt signaling pathways. Molecular docking confirmed strong binding affinity between key compounds and targets. Professor Xie's therapeutic strategy for HFS emphasizes "spleen fortification, phlegm elimination, and stasis resolution." The core herb combination likely exerts anti-HFS effects via modulation of MAPK and PI3K-Akt pathways, providing a pharmacological basis for TCM-driven HFS management.

Network Pharmacology↗

A data mining approach to characterizing medical code usage patterns.

This research describes a synthetic data mining approach to identifying diagnostic (ICD-9) and procedure (CPT) code usage patterns in two US. hospitals, with the goal of determining the adequacy and effectiveness of the current coding classification systems. We combine relative frequency measurements with measures of industry concentration borrowed from industrial economics in order to (1) ascertain the extent to which physicians utilize the available codes in classifying patients and (2) discover the factors that impinge on code usage. Our results partition the domain into areas for which the coding systems perform well and those areas for which the systems perform relatively poorly. The goal is to use this approach to understand how coding systems are used and to highlight areas for targeted improvement of the current coding

Data Interpretation, Statistical↗

Mass spectrometric genomic data mining: Novel insights into bioenergetic pathways in Chlamydomonas reinhardtii.

A new high-throughput computational strategy was established that improves genomic data mining from MS experiments. The MS/MS data were analyzed by the SEQUEST search algorithm and a combination of de novo amino acid sequencing in conjunction with an error-tolerant database search tool, operating on a 256 processor computer cluster. The error-tolerant search tool, previously established as GenomicPeptideFinder (GPF), enables detection of intron-split and/or alternatively spliced peptides from MS/MS data when deduced from genomic DNA. Isolated thylakoid membranes from the eukaryotic green alga Chlamydomonas reinhardtii were separated by 1-D SDS gel electrophoresis, protein bands were excised from the gel, digested in-gel with trypsin and analyzed by coupling nano-flow LC with MS/MS. The concerted action of SEQUEST and GPF allowed identification of 2622 distinct peptides. In total 448 peptides were identified by GPF analysis alone, including 98 intron-split peptides, resulting in the identification of novel proteins, improved annotation of gene models, and evidence of alternative splicing.

Algorithms↗

Sharing medical data for patient path analysis with data mining method.

The Agora Data project started in October 1997 in France. The objective was to share medical data between several medical institutions to analysis medical care pathways for patients that suffer from low back pain. The analysis of the medical records decomposed in three steps allowed us to produce knowledge on medical contacts of patients with the health care system. In order to study the relations between these contacts, we created medical path of patients within the framework of the possible contacts we had isolated. This work relates the implementation and the first results of the pilot study.

Confidentiality↗

Common denominator procedure: a novel approach to gene-expression data mining for identification of phenotype-specific genes.

MOTIVATION: We have established a novel data mining procedure for the identification of genes associated with pre-defined phenotypes and/or molecular pathways. Based on the observation that these genes are frequently expressed in the same place or in close proximity at about the same time, we have devised an approach termed Common Denominator Procedure. One unusual feature of this approach is that the specificity and probability to identify genes linked to the desired phenotype/pathway increase with greater diversity of the input data. RESULT: To show the feasibility of our approach, the Cancer Genome Anatomy Project expression data combined with a defined set of angiogenic factors was used to identify additional and novel angiogenesis-associated genes. A multitude of these additional genes were known to be associated with angiogenesis according to published data, verifying our approach. For some of the remaining candidate genes, application of a high-throughput functional genomics platform (XantoScreen) provided further experimental evidence for association with angiogenesis.

Angiogenic Proteins↗

Data mining and genetic algorithm based gene/SNP selection.

OBJECTIVE: Genomic studies provide large volumes of data with the number of single nucleotide polymorphisms (SNPs) ranging into thousands. The analysis of SNPs permits determining relationships between genotypic and phenotypic information as well as the identification of SNPs related to a disease. The growing wealth of information and advances in biology call for the development of approaches for discovery of new knowledge. One such area is the identification of gene/SNP patterns impacting cure/drug development for various diseases. METHODS: A new approach for predicting drug effectiveness is presented. The approach is based on data mining and genetic algorithms. A global search mechanism, weighted decision tree, decision-tree-based wrapper, a correlation-based heuristic, and the identification of intersecting feature sets are employed for selecting significant genes. RESULTS: The feature selection approach has resulted in 85% reduction of number of features. The relative increase in cross-validation accuracy and specificity for the significant gene/SNP set was 10% and 3.2%, respectively. CONCLUSION: The feature selection approach was successfully applied to data sets for drug and placebo subjects. The number of features has been significantly reduced while the quality of knowledge was enhanced. The feature set intersection approach provided the most significant genes/SNPs. The results reported in the paper discuss associations among SNPs resulting in patient-specific treatment protocols.

Algorithms↗

Application of data mining to predict the dosage of vancomycin as an outcome variable in a teaching hospital population.

OBJECTIVE: Data mining is a process used to extract potentially valuable information hidden in large volumes of raw data. The aim of this study was to explore the possibility of using easy to implement and effective supervised learning techniques to predict the dosage of vancomycin. METHODS: To reach this goal, we considered the prediction of the dosage of vancomycin as a classification problem. We chose the C4.5 decision tree technique for the dosage prediction process and supplied it with a boosting technique to enhance its performance. RESULTS: The potential predictor variables were collected from 833 patients with methicillin-resistant Staphylococcus aureus, or penicillin intolerance who were being treated with vancomycin and undergoing therapeutic drug monitoring (TDM) after attainment of steady state blood concentrations. Attributes tested as potential predictors included age, sex, weight, serum creatinine concentration, dosing interval, and variables from 1-compartment model kinetics such as Kd, Vd, and t(1/2). CONCLUSIONS: The results showed that the proposed method can utilize a variety of parameters to predict the dosage of vancomycin in the population used and that it performs well over a range of patient ages and renal function. The method may offer an alternative to existing methods used to support decision-making in clinical practice.

Adolescent↗

Antipsychotic drugs and heart muscle disorder in international pharmacovigilance: data mining study.

OBJECTIVES: To examine the relation between antipsychotic drugs and myocarditis and cardiomyopathy. DESIGN: Data mining using bayesian statistics implemented in a neural network architecture. SETTING: International database on adverse drug reactions run by the World Health Organization programme for international drug monitoring. MAIN OUTCOME MEASURES: Reports mentioning antipsychotic drugs, cardiomyopathy, or myocarditis. RESULTS: A strong signal existed for an association between clozapine and cardiomyopathy and myocarditis. An association was also seen with other antipsychotics as a group. The association was based on sufficient cases with adequate documentation and apparent lack of confounding to constitute a signal. Associations between myocarditis or cardiomyopathy and lithium, chlorpromazine, fluphenazine, haloperidol, and risperidone need further investigation. CONCLUSIONS: Some antipsychotic drugs seem to be linked to cardiomyopathy and myocarditis. The study shows the potential of bayesian neural networks in analysing data on drug safety.

Antipsychotic Agents↗

Data mining server--on-line knowledge induction tool.

The aim of this paper is to present an on-line data mining tool and illustrate its use on example of real medical data. Data from the Laboratory for in-vitro Thyroid diagnostics at the Sisters of Charity University Hospital in Zagreb were used. Preparation of the data set and one session of knowledge induction is described.

Croatia↗

Protein folding and unfolding simulations: a new challenge for data mining.

One of the unsolved paradigms in molecular biology is the protein folding problem. In recent years, with the identification of several diseases as protein folding disorders and with the explosion of genome information and the need for efficient ways to predict protein structure, protein folding became a central issue in molecular sciences research. Using molecular dynamics unfolding simulations of an amyloidogenic protein--transthyretin--as an example, we put forward a series of ideas on how simulations of this type may be used to infer rules and unfolding behavior in amyloidogenic proteins, and to extrapolate rules for protein folding in different structural classes of proteins. These, in turn, could help in the development of protein structure prediction methods. The need to analyse different proteins and to run multiple simulations creates a huge amount of data which has to be stored, managed, analyzed and shared (database and Grid technology; data mining). Once the data is captured, the next challenge is to find meaningful patterns (associations, correlations, clusters, rules, relationships) among molecular properties, or their relative importance at different stages of the folding or unfolding processes. This clearly puts new and interesting challenges to the bioinformatics community.

Computational Biology↗

Molecular determinants for ATP-binding in proteins: a data mining and quantum chemical analysis.

Adenosine 5'-triphosphate (ATP) plays an essential role in all forms of life. Molecular recognition of ATP in proteins is a subject of great importance for understanding enzymatic mechanism and for drug design. We have carried out a large-scale data mining of the Protein Data Bank (PDB) to analyze molecular determinants for recognition of the adenine moiety of ATP by proteins. Non-bonded intermolecular interactions (hydrogen bonding, pi-pi stacking interactions, and cation-pi interactions) between adenine base and surrounding residues in its binding pockets are systematically analyzed for 68 non-redundant, high-resolution crystal structures of adenylate-binding proteins. In addition to confirming the importance of the widely known hydrogen bonding, we found out that cation-pi interactions between adenine base and positively charged residues (Lys and Arg) and pi-pi stacking interactions between adenine base and surrounding aromatic residues (Phe, Tyr, Trp) are also crucial for adenine binding in proteins. On average, there exist 2.7 hydrogen bonding interactions, 1.0 pi-pi stacking interactions, and 0.8 cation-pi interactions in each adenylate-binding protein complex. Furthermore, a high-level quantum chemical analysis was performed to analyze contributions of each of the three forms of intermolecular interactions (i.e. hydrogen bonding, pi-pi stacking interactions, and cation-pi interactions) to the overall binding force of the adenine moiety of ATP in proteins. Intermolecular interaction energies for representative configurations of intermolecular complexes were analyzed using the supermolecular approach at the MP2/6-311 + G* level, which resulted in substantial interaction strengths for all the three forms of intermolecular interactions. This work represents a timely undertaking at a historical moment when a large number of X-ray crystallographic structures of proteins with bound ATP ligands have become available, and when high-level quantum chemical analysis of intermolecular interactions of large biomolecular systems becomes computationally feasible. The establishment of the molecular basis for recognition of the adenine moiety of ATP in proteins will directly impact molecular design of ATP-binding site targeted enzyme inhibitors such as kinase inhibitors.

Adenine↗

Analysis of hospitalised patient flows using data-mining.

UNLABELLED: Face to the development of hospital information system in the "Hôpital Européen George Pompidou" (HEGP), computerized patients records made medical data easier to analyse than before. We use data-mining technology to analyse intra-hospital patients' paths with one year of PMSI data (a French medical information system similar to Diagnosis Related Group). METHODS: 1. "sequential patterns mining" was used to analyse the most frequent patients' paths, 2. an integrated framework of "association rules mining" and "classification rule mining" was used to build prediction rules of patients' paths. RESULT: We construct a rule based prediction model, which gives the tendency of the patient's paths between the different medical units.

Computer Communication Networks↗

ONCOMINE: a cancer microarray database and integrated data-mining platform.

DNA microarray technology has led to an explosion of oncogenomic analyses, generating a wealth of data and uncovering the complex gene expression patterns of cancer. Unfortunately, due to the lack of a unifying bioinformatic resource, the majority of these data sit stagnant and disjointed following publication, massively underutilized by the cancer research community. Here, we present ONCOMINE, a cancer microarray database and web-based data-mining platform aimed at facilitating discovery from genome-wide expression analyses. To date, ONCOMINE contains 65 gene expression datasets comprising nearly 48 million gene expression measurements form over 4700 microarray experiments. Differential expression analyses comparing most major types of cancer with respective normal tissues as well as a variety of cancer subtypes and clinical-based and pathology-based analyses are available for exploration. Data can be queried and visualized for a selected gene across all analyses or for multiple genes in a selected analysis. Furthermore, gene sets can be limited to clinically important annotations including secreted, kinase, membrane, and known gene-drug target pairs to facilitate the discovery of novel biomarkers and therapeutic targets.

Databases, Genetic↗

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease↗

Association analysis for quantitative traits by data mining: QHPM.

Previously, we have presented a data mining-based algorithmic approach to genetic association analysis, Haplotype Pattern Mining. We have now extended the approach with the possibility of analysing quantitative traits and utilising covariates. This is accomplished by using a linear model for measuring association. We present results with the extended version, QHPM, with simulated quantitative trait data. One data set was simulated with the population simulator package Populus, and another was obtained from GAW12. In the former, there were 2-3 underlying susceptibility genes for a trait, each with several ancestral disease mutations, and 1 or 2 environmental components. We show that QHPM is capable of finding the susceptibility loci, even when there is strong allelic heterogeneity and environmental effects in the disease models. The power of finding quantitative trait loci is dependent on the ascertainment scheme of the data: collecting the study subjects from both ends of the quantitative trait distribution is more effective than using unselected individuals or individuals ascertained based on disease status, but QHPM has good power to localize the genes even with unselected individuals. Comparison with quantitative trait TDT (QTDT) showed that QHPM has better localization accuracy when the gene effect is weak.

Chromosome Mapping↗