Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Sharing medical data for patient path analysis with data mining method.

The Agora Data project started in October 1997 in France. The objective was to share medical data between several medical institutions to analysis medical care pathways for patients that suffer from low back pain. The analysis of the medical records decomposed in three steps allowed us to produce knowledge on medical contacts of patients with the health care system. In order to study the relations between these contacts, we created medical path of patients within the framework of the possible contacts we had isolated. This work relates the implementation and the first results of the pilot study.

Confidentiality↗

Antipsychotic drugs and heart muscle disorder in international pharmacovigilance: data mining study.

OBJECTIVES: To examine the relation between antipsychotic drugs and myocarditis and cardiomyopathy. DESIGN: Data mining using bayesian statistics implemented in a neural network architecture. SETTING: International database on adverse drug reactions run by the World Health Organization programme for international drug monitoring. MAIN OUTCOME MEASURES: Reports mentioning antipsychotic drugs, cardiomyopathy, or myocarditis. RESULTS: A strong signal existed for an association between clozapine and cardiomyopathy and myocarditis. An association was also seen with other antipsychotics as a group. The association was based on sufficient cases with adequate documentation and apparent lack of confounding to constitute a signal. Associations between myocarditis or cardiomyopathy and lithium, chlorpromazine, fluphenazine, haloperidol, and risperidone need further investigation. CONCLUSIONS: Some antipsychotic drugs seem to be linked to cardiomyopathy and myocarditis. The study shows the potential of bayesian neural networks in analysing data on drug safety.

Antipsychotic Agents↗

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease↗

Association analysis for quantitative traits by data mining: QHPM.

Previously, we have presented a data mining-based algorithmic approach to genetic association analysis, Haplotype Pattern Mining. We have now extended the approach with the possibility of analysing quantitative traits and utilising covariates. This is accomplished by using a linear model for measuring association. We present results with the extended version, QHPM, with simulated quantitative trait data. One data set was simulated with the population simulator package Populus, and another was obtained from GAW12. In the former, there were 2-3 underlying susceptibility genes for a trait, each with several ancestral disease mutations, and 1 or 2 environmental components. We show that QHPM is capable of finding the susceptibility loci, even when there is strong allelic heterogeneity and environmental effects in the disease models. The power of finding quantitative trait loci is dependent on the ascertainment scheme of the data: collecting the study subjects from both ends of the quantitative trait distribution is more effective than using unselected individuals or individuals ascertained based on disease status, but QHPM has good power to localize the genes even with unselected individuals. Comparison with quantitative trait TDT (QTDT) showed that QHPM has better localization accuracy when the gene effect is weak.

Chromosome Mapping↗

Selected techniques for data mining in medicine.

Widespread use of medical information systems and explosive growth of medical databases require traditional manual data analysis to be coupled with methods for efficient computer-assisted analysis. This paper presents selected data mining techniques that can be applied in medicine, and in particular some machine learning techniques including the mechanisms that make them better suited for the analysis of medical databases (derivation of symbolic rules, use of background knowledge, sensitivity and specificity of induced descriptions). The importance of the interpretability of results of data analysis is discussed and illustrated on selected medical applications.

Adult↗

XML-based visual data mining in medicine.

Medical databases in general are characterized by a high degree of complexity in terms of quantity of items, number of parameter values and data types (free text, categorical, numerical and other). Substantial domain knowledge is required for adequate formalization of medical entities. In this context we developed medical database plot (mdplot), a data mining tool to visualize both structure and quality of data in medical databases to identify items suitable for evaluation. Data models are provided in XML format. Missing data is identified to enable targeted efforts to improve data quality prior to analysis. Database items are classified as 1:1- related to the patient (i.e. variables are collected once per patient) and 1:n related. mdplot provides a list of all classes contained in a database, the number of records each and a condensed bar chart for semi-quantitative description of completeness according to four types of items: categorical, numerical, text and other. All items in a category are grouped from left to right, the height of each bar represents the proportion of non-missing values with respect to the total number of records in the class; thus the amount of content in a specific class is visualized. By selection of a specific class, a detailed description of it is provided including mean completeness in each item category as well as number of values per item. The new methodology was applied to a cardiological research database consisting of 619 items on 88 patients.

Atrial Fibrillation↗

The microarray explorer tool for data mining of cDNA microarrays: application for the mammary gland.

The Microarray Explorer (MAExplorer) is a versatile Java-based data mining bioinformatic tool for analyzing quantitative cDNA expression profiles across multiple microarray platforms and DNA labeling systems. It may be run as either a stand-alone application or as a Web browser applet over the Internet. With this program it is possible to (i) analyze the expression of individual genes, (ii) analyze the expression of gene families and clusters, (iii) compare expression patterns and (iv) directly access other genomic databases for clones of interest. Data may be downloaded as required from a Web server or in the case of the stand-alone version, reside on the user's computer. Analyses are performed in real-time and may be viewed and directly manipulated in images, reports, scatter plots, histograms, expression profile plots and cluster analyses plots. A key feature is the clone data filter for constraining a working set of clones to those passing a variety of user-specified logical and statistical tests. Reports may be generated with hypertext Web access to UniGene, GenBank and other Internet databases for sets of clones found to be of interest. Users may save their explorations on the Web server or local computer and later recall or share them with other scientists in this groupware Web environment. The emphasis on direct manipulation of clones and sets of clones in graphics and tables provides a high level of interaction with the data, making it easier for investigators to test ideas when looking for patterns. We have used the MAExplorer to profile gene expression patterns of 1500 duplicated genes isolated from mouse mammary tissue. We have identified genes that are preferentially expressed during pregnancy and during lactation. One gene we identified, carbonic anhydrase III, is highly expressed in mammary tissue from virgin and pregnant mice and in gene knock-out mice with underdeveloped mammary epithelium. Other genes, which include those encoding milk proteins, are preferentially expressed during lactation.

Animals↗

Data mining goes multidimensional.

The success of a healthcare organization depends on its ability to acquire, store, analyze and compare data across many parts of the enterprise, by many individuals. While relational databases have been around since the 1970s, their two-dimensional structure has limited--or made impossible--the kind of cross-dimensional trend analysis so necessary to healthcare today. Enter online analytical processing (OLAP), in which servers store data in multiple dimensions, opening a world of opportunity for data-mining across the enterprise. In this issue of HEALTHCARE INFORMATICS, we feature our first report from the National Software Testing Laboratories (NSTL) about technologies that will change the way healthcare does business. A division of The McGraw-Hill Companies, NSTL is an independent software and hardware testing lab offering services that include compatibility testing, bug testing, comparison testing, documentation evaluation and usability.

Computer User Training↗

Comparative genomics using data mining tools.

We have analysed the genomes of representatives of three kingdoms of life, namely, archaea, eubacteria and eukaryota using data mining tools based on compositional analyses of the protein sequences. The representatives chosen in this analysis were Methanococcus jannaschii, Haemophilus influenzae and Saccharomyces cerevisiae. We have identified the common and different features between the three genomes in the protein evolution patterns. M. jannaschii has been seen to have a greater number of proteins with more charged amino acids whereas S. cerevisiae has been observed to have a greater number of hydrophilic proteins. Despite the differences in intrinsic compositional characteristics between the proteins from the different genomes we have also identified certain common characteristics. We have carried out exploratory Principal Component Analysis of the multivariate data on the proteins of each organism in an effort to classify the proteins into clusters. Interestingly, we found that most of the proteins in each organism cluster closely together, but there are a few 'outliers'. We focus on the outliers for the functional investigations, which may aid in revealing any unique features of the biology of the respective organisms

Archaeal Proteins↗

Data mining for seeking an accurate quantitative relationship between molecular structure and GC retention indices of alkenes by projection pursuit.

Primary data mining on alkenes for seeking an accurate quantitative relationship between the molecular structure and retention indices of gas chromatography is developed in this paper. Based on the results obtained from projection pursuit, all alkenes investigated show an interesting classification. Thus, a new variable named class distance variable of alkenes, which essentially describes information about the branch, position of the double bonds, the number of double bonds, and so on for alkenes, is proposed. With the help of the new variable, both fitting and prediction accuracy of the regression model can be dramatically improved. The results obtained in this work show that the technique of projection pursuit developed in statistics is a quite promising tool for seeking an accurate quantitative structure-retention relationship (QSRR).

Journal Article↗

Data warehouse and data mining in a surgical clinic.

Hospitals and clinics have taken advantage of information systems to streamline many clinical and administrative processes. However, the potential of health care information technology as a source of data for clinical and administrative decision support has not been fully explored. In response to pressure for timely information, many hospitals are developing clinical data warehouses. This paper attempts to identify problem areas in the process of developing a data warehouse to support data mining in surgery. Based on the experience from a data warehouse in surgery several solutions are discussed.

Databases, Bibliographic↗

yMGV: a database for visualization and data mining of published genome-wide yeast expression data.

The yeast Microarray Global Viewer (yMGV) is an on-line database providing a synthetic view of the transcriptional expression profiles of Saccharomyces cerevisiae genes in most of the published expression datasets. yMGV displays a one-screen graphical representation of gene expression variations for each published genome-wide experiment, allowing quick retrieval of experimental conditions affecting expression of this gene. yMGV also provides tools to isolate groups of genes sharing similar transcription profiles in a defined subset of experiments. Additionally, yMGV furnishes a set of statistical tools for critical assessment of published data. We therefore believe that yMGV is an efficient tool that affords a quick and comprehensive overview of microarray data and generates new gene classifications. As of 20 March 2001 the yMGV database contains 6 000 000 measurements, representing genome-wide expression comparisons of 932 experiments from 39 microarray publications. The yMGV interface is available at http://transcriptome.ens.fr/ymgv/.

Computational Biology↗

A data mining technique for discovering distinct patterns of hand signs: implications in user training and computer interface design.

Hand signs are considered as one of the important ways to enter information into computers for certain tasks. Computers receive sensor data of hand signs for recognition. When using hand signs as computer inputs, we need to (1) train computer users in the sign language so that their hand signs can be easily recognized by computers, and (2) design the computer interface to avoid the use of confusing signs for improving user input performance and user satisfaction. For user training and computer interface design, it is important to have a knowledge of which signs can be easily recognized by computers and which signs are not distinguishable by computers. This paper presents a data mining technique to discover distinct patterns of hand signs from sensor data. Based on these patterns, we derive a group of indistinguishable signs by computers. Such information can in turn assist in user training and computer interface design.

Algorithms↗

Protecting patient privacy in clinical data mining.

This paper investigates whether HIPAA de-identification requirements--as well as proposed AAMC de-identification standards--were met in a large clinical data mining study (1997-2001) conducted at Duke University prior to the publication of the final rule. While HIPAA has improved de-identification standards, the study also shows that privacy issues may persist even in de-identified large clinical databases.

Biomedical Research↗

Data mining of sequences and 3D structures of allergenic proteins.

MOTIVATION: Many sequences, and in some cases structures, of proteins that induce an allergic response in atopic individuals have been determined in recent years. This data indicates that allergens, regardless of source, fall into discreet protein families. Similarities in the sequence may explain clinically observed cross-reactivities between different biological triggers. However, previously available allergy databases group allergens according to their biological sources, or observed clinical cross-reactivities, without providing data about the proteins. A computer-aided data mining system is needed to compare the sequential and structural details of known allergens. This information will aid in predicting allergenic cross-responses and eventually in determining possible common characteristics of IgE recognition. RESULTS: The new web-based Structural Database of Allergenic Proteins (SDAP) permits the user to quickly compare the sequence and structure of allergenic proteins. Data from literature sources and previously existing lists of allergens are combined in a MySQL interactive database with a wide selection of bioinformatics applications. SDAP can be used to rapidly determine the relationship between allergens and to screen novel proteins for the presence of IgE or T-cell epitopes they may share with known allergens. Further, our novel similarity search method, based on five dimensional descriptors of amino acid properties, can be used to scan the SDAP entries with a peptide sequence. For example, when a known IgE binding epitope from shrimp tropomyosin was used as a query, the method rapidly identified a similar sequence in known shellfish and insect allergens. This prediction of cross-reactivity between allergens is consistent with clinical observations. AVAILABILITY: SDAP is available on the web at http://fermi.utmb.edu/SDAP/index.html

Allergens↗

Data mining with decision trees for diagnosis of breast tumor in medical ultrasonic images.

To increase the ability of ultrasonographic (US) technology for the differential diagnosis of solid breast tumors, we describe a novel computer-aided diagnosis (CADx) system using data mining with decision tree for classification of breast tumor to increase the levels of diagnostic confidence and to provide the immediate second opinion for physicians. Cooperating with the texture information extracted from the region of interest (ROI) image, a decision tree model generated from the training data in a top-down, general-to-specific direction with 24 co-variance texture features is used to classify the tumors as benign or malignant. In the experiments, accuracy rates for a experienced physician and the proposed CADx are 86.67% (78/90) and 95.50% (86/90), respectively.

Breast Neoplasms↗

Data mining parasite genomes: haystack searching with a computer.

A number of genomes of parasitic organisms are presently being sequenced in the public domain, including Plasmodium falciparum, Leishmania major and Trypanosoma brucei with the likelihood of at least expressed sequence tag (EST) projects for several filarial and apicomplexan species. The early and timely release of sequence data to the community via the World Wide Web (www), and the public databases, (EMBL and GENBANK), forms an invaluable resource. Data mining, or 'haystack searching' this resource is becoming more fruitful to all members of the scientific community as the volume of data, diversity of genomes sampled, and accessibility increase.

Animals↗

Use of 3D QSAR methodology for data mining the National Cancer Institute Repository of Small Molecules: application to HIV-1 reverse transcriptase inhibition.

A three-dimensional (3D) stereoelectronic pharmacophore developed from a 3D quantitative structure-activity relationship (QSAR) investigation formed the basis of the development of a two-phase data-mining methodology to uncover novel leads to inhibit human immunodeficiency virus type 1 (HIV-1) reverse transcriptase at the nonnucleoside binding site. The database searching phase employed a field search for ligand requirements (such as log P, molecular volume) that were accessible from the database keys. Next, a 3D database search was performed that used an automated fitting procedure and the calculation of several binding parameters. These binding parameters were used to test the hits by a discriminant function that was previously trained to recognize active from inactive analogs. During the structural evaluation phase of the methodology, conformational properties and complementary receptor features of the hits were examined by 2D and 3D evaluations, which were followed by molecular modeling investigations. When this method was applied to a test database, an improvement from 6.4% to 100% active analogs was achieved.

Database Management Systems↗