Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Data mining with decision trees for diagnosis of breast tumor in medical ultrasonic images.

To increase the ability of ultrasonographic (US) technology for the differential diagnosis of solid breast tumors, we describe a novel computer-aided diagnosis (CADx) system using data mining with decision tree for classification of breast tumor to increase the levels of diagnostic confidence and to provide the immediate second opinion for physicians. Cooperating with the texture information extracted from the region of interest (ROI) image, a decision tree model generated from the training data in a top-down, general-to-specific direction with 24 co-variance texture features is used to classify the tumors as benign or malignant. In the experiments, accuracy rates for a experienced physician and the proposed CADx are 86.67% (78/90) and 95.50% (86/90), respectively.

Breast Neoplasms↗

Data mining parasite genomes: haystack searching with a computer.

A number of genomes of parasitic organisms are presently being sequenced in the public domain, including Plasmodium falciparum, Leishmania major and Trypanosoma brucei with the likelihood of at least expressed sequence tag (EST) projects for several filarial and apicomplexan species. The early and timely release of sequence data to the community via the World Wide Web (www), and the public databases, (EMBL and GENBANK), forms an invaluable resource. Data mining, or 'haystack searching' this resource is becoming more fruitful to all members of the scientific community as the volume of data, diversity of genomes sampled, and accessibility increase.

Animals↗

Forensic visualization of foreign matter in human tissue by near-infrared spectral imaging: methodology and data mining strategies.

BACKGROUND: Rapidity of data acquisition, high image fidelity and large field of view are of tremendous value when looking for chemical contaminants or for the proverbial "needle in the haystack" - in this case foreign inclusions in histologic sections of biopsy or autopsy tissues. Near infrared chemical imaging is one of three chemical imaging techniques (NIR, MIR and Raman) based on vibrational spectroscopy, and provides distinct technical advantages for this application. METHODS: We have chosen to utilize and evaluate near infrared (NIR) imaging for studies of foreign materials in tissue because the experimental configuration is relatively simple, data collection is rapid, and large sample areas can be screened with high image fidelity and spatial resolution. RESULTS: We have shown that NIR imaging can readily find and identify silicone gel inclusions in biological tissue samples. Additionally, preliminary results indicate that spectral signatures in the data set are also potentially sensitive to structural changes in the surrounding tissue that may be induced by the foreign body. CONCLUSIONS: NIR chemical imaging is a powerful, non-destructive tool for localization and identifying foreign contaminants in biological tissue. Preliminary results indicate that NIR imaging is also sensitive enough to differentiate tissue types (perhaps based on collagen structural differences), and provide data on the spatial localization of these components.

Animals↗

Cancer gene search with data-mining and genetic algorithms.

Cancer leads to approximately 25% of all mortalities, making it the second leading cause of death in the United States. Early and accurate detection of cancer is critical to the well being of patients. Analysis of gene expression data leads to cancer identification and classification, which will facilitate proper treatment selection and drug development. Gene expression data sets for ovarian, prostate, and lung cancer were analyzed in this research. An integrated gene-search algorithm for genetic expression data analysis was proposed. This integrated algorithm involves a genetic algorithm and correlation-based heuristics for data preprocessing (on partitioned data sets) and data mining (decision tree and support vector machines algorithms) for making predictions. Knowledge derived by the proposed algorithm has high classification accuracy with the ability to identify the most significant genes. Bagging and stacking algorithms were applied to further enhance the classification accuracy. The results were compared with that reported in the literature. Mapping of genotype information to the phenotype parameters will ultimately reduce the cost and complexity of cancer detection and classification.

Algorithms↗

Development of a clinical pathways analysis system with adaptive Bayesian nets and data mining techniques.

The use and development of software in the medical field offers tremendous opportunities for making health care delivery more efficient, more effective, and less error-prone. We discuss and explore the use of clinical pathways analysis with Adaptive Bayesian Networks and Data Mining Techniques to perform such analyses. The computation of "lift" (a measure of completed pathways improvement potential) leads us to optimism regarding the potential for this approach.

Bayes Theorem↗

Database of repetitive elements in complete genomes and data mining using transcription factor binding sites.

Approximately 43% of the human genome is occupied by repetitive elements. Even more, around 51% of the rice genome is occupied by repetitive elements. The analysis presented here indicates that repetitive elements in complete genomes may have been very important in the evolutionary genomics. In this study, a database, called the Repeat Sequence Database, is first designed and implemented to store complete and comprehensive repetitive sequences. See http://rsdb.csie.ncu.edu.tw for more information. The database contains direct, inverted and palindromic repetitive sequences, and each repetitive sequence has a variable length ranging from seven to many hundred nucleotides. The repetitive sequences in the database are explored using a mathematical algorithm to mine rules on how combinations of individual binding sites are distributed among repetitive sequences in the database. Combinations of transcription factor binding sites in the repetitive sequences are obtained and then data mining techniques are applied to mine association rules from these combinations. The discovered associations are further pruned to remove insignificant associations and obtain a set of associations. The mined association rules facilitate efforts to identify gene classes regulated by similar mechanisms and accurately predict regulatory elements. Experiments are performed on several genomes including C. elegans, human chromosome 22, and yeast.

Algorithms↗

Use of 3D QSAR methodology for data mining the National Cancer Institute Repository of Small Molecules: application to HIV-1 reverse transcriptase inhibition.

A three-dimensional (3D) stereoelectronic pharmacophore developed from a 3D quantitative structure-activity relationship (QSAR) investigation formed the basis of the development of a two-phase data-mining methodology to uncover novel leads to inhibit human immunodeficiency virus type 1 (HIV-1) reverse transcriptase at the nonnucleoside binding site. The database searching phase employed a field search for ligand requirements (such as log P, molecular volume) that were accessible from the database keys. Next, a 3D database search was performed that used an automated fitting procedure and the calculation of several binding parameters. These binding parameters were used to test the hits by a discriminant function that was previously trained to recognize active from inactive analogs. During the structural evaluation phase of the methodology, conformational properties and complementary receptor features of the hits were examined by 2D and 3D evaluations, which were followed by molecular modeling investigations. When this method was applied to a test database, an improvement from 6.4% to 100% active analogs was achieved.

Database Management Systems↗

Data-mining analysis suggests an epigenetic pathogenesis for type 2 diabetes.

The etiological origin of type 2 diabetes mellitus (T2DM) has long been controversial. The body of literature related to T2DM is vast and varied in focus, making a broad epidemiological perspective difficult, if not impossible. A data-mining approach was used to analyze all electronically available scientific literature, over 12 million Medline records, for "objects" such as genes, diseases, phenotypes, and chemical compounds linked to other objects within the T2DM literature but were not themselves within the T2DM literature. The goal of this analysis was to conduct a comprehensive survey to identify novel factors implicated in the pathology of T2DM by statistically evaluating mutually shared associations. Surprisingly, epigenetic factors were among the highest statistical scores in this analysis, strongly implicating epigenetic changes within the body as causal factors in the pathogenesis of T2DM. Further analysis implicates adipocytes as the potential tissue of origin, and cytokines or cytokine-like genes as the dysregulated factor(s) responsible for the T2DM phenotype. The analysis provides a wealth of literature supporting this hypothesis, which-if true-represents an important paradigm shift for researchers studying the pathogenesis of T2DM.

Journal Article↗

Visual management of large scale data mining projects.

This paper describes a unified framework for visualizing the preparations for, and results of, hundreds of machine learning experiments. These experiments were designed to improve the accuracy of enzyme functional predictions from sequence, and in many cases were successful. Our system provides graphical user interfaces for defining and exploring training datasets and various representational alternatives, for inspecting the hypotheses induced by various types of learning algorithms, for visualizing the global results, and for inspecting in detail results for specific training sets (functions) and examples (proteins). The visualization tools serve as a navigational aid through a large amount of sequence data and induced knowledge. They provided significant help in understanding both the significance and the underlying biological explanations of our successes and failures. Using these visualizations it was possible to efficiently identify weaknesses of the modular sequence representations and induction algorithms which suggest better learning strategies. The context in which our data mining visualization toolkit was developed was the problem of accurately predicting enzyme function from protein sequence data. Previous work demonstrated that approximately 6% of enzyme protein sequences are likely to be assigned incorrect functions on the basis of sequence similarity alone. In order to test the hypothesis that more detailed sequence analysis using machine learning techniques and modular domain representations could address many of these failures, we designed a series of more than 250 experiments using information-theoretic decision tree induction and naive Bayesian learning on local sequence domain representations of problematic enzyme function classes. In more than half of these cases, our methods were able to perfectly discriminate among various possible functions of similar sequences. We developed and tested our visualization techniques on this application.

Alcohol Dehydrogenase↗

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology↗

Neural Query System: Data-mining from within the NEURON simulator.

We have developed a simulation tool within the NEURON simulator to assist in organization, verification, and analysis of simulations. This tool, denominated Neural Query System (NQS), provides a relational database system, a query function based on the SELECT function of Structured Query Language, and data-mining tools. We show how NQS can be used to organize, manage, verify, and visualize parameters for both single cell and network simulations. We demonstrate an additional use of NQS to organize simulation output and relate outputs to parameters in a network model. The NQS software package is available at http://senselab. med.yale.edu/senselab/SimToolDB.

Animals↗

A retrospective evaluation of a data mining approach to aid finding new adverse drug reaction signals in the WHO international database.

BACKGROUND: The detection of new drug safety signals is of growing importance with ever more new drugs becoming available and exposure to medicines increasing. The task of evaluating information relating to safety lies with national agencies and, for international data, with the World Health Organization Programme for International Drug Monitoring. RATIONALE: An established approach for identifying new drug safety signals from the international database of more than 2 million case reports depends upon clinical experts from around the world. With a very large amount of information to evaluate, such an approach is open to human error. To aid the clinical review, we have developed a new signalling process using Bayesian logic, applied to data mining, within a confidence propagation neural network (Bayesian Confidence Propagation Neural Network; BCPNN). Ultimately, this will also allow the evaluation of complex variables. METHODS: The first part of this study tested the predictive value of the BCPNN in new signal detection as compared with reference literature sources (Martindale's Extra Pharmacopoeia in 1993 and July 2000, and the Physicians Desk Reference in July 2000). In the second part of the study, results with the BCPNN method were compared with those of the former signalling procedure. RESULTS: In the study period (the first quarter of 1993) 107 drug-adverse reaction combinations were highlighted as new positive associations by the BCPNN, and referred to new drugs. 15 drug-adverse reaction combinations on new drugs became negative BCPNN associations in the study period. The BCPNN method detected signals with a positive predictive value of 44% and the negative predictive value was 85%. 17 as yet unconfirmed positive associations could not be dismissed with certainty as false positive signals. Of the 10 drug-adverse reaction signals produced by the former signal detection system from data sent out for review during the study period, 6 were also identified by the BCPNN. These 6 associations have all had a more than 10-fold increase of reports and 4 of them have been included in the reference sources. The remaining 4 signals that were not identified by the BCPNN had a small, or no, increase in the number of reports, and are not listed in the reference sources. CONCLUSION: Our evaluation showed that the BCPNN approach had a high and promising predictive value in identifying early signals of new adverse drug reactions.

Algorithms↗

Knowledge discovery and data mining to assist natural language understanding.

As natural language processing systems become more frequent in clinical use, methods for interpreting the output of these programs become increasingly important. These methods require the effort of a domain expert, who must build specific queries and rules for interpreting the processor output. Knowledge discovery and data mining tools can be used instead of a domain expert to automatically generate these queries and rules. C5.0, a decision tree generator, was used to create a rule base for a natural language understanding system. A general-purpose natural language processor using this rule base was tested on a set of 200 chest radiograph reports. When a small set of reports, classified by physicians, was used as the training set, the generated rule base performed as well as lay persons, but worse than physicians. When a larger set of reports, using ICD9 coding to classify the set, was used for training the system, the rule base performed worse than the physicians and lay persons. It appears that a larger, more accurate training set is needed to increase performance of the method.

Artificial Intelligence↗

Patient-recognition data-mining model for BCG-plus interferon immunotherapy bladder cancer treatment.

Bladder cancer is the fifth most common malignant disease in the United States with an annual incidence of around 63,210 new cases and 13,180 deaths. The cost for providing care for patients with bladder cancer disease is high. Bladder cancer treatment options such as immunotherapy, chemotherapy, radiation therapy, transurethral resection, and cystectomy, are used with varying success rates. In this research, data from a nationwide bacillus Calmette-Gue rin (BCG) plus interferon-alpha (IFN-alpha) immunotherapy clinical trial was considered. Data mining algorithms were used to analyze the effectiveness of immunotherapy treatment and to understand the prominent parameters and their interactions. The extracted knowledge was used to build a patient recognition model for prediction of treatment outcomes. The data was analyzed to understand the impact of various parameters on the treatment outcome. A list of significant parameters such as cumulative tumor size, presence of residual disease, stages of prior bladder cancer, current state of bladder cancer, and the presence of current bladder cancer (T1) is provided. The decision-making approach outlined in the paper supplemented with additional knowledge bases will lead to a comprehensive analytical road map of the BCG/IFN-alpha immunotherapy treatment. It will provide individualized guidelines for each stage of the treatment as well as measure the success of the treatment.

Adjuvants, Immunologic↗

A data mining approach to simulating farmers' crop choices for integrated water resources management.

Water and land resources in Thailand are increasingly under pressure from development. In particular, there are many resource conflicts associated with agricultural production in northern Thailand. Communities in these areas are significantly constrained in the land and water management decisions they are able to make. This paper describes the application of a data mining approach to describing and simulating farmers' decision rules in a catchment in northern Thailand. This approach is being applied to simulate social, economic and biophysical constraints on farmers' decisions in these areas as part of an integrated water management model.

Agriculture↗

A data mining approach for analyzing density maps representing macromolecular structures.

Results of electron microscopy-based three-dimensional reconstructions of macromolecules or their complexes are usually stored as density maps. Each point ("voxel") in the map represents a density value and one approach for studying details of the map is to display an isosurface enclosing areas of interest. We have taken a data mining approach not only focusing on the areas of immediate interest but determining all possible separate entities ("blobs") from a density map. After the entire density map is analyzed with our mining program BLOBBER, properties of all detected blobs can be browsed and sets of blobs can be visualized using our VIZBLOB program. Since BLOBBER analyzes density maps using only density information and relates it to spatial relationships, BLOBBER can be used to analyze symmetrical or asymmetrical density maps from any source. To test our program we have analyzed published bacteriophage PRD1 reconstructions. We identified various structural details ranging from individual proteins to major complexes such as the whole capsid shell and more elaborate details of possible connections between membrane interfaces. This approach can also be a useful preprocessing tool for visualizing reconstructions.

Bacteriophages↗