Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

GeneMerge--post-genomic analysis, data mining, and hypothesis testing.

SUMMARY: GeneMerge is a web-based and standalone program written in PERL that returns a range of functional and genomic data for a given set of study genes and provides statistical rank scores for over-representation of particular functions or categories in the data set. Functional or categorical data of all kinds can be analyzed with GeneMerge, facilitating regulatory and metabolic pathway analysis, tests of population genetic hypotheses, cross-experiment comparisons, and tests of chromosomal clustering, among others. GeneMerge can perform analyses on a wide variety of genomic data quickly and easily and facilitates both data mining and hypothesis testing. AVAILABILITY: GeneMerge is available free of charge for academic use over the web and for download from: http://www.oeb.harvard.edu/hartl/lab/publications/GeneMerge.html.

Algorithms↗

Protecting patient privacy in clinical data mining.

This paper investigates whether HIPAA de-identification requirements--as well as proposed AAMC de-identification standards--were met in a large clinical data mining study (1997-2001) conducted at Duke University prior to the publication of the final rule. While HIPAA has improved de-identification standards, the study also shows that privacy issues may persist even in de-identified large clinical databases.

Biomedical Research↗

Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup.

MOTIVATION: The biological literature is a major repository of knowledge. Many biological databases draw much of their content from a careful curation of this literature. However, as the volume of literature increases, the burden of curation increases. Text mining may provide useful tools to assist in the curation process. To date, the lack of standards has made it impossible to determine whether text mining techniques are sufficiently mature to be useful. RESULTS: We report on a Challenge Evaluation task that we created for the Knowledge Discovery and Data Mining (KDD) Challenge Cup. We provided a training corpus of 862 articles consisting of journal articles curated in FlyBase, along with the associated lists of genes and gene products, as well as the relevant data fields from FlyBase. For the test, we provided a corpus of 213 new ('blind') articles; the 18 participating groups provided systems that flagged articles for curation, based on whether the article contained experimental evidence for gene expression products. We report on the evaluation results and describe the techniques used by the top performing groups.

Abstracting and Indexing↗

Data mining of sequences and 3D structures of allergenic proteins.

MOTIVATION: Many sequences, and in some cases structures, of proteins that induce an allergic response in atopic individuals have been determined in recent years. This data indicates that allergens, regardless of source, fall into discreet protein families. Similarities in the sequence may explain clinically observed cross-reactivities between different biological triggers. However, previously available allergy databases group allergens according to their biological sources, or observed clinical cross-reactivities, without providing data about the proteins. A computer-aided data mining system is needed to compare the sequential and structural details of known allergens. This information will aid in predicting allergenic cross-responses and eventually in determining possible common characteristics of IgE recognition. RESULTS: The new web-based Structural Database of Allergenic Proteins (SDAP) permits the user to quickly compare the sequence and structure of allergenic proteins. Data from literature sources and previously existing lists of allergens are combined in a MySQL interactive database with a wide selection of bioinformatics applications. SDAP can be used to rapidly determine the relationship between allergens and to screen novel proteins for the presence of IgE or T-cell epitopes they may share with known allergens. Further, our novel similarity search method, based on five dimensional descriptors of amino acid properties, can be used to scan the SDAP entries with a peptide sequence. For example, when a known IgE binding epitope from shrimp tropomyosin was used as a query, the method rapidly identified a similar sequence in known shellfish and insect allergens. This prediction of cross-reactivity between allergens is consistent with clinical observations. AVAILABILITY: SDAP is available on the web at http://fermi.utmb.edu/SDAP/index.html

Allergens↗

Application of a data-mining method based on Bayesian networks to lesion-deficit analysis.

Although lesion-deficit analysis (LDA) has provided extensive information about structure-function associations in the human brain, LDA has suffered from the difficulties inherent to the analysis of spatial data, i.e., there are many more variables than subjects, and data may be difficult to model using standard distributions, such as the normal distribution. We herein describe a Bayesian method for LDA; this method is based on data-mining techniques that employ Bayesian networks to represent structure-function associations. These methods are computationally tractable, and can represent complex, nonlinear structure-function associations. When applied to the evaluation of data obtained from a study of the psychiatric sequelae of traumatic brain injury in children, this method generates a Bayesian network that demonstrates complex, nonlinear associations among lesions in the left caudate, right globus pallidus, right side of the corpus callosum, right caudate, and left thalamus, and subsequent development of attention-deficit hyperactivity disorder, confirming and extending our previous statistical analysis of these data. Furthermore, analysis of simulated data indicates that methods based on Bayesian networks may be more sensitive and specific for detecting associations among categorical variables than methods based on chi-square and Fisher exact statistics.

Algorithms↗

Data mining with decision trees for diagnosis of breast tumor in medical ultrasonic images.

To increase the ability of ultrasonographic (US) technology for the differential diagnosis of solid breast tumors, we describe a novel computer-aided diagnosis (CADx) system using data mining with decision tree for classification of breast tumor to increase the levels of diagnostic confidence and to provide the immediate second opinion for physicians. Cooperating with the texture information extracted from the region of interest (ROI) image, a decision tree model generated from the training data in a top-down, general-to-specific direction with 24 co-variance texture features is used to classify the tumors as benign or malignant. In the experiments, accuracy rates for a experienced physician and the proposed CADx are 86.67% (78/90) and 95.50% (86/90), respectively.

Breast Neoplasms↗

Data mining parasite genomes: haystack searching with a computer.

A number of genomes of parasitic organisms are presently being sequenced in the public domain, including Plasmodium falciparum, Leishmania major and Trypanosoma brucei with the likelihood of at least expressed sequence tag (EST) projects for several filarial and apicomplexan species. The early and timely release of sequence data to the community via the World Wide Web (www), and the public databases, (EMBL and GENBANK), forms an invaluable resource. Data mining, or 'haystack searching' this resource is becoming more fruitful to all members of the scientific community as the volume of data, diversity of genomes sampled, and accessibility increase.

Animals↗

Database of repetitive elements in complete genomes and data mining using transcription factor binding sites.

Approximately 43% of the human genome is occupied by repetitive elements. Even more, around 51% of the rice genome is occupied by repetitive elements. The analysis presented here indicates that repetitive elements in complete genomes may have been very important in the evolutionary genomics. In this study, a database, called the Repeat Sequence Database, is first designed and implemented to store complete and comprehensive repetitive sequences. See http://rsdb.csie.ncu.edu.tw for more information. The database contains direct, inverted and palindromic repetitive sequences, and each repetitive sequence has a variable length ranging from seven to many hundred nucleotides. The repetitive sequences in the database are explored using a mathematical algorithm to mine rules on how combinations of individual binding sites are distributed among repetitive sequences in the database. Combinations of transcription factor binding sites in the repetitive sequences are obtained and then data mining techniques are applied to mine association rules from these combinations. The discovered associations are further pruned to remove insignificant associations and obtain a set of associations. The mined association rules facilitate efforts to identify gene classes regulated by similar mechanisms and accurately predict regulatory elements. Experiments are performed on several genomes including C. elegans, human chromosome 22, and yeast.

Algorithms↗

Use of 3D QSAR methodology for data mining the National Cancer Institute Repository of Small Molecules: application to HIV-1 reverse transcriptase inhibition.

A three-dimensional (3D) stereoelectronic pharmacophore developed from a 3D quantitative structure-activity relationship (QSAR) investigation formed the basis of the development of a two-phase data-mining methodology to uncover novel leads to inhibit human immunodeficiency virus type 1 (HIV-1) reverse transcriptase at the nonnucleoside binding site. The database searching phase employed a field search for ligand requirements (such as log P, molecular volume) that were accessible from the database keys. Next, a 3D database search was performed that used an automated fitting procedure and the calculation of several binding parameters. These binding parameters were used to test the hits by a discriminant function that was previously trained to recognize active from inactive analogs. During the structural evaluation phase of the methodology, conformational properties and complementary receptor features of the hits were examined by 2D and 3D evaluations, which were followed by molecular modeling investigations. When this method was applied to a test database, an improvement from 6.4% to 100% active analogs was achieved.

Database Management Systems↗

Visual management of large scale data mining projects.

This paper describes a unified framework for visualizing the preparations for, and results of, hundreds of machine learning experiments. These experiments were designed to improve the accuracy of enzyme functional predictions from sequence, and in many cases were successful. Our system provides graphical user interfaces for defining and exploring training datasets and various representational alternatives, for inspecting the hypotheses induced by various types of learning algorithms, for visualizing the global results, and for inspecting in detail results for specific training sets (functions) and examples (proteins). The visualization tools serve as a navigational aid through a large amount of sequence data and induced knowledge. They provided significant help in understanding both the significance and the underlying biological explanations of our successes and failures. Using these visualizations it was possible to efficiently identify weaknesses of the modular sequence representations and induction algorithms which suggest better learning strategies. The context in which our data mining visualization toolkit was developed was the problem of accurately predicting enzyme function from protein sequence data. Previous work demonstrated that approximately 6% of enzyme protein sequences are likely to be assigned incorrect functions on the basis of sequence similarity alone. In order to test the hypothesis that more detailed sequence analysis using machine learning techniques and modular domain representations could address many of these failures, we designed a series of more than 250 experiments using information-theoretic decision tree induction and naive Bayesian learning on local sequence domain representations of problematic enzyme function classes. In more than half of these cases, our methods were able to perfectly discriminate among various possible functions of similar sequences. We developed and tested our visualization techniques on this application.

Alcohol Dehydrogenase↗

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology↗

A retrospective evaluation of a data mining approach to aid finding new adverse drug reaction signals in the WHO international database.

BACKGROUND: The detection of new drug safety signals is of growing importance with ever more new drugs becoming available and exposure to medicines increasing. The task of evaluating information relating to safety lies with national agencies and, for international data, with the World Health Organization Programme for International Drug Monitoring. RATIONALE: An established approach for identifying new drug safety signals from the international database of more than 2 million case reports depends upon clinical experts from around the world. With a very large amount of information to evaluate, such an approach is open to human error. To aid the clinical review, we have developed a new signalling process using Bayesian logic, applied to data mining, within a confidence propagation neural network (Bayesian Confidence Propagation Neural Network; BCPNN). Ultimately, this will also allow the evaluation of complex variables. METHODS: The first part of this study tested the predictive value of the BCPNN in new signal detection as compared with reference literature sources (Martindale's Extra Pharmacopoeia in 1993 and July 2000, and the Physicians Desk Reference in July 2000). In the second part of the study, results with the BCPNN method were compared with those of the former signalling procedure. RESULTS: In the study period (the first quarter of 1993) 107 drug-adverse reaction combinations were highlighted as new positive associations by the BCPNN, and referred to new drugs. 15 drug-adverse reaction combinations on new drugs became negative BCPNN associations in the study period. The BCPNN method detected signals with a positive predictive value of 44% and the negative predictive value was 85%. 17 as yet unconfirmed positive associations could not be dismissed with certainty as false positive signals. Of the 10 drug-adverse reaction signals produced by the former signal detection system from data sent out for review during the study period, 6 were also identified by the BCPNN. These 6 associations have all had a more than 10-fold increase of reports and 4 of them have been included in the reference sources. The remaining 4 signals that were not identified by the BCPNN had a small, or no, increase in the number of reports, and are not listed in the reference sources. CONCLUSION: Our evaluation showed that the BCPNN approach had a high and promising predictive value in identifying early signals of new adverse drug reactions.

Algorithms↗

Knowledge discovery and data mining to assist natural language understanding.

As natural language processing systems become more frequent in clinical use, methods for interpreting the output of these programs become increasingly important. These methods require the effort of a domain expert, who must build specific queries and rules for interpreting the processor output. Knowledge discovery and data mining tools can be used instead of a domain expert to automatically generate these queries and rules. C5.0, a decision tree generator, was used to create a rule base for a natural language understanding system. A general-purpose natural language processor using this rule base was tested on a set of 200 chest radiograph reports. When a small set of reports, classified by physicians, was used as the training set, the generated rule base performed as well as lay persons, but worse than physicians. When a larger set of reports, using ICD9 coding to classify the set, was used for training the system, the rule base performed worse than the physicians and lay persons. It appears that a larger, more accurate training set is needed to increase performance of the method.

Artificial Intelligence↗

A data mining approach for analyzing density maps representing macromolecular structures.

Results of electron microscopy-based three-dimensional reconstructions of macromolecules or their complexes are usually stored as density maps. Each point ("voxel") in the map represents a density value and one approach for studying details of the map is to display an isosurface enclosing areas of interest. We have taken a data mining approach not only focusing on the areas of immediate interest but determining all possible separate entities ("blobs") from a density map. After the entire density map is analyzed with our mining program BLOBBER, properties of all detected blobs can be browsed and sets of blobs can be visualized using our VIZBLOB program. Since BLOBBER analyzes density maps using only density information and relates it to spatial relationships, BLOBBER can be used to analyze symmetrical or asymmetrical density maps from any source. To test our program we have analyzed published bacteriophage PRD1 reconstructions. We identified various structural details ranging from individual proteins to major complexes such as the whole capsid shell and more elaborate details of possible connections between membrane interfaces. This approach can also be a useful preprocessing tool for visualizing reconstructions.

Bacteriophages↗

Data mining to support simulation modeling of patient flow in hospitals.

Spiraling health care costs in the United States are driving institutions to continually address the challenge of optimizing the use of scarce resources. One of the first steps towards optimizing resources is to utilize capacity effectively. For hospital capacity planning problems such as allocation of inpatient beds, computer simulation is often the method of choice. One of the more difficult aspects of using simulation models for such studies is the creation of a manageable set of patient types to include in the model. The objective of this paper is to demonstrate the potential of using data mining techniques, specifically clustering techniques such as K-means, to help guide the development of patient type definitions for purposes of building computer simulation or analytical models of patient flow in hospitals. Using data from a hospital in the Midwest this study brings forth several important issues that researchers need to address when applying clustering techniques in general and specifically to hospital data.

Algorithms↗

MtDB: a database for personalized data mining of the model legume Medicago truncatula transcriptome.

In order to identify the genes and gene functions that underlie key aspects of legume biology, researchers have selected the cool season legume Medicago truncatula (Mt) as a model system for legume research. A set of >170 000 Mt ESTs has been assembled based on in-depth sampling from various developmental stages and pathogen-challenged tissues. MtDB is a relational database that integrates Mt transcriptome data and provides a wide range of user-defined data mining options. The database is interrogated through a series of interfaces with 58 options grouped into two filters. In addition, the user can select and compare unigene sets generated by different assemblers: Phrap, Cap3 and Cap4. Sequence identifiers from all public Mt sites (e.g. IDs from GenBank, CCGB, TIGR, NCGR, INRA) are fully cross-referenced to facilitate comparisons between different sites, and hypertext links to the appropriate database records are provided for all queries' results. MtDB's goal is to provide researchers with the means to quickly and independently identify sequences that match specific research interests based on user-defined criteria. The underlying database and query software have been designed for ease of updates and portability to other model organisms. Public access to the database is at http://www.medicago.org/MtDB.

Chromosome Mapping↗

Databases and data mining for computational vaccinology.

Drugs and vaccines are keys to the effective fight against disease. While the pharmaceutical industry has developed an awesome array of real and virtual approaches to rational drug discovery, the complexity of the immune system hampers attempts to design and develop vaccines in a rational manner. The goal of immunoinformatics (the application of informatics techniques to immunological macromolecules), an emergent sub-discipline of bioinformatics, is to develop computational vaccinology as a potent tool in the quest for new vaccines. Databases and data mining, the two principal weapons at the disposal of the in silico vaccinologist, will be presented in the light of current developments.

Computational Biology↗

Data mining crystallization databases: knowledge-based approaches to optimize protein crystal screens.

Protein crystallization is a major bottleneck in protein X-ray crystallography, the workhorse of most structural proteomics projects. Because the principles that govern protein crystallization are too poorly understood to allow them to be used in a strongly predictive sense, the most common crystallization strategy entails screening a wide variety of solution conditions to identify the small subset that will support crystal nucleation and growth. We tested the hypothesis that more efficient crystallization strategies could be formulated by extracting useful patterns and correlations from the large data sets of crystallization trials created in structural proteomics projects. A database of crystallization conditions was constructed for 755 different proteins purified and crystallized under uniform conditions. Forty-five percent of the proteins formed crystals. Data mining identified the conditions that crystallize the most proteins, revealed that many conditions are highly correlated in their behavior, and showed that the crystallization success rate is markedly dependent on the organism from which proteins derive. Of the proteins that crystallized in a 48-condition experiment, 60% could be crystallized in as few as 6 conditions and 94% in 24 conditions. Consideration of the full range of information coming from crystal screening trials allows one to design screens that are maximally productive while consuming minimal resources, and also suggests further useful conditions for extending existing screens.

Archaeal Proteins↗