Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

An approach to the perceptual optimization of complex visualizations.

This paper proposes a new experimental framework within which evidence regarding the perceptual characteristics of a visualization method can be collected, and describes how this evidence can be explored to discover principles and insights to guide the design of perceptually near-optimal visualizations. We make the case that each of the current approaches for evaluating visualizations is limited in what it can tell us about optimal tuning and visual design. We go on to argue that our new approach is better suited to optimizing the kinds of complex visual displays that are commonly created in visualization. Our method uses human-in-the-loop experiments to selectively search through the parameter space of a visualization method, generating large databases of rated visualization solutions. Data mining is then used to extract results from the database, ranging from highly specific exemplar visualizations for a particular data set, to more broadly applicable guidelines for visualization design. We illustrate our approach using a recent study of optimal texturing for layered surfaces viewed in stereo and in motion. We show that a genetic algorithm is a valuable way of guiding the human-in-the-loop search through visualization parameter space. We also demonstrate several useful data mining methods including clustering, principal component analysis, neural networks, and statistical comparisons of functions of parameters.

Algorithms↗

Mining chemical structural information from the drug literature.

It is easier to find too many documents on a life science topic than to find the right information inside these documents. With the application of text data mining to biological documents, it is no surprise that researchers are starting to look at applications that mine out chemical information. The mining of chemical entities--names and structures--brings with it some unique challenges, which commercial and academic efforts are beginning to address. Ultimately, life science text data mining applications need to focus on the marriage of biological and chemical information.

Chemistry, Pharmaceutical↗

Assessing association rules and decision trees on analysis of diabetes data from the DiabCare program in France.

Recent advances in information technology have made it possible to solve increasingly complex problems, and also to collect and store huge amounts of information. These vast quantities of data further have to be transformed into relevant value-added and "decision-quality" knowledge. It is against this background that the KDD (Knowledge Discovery in Databases), a multidisciplinary field using computer learning, artificial intelligence, statistics, database technology, expert systems, and data visualization, appeared in the early 90's. In order to assess these technologies in the medical field, we have tested some of these techniques on a large database at our disposal, named DiabCare stemming from the WHO - DiabCare program for the application of the Saint-Vincent Declaration. It contains evaluation data on the health care of patients with diabetes, and in particular, its complications. So far, data analysis has been done using classical statistical methods, and we now intend to make use of such data-mining tools as Associations Rules and Decision and Classification Trees for further exploration of this database. The results presented here show that data mining techniques can be used successfully to extract knowledge from medical databases. The results obtained using Association Rules and especially Decision Trees are very promising.

Algorithms↗

Reports of hyperkalemia after publication of RALES--a pharmacovigilance study.

PURPOSE: A population-based study and anecdotal reports have indicated that the publication of the Randomized Aldactone Evaluation Study (RALES) was associated with not merely a broader use of spironolactone in the treatment of heart failure, but also with a coinciding sharp increase in hyperkalemia-associated morbidity/mortality in patients also being treated with ACE-inhibitors. Data mining algorithms (DMAs) are being applied to spontaneous reporting system (SRS) databases in hopes of obtaining early warnings/additional insights into post-licensure safety data. We applied two DMAs (i.e. multi-item gamma Poisson shrinker [MGPS] and proportional reporting ratios [PRRs]) to spontaneous reporting system (SRS) data to determine if these DMAs could have provided an earlier indication of a possible hyperkalemia safety issue. METHODS: MGPS and PRRs were retrospectively applied to US FDA-AERS, an SRS database. Year-by-year analysis and analysis of increasing cumulative time intervals were performed on cases in which both spironolactone and hyperkalemia and possibly related cardiac events had been reported. RESULTS: Neither of the DMAs initially provided a compelling signal of disproportionate reporting (SDR) for hyperkalemia after publication of RALES. However, using events consistent with clinical sequelae of hyperkalemia (e.g,. sudden death), SDRs were identified with PRRs. CONCLUSIONS: The quality and usefulness of data mining analysis is highly situation dependent and may vary with the knowledge and experience of the drug safety reviewer. Our analysis suggests that contemporary DMAs may have significant limitations in detecting increased frequency of labeled events in real-life prospective pharmacovigilance. There is a paucity of research in this area and we recommend further research for new approaches to detecting increased frequency of labeled events.

Adverse Drug Reaction Reporting Systems↗

Predicting the likelihood of falls among the elderly using likelihood basis pursuit technique.

This study reports on the application of the knowledge discovery in database process to generate models that can predict the likelihood of falls among the elderly who reside in long-term care facilities. This process was applied to data held in the Minimum Data Set, a comprehensive resident assessment instrument being used in all Medicare and Medicaid supported nursing homes in the United States. For this study, we incorporated a new data mining technique, Likelihood Basis Pursuit, into the process. Using this technique, we were able to correctly identify which of the variables in this data set were associated with falls and generate models that could make fall likelihood predictions based upon those variables. Because the model provides probabilities based upon the exact combination of variables present in a particular resident, models constructed using this new data mining technique have the potential to be more useful for assessing fall risk.

Accidental Falls↗

In silico gene function prediction using ontology-based pattern identification.

MOTIVATION: With the emergence of genome-wide expression profiling data sets, the guilt by association (GBA) principle has been a cornerstone for deriving gene functional interpretations in silico. Given the limited success of traditional methods for producing clusters of genes with great amounts of functional similarity, new data-mining algorithms are required to fully exploit the potential of high-throughput genomic approaches. RESULTS: Ontology-based pattern identification (OPI) is a novel data-mining algorithm that systematically identifies expression patterns that best represent existing knowledge of gene function. Instead of relying on a universal threshold of expression similarity to define functionally related groups of genes, OPI finds the optimal analysis settings that yield gene expression patterns and gene lists that best predict gene function using the principle of GBA. We applied OPI to a publicly available gene expression data set on the life cycle of the malarial parasite Plasmodium falciparum and systematically annotated genes for 320 functional categories based on current Gene Ontology annotations. An ontology-based hierarchical tree of the 320 categories provided a systems-wide biological view of this important malarial parasite.

Algorithms↗

[Rule induction algorithm for brain glioma using support vector machine].

A new proposed data mining technique, support vector machine (SVM), is used to predict the degree of malignancy in brain glioma. Based on statistical learning theory, SVM realizes the principle of data dependent structure risk minimization, so it can depress the overfitting with better generalization performance, since the prediction in medical diagnosis often deals with a small sample. SVM based rule induction algorithm is implemented in comparison with other data mining techniques such as artificial neural networks, rule induction algorithm and fuzzy rule extraction algorithm based on fuzzy max-min neural networks (FRE-FMMNN) proposed recently. Computation results by 10 fold cross validation method show that SVM can get higher prediction accuracy than artificial neural networks and FRE-FMMNN, which implies SVM can get higher accuracy and more reliability. On the whole data sets, SVM gets one rule with the classification accuracy of 89.29%, while FRE-FMMNN gets two rules of 84. 64%, in which the rule got by SVM is of quantity relation and contains more information than the two rules by FRE-FMMNN. All the above show SVM is a potential algorithm for the medical diagnosis such as the prediction of the degree of malignancy in brain glioma.

Algorithms↗

Determining a detectable threshold of signal intensity in cDNA microarray based on accumulated distribution.

In microarray data mining, one of the key problems is how to handle weak signals. Based on a bent piecewise linear accumulated distribution generally found in the microarray data, a new detectable threshold finding method is proposed to filter genes with unreliable information in this paper. More reliable and reproducible data is produced for the subsequent data mining.

DNA, Complementary↗

From association to alert--a revised approach to international signal analysis.

From the inception of the WHO international drug monitoring programme, the main aim has been to detect signals of adverse reaction problems as early as possible. The Uppsala Monitoring Centre (UMC), is now in a better position to fulfil this mission. Using the latest technology, new tools have been developed which allow for rapid, robust and comprehensive data mining of the WHO database. Based on retrospective time scans made during the pilot phase the current threshold used is the 97.5% confidence level of difference from the generality of the database. To maximize the capacity for picking up signals, we intend to extend today's panel of expert consultants, as well as doing our own review. The new system includes an enhanced follow-up list of signals, a 're-signalling' procedure and a cumulative historical file of all drug-ADR associations. Already we produce some 50 signals per year, cisapride and tachycardia being an example of a controversial signal only recently accepted. With the addition of new tools for follow-up of important signals such as complex variable data mining techniques, and the combination of WHO ADR data with sales and prescription figures from the IMS, we will be able to provide more information that should benefit regulators, producers, prescribers, and most importantly, the users of medicines.

Journal Article↗

BRENDA, AMENDA and FRENDA: the enzyme information system in 2007.

The BRENDA (BRaunschweig ENzyme DAtabase) enzyme information system (http://www.brenda.uni-koeln.de) is the largest publicly available enzyme information system worldwide. The major parts of its contents are manually extracted from primary literature. It is not restricted to specific groups of enzymes, but includes information on all identified enzymes irrespective of the enzyme's source. The range of data encompasses functional, structural, sequence, localisation, disease-related, isolation, stability information on enzyme and ligand-related data. Each single entry is linked to the enzyme source and to a literature reference. Recently the data repository was complemented by text-mining data in AMENDA (Automatic Mining of ENzyme DAta) and FRENDA (Full Reference ENzyme DAta). A genome browser, membrane protein prediction and full-text search capacities were added. The newly implemented web service provides instant access to the data for programmers via a SOAP (Simple Object Access Protocol) interface. The BRENDA data can be downloaded in the form of a text file from the beginning of 2007.

Animals↗

Mining mass spectra for diagnosis and biomarker discovery of cerebral accidents.

In this paper we try to identify potential biomarkers for early stroke diagnosis using surface-enhanced laser desorption/ionization mass spectrometry coupled with analysis tools from machine learning and data mining. Data consist of 42 specimen samples, i.e., mass spectra divided in two big categories, stroke and control specimens. Among the stroke specimens two further categories exist that correspond to ischemic and hemorrhagic stroke; in this paper we limit our data analysis to discriminating between control and stroke specimens. We performed two suites of experiments. In the first one we simply applied a number of different machine learning algorithms; in the second one we have chosen the best performing algorithm as it was determined from the first phase and coupled it with a number of different feature selection methods. The reason for this was 2-fold, first to establish whether feature selection can indeed improve performance, which in our case it did not seem to confirm, but more importantly to acquire a small list of potentially interesting biomarkers. Of the different methods explored the most promising one was support vector machines which gave us high levels of sensitivity and specificity. Finally, by analyzing the models constructed by support vector machines we produced a small set of 13 features that could be used as potential biomarkers, and which exhibited good performance both in terms of sensitivity, specificity and model stability.

Adult↗

Building knowledge in a complex preterm birth problem domain.

Data mining methods used a racially diverse sample (n = 19,970) of pregnant women and 1,622 variables that were collected in Duke's TMR electronic patient record over a 10-year period. Different statistical and data mining methods were similar when compared using receiver operating characteristic (ROC) curves. Best results found that seven demographic variables yielded .72 and addition of hundreds of other clinical variables added only .03 to the area under the curve (AUC). Similar results across methods suggest that results were data-driven and not method-dependent, and that demographic variables may offer a small set of parsimonious variables with predictive accuracy in a racially diverse population. Work to determine relevant variables for improved predictive accuracy is ongoing.

Area Under Curve↗

Artificial intelligence techniques for monitoring dangerous infections.

The monitoring and detection of nosocomial infections is a very important problem arising in hospitals. A hospital-acquired or nosocomial infection is a disease that develops after admission into the hospital and it is the consequence of a treatment, not necessarily a surgical one, performed by the medical staff. Nosocomial infections are dangerous because they are caused by bacteria which have dangerous (critical) resistance to antibiotics. This problem is very serious all over the world. In Italy, almost 5-8% of the patients admitted into hospitals develop this kind of infection. In order to reduce this figure, policies for controlling infections should be adopted by medical practitioners. In order to support them in this complex task, we have developed a system, called MERCURIO, capable of managing different aspects of the problem. The objectives of this system are the validation of microbiological data and the creation of a real time epidemiological information system. The system is useful for laboratory physicians, because it supports them in the execution of the microbiological analyses; for clinicians, because it supports them in the definition of the prophylaxis, of the most suitable antibi-otic therapy and in monitoring patients' infections; and for epidemiologists, because it allows them to identify outbreaks and to study infection dynamics. In order to achieve these objectives, we have adopted expert system and data mining techniques. We have also integrated a statistical module that monitors the diffusion of nosocomial infections over time in the hospital, and that strictly interacts with the knowledge based module. Data mining techniques have been used for improving the system knowledge base. The knowledge discovery process is not antithetic, but complementary to the one based on manual knowledge elicitation. In order to verify the reliability of the tasks performed by MERCURIO and the usefulness of the knowledge discovery approach, we performed a test based on a dataset of real infection events. In the validation task MERCURIO achieved an accuracy of 98.5%, a sensitivity of 98.5% and a specificity of 99%. In the therapy suggestion task, MERCURIO achieved very high accuracy and specificity as well. The executed test provided many insights to experts, too (we discovered some of their mistakes). The knowledge discovery approach was very effective in validating part of the MERCURIO knowledge base, and also in extending it with new validation rules, confirmed by interviewed microbiologists and specific to the hospital laboratory under consideration.

Artificial Intelligence↗

Altered gene expression profiles reveal similarities and differences between Parkinson disease and model systems.

Parkinson disease (PD) targets dopaminergic neurons in the substantia nigra, resulting in motor disturbances such as resting tremor, bradykinesia, and rigidity. Pathogenic processes likely occur over several decades, in that an overwhelming percentage of neurons are already dead at the time of clinical diagnosis. For this reason, the usage of animal model systems to discover the early steps in the pathologic cascade is required. These include exposure to the neurotoxin 1-methyl-4-phenyl-1,2,3,6-tetrahydropyridine (MPTP), which selectively kills dopamine neurons in the substantia nigra, and genetic models incorporating mutations in the alpha-synuclein gene that cause disease in human patients. Through the evaluation of these models at multiple time points, it is possible to discover novel gene expression changes that may underlie disease pathogenesis. Specifically, the authors hypothesize that animal models of PD and human PD brains share a gene expression profile that signifies certain aspects of pathogenesis and/or recovery-resistance. To test this and similar hypotheses, the authors and others have utilized new microarray technology that enables the sampling of thousands of genes' expression level in one assay. Because the technology is fairly new and results can vary depending on methods used, results must be evaluated with care. Multiple array and data-mining options can be used to make the most accurate inferences as to differentially expressed genes in each set of samples. The authors developed a fusion classifier approach whereby individual data-mining algorithms generate lists of significant genes. The lists are subsequently queried, and only genes unanimously called significant are retained for further validation. Although the authors' approach identified hundreds of differentially expressed genes in each of three PD systems, only a few were common between the human and animal substantia nigra. These were related to dopamine phenotype, synaptic function, and the mitochondrial metabolism, implicating the presynaptic terminal as a primary site of injury. The time course of the authors' experiments indicates that if the synaptic changes could be prevented, this may alleviate some cell death, in that these changes precede neuronal loss.

1-Methyl-4-phenyl-1,2,3,6-tetrahydropyridine↗

Investigating trends in acoustics research from 1970-1999.

Text data mining is a burgeoning field in which new information is extracted from existing text databases. Computational methods are used to compare relationships between database elements to yield new information about the existing data. Text data mining software was used to determine research trends in acoustics for the years 1970, 1980, 1990, and 1999. Trends were indicated by the number of published articles in the categories of acoustics using the Journal of the Acoustical Society of America (JASA) as the article source. Research was classified using a method based on the Physics and Astronomy Classification Scheme (PACS). Research was further subdivided into world regions, including North and South America, Eastern and Western Europe, Asia, Africa, Middle East, and Australia/New Zealand. In order to gauge the use of JASA as an indicator of international acoustics research, three subjects, underwater sound, nonlinear acoustics, and bioacoustics, were further tracked in 1999, using all journals in the INSPEC database. Research trends indicated a shift in emphasis of certain areas, notably underwater sound, audition, and speech. JASA also showed steady growth, with increasing participation by non-US authors, from about 20% in 1970 to nearly 50% in 1999.

Journal Article↗

ResourceLog: an embeddable tool for dynamically monitoring the usage of web-based bioscience resources.

The present study described an open source application, ResourceLog, that allows website administrators to record and analyze the usage of online resources. The application includes four components: logging, data mining, administrative interface, and back-end database. The logging component is embedded in the host website. It extracts and streamlines information about the Web visitors, the scripts, and dynamic parameters from each page request. The data mining component runs as a set of scheduled tasks that identify visitors of interest, such as those who have heavily used the resources. The identified visitors will be automatically subjected to a voluntary user survey. The usage of the website content can be monitored through the administrative interface and subjected to statistical analyses. As a pilot project, ResourceLog has been implemented in SenseLab, a Web-based neuroscience database system. ResourceLog provides a robust and useful tool to aid system evaluation of a resource-driven Web application, with a focus on determining the effectiveness of data sharing in the field and with the general public.

Bibliometrics↗

A novel and accurate diagnostic test for human African trypanosomiasis.

INTRODUCTION: Human African trypanosomiasis (sleeping sickness) affects up to half a million people every year in sub-Saharan Africa. Because current diagnostic tests for the disease have low accuracy, we sought to develop a novel test that can diagnose human African trypanosomiasis with high sensitivity and specificity. METHODS: We applied serum samples from 85 patients with African trypanosomiasis and 146 control patients who had other parasitic and non-parasitic infections to a weak cation exchange chip, and analysed with surface-enhanced laser desorption-ionisation time-of-flight mass spectrometry. Mass spectra were then assessed with three powerful data-mining tools: a tree classifier, a neural network, and a genetic algorithm. FINDINGS: Spectra (2-100 kDa) were grouped into training (n=122) and testing (n=109) sets. The training set enabled data-mining software to identify distinct serum proteomic signatures characteristic of human African trypanosomiasis among 206 protein clusters. Sensitivity and specificity, determined with the testing set, were 100% and 98.6%, respectively, when the majority opinion of the three algorithms was considered. This novel approach is much more accurate than any other diagnostic test. INTERPRETATION: Our report of the accurate diagnosis of an infection by use of proteomic signature analysis could form the basis for diagnostic tests for the disease, monitoring of response to treatment, and for improving the accuracy of patient recruitment in large-scale epidemiological studies.

Adolescent↗

The future of epidemiology: methodological challenges and multilevel inference.

A decade ago there was considerable debate about the appropriate objectives and paradigms of modern epidemiologic research. One concern put forth in these debates was that "risk factor epidemiology" might be forcing our field to focus more on individuals and less on populations and public health. Today, most epidemiologists acknowledge that public health is influenced by both population-level and individual-level determinants. Ecologic studies are valuable tools for generating hypotheses and addressing group-level determinants of disease risk. Traditional risk factor studies and genomic studies have helped establish the multifactorial concept of disease causation. Individual-level studies also have provided the biomedical community with hypotheses that have stimulated research into disease mechanisms that have led to reductions in morbidity and mortality for diseases such as HIV/AIDS, cardiovascular disease, and cancer. Current debates about the role of genomic data in epidemiology and public health mirror the debates about risk factor epidemiology one decade ago. Genomic variation is measured at the individual level, but how this variation is maintained in human populations is a group-level (population) phenomenon that is worthy of epidemiologic investigation in its own right. Multilevel epidemiology seeks to understand multiple levels of inference, from genes to individuals to populations and could combine hypothesis-driven research with aspects of data mining. Multilevel epidemiology calls for the study of health and disease determinants defined at the population level and individual level for a more comprehensive strategy to understanding human disease etiology. With the continued development of multilevel statistical methods and the advent of data mining, the technical constraints of the past will become less relevant to the next generation of epidemiologists who wish to embrace a more multilevel epidemiology.

Data Interpretation, Statistical↗