Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Clinical and pharmacogenomic data mining: 1. Generalized theory of expected information and application to the development of tools.

New scientific problems, arising from the human genome project, are challenging the classical means of using statistics. Yet quantified knowledge in the form of rules and rule strengths based on real relationships in data, as opposed to expert opinion, is urgently required for researcher and physician decision support. The problem is that with many parameters, the space to be analyzed is highly dimensional. That is, the combinations of data to examine are subject to a combinatorial explosion as the number of possible events (entries, items, sub-records) (a),(b),(c),... per record (a,b,c,..) increases, and hence much of the space is sparsely populated. These combinatorial considerations are particularly problematic for identifying those associations called "Unicorn Events" which occur significantly less than expected to the extent that they are never seen to be counted. To cope with the combinatorial explosion, a novel numerical "book keeping" approach is taken to generate information terms relating to the combinatorial subsets of events (a,b,c,..), and, most importantly, the zeta (Zeta) function is employed. The incomplete Zeta function zeta(s,n) with s = 1, in which frequencies of occurrence such as n = n(a,b,c,...) determine the range of summation n, is argued to be the natural choice of information function. It emerges from Bayesian integration, taken over the distribution of possible values of information measures for sparse and ample data alike. Expected mutual information l(a;b;c) in nats (i.e., natural units analogous to bits but based on the natural logarithm), such as is available to the observer, is measured as e.g., the difference zeta(s,o(a,b,c..)) - zeta(s,e(a,b,c..)) where o(a,b,c,..) and e(a,b,c,..) are, or relate to, the observed and expected frequencies of occurrence, respectively. For real values of s > 1 the qualitative impact of strongly (positively or negatively) ranked data is preserved despite several numerical approximations. As real s increases, and the output of the information functions converge into three values +1, 0, and -1 nats representing a trinary logic system. For quantitative data, a useful ad hoc method, to report sigma-normalized covariations in an analogous manner to mutual information for significance comparison purposes, is demonstrated. Finally, the potential ability to make use of mutual information in a complex biomedical study, and to include Bayesian prior information derived from statistical, tabular, anecdotal, and expert opinion is briefly illustrated.

Clinical Trials as Topic↗

Clinical and pharmacogenomic data mining: 2. A simple method for the combination of information from associations and multivariances to facilitate analysis, decision, and design in clinical research and practice.

The physician and researcher must ultimately be able to combine qualitative and quantitative features from a variety of combinations of observations on data of many component items (i.e., many dimensions), and hence reach simple conclusions about interpretation, rational courses of action, and design. In the first paper of this series, it was noted that such needs are challenging the classical means of using statistics. Hence, the paper proposed the use of a Generalized Theory of Expected Information or "Zeta Theory". The conjoint event [a,b,c,..] is seen as a rule of association for a,b,c,.. associated with a rule strength I(a;b;c;...) = xi(s,o[a,b,c,..]) - xi (s,e[a,b,c,...]), where xi is the incomplete Zeta Function. Here, o[a,b,c,...] is the observed, and e[a,b,c,..] the expected, frequency of occurrence of conjoint event [a,b,c,...]. The present paper explores how output from this approach might be assembled in a form better suited for decision support. Related to this is the difficulty that the treatment of covariance and multivariance was previously rendered as a "fuzzy association" so that the output would fall into a similar form as the true associations, but this was a somewhat ad hoc approach in which only the final I( ) had any meaning. Users at clinical research sites had subsequently requested an alternative approach in which "effective frequencies" o[ ] and e[ ] calculated from the above variances and used to evaluate I( ) give some intuitive feeling analogous to the association treatment, and this is explored here. Though the present paper is theoretical, real examples are used to illustrate application. One clinical-genomic example illustrates experimental design by identifying data which is, or is not, statistically germane to the study. We also report on some impressions based on applying these techniques in studies of real, extensive patient record data which are now emerging, as well as on molecular design data originally studied in part to test the ability to deduce the effects of simple natural patient sequence variations ("SNPs") on patient protein activity. On the basis of these study experiences, methods of rationalizing and condensing the rules implied by associations and variances between data, as well as discussion of the difficulty of what is meant by "condensed", are presented in the Appendix.

Biomedical Research↗

Dynamic and static approaches to clinical data mining.

In sequential diagnosis, the usefulness of a test can be assessed only in the context of a chosen diagnostic strategy, and depends on the evidence provided by previous test results. Choosing the most useful test at each stage of the evidence-gathering process therefore requires a dynamic approach to data analysis. An implementation of such an approach in an intelligent program for sequential diagnosis based on the evidence-gathering strategies used by doctors is described. On the other hand, a static approach to data analysis is appropriate in the discovery of knowledge required, for example, to explain or justify a diagnosis by identifying the most important findings, both positive and negative, on which the diagnosis is based. An algorithm for the discovery of features which always provide evidence in favour of, or against, a diagnosis selected by the data miner is presented. Dominance relationships among features in the data set are also discovered such that if one feature dominates another, it always provides more evidence in favour of the diagnosis, or less evidence against it.

Abdominal Pain↗

Gene expression databases and data mining.

The DNA microarray technology has arguably caught the attention of the worldwide life science community and is now systematically supporting major discoveries in many fields of study. The majority of the initial technical challenges of conducting experiments are being resolved, only to be replaced with new informatics hurdles, including statistical analysis, data visualization, interpretation, and storage. Two systems of databases, one containing expression data and one containing annotation data are quickly becoming essential knowledge repositories of the research community. This present paper surveys several databases, which are considered "pillars" of research and important nodes in the network. This paper focuses on a generalized workflow scheme typical for microarray experiments using two examples related to cancer research. The workflow is used to reference appropriate databases and tools for each step in the process of array experimentation. Additionally, benefits and drawbacks of current array databases are addressed, and suggestions are made for their improvement.

Breast Neoplasms↗

A data-mining approach to spacer oligonucleotide typing of Mycobacterium tuberculosis.

MOTIVATION: The Direct Repeat (DR) locus of Mycobacterium tuberculosis is a suitable model to study (i) molecular epidemiology and (ii) the evolutionary genetics of tuberculosis. This is achieved by a DNA analysis technique (genotyping), called sp acer oligo nucleotide typing (spoligotyping ). In this paper, we investigated data analysis methods to discover intelligible knowledge rules from spoligotyping, that has not yet been applied on such representation. This processing was achieved by applying the C4.5 induction algorithm and knowledge rules were produced. Finally, a Prototype Selection (PS) procedure was applied to eliminate noisy data. This both simplified decision rules, as well as the number of spacers to be tested to solve classification tasks. In the second part of this paper, the contribution of 25 new additional spacers and the knowledge rules inferred were studied from a machine learning point of view. From a statistical point of view, the correlations between spacers were analyzed and suggested that both negative and positive ones may be related to potential structural constraints within the DR locus that may shape its evolution directly or indirectly. RESULTS: By generating knowledge rules induced from decision trees, it was shown that not only the expert knowledge may be modeled but also improved and simplified to solve automatic classification tasks on unknown patterns. A practical consequence of this study may be a simplification of the spoligotyping technique, resulting in a reduction of the experimental constraints and an increase in the number of samples processed.

Algorithms↗

Data mining the Arabidopsis genome reveals fifteen 14-3-3 genes. Expression is demonstrated for two out of five novel genes.

In plants, 14-3-3 proteins are key regulators of primary metabolism and membrane transport. Although the current dogma states that 14-3-3 isoforms are not very specific with regard to target proteins, recent data suggest that the specificity may be high. Therefore, identification and characterization of all 14-3-3 (GF14) isoforms in the model plant Arabidopsis are important. Using the information now available from The Arabidopsis Information Resource, we found three new GF14 genes. The potential expression of these three genes, and of two additional novel GF14 genes (Rosenquist et al., 2000), in leaves, roots, and flowers was examined using reverse transcriptase-polymerase chain reaction and cDNA library polymerase chain reaction screening. Under normal growth conditions, two of these genes were found to be transcribed. These genes were named grf11and grf12, and the corresponding new 14-3-3 isoforms were named GF14omicron and GF14iota, respectively. The gene coding for GF14omicron was expressed in leaves, roots, and flowers, whereas the gene coding for GF14iota was only expressed in flowers. Gene structures and relationships between all members of the GF14 gene family were deduced from data available through The Arabidopsis Information Resource. The data clearly support the theory that two 14-3-3 genes were present when eudicotyledons diverged from monocotyledons. In total, there are 15 14-3-3 genes (grfs 1-15) in Arabidopsis, of which 12 (grfs 1-12) now have been shown to be expressed.

14-3-3 Proteins↗

[Comparative analysis via data mining on the clinical features of Western medicine and Chinese medicine in diagnosing rheumatoid arthritis].

OBJECTIVE: To compare the clinical characteristics of traditional Chinese medicine (TCM) and Western medicine (WM) in diagnosing rheumatoid arthritis (RA). METHODS: A total of 85 clinical RA related messages were enacted and classified into 5 sets as pathological locations, quantitative diagnosis, symptomatic descriptions, general status and environment factors. The respective frequency of their presence in the TCM or WM data sets for RA diagnosis collected from MEDLINE and China National Knowledge Infrastructure (CNKI) were analyzed statistically by Chi-square test, and the relationship of some TCM diagnostic factors/conditions with the RA related biological factors was analyzed by co-occurrence-based literature approach of mining. RESULTS: There was significant difference between the diagnostic pattern of WM and TCM (P < 0.01). Compared with that of WM, TCM diagnosis on RA paid more attention to environmental factors and symptomatic descriptions, which showed definite association with the cytokines and neuro-endocrine factors in RA. CONCLUSION: Examination of environmental factors and symptomatic descriptions for RA diagnosis is one of the important characteristics of TCM treatment in accordance to syndrome differentiation, it has its potential biologic basis. A novel approach is proposed for exploring the characteristics of TCM in diagnosis and observation.

Arthritis, Rheumatoid↗

Data mining for proteins characteristic of clades.

A synapomorphy is a phylogenetic character that provides evidence of shared descent. Ideally a synapomorphy is ubiquitous within the clade of related organisms and nonexistent outside the clade, implying that it arose after divergence from other extant species and before the last common ancestor of the clade. With the recent proliferation of genetic sequence data, molecular synapomorphies have assumed great importance, yet there is no convenient means to search for them over entire genomes. We have developed a new program called Conserv, which can rapidly assemble orthologous sequences and rank them by various metrics, such as degree of conservation or divergence from out-group orthologs. We have used Conserv to conduct a largescale search for molecular synapomorphies for bacterial clades. The search discovered sequences unique to clades, such as Actinobacteria, Firmicutes and gamma-Proteobacteria, and shed light on several open questions, such as whether Symbiobacterium thermophilum belongs with Actinobacteria or Firmicutes. We conclude that Conserv can quickly marshall evidence relevant to evolutionary questions that would be much harder to assemble with other tools.

Amino Acid Sequence↗

Knowledge discovery in gene-expression-microarray data: mining the information output of the genome.

A key aspect of the genomics revolution is the transformation of large amounts of biological information into an electronic format, leading to an information-based approach to biomedical problems. Large-scale RNA assays and gene-expression-microarray studies, in particular, represent the second wave of the genomics revolution, providing gene-expression data that complement gene-sequence data and help our understanding of the molecular basis of health and disease. They are being applied at several stages in the drug-development process and could ultimately have broad applications in disease diagnosis and patient prognosis.

Animals↗

The NEIBank project for ocular genomics: data-mining gene expression in human and rodent eye tissues.

NEIBank is a project to gather and organize genomic resources for eye research. The first phase of this project covers the construction and sequence analysis of cDNA libraries from human and animal model eye tissues to develop an overview of the repertoire of genes expressed in the eye and a resource of cDNA clones for further studies. The sequence data are grouped and identified using the tools of bioinformatics and the results are displayed through a web site where they can be interrogated by keyword search, chromosome location, by Blast (sequence comparison) or by alignment on completed genomes. Many novel proteins and novel splice forms of known genes have already emerged from analysis of the accumulating data. This review provides an overview of the current state of the database for human eye tissues, with specific comparisons to some parallel data from mouse and rat, and with illustrative examples of the kinds of insights and discoveries these data can produce. One of the major themes that emerges is that at the molecular level human eye tissues have significant differences from those of rodents, encompassing species specific genes, alternative splice forms and great variation in levels of gene expression. These point to specific adaptations and mechanisms in the human eye and emphasize that care needs to be taken in the application of appropriate animal model systems.

Amino Acid Sequence↗

Functional genomics and proteomics in the clinical neurosciences: data mining and bioinformatics.

The goal of this chapter is to introduce some of the available computational methods for expression analysis. Genomic and proteomic experimental techniques are briefly discussed to help the reader understand these methods and results better in context with the biological significance. Furthermore, a case study is presented that will illustrate the use of these analytical methods to extract significant biomarkers from high-throughput microarray data. Genomic and proteomic data analysis is essential for understanding the underlying factors that are involved in human disease. Currently, such experimental data are generally obtained by high-throughput microarray or mass spectrometry technologies among others. The sheer amount of raw data obtained using these methods warrants specialized computational methods for data analysis. Biomarker discovery for neurological diagnosis and prognosis is one such example. By extracting significant genomic and proteomic biomarkers in controlled experiments, we come closer to understanding how biological mechanisms contribute to neural degenerative diseases such as Alzheimers' and how drug treatments interact with the nervous system. In the biomarker discovery process, there are several computational methods that must be carefully considered to accurately analyze genomic or proteomic data. These methods include quality control, clustering, classification, feature ranking, and validation. Data quality control and normalization methods reduce technical variability and ensure that discovered biomarkers are statistically significant. Preprocessing steps must be carefully selected since they may adversely affect the results of the following expression analysis steps, which generally fall into two categories: unsupervised and supervised. Unsupervised or clustering methods can be used to group similar genomic or proteomic profiles and therefore can elucidate relationships within sample groups. These methods can also assign biomarkers to sub-groups based on their expression profiles across patient samples. Although clustering is useful for exploratory analysis, it is limited due to its inability to incorporate expert knowledge. On the other hand, classification and feature ranking are supervised, knowledge-based machine learning methods that estimate the distribution of biological expression data and, in doing so, can extract important information about these experiments. Classification is closely coupled with feature ranking, which is essentially a data reduction method that uses classification error estimation or other statistical tests to score features. Biomarkers can subsequently be extracted by eliminating insignificantly ranked features. These analytical methods may be equally applied to genetic and proteomic data. However, because of both biological differences between the data sources and technical differences between the experimental methods used to obtain these data, it is important to have a firm understanding of the data sources and experimental methods. At the same time, regardless of the data quality, it is inevitable that some discovered biomarkers are false positives. Thus, it is important to validate discovered biomarkers. The validation process may be slow; yet, the overall biomarker discovery process is significantly accelerated due to initial feature ranking and data reduction steps. Information obtained from the validation process may also be used to refine data analysis procedures for future iteration. Biomarker validation may be performed in a number of ways - bench-side in traditional labs, web-based electronic resources such as gene ontology and literature databases, and clinical trials.

Animals↗

Inside the data mine: showcasing UR/QA.

The latest look inside San Ramon, CA-based health care consulting company GE Medical Systems Health Care Solutions (HCS--formerly MECON)Ddata Mine shows that health care facilities spend an average of about $29 per adjusted discharge for services rendered by their utilization review and quality assurance (UR/QA) departments. How do you compare?

California↗

A 3D interactive multimodal viewer as data mining tool for the Visible Human Dataset color image histograms.

An on-line virtual three-dimensional immersive environment to navigate through colorimetric characterization of the Visible Human Dataset (VHD) cryosectional cross-section color images is introduced. Real-time analysis of color component characteristics of a user defined set of VHD images is now possible. This is a potentially useful resource to many developers working on the VHD raw data, however it could be used in medical education.

Female↗

Topology of gene expression networks as revealed by data mining and modeling.

MOTIVATION: Interpretation of high-throughput gene expression profiling requires a knowledge of the design principles underlying the networks that sustain cellular machinery. Recently a novel approach based on the study of network topologies has been proposed. This methodology has proven to be useful for the analysis of a variety of biological systems, including metabolic networks, networks of protein-protein interactions, and gene networks that can be derived from gene expression data. In the present paper, we focus on several important issues related to the topology of gene expression networks that have not yet been fully studied. RESULTS: The networks derived from gene expression profiles for both time series experiments in yeast and perturbation experiments in cell lines are studied. We demonstrate that independent from the experimental organism (yeast versus cell lines) and the type of experiment (time courses versus perturbations) the extracted networks have similar topological characteristics suggesting together with the results of other common principles of the structural organization of biological networks. A novel computational model of network growth that reproduces the basic design principles of the observed networks is presented. Advantage of the model is that it provides a general mechanism to generate networks with different types of topology by a variation of a few parameters. We investigate the robustness of the network structure to random damages and to deliberate removal of the most important parts of the system and show a surprising tolerance of gene expression networks to both kinds of disturbance.

Algorithms↗

TreeGeneBrowser: phylogenetic data mining of gene sequences from public databases.

MOTIVATION: Sequence databases represent an enormous resource of phylogenetic information, but there is a lack of tools for accessing that information in order to assess the amount of evolutionary information in these databases that may be suitable for phylogenetic reconstruction and for identifying areas of the taxonomy that are under-represented for specific gene sequences. RESULTS: We have developed TreeGeneBrowser which allows inspection and evaluation of gene sequence data for phylogenetic reconstruction. This program improves the efficiency of identification of genes that may be useful for particular phylogenetic studies and identifies taxa and taxonomic branches that are under-represented in sequence databases.

Algorithms↗

Antioxidant defense in Plasmodium falciparum--data mining of the transcriptome.

The intraerythrocytic malaria parasite is under constant oxidative stress originating both from endogenous and exogenous processes. The parasite is endowed with a complete network of enzymes and proteins that protect it from those threats, but also uses redox activities to regulate enzyme activities. In the present analysis, the transcription of the genes coding for the antioxidant defense elements are viewed in the time-frame of the intraerythrocytic cycle. Time-dependent transcription data were taken from the transcriptome of the human malaria parasite Plasmodium falciparum. Whereas for several processes the transcription of the many participating genes is coordinated, in the present case there are some outstanding deviations where gene products that utilize glutathione or thioredoxin are transcribed before the genes coding for elements that control the levels of those substrates are transcribed. Such insights may hint to novel, non-classical pathways that necessitate further investigations.

Animals↗