Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Bayesian networks for knowledge discovery in large datasets: basics for nurse researchers.

The growth of nursing databases necessitates new approaches to data analyses. These databases, which are known to be massive and multidimensional, easily exceed the capabilities of both human cognition and traditional analytical approaches. One innovative approach, knowledge discovery in large databases (KDD), allows investigators to analyze very large data sets more comprehensively in an automatic or a semi-automatic manner. Among KDD techniques, Bayesian networks, a state-of-the art representation of probabilistic knowledge by a graphical diagram, has emerged in recent years as essential for pattern recognition and classification in the healthcare field. Unlike some data mining techniques, Bayesian networks allow investigators to combine domain knowledge with statistical data, enabling nurse researchers to incorporate clinical and theoretical knowledge into the process of knowledge discovery in large datasets. This tailored discussion presents the basic concepts of Bayesian networks and their use as knowledge discovery tools for nurse researchers.

Artificial Intelligence↗

Text mining and its potential applications in systems biology.

With biomedical literature increasing at a rate of several thousand papers per week, it is impossible to keep abreast of all developments; therefore, automated means to manage the information overload are required. Text mining techniques, which involve the processes of information retrieval, information extraction and data mining, provide a means of solving this. By adding meaning to text, these techniques produce a more structured analysis of textual knowledge than simple word searches, and can provide powerful tools for the production and analysis of systems biology models.

Artificial Intelligence↗

Techniques: application of systems biology to absorption, distribution, metabolism, excretion and toxicity.

It is widely recognized that either predicting or determining the absorption, distribution, metabolism, excretion and toxicity (ADME/Tox) properties of molecules helps to prevent the failure of some compounds before they reach the clinic. Consequently, there has been considerable research into developing better in silico, in vitro and in vivo methods and models. Toxicogenomics, proteomics, metabonomics and pharmacogenomics represent the latest experimental approaches that can be combined with high-throughput molecular screening of targets to provide a view of the complete biological system that is modulated by a compound. The functional interpretation and relevance of these complex multidimensional data to the phenotype observed in humans is the focus of current research in toxicology. Multiple content databases, data mining and predictive modeling algorithms, visualization tools, and high-throughput data-analysis solutions are being integrated to form systems-ADME/Tox. In this review, we focus on the most recent advances and applications in this area.

Animals↗

Recursive partitioning models for linkage in COGA data.

We have developed a recursive-partitioning (RP) algorithm for identifying phenotype and covariate groupings that interact with the evidence for linkage. This data-mining approach for detecting gene x environment interactions uses genotype and covariate data on affected relative pairs to find evidence for linkage heterogeneity across covariate-defined subgroups. We adapted a likelihood-ratio based test of linkage parameterized with relative risks to a recursive partitioning framework, including a cross-validation based deviance measurement for choosing optimal tree size and a bootstrap sampling procedure for choosing robust tree structure. ALDX2 category 5 individuals were considered affected, categories 1 and 3 unaffected, and all others unknown. We sampled non-overlapping affected relative pairs from each family; therefore, we used 144 affected pairs in the RP model. Twenty pair-level covariates were defined from smoking status, maximum drinks, ethnicity, sex, and age at onset. Using the all-pairs score in GENEHUNTER, the nonparametric linkage tests showed no regions with suggestive linkage evidence. However, using the RP model, several suggestive regions were found on chromosomes 2, 4, 6, 14, and 20, with detection of associated covariates such as sex and age at onset.

Alcoholism↗

On combining multiple microarray studies for improved functional classification by whole-dataset feature selection.

As microarray technologies become routinely applied in genome laboratories for studying gene expression, it is not uncommon that experiments on identical or similar sets of genes are conducted by multiple laboratories for various functional studies of these genes. Much of such data are often available to researchers for their data analysis, either through collaborators or from online gene expression databases. It will be useful to combine data from different microarray studies to improve the microarray data mining results. We show that the functional classification of genes from microarray data can be improved further by combining gene expression data from multiple microarray studies, even if the experimental focus or conditions for each experimental study may differ. However, blindly combining all available datasets may not always improve the analysis results---it is important to be selective of the datasets for inclusion. In our approach, we consider each dataset to be one feature, and then apply feature selection strategies to select appropriate datasets for training. With a simple hill-climbing method, we show that gene classification performances can be improved by whole-dataset feature selection.

Algorithms↗

[Chemical evaluation by cancer cell line panel and its role in molecular target-based anticancer drug screening].

Mechanism-based or target-based evaluation of chemicals is important in the discovery and development of anticancer drugs. A new scale for the mechanism-oriented evaluation is acquired by creating a database of drugs that includes their activities concerning growth inhibition against a set of cancer cell lines and by employing a specific data-mining method (Paull KD, et al: J Natl Cancer Inst 81: 1088-1092, 1989). According to this principle, we have established a new system for drug evaluation using a panel of 39 human cancer cell lines, JFCR-39 Cell Line Panel. The JFCR-39 Cell Line Panel is a system combining a wet system (drug-sensitivity test) and a dry system (data base and its mining), and can predict the mechanism of action of chemicals. Therefore, it is useful in anticancer drug discovery. The JFCR-39 Cell Line Panel now plays an important role as a core drug evaluation system in the molecular target-based drug screening conducted by Screening Committee of New Anticancer Agents supported by Grant-in-Aid for Scientific Research on Priority Area "Cancer" from The Ministry of Education, Culture, Sports, Science and Technology, Japan. I outlined the JFCR-39 Cell Line Panel in this review.

Antineoplastic Agents↗

High throughput multiple combination extraction from large scale polymorphism data by exact tree method.

Single nucleotide polymorphisms (SNPs) are increasingly becoming important in clinical settings as useful genetic markers. For the evaluation of genetic risk factors of multifactorial diseases, it is not sufficient to focus on individual SNPs. It is preferable to evaluate combinations of multiple markers, because it allows us to examine the interactions between multiple factors. If all the combinations possible were evaluated round-robin, the number of calculations would rapidly explode as the number of markers analyzed increased. To overcome this limitation, we devised the exact tree method based on decision tree analysis and applied it to 14 SNP data from 68 Japanese stroke patients and 189 healthy controls. From the obtained tree models, we succeeded in extracting multiple statistically significant combinations that elevate the risk of stroke. From this result, we inferred that this method would work more efficiently in the whole genome study, which handles thousands of genetic markers. This exploratory data mining method will facilitate the extraction of combinations from large-scale genetic data and provide a good foothold for further verificatory research.

Adult↗

Comparisons of gene colinearity in genomes using GeneOrder2.0.

Comparative genomics is enhanced by data mining the rapidly expanding DNA sequence databases. Because of the immense amount of data, computational tools and methods are needed to augment traditional manual visualizations and manipulations of these data. GeneOrder2.0, a Java-based interactive software programme, organizes genome sequence data into tabular and graphical visualizations of the extent of colinearity of genes between any two chromosome genomes of < or =250 kilobases. Both GenBank and proprietary data can be analyzed with this tool.

Computational Biology↗

The allometry of avian basal metabolic rate: good predictions need good data.

Basal metabolic rate (BMR) is often predicted by allometric interpolation, but such predictions are critically dependent on the quality of the data used to derive allometric equations relating BMR to body mass (Mb). An examination of the metabolic rates used to produce conventional and phylogenetically independent allometries for avian BMR in a recent analysis revealed that only 67 of 248 data unambiguously met the criteria for BMR and had sample sizes with n>/=3. The metabolic rates that represented BMR were significantly lower than those that did not meet the criteria for BMR or were measured under unspecified conditions. Moreover, our conventional allometric estimates of BMR (W; logBMR=-1.461+0.669logMb) using a more constrained data set that met the conditions that define BMR and had n>/=3 were 10%-12% lower than those obtained in the earlier analysis. The inclusion of data that do not represent BMR results in the overestimation of predicted BMR and can potentially lead to incorrect conclusions concerning metabolic adaptation. Our analyses using a data set that included only BMR with n>/=3 were consistent with the conclusion that BMR does not differ between passerine and nonpasserine birds after taking phylogeny into account. With an increased focus on data mining and synthetic analyses, our study suggests that a thorough knowledge of how data sets are generated and the underlying constraints on their interpretation is a necessary prerequisite for such exercises.

Analysis of Variance↗

Progenetix.net: an online repository for molecular cytogenetic aberration data.

UNLABELLED: Through sequencing projects and, more recently, array-based expression analysis experiments, a wealth of genetic data has become accessible via online resources. In contrast, few of the (molecular-) cytogenetic aberration data collected in the last decades are available in a format suitable for data mining procedures. www.progenetix.net is a new online repository for previously published chromosomal aberration data, allowing the addition of band-specific information about chromosomal imbalances to oncologic data analysis efforts. AVAILABILITY: http://www.progenetix.net CONTACT: mbaudis@stanford.edu

Chromosome Aberrations↗

Using data warehousing and OLAP in public health care.

The paper describes the possibilities of using data warehousing and OLAP technologies in public health care in general and then our own experience with these technologies gained during the implementation of a data warehouse of outpatient data at the national level. Such a data warehouse serves as a basis for advanced decision support systems based on statistical, OLAP or data mining methods. We used OLAP to enable interactive exploration and analysis of the data. We found out that data warehousing and OLAP are suitable for the domain of public health and that they enable new analytical possibilities in addition to the traditional statistical approaches.

Ambulatory Care↗

Cell cycle, DNA replication, repair, and recombination in the dimorphic human pathogenic fungus Paracoccidioides brasiliensis.

DNA replication, together with repair mechanisms and cell cycle control, are the most important cellular processes necessary to maintain correct transfer of genetic information to the progeny. These processes are well conserved throughout the Eukarya, and the genes that are involved provide essential information for understanding the life cycle of an organism. We used computational tools for data mining of genes involved in these processes in the pathogenic fungus Paracoccidiodes brasiliensis. Data derived from transcriptome analysis revealed that the cell cycle of this fungus, as well as DNA replication and repair, and the recombination machineries, are highly similar to those of the yeast Saccharomyces cerevisiae. Among orthologs detected in both species, there are genes related to cytoskeleton structure and assembly, chromosome segregation, and cell cycle control genes. We identified at least one representative gene from each step of the initiation of DNA replication. Major players in the process of DNA damage and repair were also identified.

Cell Cycle↗

EXAMINE: a computational approach to reconstructing gene regulatory networks.

Reverse-engineering of gene networks using linear models often results in an underdetermined system because of excessive unknown parameters. In addition, the practical utility of linear models has remained unclear. We address these problems by developing an improved method, EXpression Array MINing Engine (EXAMINE), to infer gene regulatory networks from time-series gene expression data sets. EXAMINE takes advantage of sparse graph theory to overcome the excessive-parameter problem with an adaptive-connectivity model and fitting algorithm. EXAMINE also guarantees that the most parsimonious network structure will be found with its incremental adaptive fitting process. Compared to previous linear models, where a fully connected model is used, EXAMINE reduces the number of parameters by O(N), thereby increasing the chance of recovering the underlying regulatory network. The fitting algorithm increments the connectivity during the fitting process until a satisfactory fit is obtained. We performed a systematic study to explore the data mining ability of linear models. A guideline for using linear models is provided: If the system is small (3-20 elements), more than 90% of the regulation pathways can be determined correctly. For a large-scale system, either clustering is needed or it is necessary to integrate information in addition to expression profile. Coupled with the clustering method, we applied EXAMINE to rat central nervous system development (CNS) data with 112 genes. We were able to efficiently generate regulatory networks with statistically significant pathways that have been predicted previously.

Algorithms↗

Using discordance to improve classification in narrative clinical databases: an application to community-acquired pneumonia.

Data mining in electronic medical records may facilitate clinical research, but much of the structured data may be miscoded, incomplete, or non-specific. The exploitation of narrative data using natural language processing may help, although nesting, varying granularity, and repetition remain challenges. In a study of community-acquired pneumonia using electronic records, these issues led to poor classification. Limiting queries to accurate, complete records led to vastly reduced, possibly biased samples. We exploited knowledge latent in the electronic records to improve classification. A similarity metric was used to cluster cases. We defined discordance as the degree to which cases within a cluster give different answers for some query that addresses a classification task of interest. Cases with higher discordance are more likely to be incorrectly classified, and can be reviewed manually to adjust the classification, improve the query, or estimate the likely accuracy of the query. In a study of pneumonia--in which the ICD9-CM coding was found to be very poor--the discordance measure was statistically significantly correlated with classification correctness (.45; 95% CI .15-.62).

Adult↗

Applications of qualitative multi-attribute decision models in health care.

Hierarchical decision models are a general decision support methodology aimed at the classification or evaluation of options that occur in decision-making processes. They are also important for the analysis, simulation and explanation of options. Decision models are typically developed through the decomposition of complex decision problems into smaller and less complex subproblems; the result of such decomposition is a hierarchical structure that consists of attributes and utility functions. This article presents an approach to the development and application of qualitative hierarchical decision models that is based on DEX, an expert system shell for multi-attribute decision support. The distinguishing characteristics of DEX are the use of qualitative (symbolic) attributes, and 'if-then' decision rules. Also, DEX provides a number of methods for the analysis of models and options, such as selective explanation and what-if analysis. We demonstrate the applicability and flexibility of the approach presenting four real-life applications of DEX in health care: assessment of breast cancer risk, assessment of basic living activities in community nursing, risk assessment in diabetic foot care, and technical analysis of radiogram errors. In particular, we highlight and justify the importance of knowledge presentation and option analysis methods for practical decision-making. We further show that, using a recently developed data mining method called HINT, such hierarchical decision models can be discovered from retrospective patient data.

Breast Neoplasms↗

Cytochrome P450 mono-oxygenases in conifer genomes: discovery of members of the terpenoid oxygenase superfamily in spruce and pine.

Diterpene resin acids, together with monoterpenes and sesquiterpenes, are the most prominent defence chemicals in conifers. These compounds belong to the large group of structurally diverse terpenoids formed by enzymes known as terpenoid synthases. CYPs (cytochrome P450-dependent mono-oxygenases) can further increase the structural diversity of these terpenoids. While most terpenoids are characterized as specialized or secondary metabolites, some terpenoids, such as the phytohormones GA (gibberellic acid), BRs (brassinosteroids) and ABA (abscisic acid), have essential functions in plant growth and development. To date, very few CYP genes involved in conifer terpenoid metabolism have been functionally characterized and were limited to two systems, yew (Taxus) and loblolly pine (Pinus taeda). The characterized yew CYP genes are involved in taxol diterpene biosynthesis, while the only characterized pine terpenoid CYP gene is part of DRA (diterpene resin acid) biosynthesis. These CYPs from yew and pine are members of two apparently conifer-specific CYP families within the larger CYP85 clan, one of four plant CYP multifamily clans. Other CYP families within the CYP85 clan were characterized from a variety of angiosperms with functions in terpenoid phytohormone metabolism of GA, BR, and ABA. The recent development of EST (expressed sequence tag) and FLcDNA (where FL is full-length) sequence databases and cDNA collections for species of two conifers, spruce (Picea) and pine, allows for the discovery of new terpenoid CYPs in gymnosperms by means of large-scale sequence mining, phylogenetic analysis and functional characterization. Here, we present a snapshot of conifer CYP data mining, discovery of new conifer CYPs in all but one family within the CYP85 clan, and suggestions for their functional characterization. This paper will focus on the discovery of conifer CYPs associated with diterpene metabolism and CYP with possible functions in the formation of GA, BR, and ABA in conifers.

Algorithms↗

Systematic functional evaluation of CNGA1 missense variants associated with retinitis pigmentosa.

BACKGROUND: Missense variants are frequently classified as variants of uncertain significance (VUS) according to the guidelines of the American College of Medical Genetics and Genomics and the Association of Molecular Pathology (ACMG/AMP). Consequently, disease relevance remains elusive, impeding molecular genetic diagnostics, patients` and family genetic counseling, and identification of patients eligible for clinical trials. Functional studies are critical for resolving the clinical significance of VUS. CNGA1 encodes the main subunit of the rod cyclic nucleotide-gated (CNG) channel, a vital component of the phototransduction cascade. Variants in CNGA1 are a rare cause of autosomal recessive retinitis pigmentosa and a phase I/II gene augmentation trial (NCT06291935) is currently ongoing highlighting the necessity to differentiate benign from pathogenic variants. METHODS: CNGA1 missense variants compiled from retinal disease patient cohorts, public databases and literature were functionally investigated using a medium-throughput aequorin-based assay and in vitro minigene splice assays for predicted exonic spliceogenic variants. Functional data were correlated with the in silico prediction of five variant effect predictors (VEPs) and applied to support or revise variants' ACMG/AMP classification. RESULTS: Data mining revealed 86 missense CNGA1 variants - including three novel - most of them lacking functional data; 65.1% of the variants were initially classified as VUS. The aequorin-based assay showed that 72.1% of tested variants significantly impaired CNG channel function and were classified as functionally abnormal, while 23.3% were functionally normal and 5% remained functionally uncertain. Correlation of the functional data with in silico predictions identified AlphaMissense and CPT-1 to be the most suitable tools for assessing CNGA1 missense variants. Using in vitro minigene splice assays, two putative missense variants were shown to induce missplicing. Based on the functional findings, 62.1% of the variants initially classified as VUS were re-categorized as likely pathogenic or likely benign. Furthermore, 93.3% of the variants initially classified as likely pathogenic showed an effect on CNGA1 channel function, confirming their disease relevance and supporting their reclassification as pathogenic. CONCLUSION: This study represents the first comprehensive functional assessment of disease-associated CNGA1 missense variants, thus significantly advancing the understanding of their disease relevance and improving molecular genetic diagnostics in patients.

Humans↗

Automated tissue analysis--a bioinformatics perspective.

OBJECTIVES: Recent progress in automated tissue analysis (tissomics) provides reproducible phenotypical characterization of histological specimens. We introduce informatics tools to cluster and correlate quantitative tissue profiles with gene expression data. The great potential of synergies between tissue analysis and bioinformatics and its perspectives in medical research and computational diagnostics are discussed. METHODS: Key enablers in microscopic imaging and machine vision are reviewed to perform a high-throughput tissue analysis. Methodologies are described and results are demonstrated that support a combined analysis of tissue with gene expression profiles whereby the consideration of individual responses is key. RESULTS: Comprehensive histomorphometric profiles, extracted using machine vision, provide information regarding the components and heterogeneity of a tissue in a reproducible format amenable to data mining and analysis. Tissue quantitative information can be placed in synergetic context with bioinformatics data, such as gene expression profiles, for a more comprehensive stratification of individual responses. From a bioinformatics point of view tissue data are co-variants that support the identification of candidate genes relevant in tissue injury or disease. CONCLUSIONS: Progress in automated analytics enables the generation of quantitative data about tissue previously limited to visual histopathology. Such reproducible data sets can be statistically correlated and clustered throughout the continuum of bioinformatics. The combined approach supports a system-wide view of biology and has a potential to accelerate developments for a personalized computational diagnosis.

Automation↗