Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Analysis of complex brain disorders with gene expression microarrays: schizophrenia as a disease of the synapse.

The level of cellular and molecular complexity of the nervous system creates unique problems for the neuroscientist in the design and implementation of functional genomic studies. Microarray technologies can be powerful, with limitations, when applied to the analysis of human brain disorders. Recently, using cDNA microarrays, altered gene expression patterns between subjects with schizophrenia and controls were shown. Functional data mining led to two novel discoveries: a consistent decrease in the group of transcripts encoding proteins that regulate presynaptic function; and the most changed gene, which has never been previously associated with schizophrenia, regulator of G-protein signaling 4. From these and other findings, a hypothesis has been formulated to suggest that schizophrenia is a disease of the synapse. In the context of a neurodevelopmental model, it is proposed that impaired mechanics of synaptic transmission in specific neural circuits during childhood and adolescence ultimately results in altered synapse formation or pruning, or both, which manifest in the clinical onset of the disease.

Brain Chemistry↗

The human proteomics initiative (HPI).

The availability of the human genome sequence has enabled the exploration and exploitation of the human genome and proteome to begin. Research has now focussed on the annotation of the genome and in particular of the proteome. With expert annotation extracted from the literature by biologists as the foundation, it has been possible to expand into the areas of data mining and automatic annotation. With further development and integration of pattern recognition methods and the application of alignments clustering, proteome analysis can now be provided in a meaningful way. These various approaches have been integrated to attach, extract and combine as much relevant information as possible to the proteome. This resource should be valuable to users from both research and industry.

Algorithms↗

Knowledge discovery with classification rules in a cardiovascular dataset.

In this paper we study an evolutionary machine learning approach to data mining and knowledge discovery based on the induction of classification rules. A method for automatic rules induction called AREX using evolutionary induction of decision trees and automatic programming is introduced. The proposed algorithm is applied to a cardiovascular dataset consisting of different groups of attributes which should possibly reveal the presence of some specific cardiovascular problems in young patients. A case study is presented that shows the use of AREX for the classification of patients and for discovering possible new medical knowledge from the dataset. The defined knowledge discovery loop comprises a medical expert's assessment of induced rules to drive the evolution of rule sets towards more appropriate solutions. The final result is the discovery of a possible new medical knowledge in the field of pediatric cardiology.

Algorithms↗

Bioinformatics in glycobiology.

In comparison with genes and proteins, attention paid to oligosaccharides that modify proteins is still marginal. Accordingly, bioinformatics is so far poorly involved in glycobiology. Some initiatives have been taken, however, to collect in databases all glycobiology-relevant information or to design specific data mining algorithms to infer predictions or identify oligosaccharide structures. In this review, we make a non-exhaustive survey of the available glycobiology-related bioinformatic resources, focussing mainly on those resources that are available through the World Wide Web. Some well-curated databases are identified, but the development of specialised algorithms appears to be limited.

Algorithms↗

Genomic analysis of stress response genes.

Mammalian cells respond to a wide range of external stimuli including growth factors, peptide hormones, cytokines, osmotic stress, heat shock, pharmacological agents and toxicants via multiple signalling pathways. Genome-wide transcript profiling simultaneously monitors the gene expression programs downstream of all signal transduction pathways and can identify novel molecular targets for stress-inducing signals. Our laboratory has combined transcript profiling of cytotoxic compounds with experimental systems in which signalling components are disrupted (e.g. small molecule protein kinase inhibitors) to reveal the contribution of specific signalling pathways to the transcriptional response to toxicant-induced stress. A complementary approach for elucidating the molecular mechanisms that regulate transcriptional responses to toxicants involves DNA sequence analysis of gene regulatory regions obtained via data mining of recently completed mammalian genome sequences. Together, these approaches reveal the molecular mechanisms used to finely tune alterations in gene expression, enabling cells to react in an appropriate manner to external stress-inducing stimuli.

Gene Expression Regulation, Enzymologic↗

Quick fuzzy backpropagation algorithm.

A modification of the fuzzy backpropagation (FBP) algorithm called QuickFBP algorithm is proposed, where the computation of the net function is significantly quicker. It is proved that the FBP algorithm is of exponential time complexity, while the QuickFBP algorithm is of polynomial time complexity. Convergence conditions of the QuickFBP, resp. the FBP algorithm are defined and proved for: (1) single output neural networks in case of training patterns with different targets; and (2) multiple output neural networks in case of training patterns with equivalued target vector. They support the automation of the weights training process (quasi-unsupervised learning) establishing the target value(s) depending on the network's input values. In these cases the simulation results confirm the convergence of both algorithms. An example with a large-sized neural network illustrates the significantly greater training speed of the QuickFBP rather than the FBP algorithm. The adaptation of an interactive web system to users on the basis of the QuickFBP algorithm is presented. Since the QuickFBP algorithm ensures quasi-unsupervised learning, this implies its broad applicability in areas of adaptive and adaptable interactive systems, data mining, etc. applications.

Algorithms↗

Knowledge discovery approach to automated cardiac SPECT diagnosis.

The paper describes a computerized process of myocardial perfusion diagnosis from cardiac single proton emission computed tomography (SPECT) images using data mining and knowledge discovery approach. We use a six-step knowledge discovery process. A database consisting of 267 cleaned patient SPECT images (about 3000 2D images), accompanied by clinical information and physician interpretation was created first. Then, a new user-friendly algorithm for computerizing the diagnostic process was designed and implemented. SPECT images were processed to extract a set of features, and then explicit rules were generated, using inductive machine learning and heuristic approaches to mimic cardiologist's diagnosis. The system is able to provide a set of computer diagnoses for cardiac SPECT studies, and can be used as a diagnostic tool by a cardiologist. The achieved results are encouraging because of the high correctness of diagnoses.

Artificial Intelligence↗

The Colorectal Cancer Recurrence Support (CARES) System.

Colorectal cancer has risen in incidence to become the second commonest form of cancer in Singapore. The primary treatment is surgery but up to 50% of patients still suffer from recurrence of the cancer after surgery. Early identification of recurrence will increase the effectiveness of therapy and the survival of patients. This paper describes the CARES (Cancer Recurrence Support) System, whose objective is to predict the recurrence of colorectal cancer, using Case-based Reasoning (CBR), and supported by other techniques such as data mining and natural language processing. The CARES System employs CBR to compare and contrast between the new and past colorectal cancer patient cases, and makes inferences based on those comparisons to determine the high risk patient groups. The features and functionality of the system are described.

Artificial Intelligence↗

Two-Stage Machine Learning model for guideline development.

We present a Two-Stage Machine Learning (ML) model as a data mining method to develop practice guidelines and apply it to the problem of dementia staging. Dementia staging in clinical settings is at present complex and highly subjective because of the ambiguities and the complicated nature of existing guidelines. Our model abstracts the two-stage process used by physicians to arrive at the global Clinical Dementia Rating Scale (CDRS) score. The model incorporates learning intermediate concepts (CDRS category scores) in the first stage that then become the feature space for the second stage (global CDRS score). The sample consisted of 678 patients evaluated in the Alzheimer's Disease Research Center at the University of California, Irvine. The demographic variables, functional and cognitive test results used by physicians for the task of dementia severity staging were used as input to the machine learning algorithms. Decision tree learners and rule inducers (C4.5, Cart, C4.5 rules) were selected for our study as they give expressive models, and Naive Bayes was used as a baseline algorithm for comparison purposes. We first learned the six CDRS category scores (memory, orientation, judgement and problem solving, personal care, home and hobbies, and community affairs). These learned CDRS category scores were then used to learn the global CDRS scores. The Two-Stage ML model classified as well as or better than the published inter-rater agreements for both the category and global CDRS scoring by dementia experts. Furthermore, for the most critical distinction, normal versus very mildly impaired, the Two-Stage ML model was 28.1 and 6.6% more accurate than published performances by domain experts. Our study of the CDRS examined one of the largest, most diverse samples in the literature, suggesting that our findings are robust. The Two-Stage ML model also identified a CDRS category, Judgment and Problem Solving, which has low classification accuracy similar to published reports. Since this CDRS category appears to be mainly responsible for misclassification of the global CDRS score when it occurs, further attribute and algorithm research on the Judgment and Problem Solving CDRS score could improve its accuracy as well as that of the global CDRS score.

Algorithms↗

Putting engineering back into protein engineering: bioinformatic approaches to catalyst design.

Complex multivariate engineering problems are commonplace and not unique to protein engineering. Mathematical and data-mining tools developed in other fields of engineering have now been applied to analyze sequence-activity relationships of peptides and proteins and to assist in the design of proteins and peptides with specified properties. Decreasing costs of DNA sequencing in conjunction with methods to quickly synthesize statistically representative sets of proteins allow modern heuristic statistics to be applied to protein engineering. This provides an alternative approach to expensive assays or unreliable high-throughput surrogate screens.

Algorithms↗

Maximum-likelihood crystallization.

The crystallization facility of the TB Structural Genomics Consortium, one of nine NIH-sponsored structural genomics pilot projects, employs a combinatorial random sampling technique in high-throughput crystallization screening. Although data are still sparse and a comprehensive analysis cannot be performed at this stage, preliminary results appear to validate the random-screening concept. A discussion of statistical crystallization data analysis aims to draw attention to the need for comprehensive and valid sampling protocols. In view of limited overlap in techniques and sampling parameters between the publicly funded high-throughput crystallography initiatives, exchange of information should be encouraged, aiming to effectively integrate data mining efforts into a comprehensive predictive framework for protein crystallization.

Crystallization↗

Use of recursive partitioning in the sequential screening of G-protein-coupled receptors.

High-throughput screening (HTS) is changing as more compounds and better assay techniques become available. HTS is also generating a large amount of data. There is a need to rationalize the HTS process, because, in some cases, the screening of all available compounds is not economically feasible. In addition to the selection of promising compounds, there is a need to learn from the data that we collect. In this paper, we use a data-mining method, recursive partitioning, to help uncover and understand structure-activity relations and to help biology and chemistry experts make better decisions on which compounds to screen next and better characterize. The sequential-screening process is presented and the results of applying that process to 14 G-protein-coupled receptor assays are reported.

Animals↗

Computer-aided diagnosis of breast tumors with different US systems.

RATIONALE AND OBJECTIVES: The authors performed this study to determine whether a computer-aided diagnostic (CAD) system was suitable from one ultrasound (US) unit to another after parameters were adjusted by using intelligent selection algorithms. MATERIALS AND METHODS: The authors used texture analysis and data mining with a decision tree model to classify breast tumors with different US systems. The databases of training cases from one unit and testing cases from another were collected from different countries. Regions of interest on US scans and co-variance texture parameters were used in the diagnosis system. Proposed adjustment schemes for different US systems were used to transform the information needed for a differential diagnosis. RESULTS: Comparison of the diagnostic system with and without adjustment, respectively, yielded the following results: accuracy, 89.9% and 82.2%; sensitivity, 94.6% and 92.2%; specificity, 85.4% and 72.3%; positive predictive value, 86.5% and 76.8%; and negative predictive value, 94.1% and 90.4%. The improvement in accuracy, specificity, and positive predictive value was statistically significant. Diagnostic performance was improved after the adjustment. CONCLUSION: After parameters were adjusted by using intelligent selection algorithms, the performance of the proposed CAD system was better both with the same and with different systems. Different resolutions, different setting conditions, and different scanner ages are no longer obstacles to the application of such a CAD system.

Algorithms↗

A cardiovascular EST repertoire: progress and promise for understanding cardiovascular disease.

The application of expressed sequence tag (EST) technology has proven to be an effective tool for gene discovery and the generation of gene expression profiles. The generation of an EST resource for the cardiovascular system has revealed significant insights into the changes in gene expression that guide heart development and disease. Furthermore, an important genetic resource has been developed for cardiovascular biology that is valuable for data mining and disease gene discovery.

Animals↗

High-throughput and virtual screening: core lead discovery technologies move towards integration.

In addition to high-throughput screening (HTS), the main lead discovery technology employed by most pharmaceutical companies today is virtual screening (VS). Although the two techniques have somewhat different philosophical origins, they contain many synergies that can potentially enhance the lead discovery process. Here, we describe many of the latest developments in VS technology with particular emphasis on their potential impact on HTS in, for example, focussed screening and data mining. In addition, we highlight key issues that need to be addressed before the potential of such efforts can be fully realized.

Journal Article↗

In need of high-throughput behavioral systems.

One of the current major bottlenecks in drug discovery is in vivo testing of candidate drugs in behavioral paradigms in normal or genetically altered mice. This testing is essential in discovering gene function and predicting potential efficacy of CNS drugs in humans. New efforts in the biotech community aim to alleviate this bottleneck by developing higher-throughput systems of behavioral, neurological and physiological analyses. Together with large pharmacological databases, equipped with state-of-the-art bioinformatic and/or data-mining algorithms, these systems will provide rapid and accurate indices of the therapeutic potential of novel drugs. By providing a substantial increase in the speed of behavioral testing, new high-throughput systems will facilitate current behavioral research with faster, more reliable approaches. Furthermore, screening whole drug-libraries and comparing the profiles of novel compounds to those of known compounds will facilitate the discovery of novel drugs. Target validation will also become more efficient with the fast characterization of novel mutant mice.

Animals↗

Visual and computational analysis of structure--activity relationships in high-throughput screening data.

Novel analytic methods are required to assimilate the large volumes of structural and bioassay data generated by combinatorial chemistry and high-throughput screening programmes in the pharmaceutical and agrochemical industries. Recent work in visualisation and data mining has been used to develop structure--activity relationships from such chemical-biological datasets.

Computational Biology↗

Functional and comparative genomics of pathogenic bacteria.

Microarray expression profiling and the development of data-mining tools and new statistical instruments affords an unprecedented opportunity for the genome-scale study of bacterial pathogenicity. Expression profiles obtained from bacteria grown in media simulating host microenvironments yield a portrait of interacting metabolic pathways and multistage developmental programs and disclose regulatory networks. The analysis of closely related strains and species by microarray-based comparative genomics provides a measure of genetic variability within natural populations and identifies crucial differences between pathogen and commensal. In the near future, the combined use of bacterial and host microarrays to study the same infected tissue will reveal the host-pathogen dialogue in a gene-by-gene and site- and time-specific manner. This review discusses the use of microarray-based expression profiling to identify genes of pathogenic bacteria that are differentially regulated in response to host-specific signals. Additionally, the review describes the application of microarray methods to disclose differences in gene content between taxonomically related strains that vary with respect to pathogenic phenotype.

Animals↗