Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Knowledge discovery for advanced clinical data management and analysis.

Knowledge discovery is a broad research field in which methods are developed to support discovery of novel and potentially useful knowledge from clinical databases and registers in systems for patient care. However, the techniques available are not readily applicable in medical domains, due to, among other reasons, low user friendliness and lack of proper methodological background. Data mining approaches to be explored and improved are predictive modelling, segmentation, dependency modelling, summarization, and change and deviation detection/modelling (in data or knowledge). Another and original contribution of the research is to build up efficient feedback loops. Human experts and available domain expert systems could provide suggestions as how to improve all major steps in the knowledge discovery process such as evaluation of knowledge, choice of data mining methods and data input. A long tradition of collecting and maintaining clinical and administrative data could be found in fields of oncology, cardiology, coronary surgery, social and primary health care medicine. All these areas, that gather data over long periods of time, could benefit from knowledge discovery.

Artificial Intelligence↗

Implementation of automated signal generation in pharmacovigilance using a knowledge-based approach.

Automated signal generation is a growing field in pharmacovigilance that relies on data mining of huge spontaneous reporting systems for detecting unknown adverse drug reactions (ADR). Previous implementations of quantitative techniques did not take into account issues related to the medical dictionary for regulatory activities (MedDRA) terminology used for coding ADRs. MedDRA is a first generation terminology lacking formal definitions; grouping of similar medical conditions is not accurate due to taxonomic limitations. Our objective was to build a data-mining tool that improves signal detection algorithms by performing terminological reasoning on MedDRA codes described with the DAML+OIL description logic. We propose the PharmaMiner tool that implements quantitative techniques based on underlying statistical and bayesian models. It is a JAVA application displaying results in tabular format and performing terminological reasoning with the Racer inference engine. The mean frequency of drug-adverse effect associations in the French database was 2.66. Subsumption reasoning based on MedDRA taxonomical hierarchy produced a mean number of occurrence of 2.92 versus 3.63 (p < 0.001) obtained with a combined technique using subsumption and approximate matching reasoning based on the ontological structure. Semantic integration of terminological systems with data mining methods is a promising technique for improving machine learning in medical databases.

Adverse Drug Reaction Reporting Systems↗

Mining association rules from a pediatric primary care decision support system.

The purpose of this study was to apply an unsupervised data mining algorithm to a database containing data collected at the point of care for clinical decision support. The data set was taken from the Child Health Improvement Program (CHIP), a preventive services tracking and reminder system in use at the University of North Carolina. The database contains over 30,000 visits. We used a previously described pattern discovery algorithm to extract 2nd and 3rd order association rules from the data and reviewed the literature two see if the associations had been described before. The algorithm discovered 16 2nd order associations and 103 3rd order associations. The 3rd order associations contained no new information. The 2nd order associations demonstrated a covariance among a range of health risk behaviors. Additionally, the algorithm discovered that both tobacco smoke exposure and chronic cardiopulmonary disease are associated with failure on developmental screens. These relationships have been described before and have been attributed to underlying poverty. The work demonstrates the ability of unsupervised data mining by rule association on sparse clinical data to discover clinically important associations. However, many associations may be previously known or explained by confounding variables.

Algorithms↗

[Study on the application of classification tree model in screening the risk factors of malignant tumor].

OBJECTIVE: To introduce the partitioning algorithm of classification tree model, and to explore the value of this data mining technique applied in data analysis of multifactorial diseases as malignant tumors. METHODS: Data was analyzed from a survey that conducted on 84 breast cancer patients and 273 cancer-free controls selected randomly in Jiashan county. The classification tree model was constructed using Exhaustive CHAID method and evaluated by the Risk statistics and the area under the ROC curve. RESULTS: 9 out of 105 effect risks factors were selected, in which career was the most important factor indicating that workers, teachers and retirees suffered much more risks than others. Nevertheless, the number of pregnancies, breast examination, reasons for menopause, age at menarche, intake of shrimp, crab, kipper, kelp and laver etc were also risk factors on breast cancer. However, physical exercise played different roles on different people. The Risk statistics of model was 0.174, and the area under the ROC curve was 0.872 which was significantly different from 0.5, suggesting that the classification tree model fit the actuality very well. CONCLUSION: The classification tree model could screen out the major affecting factors quickly and effectively and could also identify the cutting-points for continuous and ordinal variables, as well as revealing the complex interaction among the factors at many levels. This model might become a powerful tool to explore the complexities of the risks on diseases.

Algorithms↗

Mechanism of Shaofu Zhuyu decoction in improving diabetic mellitus erectile dysfunction inhibition of ferroptosis based on network pharmacology and experimental validation.

OBJECTIVE: To explore the medication patterns and mechanisms of action of Shaofu Zhuyu decoction (, SFZYD) in inhibiting ferroptosis through the nuclear factor erythroid 2-related factor 2 (Nrf2)/heme oxygenase 1 (HO-1)/glutathione peroxidase 4 (GPX4) pathway to improve diabetes mellitus-induced erectile dysfunction (DMED). METHODS: Firstly, data mining was employed to identify the medication patterns of Traditional Chinese Medicine (TCM) in treating DMED. Secondly, network pharmacology combined with a ferroptosis database was used to predict the targets. Subsequently, cell counting kit-8, 4',6-diamidino-2-phenylindole staining, reverse transcription-polymerase chain reaction (RT-PCR), and reagent kits were utilized to assess the repair effects of SFZYD on corpus cavernosum endothelial cells (CCECs) induced by high glucose (HG). Metabolic indicators, hematoxylin-eosin staining, and Masson staining were performed to observe the restorative effects of SFZYD on erectile function and penile tissue in diabetic rats. Finally, using Nrf2 inhibitors, the expression of related proteins and mRNAs was detected through Western blotting and RT-PCR. Reactive oxygen species levels and mitochondrial membrane potential were detected by flow cytometry. RESULTS: Data mining revealed that the prescription rules for blood stasis-type DMED coincide with the treatment principles of SFZYD. Network pharmacology identified 48 ferroptosis-related targets, primarily heme oxygenase 1 (HMOX1) and GPX4. Kyoto Encyclopedia of Genes and Genomes enrichment analysis associated these targets with the ferroptosis pathway. SFZYD repaired HG-induced CCECs damage and restored HMOX1 and GPX4 mRNA levels. in vivo, SFZYD effectively alleviated erectile dysfunction and repaired blood sinuses and fibrosis in diabetic rats. Following Nrf2 inhibition, the expression of Nrf2, HMOX1, and GPX4 decreased, while SFZYD intervention reversed these effects, improving ferroptosis and oxidative stress indicators. CONCLUSION: This study explored the potential mechanisms and efficacy of the TCM prescription SFZYD in treating DMED through data mining, network pharmacology analysis, cellular experiments, and animal experiments. It verified its effectiveness in repairing HG-induced CCECs damage, improving the pathological state of penile tissue in diabetic rats, and restoring erectile function by regulating the Nrf2/HO-1/GPX4 signaling pathway. This provides new insights and scientific evidence for treating DMED with TCM.

Male↗

Megavariate data analysis of mass spectrometric proteomics data using latent variable projection method.

There are many data mining techniques for processing and general learning of multivariate data. However, we believe the wavelet transformation and latent variable projection method are particularly useful for spectroscopic and chromatographic data. Projection based methods are designed to handle hugely multivariate nature of such data effectively. For the actual analysis of the data we have used latent variable projection methods such as principal component analysis (PCA) and partial least squares projection to latent structures based discriminant analysis (PLS-DA) to analyze the raw data presented to the participants of the First Duke Proteomics Data Mining Conference. PCA was used to solve problem #1 (clustering problem) and the PLS-DA was used to solve problem #2 (classification problem). The idea of internal and external cross-validation was used to validate the model obtained from the classification analysis. The simple two-component PLS-DA model obtained from the analysis performed well. The model has completely separated the two groups from all the data. The same model applied on two-thirds of the data showed good performance by external validation with independent test set of remaining 13 specimens obtained by setting aside the spectra of every third specimen (accuracy of 85%).

Artificial Intelligence↗

Mining microarray datasets aided by knowledge stored in literature.

DNA microarray technology produces large amounts of data. For data mining of these datasets, background information on genes can be helpful. Unfortunately most information is stored in free text. Here, we present an approach to use this information for DNA microarray data mining.

Databases, Genetic↗

Tools for statistical analysis with missing data: application to a large medical database.

Missing data is a common feature of large data sets in general and medical data sets in particular. Depending on the goal of statistical analysis, various techniques can be used to tackle this problem. Imputation methods consist in substituting the missing values with plausible or predicted values so that the completed data can then be analysed with any chosen data mining procedure. In this work, we study imputation in the context of multivariate data and we evaluate a number of methods which can be used by today's standard statistical software packages. Imputation using multivariate classification, multiple imputation and imputation by factorial analysis are compared using simulated data and a large medical database (from the diabetes field) with numerous missing values. Our main result is to provide a control chart for assessing data quality after the imputation process. To this end, we developed an algorithm for which the input is a set of parameters describing the underlying data (e.g., covariance matrix, distribution) and the output is a chart which plots the change in the prediction error with respect to the proportion of missing values. The chart is built by means of an iterative algorithm involving four steps: (1) a sample of simulated data is drawn by using the input parameters; (2) missing values are randomly generated; (3) an imputation method is used to fill in the missing data and (4) the prediction error is computed. Steps 1 to 4 are repeated in order to estimate the distribution of the prediction error. The control chart was established for the 3 imputation methods studied here, assuming a multivariate normal distribution of data. The use of this tool on a large medical database was then investigated. We show how the control chart can be used to assess the quality of the imputation process in the pre-processing step upstream of data mining procedures.

Algorithms↗

An integrative genomic approach to uncover molecular mechanisms of prokaryotic traits.

With mounting availability of genomic and phenotypic databases, data integration and mining become increasingly challenging. While efforts have been put forward to analyze prokaryotic phenotypes, current computational technologies either lack high throughput capacity for genomic scale analysis, or are limited in their capability to integrate and mine data across different scales of biology. Consequently, simultaneous analysis of associations among genomes, phenotypes, and gene functions is prohibited. Here, we developed a high throughput computational approach, and demonstrated for the first time the feasibility of integrating large quantities of prokaryotic phenotypes along with genomic datasets for mining across multiple scales of biology (protein domains, pathways, molecular functions, and cellular processes). Applying this method over 59 fully sequenced prokaryotic species, we identified genetic basis and molecular mechanisms underlying the phenotypes in bacteria. We identified 3,711 significant correlations between 1,499 distinct Pfam and 63 phenotypes, with 2,650 correlations and 1,061 anti-correlations. Manual evaluation of a random sample of these significant correlations showed a minimal precision of 30% (95% confidence interval: 20%-42%; n = 50). We stratified the most significant 478 predictions and subjected 100 to manual evaluation, of which 60 were corroborated in the literature. We furthermore unveiled 10 significant correlations between phenotypes and KEGG pathways, eight of which were corroborated in the evaluation, and 309 significant correlations between phenotypes and 166 GO concepts evaluated using a random sample (minimal precision = 72%; 95% confidence interval: 60%-80%; n = 50). Additionally, we conducted a novel large-scale phenomic visualization analysis to provide insight into the modular nature of common molecular mechanisms spanning multiple biological scales and reused by related phenotypes (metaphenotypes). We propose that this method elucidates which classes of molecular mechanisms are associated with phenotypes or metaphenotypes and holds promise in facilitating a computable systems biology approach to genomic and biomedical research.

Algorithms↗

[The ideal form of laboratory information management].

In a clinical laboratory, not many staff can point out the problems of laboratory information management. Although the clinical laboratory introduced information systems in early stage, no organization supplies specialists to this field. Much knowledge is hidden in the clinical laboratory data, which can be discovered by data-mining technology. We can contribute to medical development with this technology. Moreover, the cost of routine work and research work may also be mitigated. However, data-mining technology including structurally recorded data and diversified analytic systems are required to build such capability. The laboratory information management division should make sufficient use of the formal information with non-fixed data base searching. This section should become an important section in the hospital by supplying advanced knowledge discovery and strategic decision-making. In this paper, we discuss the necessity of the information education in the clinical laboratory field and describe the importance of information management in a clinical laboratory.

Clinical Laboratory Information Systems↗

Association rule mining in peer-to-peer systems.

We extend the problem of association rule mining--a key data mining problem--to systems in which the database is partitioned among a very large number of computers that are dispersed over a wide area. Such computing systems include grid computing platforms, federated database systems, and peer-to-peer computing environments. The scale of these systems poses several difficulties, such as the impracticality of global communications and global synchronization, dynamic topology changes of the network, on-the-fly data updates, the need to share resources with other applications, and the frequent failure and recovery of resources. We present an algorithm by which every node in the system can reach the exact solution, as if it were given the combined database. The algorithm is entirely asynchronous, imposes very little communication overhead, transparently tolerates network topology changes and node failures, and quickly adjusts to changes in the data as they occur. Simulation of up to 10,000 nodes show that the algorithm is local: all rules, except for those whose confidence is about equal to the confidence threshold, are discovered using information gathered from a very small vicinity, whose size is independent of the size of the system.

Algorithms↗

The predictive power of the CluSTr database.

SUMMARY: The CluSTr database employs a fully automatic single-linkage hierarchical clustering method based on a similarity matrix. In order to compute the matrix, first all-against-all pair-wise comparisons between protein sequences are computed using the Smith-Waterman algorithm. The statistical significance of the similarity scores is then assessed using a Monte Carlo analysis, yielding Z-values, which are used to populate the matrix. This paper describes automated annotation experiments that quantify the predictive power and hence the biological relevance of the CluSTr data. The experiments utilized the UniProt data-mining framework to derive annotation predictions using combinations of InterPro and CluSTr. We show that this combination of data sources greatly increases the precision of predictions made by the data-mining framework, compared with the use of InterPro data alone. We conclude that the CluSTr approach to clustering proteins makes a valuable contribution to traditional protein classifications. AVAILABILITY: http://www.ebi.ac.uk/clustr/.

Algorithms↗

BioAfrica's HIV-1 proteomics resource: combining protein data with bioinformatics tools.

Most Internet online resources for investigating HIV biology contain either bioinformatics tools, protein information or sequence data. The objective of this study was to develop a comprehensive online proteomics resource that integrates bioinformatics with the latest information on HIV-1 protein structure, gene expression, post-transcriptional/post-translational modification, functional activity, and protein-macromolecule interactions. The BioAfrica HIV-1 Proteomics Resource http://bioafrica.mrc.ac.za/proteomics/index.html is a website that contains detailed information about the HIV-1 proteome and protease cleavage sites, as well as data-mining tools that can be used to manipulate and query protein sequence data, a BLAST tool for initiating structural analyses of HIV-1 proteins, and a proteomics tools directory. The Proteome section contains extensive data on each of 19 HIV-1 proteins, including their functional properties, a sample analysis of HIV-1HXB2, structural models and links to other online resources. The HIV-1 Protease Cleavage Sites section provides information on the position, subtype variation and genetic evolution of Gag, Gag-Pol and Nef cleavage sites. The HIV-1 Protein Data-mining Tool includes a set of 27 group M (subtypes A through K) reference sequences that can be used to assess the influence of genetic variation on immunological and functional domains of the protein. The BLAST Structure Tool identifies proteins with similar, experimentally determined topologies, and the Tools Directory provides a categorized list of websites and relevant software programs. This combined database and software repository is designed to facilitate the capture, retrieval and analysis of HIV-1 protein data, and to convert it into clinically useful information relating to the pathogenesis, transmission and therapeutic response of different HIV-1 variants. The HIV-1 Proteomics Resource is readily accessible through the BioAfrica website at: http://bioafrica.mrc.ac.za/proteomics/index.html.

Africa↗

On combining recursive partitioning and simulated annealing to detect groups of biologically active compounds.

Statistical data mining methods have proven to be powerful tools for investigating correlations between molecular structure and biological activity. Recursive partitioning (RP), in particular, offers several advantages in mining large, diverse data sets resulting from high throughput screening. When used with binary molecular descriptors, the standard implementation of RP splits on single descriptors. We use simulated annealing (SA) to find combinations of molecular descriptors whose simultaneous presence best separates off the most active, chemically similar group of compounds. The search is incorporated into a recursive partitioning design to produce a regression tree for biological activity on the space of structural fingerprints. Each node is characterized by a specific combination of structural features, and the terminal nodes with high average activities correspond, roughly, to different classes of compounds. Using LeadScope structural features as descriptors to mine a database from the National Cancer Institute, the merging of RP and SA consistently identifies structurally homogeneous classes of highly potent anticancer agents.

Algorithms↗

Key enzymes of the protocatechuate branch of the beta-ketoadipate pathway for aromatic degradation in Corynebacterium glutamicum.

Although the protocatechuate branch of the beta-ketoadipate pathway in Gram+ bacteria has been well studied, this branch is less understood in Gram+ bacteria. In this study, Corynebacterium glutamicum was cultivated with protocatechuate, p-cresol, vanillate and 4-hydroxybenzoate as sole carbon and energy sources for growth. Enzymatic assays indicated that growing cells on these aromatic compounds exhibited protocatechuate 3,4-dioxygenase activities. Data-mining of the genome of this bacterium revealed that the genetic locus ncg12314-ncg12315 encoded a putative protocatechuate 3,4-dioxygenase. The genes, ncg12314 and ncg12315, were amplified by PCR technique and were cloned into plasmid (pET21aP34D). Recombinant Escherichia coli strain harboring this plasmid actively expressed protocatechuate 3,4-dioxygenase activity. Further, when this locus was disrupted in C. glutamicum, the ability to degrade and assimilate protocatechuate, p-cresol, vanillate or 4-hydroxybenzoate was lost and protocatechuate 3,4-dioxygenase activity was disappeared. The ability to grow with these aromatic compounds and protocatechuate 3,4-dioxygenase activity of C. glutamicum mutant could be restored by gene complementation. Thus, it is clear that the key enzyme for ring-cleavage, protocatechuate 3,4-dioxygenase, was encoded by ncg12314 and ncg12315. The additional genes involved in the protocatechuate branch of the beta-ketoadipate pathway were identified by mining the genome data publically available in the GenBank. The functional identification of genes and their unique organization in C. glutamicum provided new insight into the genetic diversity of aromatic compound degradation.

Adipates↗

Assessment of approximate string matching in a biomedical text retrieval problem.

Text-based search is widely used for biomedical data mining and knowledge discovery. Character errors in literatures affect the accuracy of data mining. Methods for solving this problem are being explored. This work tests the usefulness of the Smith-Waterman algorithm with affine gap penalty as a method for biomedical literature retrieval. Names of medicinal herbs collected from herbal medicine literatures are matched with those from medicinal chemistry literatures by using this algorithm at different string identity levels (80-100%). The optimum performance is at string identity of 88%, at which the recall and precision are 96.9% and 97.3%, respectively. Our study suggests that the Smith-Waterman algorithm is useful for improving the success rate of biomedical text retrieval.

Algorithms↗

Structural and functional characterization of pi bulges and other short intrahelical deformations.

We data-mined the Protein Data Bank for short intrahelical deformations, including pi bulges. These are defined as a contiguous stretch of intrahelical residues deviating from the standard alpha-helical i-->i-4 hydrogen bonding pattern, bilaterally flanked by at least one alpha-helical turn resulting in a helix kink of less than 40 degrees. We find that such motifs exist in 4.7% of a PDB subset filtered by quality metrics (resolution <2.5 A, R-factor <0.25, sequence identity <35%). These are typically characterized by at least one i-->i-5 main chain hydrogen bond, with energetically favorable main chain dihedral angles, followed by a variable number of main chain carbonyl groups that do not accept intrahelical main chain hydrogen bonds. Their stabilization commonly occurs via hydrogen bonding to water molecules or polar groups. Numerous deformations are implicated in basic yet vital functional roles, commonly as ligand binding site contributors.

Amino Acid Motifs↗

Prediction of stone disease by discriminant analysis and artificial neural networks in genetic polymorphisms: a new method.

OBJECTIVE: To use information from genetic polymorphisms and from patients (drinking/exercise habits) to identify their association with stone disease, the main analytical and predictive tools being discriminant analysis (DA) and artificial neural networks (ANNs). PATIENTS, SUBJECTS AND METHODS: Urinary stone disease is common in Taiwan; the formation of calcium oxalate stone is reportedly associated with genetic polymorphisms but there are many of these. Genotyping requires many individuals and markers because of the complexity of gene-gene and gene-environmental factor interactions. With the development of artificial intelligence, data-mining tools like ANNs can be used to derive more from patient data in predicting disease. Thus we compared 151 patients with calcium oxalate stones and 105 healthy controls for the presence of four genetic polymorphisms; cytochrome p450c17, E-cadherin, urokinase and vascular endothelial growth factor (VEGF). Information about environmental factors, e.g. water, milk and coffee consumption, and outdoor activities, was also collected. Stepwise DA and ANNs were used as classification methods to obtain an effective discriminant model. RESULTS: With only the genetic variables, DA successfully classified 64% of the participants, but when all related factors (gene and environmental factors) were considered simultaneously, stepwise DA was successful in classifying 74%. The results for DA were best when six variables (sex, VEGF, stone number, coffee, milk, outdoor activities), found by iterative selection, were used. The ANN successfully classified 89% of participants and was better than DA when considering all factors in the model. A sensitivity analysis of the input parameters for ANN was conducted after the ANN program was trained; the most important inputs affecting stone disease were genetic (VEGF), while the second and third were water and milk consumption. CONCLUSIONS: While data-mining tools such as DA and ANN both provide accurate results for assessing genetic markers of calcium stone disease, the ANN provides a better prediction than the DA, especially when considering all (genetic and environmental) related factors simultaneously. This model provides a new way to study stone disease in combination with genetic polymorphisms and environmental factors.

Adult↗