Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

Knowledge discovery in databases of biomechanical variables: application to the sit to stand motor task.

ABSTRACT : BACKGROUND : The interpretation of data obtained in a movement analysis laboratory is a crucial issue in clinical contexts. Collection of such data in large databases might encourage the use of modern techniques of data mining to discover additional knowledge with automated methods. In order to maximise the size of the database, simple and low-cost experimental set-ups are preferable. The aim of this study was to extract knowledge inherent in the sit-to-stand task as performed by healthy adults, by searching relationships among measured and estimated biomechanical quantities. An automated method was applied to a large amount of data stored in a database. The sit-to-stand motor task was already shown to be adequate for determining the level of individual motor ability. METHODS : The technique of search for association rules was chosen to discover patterns as part of a Knowledge Discovery in Databases (KDD) process applied to a sit-to-stand motor task observed with a simple experimental set-up and analysed by means of a minimum measured input model. Selected parameters and variables of a database containing data from 110 healthy adults, of both genders and of a large range of age, performing the task were considered in the analysis. RESULTS : A set of rules and definitions were found characterising the patterns shared by the investigated subjects. Time events of the task turned out to be highly interdependent at least in their average values, showing a high level of repeatability of the timing of the performance of the task. CONCLUSIONS : The distinctive patterns of the sit-to-stand task found in this study, associated to those that could be found in similar studies focusing on subjects with pathologies, could be used as a reference for the functional evaluation of specific subjects performing the sit-to-stand motor task.

Journal Article↗

Reconsidering medical malpractice reform: the case for arbitration and transparency in non-emergent contexts.

This Article proposes a two-pronged legislative response to the current debate over medical malpractice insurance. The author does not advocate mandatory caps on malpractice damages, nor the imposition of a uniform regime on the field of medicine. Rather, he articulates some of the important legal, medical, and societal benefits that would come from embracing arbitration in the non-emergent medical malpractice context. The author also calls for the reformulation of the National Practitioner Data Bank to achieve greater transparentcy and to leverage advances in information technology and data-mining software to measure the risk levels of individual practitioners. This reform, in turn, would open up the possibility of greater subcategorization of premiums and more effective deterrence in medical malpractice insurance.

Federal Government↗

Mining for diagnostic information in body surface potential maps: a comparison of feature selection techniques.

BACKGROUND: In body surface potential mapping, increased spatial sampling is used to allow more accurate detection of a cardiac abnormality. Although diagnostically superior to more conventional electrocardiographic techniques, the perceived complexity of the Body Surface Potential Map (BSPM) acquisition process has prohibited its acceptance in clinical practice. For this reason there is an interest in striking a compromise between the minimum number of electrocardiographic recording sites required to sample the maximum electrocardiographic information. METHODS: In the current study, several techniques widely used in the domains of data mining and knowledge discovery have been employed to mine for diagnostic information in 192 lead BSPMs. In particular, the Single Variable Classifier (SVC) based filter and Sequential Forward Selection (SFS) based wrapper approaches to feature selection have been implemented and evaluated. Using a set of recordings from 116 subjects, the diagnostic ability of subsets of 3, 6, 9, 12, 24 and 32 electrocardiographic recording sites have been evaluated based on their ability to correctly asses the presence or absence of Myocardial Infarction (MI). RESULTS: It was observed that the wrapper approach, using sequential forward selection and a 5 nearest neighbour classifier, was capable of choosing a set of 24 recording sites that could correctly classify 82.8% of BSPMs. Although the filter method performed slightly less favourably, the performance was comparable with a classification accuracy of 79.3%. In addition, experiments were conducted to show how (a) features chosen using the wrapper approach were specific to the classifier used in the selection model, and (b) lead subsets chosen were not necessarily unique. CONCLUSION: It was concluded that both the filter and wrapper approaches adopted were suitable for guiding the choice of recording sites useful for determining the presence of MI. It should be noted however that in this study recording sites have been suggested on their ability to detect disease and such sites may not be optimal for estimating body surface potential distributions.

Algorithms↗

Symbolic, neural, and Bayesian machine learning models for predicting carcinogenicity of chemical compounds.

Experimental programs have been underway for several years to determine the environmental effects of chemical compounds, mixtures, and the like. Among these programs is the National Toxicology Program (NTP) on rodent carcinogenicity. Because these experiments are costly and time-consuming, the rate at which test articles (i.e., chemicals) can be tested is limited. The ability to predict the outcome of the analysis at various points in the process would facilitate informed decisions about the allocation of testing resources. To assist human experts in organizing an empirical testing regime, and to try to shed light on mechanisms of toxicity, we constructed toxicity models using various machine learning and data mining methods, both existing and those of our own devising. These models took the form of decision trees, rule sets, neural networks, rules extracted from trained neural networks, and Bayesian classifiers. As a training set, we used recent results from rodent carcinogenicity bioassays conducted by the NTP on 226 test articles. We performed 10-way cross-validation on each of our models to approximate their expected error rates on unseen data. The data set consists of physical-chemical parameters of test articles, alerting chemical substructures, salmonella mutagenicity assay results, subchronic histopathology data, and information on route, strain, and sex/species for 744 individual experiments. These results contribute to the ongoing process of evaluating and interpreting the data collected from chemical toxicity studies.

Animals↗

Chemical effects in biological systems (CEBS) object model for toxicology data, SysTox-OM: design and application.

MOTIVATION: The CEBS data repository is being developed to promote a systems biology approach to understand the biological effects of environmental stressors. CEBS will house data from multiple gene expression platforms (transcriptomics), protein expression and protein-protein interaction (proteomics), and changes in low molecular weight metabolite levels (metabolomics) aligned by their detailed toxicological context. The system will accommodate extensive complex querying in a user-friendly manner. CEBS will store toxicological contexts including the study design details, treatment protocols, animal characteristics and conventional toxicological endpoints such as histopathology findings and clinical chemistry measures. All of these data types can be integrated in a seamless fashion to enable data query and analysis in a biologically meaningful manner. RESULTS: An object model, the SysBio-OM (Xirasagar et al., 2004) has been designed to facilitate the integration of microarray gene expression, proteomics and metabolomics data in the CEBS database system. We now report SysTox-OM as an open source systems toxicology model designed to integrate toxicological context into gene expression experiments. The SysTox-OM model is comprehensive and leverages other open source efforts, namely, the Standard for Exchange of Nonclinical Data (http://www.cdisc.org/models/send/v2/index.html) which is a data standard for capturing toxicological information for animal studies and Clinical Data Interchange Standards Consortium (http://www.cdisc.org/models/sdtm/index.html) that serves as a standard for the exchange of clinical data. Such standardization increases the accuracy of data mining, interpretation and exchange. The open source SysTox-OM model, which can be implemented on various software platforms, is presented here. AVAILABILITY: A universal modeling language (UML) depiction of the entire SysTox-OM is available at http://cebs.niehs.nih.gov and the Rational Rose object model package is distributed under an open source license that permits unrestricted academic and commercial use and is available at http://cebs.niehs.nih.gov/cebsdownloads. Currently, the public toxicological data in CEBS can be queried via a web application based on the SysTox-OM at http://cebs.niehs.nih.gov CONTACT: xirasagars@saic.com SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Computational Biology↗

CryptoDB: the Cryptosporidium genome resource.

CryptoDB (http://CryptoDB.org) represents a collaborative effort to locate all genome data for the apicomplexan parasite Cryptosporidium parvum in a single user-friendly database. CryptoDB currently houses the genomic sequence data for both the human type 1 H strain and the bovine type 2 IOWA strain in addition to all other available EST and GSS sequences obtained from public repositories. All data are available for data mining via BLAST, keyword searches of pre-computed BLASTX results and user-defined or PROSITE motif pattern searches. Release 1.0 of CryptoDB contains approximately 19 million bases of genome sequence for the H and IOWA strains and an additional approximately 24 million bases of GSS and EST sequence obtained from other sources. Open reading frames greater than 50 and 100 amino acids have been generated for all sequences and all data are available for bulk download. This database, like other apicomplexan parasite databases, has been built utilizing the PlasmoDB model.

Animals↗

Mining fatty acid databases for detection of novel compounds in aerobic bacteria.

This study examines how the discriminatory power of an automated bacterial whole-cell fatty acid identification system can be significantly enhanced by exploring the vast amounts of information accumulated during 15 years of routine gas chromatographic analysis of the fatty acid content of aerobic bacteria. Construction of a global peak occurrence histogram based upon a large fatty acid database is shown to serve as a highly informative tool for assessing the delineation of the naming windows used during the automatic recognition of fatty acid compounds. Along the lines of this data mining application, it is suggested that several naming windows of the Sherlock MIS TSBA50 peak naming method may need to be re-evaluated in order to fit more closely with the bulk of observed fatty acid profiles. At the same time, the global peak occurrence histogram has put forward the delineation of 32 new peak naming windows, accounting for a 26% increase in the total number of fatty acid features taken into account for bacterial identification. By scrutinizing the relationships between the newly delineated naming windows and the many taxonomic units covered within a proprietary fatty acid database, all new naming windows were proven to correspond with stable features of some specific groups of microorganisms. This latter analysis clearly underscores the impact of incorporating the new fatty acid compounds for improving the resolution of the bacterial identification system and endorses the applicability of knowledge discovery in databases within the field of microbiology.

Bacteria, Aerobic↗

Congruent strategies for carbohydrate sequencing. 1. Mining structural details by MSn.

This report is the first in a series of three focused on establishing congruent strategies for carbohydrate sequencing. The reports are divided into (i) analytical considerations that account for all aspects of small oligomer structure by MSn disassembly, (ii) database support using an ion fragment library and associated tools for high-throughput analysis, and (iii) a concluding algorithm for defining oligosaccharide topology from MSn disassembly pathways. The analytical contribution of this first report explores the limits of structural detail exposed by ion trap mass spectrometry with samples prepared as methyl derivatives and analyzed as metal ion adducts. This data mining effort focuses on correlating the fragments of small oligomers to stereospecific glycan structures, an outcome attributed to a combination of metal ion adduction and analyte conformation. Facile glycosidic cleavage introduces a point of lability (pyranosyl-1-ene) that upon collisional activation initiates subsequent ring fragmentation. Product masses and ion intensities vary with interresidue linkage, branching position, and monomer stereochemistry. Excessive fragmentation is the property of small oligomers where collisional energy within a smaller number of oscillators dissipates through extensive fragmentation. The procedures discussed in this report are unified into a singular strategy using an ion trap mass spectrometer with the sensitivity expected for electron multiplier detection. Although a small set of structures have been discussed, the basic principles considered are fully congruent, with ample opportunities for expansion.

Carbohydrate Sequence↗

Mining reproducible activation patterns in epileptic intracerebral EEG signals: application to interictal activity.

The study of interictal transient events may substantially complement the analysis of seizures in the presurgical evaluation of intractable epilepsy. A comprehensive methodology of quantifying reproducibility of activation patterns in intracerebral electroencephalography signals is presented. It may be applied to various forms of transient epileptic events under the assumption that a time of occurrence may be assigned to them. In this paper, the method is used on two different forms of interictal events (interictal spikes or sharpwaves and transient bursts of fast activity). The methodology is based on signal processing and data mining algorithms and proceeds in three steps: 1) detection of transient paroxysmal events (monochannel event); 2) identification of quasisynchronous transient paroxysmal events (multichannel events); and 3) automatic extraction of similar activation patterns. Results show that the methodology allows reproducible sequential activation sets to be identified from signals recorded in four patients. Potential advantages of the method are discussed with respect to other approaches.

Algorithms↗

Analysis of randomized controlled trials.

Although the sophistication and flexibility of the statistical technology available to the data analyst have increased, some durable, simple principles remain valid. Hypothesis-driven analyses, which were anticipated and specified in the protocol, must still be kept separate and privileged relative to the important, but risky data mining made possible by modern computers. Analyses that have a firm basis in the randomization are interpreted more easily than those that rely heavily on statistical models. Outcomes--such as quality of life, symptoms, and behaviors--that require the cooperation of subjects to be measured will come to be more and more important as trials move away from mortality as the main outcome. Inevitably, such trials will have to deal with more missing data, especially because of dropout and noncompliance. There are fundamental limits on the ability of statistical methods to compensate for such problems, so they must be considered when studies are designed. Finally, it must be emphasized that the availability of software is not a substitute for experience and statistical expertise.

Data Interpretation, Statistical↗

Distribution patterns of over-represented k-mers in non-coding yeast DNA.

MOTIVATION: Over-represented k-mers in genomic DNA regions are often of particular biological interest. For example, over-represented k-mers in co-regulated families of genes are associated with the DNA binding sites of transcription factors. To measure over-representation, we introduce a statistical background model based on single-mismatches, and apply it to the pooled 500 bp ORF Upstream Regions (USRs) of yeast. More importantly, we investigate the context and spatial distribution of over-represented k-mers in yeast USRs. RESULTS: Single and double-stranded spatial distributions of most over-represented k-mers are highly non-random, and predominantly cluster into a small number of classes that are robust with respect to over-representation measures. Specifically, we show that the three most common distribution patterns can be related to DNA structure, function, and evolution and correspond to: (a) homologous ORF clusters associated with sharply localized distributions; (b) regulatory elements associated with a symmetric broad hill-shaped distribution in the 50-200 bp USR; and (c) runs of As, Ts, and ATs associated with a broad hill-shaped distribution also in the 50-200 bp USR, with extreme structural properties. Analysis of over-representation, homology, localization, and DNA structure are essential components of a general data-mining approach to finding biologically important k-mers in raw genomic DNA and understanding the 'lexicon' of regulatory regions.

Amino Acid Motifs↗

Computer system "gene discovery" for promoter structure analysis.

This paper presents implementation of Data Mining and Knowledge Discovery techniques for searching for regularities in tables of context features of DNA sequences involved in regulation of transcription. The goal is to discover regularities that relate nucleotide sequences to the functional classes of these sequences. The search patterns for regularities have been constructed in the first-order logic augmented by probabilistic estimates. To this aim. the PC software system "Gene Discovery" has been designed. This system accepts molecular-genetical data retrieved from a database by using SQL queries. Nucleotide sequences of promoters of several functional systems were extracted from the TRRD database (http://wwwmgs.bionet.nsc.ru/mgs/gnw/trrd/) and analysed. The data include nucleotide sequences of erythroid-specific gene promoters, endocrine system gene promoters, promoter regions of the genes controlling cell cycle, promoter of genes regulating lipid metabolism, and muscle-specific gene promoters. Several regularities that relate the nucleotide sequences in the regulatory DNA and their location relative to the transcription start with each functional class have been found.

Base Sequence↗

A complete inventory of fungal kinesins in representative filamentous ascomycetes.

Complete inventories of kinesins from three pathogenic filamentous ascomycetes, Botryotinia fuckeliana, Cochliobolus heterostrophus, and Gibberella moniliformis, are described. These protein sequences were compared with those of the filamentous saprophyte, Neurospora crassa and the two yeasts Saccharomyces cerevisiae and Schizosaccharomyces pombe. Data mining and phylogenetic analysis of the motor domain yielded a constant set of 10 kinesins in the filamentous fungal species, compared with a smaller set in S. cerevisiae and S. pombe. The filamentous fungal kinesins fell into nine subfamilies when compared with well-characterized kinesins from other eukaryotes. A few putative kinesins (one in B. fuckeliana and two in C. heterostrophus) could not be defined as functional, due to unorthodox organization and lack of experimental data. The broad representation of filamentous fungal kinesins across most of the known subfamilies and the ease of gene manipulation make fungi ideal models for functional and evolutionary investigation of these proteins.

Ascomycota↗

High-throughput metabolic state analysis: the missing link in integrated functional genomics of yeasts.

The lack of comparable metabolic state assays severely limits understanding the metabolic changes caused by genetic or environmental perturbations. The present study reports the application of a novel derivatization method for metabolome analysis of yeast, coupled to data-mining software that achieve comparable throughput, effort and cost compared with DNA arrays. Our sample workup method enables simultaneous metabolite measurements throughout central carbon metabolism and amino acid biosynthesis, using a standard GC-MS platform that was optimized for this purpose. As an implementation proof-of-concept, we assayed metabolite levels in two yeast strains and two different environmental conditions in the context of metabolic pathway reconstruction. We demonstrate that these differential metabolite level data distinguish among sample types, such as typical metabolic fingerprinting or footprinting. More importantly, we demonstrate that this differential metabolite level data provides insight into specific metabolic pathways and lays the groundwork for integrated transcription-metabolism studies of yeasts.

Aerobiosis↗

Identification of ligands for two human bitter T2R receptors.

Earlier, a family of G protein-coupled receptors, termed T2Rs, was identified in the rodent and human genomes through data mining. It was suggested that these receptors mediate bitter taste perception. Analysis of the human genome revealed that the hT2R family is composed of 25 members. However, bitter ligands have been identified for only three human receptors so far. Here we report identification of two novel ligand-receptor pairs. hT2R61 is activated by 6-nitrosaccharin, a bitter derivative of saccharin. hT2R44 is activated by denatonium and 6-nitrosaccharin. Activation profiles for these receptors correlate with psychophysical data determined for the bitter compounds in human studies. Functional analysis of hT2R chimeras allowed us to identify residues in extracellular loops critical for receptor activation by ligands. The discovery of two novel bitter ligand-receptor pairs provides additional support for the hypothesis that hT2Rs mediate a bitter taste response in humans.

Amino Acid Sequence↗

Spermatozoal RNA profiles of normal fertile men.

BACKGROUND: Findings from several studies support the conclusion that spermatozoa contain a complex repertoire of mRNAs. Even though these mRNAs are thought to provide an insight into past events of spermatogenesis, their complexity and function have yet to be established. Our aim was to determine whether we could use spermatozoal mRNAs to generate a genetic fingerprint of normal fertile men. METHODS: We used a suite of microarrays containing 27016 unique expressed sequence tags (ESTs) to investigate cDNAs from a pool of 19 testes, cDNAs from a pool of nine individual ejaculate spermatozoal mRNAs, and cDNAs constructed from spermatozoal mRNAs from a single ejaculate. We also used ontological data mining to determine the function of the genes identified in each EST profile. FINDINGS: The cDNAs from the testes, pooled ejaculate, and single ejaculate hybridised to 7157, 3281, and 2780 ESTs, respectively. The testicular population contained all of the ESTs identified by the cDNAs from the pooled and individual ejaculate. The pooled ejaculate population contained all but four ESTs identified from the individual ejaculate. A subset of the spermatozoal mRNAs was associated with embryo development. INTERPRETATION: The microarray data from testes and spermatozoa (pooled and individual) were concordant, supporting the view that a spermatozoal mRNA fingerprint can be obtained from normal fertile men. Thus, profiling can be used to monitor past events-ie, gene expression of spermatogenesis. Moreover, the data suggest that, in addition to delivering the paternal genome, spermatozoa provide the zygote with a unique suite of paternal mRNAs. Ejaculate spermatozoa can now be used as a non-invasive proxy for investigations of testis-specific infertility.

Adolescent↗

Microarray expression profiling identifies early signaling transcripts associated with 6-OHDA-induced dopaminergic cell death.

The parkinsonian mimetic 6-hydroxydopamine (6-OHDA) has been shown to cause transcriptional changes associated with cellular stress and the unfolded protein response. As these cellular sequelae depend on upstream signaling events, the present study used functional genomics and proteomic approaches to aid in deciphering toxin-mediated regulatory pathways. Microarray analysis of RNA collected from multiple time points following 6-OHDA treatment was combined with data mining and clustering techniques to identify distinct functional subgroups of genes. Notably, stress-induced transcription factors such as ATF3, ATF4, CHOP, and C/EBP beta were robustly up-regulated, yet exhibited unique kinetic patterns. Genes involved in the synthesis and modification of proteins (various tRNA synthetases), protein degradation (e.g., ubiquitin, Herpud1, Sqstm1), and oxidative stress (Hmox1, Por) could be subgrouped into distinct kinetic profiles as well. Realtime PCR and/or two-dimensional electrophoresis combined with western blotting validated data derived from microarray analyses. Taken together, these data support the notion that oxidative stress and protein dysfunction play a role in Parkinson's disease, as well as provide a time course for many of the molecular events associated with 6-OHDA neurotoxicity.

Animals↗