Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

The mouse genome database (MGD): new features facilitating a model system.

The mouse genome database (MGD, http://www.informatics.jax.org/), the international community database for mouse, provides access to extensive integrated data on the genetics, genomics and biology of the laboratory mouse. The mouse is an excellent and unique animal surrogate for studying normal development and disease processes in humans. Thus, MGD's primary goals are to facilitate the use of mouse models for studying human disease and enable the development of translational research hypotheses based on comparative genotype, phenotype and functional analyses. Core MGD data content includes gene characterization and functions, phenotype and disease model descriptions, DNA and protein sequence data, polymorphisms, gene mapping data and genome coordinates, and comparative gene data focused on mammals. Data are integrated from diverse sources, ranging from major resource centers to individual investigator laboratories and the scientific literature, using a combination of automated processes and expert human curation. MGD collaborates with the bioinformatics community on the development of data and semantic standards, and it incorporates key ontologies into the MGD annotation system, including the Gene Ontology (GO), the Mammalian Phenotype Ontology, and the Anatomical Dictionary for Mouse Development and the Adult Anatomy. MGD is the authoritative source for mouse nomenclature for genes, alleles, and mouse strains, and for GO annotations to mouse genes. MGD provides a unique platform for data mining and hypothesis generation where one can express complex queries simultaneously addressing phenotypic effects, biochemical function and process, sub-cellular location, expression, sequence, polymorphism and mapping data. Both web-based querying and computational access to data are provided. Recent improvements in MGD described here include the incorporation of single nucleotide polymorphism data and search tools, the addition of PIR gene superfamily classifications, phenotype data for NIH-acquired knockout mice, images for mouse phenotypic genotypes, new functional graph displays of GO annotations, and new orthology displays including sequence information and graphic displays.

Animals↗

Dissecting tBHQ induced ARE-driven gene expression through long and short oligonucleotide arrays.

This paper compares the gene expression profiles identified by short (Affymetrix U95AV2) or long (Agilent Hu1A) oligonucleotide arrays on a model for upregulation of a cluster of antioxidant responsive element-driven genes by treatment with tert-butylhydroquinone. MAS 5.0, dCHIP, and RMA were applied to normalize the Affymetrix data, and Lowess regression was considered for Agilent data. SAM was used to identify the differential gene expression. A set of biological markers and housekeeping genes were chosen to evaluate the performance of multiple normalization approaches. Both arrays illustrated a definite set of overlapping genes between the data sets regardless of data mining tools used. However, unique gene expression profiles based on the platform used were also revealed and confirmed by quantitative RT-PCR. Further analysis of the data revealed by alternative approaches suggested that alternative splicing, multiple vs. single probe(s) measurement, and use or nonuse of mismatch probes may account for the discrepant data. Therefore, these two microarray technologies offer relatively reliable data. Integration of the gene expression profiles from different array platforms may not only help for cross-validation but also provide a more complete view of the transcriptional scenario.

Alternative Splicing↗

Investigation of metabolic objectives in cultured hepatocytes.

Using optimization based methods to predict fluxes in metabolic flux balance models has been a successful approach for some microorganisms, enabling construction of in silico models and even inference of some regulatory motifs. However, this success has not been translated to mammalian cells. The lack of knowledge about metabolic objectives in mammalian cells is a major obstacle that prevents utilization of various metabolic engineering tools and methods for tissue engineering and biomedical purposes. In this work, we investigate and identify possible metabolic objectives for hepatocytes cultured in vitro. To achieve this goal, we present a special data-mining procedure for identifying metabolic objective functions in mammalian cells. This multi-level optimization based algorithm enables identifying the major fluxes in the metabolic objective from MFA data in the absence of information about critical active constraints of the system. Further, once the objective is determined, active flux constraints can also be identified and analyzed. This information can be potentially used in a predictive manner to improve cell culture results or clinical metabolic outcomes. As a result of the application of this method, it was found that in vitro cultured hepatocytes maximize oxygen uptake, coupling of urea and TCA cycles, and synthesis of serine and urea. Selection of these fluxes as the metabolic objective enables accurate prediction of the flux distribution in the system given a limited amount of flux data; thus presenting a workable in silico model for cultured hepatocytes. It is observed that an overall homeostasis picture is also emergent in the findings.

Cell Culture Techniques↗

CLU: a new algorithm for EST clustering.

BACKGROUND: The continuous flow of EST data remains one of the richest sources for discoveries in modern biology. The first step in EST data mining is usually associated with EST clustering, the process of grouping of original fragments according to their annotation, similarity to known genomic DNA or each other. Clustered EST data, accumulated in databases such as UniGene, STACK and TIGR Gene Indices have proven to be crucial in research areas from gene discovery to regulation of gene expression. RESULTS: We have developed a new nucleotide sequence matching algorithm and its implementation for clustering EST sequences. The program is based on the original CLU match detection algorithm, which has improved performance over the widely used d2_cluster. The CLU algorithm automatically ignores low-complexity regions like poly-tracts and short tandem repeats. CONCLUSION: CLU represents a new generation of EST clustering algorithm with improved performance over current approaches. An early implementation can be applied in small and medium-size projects. The CLU program is available on an open source basis free of charge. It can be downloaded from http://compbio.pbrc.edu/pti.

Algorithms↗

Future of benchmarking: more data, more sharing, and better patient care.

Automated systems that provide whatever regulatory information is needed when it is needed; sharing of data to improve quality; data mined for specific groups of patients: Those are just a few of the trends predicted by health care experts asked to comment on the future of benchmarking and data strategies. Such improvements are needed; many hospitals continually run into problems when it comes to finding the right data sets for targeted patient groups.

Benchmarking↗

Explaining the diffusion of Medicaid home care waiver programs using VPRS decision rules.

While public preferences and legal decisions require extended Medicaid Home and Community-Based Services (HCBS), uneven program development is a major concern of policy makers and consumers. This paper presents the first Variable Precision Rough Sets (VPRS) analysis of national data to examine inter-state variation in the Medicaid HCBS waiver program. The exposition provides a detailed discussion of the methodological options and processes, tests the generated rules using a leave-one-out cross-validation, and compares VPRS classification accuracy with regression analyses of the same dataset. The results demonstrate that VPRS offers a robust method with two distinctive features for health care research. First, for policy makers and their audiences, VPRS results are presented as "if...then..." decision rules with likelihoods stated as percentages. Because this output form can be easily understood by non-specialists, the potential impact of research is enhanced. Second, for analysts generating evidence for health policy, VPRS provides a rigorous data mining tool and acknowledges inherent analytical uncertainty in the field.

Community Health Services↗

From microarrays to networks: mining expression time series.

Over the past few years, powerful new methods have been devised that enable researchers to study the expression dynamics of many genes simultaneously (e.g. gene expression profiles using cDNA microarrays). In principle, this potentially vast quantity of data enables the dissection of the complex genetic networks that control the patterns and rhythms of gene expression in the cell. Finding the patterns in those data represents the next major phase in our understanding of the programming and functioning of the living cell. Simple dynamic models can be used to generate gene expression networks. These networks reveal the phenomenological link between the expression of different genes. This review discuss how these networks are generated and outlines several data-mining techniques for extracting relationships and hypotheses in gene expression. These emerging methods can be applied to a range of biological problems.

Computational Biology↗

Web portal to an image database for high-resolution three-dimensional reconstruction.

The exponential increase of image data in high-resolution reconstructions by electron cryomicroscopy (cryoEM) has posed a need for efficient data management solutions in addition to powerful data processing procedures. Although relational databases and web portals are commonly used to manage sequences and structures in biological research, their application in cryoEM has been limited due to the complexity in accomplishing the dual tasks of interacting with proprietary software and simultaneously providing data access to users without database knowledge. Here, we report our results in developing web portal to SQL image databases used by the Image Management and Icosahedral Reconstruction System (IMIRS) to manage cryoEM images for subnanometer-resolution reconstructions. Fundamental issues related to the design and deployment of web portals to image databases are described. A web browser-based user interface was designed to accomplish data reporting and other database-related services, including user authentication, data entry, graph-based data mining, and various query and reporting tasks with interactive image manipulation capabilities. With an integrated web portal, IMIRS represents the first cryoEM application that incorporates both web-based data reporting tools and a complete set of data processing modules. Our examples should thus provide general guidelines applicable to other cryoEM technology development efforts.

Cryoelectron Microscopy↗

Comparison of different microarray data analysis programs and description of a database for microarray data management.

Data analysis and management represent a major challenge for gene expression studies using microarrays. Here, we compare different methods of analysis and demonstrate the utility of a personal microarray database. Gene expression during HIV infection of cell lines was studied using Affymetrix U-133 A and B chips. The data were analyzed using Affymetrix Microarray Suite and Data Mining Tool, Silicon Genetics GeneSpring, and dChip from Harvard School of Public Health. A small-scale database was established with FileMaker Pro Developer to manage and analyze the data. There was great variability among the programs in the lists of significantly changed genes constructed from the same data. Similarly choices of different parameters for normalization, comparison, and standardization greatly affected the outcome. As many probe sets on the U133 chip target the same Unigene clusters, the Unigene information can be used as an internal control to confirm and interpret the probe set results. Algorithms used for the determination of changes in gene expression require further refinement and standardization. The use of a personal database powered with Unigene information can enhance the analysis of gene expression data.

Database Management Systems↗

Mayday--a microarray data analysis workbench.

UNLABELLED: Mayday is a workbench for visualization, analysis and storage of microarray data. It features a graphical user interface and supports the development and integration of existing and new analysis methods. Besides the infrastructural core functionality, Mayday offers a variety of plug-ins, such as various interactive viewers, a connection to the R statistical environment, a connection to SQL-based databases and different data mining methods, including WEKA-library based methods for classification and various clustering methods. In addition, so-called meta information objects are provided for annotation of the microarray data allowing integration of data from different sources, which is a feature that, for instance, is employed in the enhanced heatmap visualization. SUPPLEMENTARY INFORMATION: The software and more detailed information including screenshots and a user guide as well as test data can be found on the Mayday home page http://www.zbit.uni-tuebingen.de/pas/mayday. The core is published under the GPL (GNU Public License) and the associated plug-ins under the LGPL (Lesser GNU Public License).

Computer Graphics↗

On the consistency of Bayesian variable selection for high dimensional binary regression and classification.

Modern data mining and bioinformatics have presented an important playground for statistical learning techniques, where the number of input variables is possibly much larger than the sample size of the training data. In supervised learning, logistic regression or probit regression can be used to model a binary output and form perceptron classification rules based on Bayesian inference. We use a prior to select a limited number of candidate variables to enter the model, applying a popular method with selection indicators. We show that this approach can induce posterior estimates of the regression functions that are consistently estimating the truth, if the true regression model is sparse in the sense that the aggregated size of the regression coefficients are bounded. The estimated regression functions therefore can also produce consistent classifiers that are asymptotically optimal for predicting future binary outputs. These provide theoretical justifications for some recent empirical successes in microarray data analysis.

Bayes Theorem↗

ToxoDB: accessing the Toxoplasma gondii genome.

ToxoDB (http://ToxoDB.org) provides a genome resource for the protozoan parasite Toxoplasma gondii. Several sequencing projects devoted to T. gondii have been completed or are in progress: an EST project (http://genome.wustl.edu/est/index.php?toxoplasma=1), a BAC clone end-sequencing project (http://www.sanger.ac.uk/Projects/T_gondii/) and an 8X random shotgun genomic sequencing project (http://www.tigr.org/tdb/e2k1/tga1/). ToxoDB was designed to provide a central point of access for all available T. gondii data, and a variety of data mining tools useful for the analysis of unfinished, un-annotated draft sequence during the early phases of the genome project. In later stages, as more and different types of data become available (microarray, proteomic, SNP, QTL, etc.) the database will provide an integrated data analysis platform facilitating user-defined queries across the different data types.

Animals↗

DNannotator: Annotation software tool kit for regional genomic sequences.

Sequence annotation is essential for genomics-based research. Investigators of a specific genomic region who have developed abundant local discoveries such as genes and genetic markers, or have collected annotations from multiple resources, can be overwhelmed by the difficulty in creating local annotation and the complexity of integrating all the annotations. Presenting such integrated data in a form suitable for data mining and high-throughput experimental design is even more daunting. DNannotator, a web application, was designed to perform batch annotation on a sizeable genomic region. It takes annotation source data, such as SNPs, genes, primers, and so on, prepared by the end-user and/or a specified target of genomic DNA, and performs de novo annotation. DNannotator can also robustly migrate existing annotations in GenBank format from one sequence to another. Annotation results are provided in GenBank format and in tab-delimited text, which can be imported and managed in a database or spreadsheet and combined with existing annotation as desired. Graphic viewers, such as Genome Browser or Artemis, can display the annotation results. Reference data (reports on the process) facilitating the user's evaluation of annotation quality are optionally provided. DNannotator can be accessed at http://sky.bsd.uchicago.edu/DNannotator.htm.

Chromosomes, Human, Pair 13↗

A practical method of predicting client revisit intention in a hospital setting.

Data mining (DM) models are an alternative to traditional statistical methods for examining whether higher customer satisfaction leads to higher revisit intention. This study used a total of 906 outpatients' satisfaction data collected from a nationwide survey interviews conducted by professional interviewers on a face-to-face basis in South Korea, 1998. Analyses showed that the relationship between overall satisfaction with hospital services and outpatients' revisit intention, along with word-of-mouth recommendation as intermediate variables, developed into a nonlinear relationship. The five strongest predictors of revisit intention were overall satisfaction, intention to recommend to others, awareness of hospital promotion, satisfaction with physician's kindness, and satisfaction with treatment level.

Adult↗

Effect of data normalization on fuzzy clustering of DNA microarray data.

BACKGROUND: Microarray technology has made it possible to simultaneously measure the expression levels of large numbers of genes in a short time. Gene expression data is information rich; however, extensive data mining is required to identify the patterns that characterize the underlying mechanisms of action. Clustering is an important tool for finding groups of genes with similar expression patterns in microarray data analysis. However, hard clustering methods, which assign each gene exactly to one cluster, are poorly suited to the analysis of microarray datasets because in such datasets the clusters of genes frequently overlap. RESULTS: In this study we applied the fuzzy partitional clustering method known as Fuzzy C-Means (FCM) to overcome the limitations of hard clustering. To identify the effect of data normalization, we used three normalization methods, the two common scale and location transformations and Lowess normalization methods, to normalize three microarray datasets and three simulated datasets. First we determined the optimal parameters for FCM clustering. We found that the optimal fuzzification parameter in the FCM analysis of a microarray dataset depended on the normalization method applied to the dataset during preprocessing. We additionally evaluated the effect of normalization of noisy datasets on the results obtained when hard clustering or FCM clustering was applied to those datasets. The effects of normalization were evaluated using both simulated datasets and microarray datasets. A comparative analysis showed that the clustering results depended on the normalization method used and the noisiness of the data. In particular, the selection of the fuzzification parameter value for the FCM method was sensitive to the normalization method used for datasets with large variations across samples. CONCLUSION: Lowess normalization is more robust for clustering of genes from general microarray data than the two common scale and location adjustment methods when samples have varying expression patterns or are noisy. In particular, the FCM method slightly outperformed the hard clustering methods when the expression patterns of genes overlapped and was advantageous in finding co-regulated genes. Thus, the FCM approach offers a convenient method for finding subsets of genes that are strongly associated to a given cluster.

Algorithms↗

Preoperative prediction of pediatric patients with effusions and edema following cardiopulmonary bypass surgery by serological and routine laboratory data.

AIM: Postoperative effusions and edema and capillary leak syndrome in children after cardiac surgery with cardiopulmonary bypass constitute considerable clinical problems. Overshooting immune response is held to be the cause. In a prospective study we investigated whether preoperative immune status differences exist in patients at risk for postsurgical effusions and edema, and to what extent these differences permit prediction of the postoperative outcome. METHODS: One-day preoperative serum levels of immunoglobulins, complement, cytokines and chemokines, soluble adhesion molecules and receptors as well as clinical chemistry parameters such as differential counts, creatinine, blood coagulation status (altogether 56 parameters) were analyzed in peripheral blood samples of 75 children (aged 3-18 years) undergoing cardiopulmonary bypass surgery (29 with postoperative effusions and edema within the first postoperative week). RESULTS: Preoperative elevation of the serum level of C3 and C5 complement components, tumor necrosis factor-alpha, percentage of leukocytes that are neutrophils, body weight and decreased percentage of lymphocytes (all P < 0.03) occurred in children developing postoperative effusions and edema. While single parameters did not predict individual outcome, >86% of the patients with postoperative effusions and oedema were correctly predicted using two different classification algorithms. Data mining by both methods selected nine partially overlapping parameters. The prediction quality was independent of the congenital heart defect. CONCLUSION: Indicators of inflammation were selected as risk indicators by explorative data analysis. This suggests that preoperative differences in the immune system and capillary permeability status exist in patients at risk for postoperative effusions. These differences are suitable for preoperative risk assessment and may be used for the benefit of the patient and to improve cost effectiveness.

Adolescent↗

Rationale and design of a large-scale trial using nicorandil as an adjunct to percutaneous coronary intervention for ST-segment elevation acute myocardial infarction: Japan-Working groups of acute myocardial infarction for the reduction of Necrotic Damage by a K-ATP channel opener (J-WIND-KATP).

BACKGROUND: The benefits of percutaneous coronary intervention (PCI) in acute myocardial infarction (AMI) are limited by reperfusion injury. In animal models, nicorandil, a hybrid of an ATP-sensitive K(+) (KATP) channel opener and nitrates, reduces infarct size, so the Japan-Working groups of acute myocardial Infarction for the reduction of Necrotic Damage by a K-ATP channel opener (J-WIND-KATP) designed a prospective, randomized, multicenter study to evaluate whether nicorandil reduces myocardial infarct size and improves regional wall motion when used as an adjunctive therapy for AMI. METHODS AND RESULTS: Twenty-six hospitals in Japan are participating in the J-WIND-KATP study. Patients with AMI who are candidates for PCI are randomly allocated to receive either intravenous nicorandil or placebo. The primary end-points are (1) estimated infarct size and (2) left ventricular function. Single nucleotide polymorphisms (SNPs) that may be associated with the function of KATP-channel and the susceptibility of AMI to the drug will be examined. Furthermore, a data mining method will be used to design the optimal combined therapy for post-myocardial infarction (MI) patients. CONCLUSIONS: It is intended that J-WIND-KATP will provide important data on the effects of nicorandil as an adjunct to PCI for AMI and that the SNPs information that will open the field of tailor-made therapy. The optimal therapeutic drug combination will also be determined for post-MI patients.

Adult↗

Corticosteroid-regulated genes in rat kidney: mining time series array data.

Kidney is a major target for adverse effects associated with corticosteroids. A microarray dataset was generated to examine changes in gene expression in rat kidney in response to methylprednisolone. Four control and 48 drug-treated animals were killed at 16 times after drug administration. Kidney RNA was used to query 52 individual Affymetrix chips, generating data for 15,967 different probe sets for each chip. Mining techniques applicable to time series data that identify drug-regulated changes in gene expression were applied. Four sequential filters eliminated probe sets that were not expressed in the tissue, not regulated by drug, or did not meet defined quality control standards. These filters eliminated 14,890 probe sets (94%) from further consideration. Application of judiciously chosen filters is an effective tool for data mining of time series datasets. The remaining data can then be further analyzed by clustering and mathematical modeling. Initial analysis of this filtered dataset identified a group of genes whose pattern of regulation was highly correlated with prototype corticosteroid enhanced genes. Twenty genes in this group, as well as selected genes exhibiting either downregulation or no regulation, were analyzed for 5' GRE half-sites conserved across species. In general, the results support the hypothesis that the existence of conserved DNA binding sites can serve as an important adjunct to purely analytic approaches to clustering genes into groups with common mechanisms of regulation. This dataset, as well as similar datasets on liver and muscle, are available online in a format amenable to further analysis by others.

Animals↗