Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Transcriptional co-regulation of secondary metabolism enzymes in Arabidopsis: functional and evolutionary implications.

The combined knowledge of the Arabidopsis genome and transcriptome now allows to get an integrated view of the dynamics and evolution of metabolic pathways in plants. We used publicly available sets of microarray data obtained in a wide range of different stress and developmental conditions to investigate the co-expression of genes encoding enzymes of secondary metabolism pathways, in particular indoles, phenylpropanoids, and flavonoids. We performed hierarchical clustering of gene expression profiles and found that major enzymes of each pathway display a clear and robust co-expression throughout all the conditions studied. Moreover, detailed analysis evidenced that some genes display co-regulation in particular physiological conditions only, certainly reflecting their modular recruitment into stress- or developmentally regulated biosynthetic pathways. The combination of these microarray data with sequence analysis allows to draw very precise hypotheses on the function of otherwise uncharacterized genes. To illustrate this approach, we focused our analysis on secondary metabolism glycosyltransferases (UGTs), a multigenic family involved in the conjugation of small molecules to sugars like glucose. We propose that UGT74B1 and UGT74C1 may be involved in aromatic and aliphatic glucosinolates synthesis, respectively. We also suggest that UGT75C1 may function as an anthocyanin-5-O-glucosyltransferase in planta. Therefore, this data-mining approach appears very powerful for the functional prediction of unknown genes, and could be transposed to virtually any other gene family. Finally, we suggest that analysis of expression pattern divergence of duplicated genes also provides some insight into the mechanisms of metabolic pathway evolution.

Arabidopsis↗

Response diversity of Arabidopsis thaliana ecotypes in elevated [CO2] in the field.

Free Air [CO(2)] Enrichment (FACE) allows for plant growth under fully open-air conditions of elevated [CO(2)] at concentrations expected to be reached by mid-century. We used Arabidopsis thaliana ecotypes Col-0, Cvi-0, and WS to analyze changes in gene expression and metabolite profiles of plants grown in "SoyFACE" (http://www.soyface.uiuc.edu/), a system of open-air rings within which [CO(2)] is elevated to approximately 550 ppm. Data from multiple rings, comparing plants in ambient air and elevated [CO(2)], were analyzed by mixed model ANOVA, linear discriminant analysis (LDA) and data-mining tools. In elevated [CO(2)], decreases in the expression of genes related to chloroplast functions characterized all lines but individual members of distinct multi-gene families were regulated differently between lines. Also, different strategies distinguished the lines with respect to the regulation of genes related to carbohydrate biosynthesis and partitioning, N-allocation and amino acid metabolism, cell wall biosynthesis, and hormone responses, irrespective of the plants' developmental status. Metabolite results paralleled reactions seen at the level of transcript expression. Evolutionary adaptation of species to their habitat and intrinsic genetic plasticity seem to determine the nature of responses to elevated [CO(2)]. Irrespective of their underlying genetic diversity, and evolutionary adaptation to different habitats, a small number of common, predominantly stress-responsive, signature transcripts appear to characterize responses of the Arabidopsis ecotypes in FACE.

Analysis of Variance↗

DNA microarrays in prostate cancer.

DNA microarray technology provides a means to examine large numbers of molecular changes related to a biological process in a high throughput manner. This review discusses plausible utilities of this technology in prostate cancer research, including definition of prostate cancer predisposition, global profiling of gene expression patterns associated with cancer initiation and progression, identification of new diagnostic and prognostic markers, and discovery of novel patient classification schemes. The technology, at present, has only been explored in a limited fashion in prostate cancer research. Some hurdles to be overcome are the high cost of the technology, insufficient sample size and repeated experiments, and the inadequate use of bioinformatics. With the completion of the Human Genome Project and the advance of several highly complementary technologies, such as laser capture microdissection, unbiased RNA amplification, customized functional arrays (eg, single-nucleotide polymorphism chips), and amenable bioinformatics software, this technology will become widely used by investigators in the field. The large amount of novel, unbiased hypotheses and insights generated by this technology is expected to have a significant impact on the diagnosis, treatment, and prevention of prostate cancer. Finally, this review emphasizes existing, but currently underutilized, data-mining tools, such as multivariate statistical analyses, neural networking, and machine learning techniques, to stimulate wider usage.

Algorithms↗

[How will urology cope with the new health care system?].

According to the report 2000 issued by the expert advisory committee on the health care system, improvement in performance requires, among other things, awareness of the need for more economic efficiency, better quality of medical care, and definition of targets. Disease prevention and early diagnosis are the main objectives of health policy and practice. Data mining becomes an important tool in epidemiological and clinical research. Urologists are called upon to engage in clinical studies to establish therapeutic standards and guidelines. Results of research must be put in to practice faster. Doctors ought to respond to changes in patients' attitudes.

Cost Control↗

Assessment of freeway traffic parameters leading to lane-change related collisions.

This study aims at 'predicting' the occurrence of lane-change related freeway crashes using the traffic surveillance data collected from a pair of dual loop detectors. The approach adopted here involves developing classification models using the historical crash data and corresponding information on real-time traffic parameters obtained from loop detectors. The historical crash and loop detector data to calibrate the neural network models (corresponding to crash and non-crash cases to set up a binary classification problem) were collected from the Interstate-4 corridor in Orlando (FL) metropolitan area. Through a careful examination of crash data, it was concluded that all sideswipe collisions and the angle crashes that occur on the inner lanes (left most and center lanes) of the freeway may be attributed to lane-changing maneuvers. These crashes are referred to as lane-change related crashes in this study. The factors explored as independent variables include the parameters formulated to capture the overall measure of lane-changing and between-lane variations of speed, volume and occupancy at the station located upstream of crash locations. Classification tree based variable selection procedure showed that average speeds upstream and downstream of crash location, difference in occupancy on adjacent lanes and standard deviation of volume and speed downstream of the crash location were found to be significantly associated with the binary variable (crash versus non-crash). The classification models based on data mining approach achieved satisfactory classification accuracy over the validation dataset. The results indicate that these models may be applied for identifying real-time traffic conditions prone to lane-change related crashes.

Accidents, Traffic↗

Development of radiology prediction models using feature analysis.

RATIONALE AND OBJECTIVES: This article provides an introduction to prediction models and their application in diagnostic imaging research. Prediction models capitalize on the different degrees of association among variables to make a prediction of a health state, formulate a rule, or quantify individual contributions of various predictor variables. The purpose of this article is to elucidate the rationale, implication, and interpretation of prediction models using imaging features. MATERIALS AND METHODS: The techniques and challenges of developing, testing, and implementing prediction models are described. Prediction model development methods are similar to data-mining techniques. RESULTS: Learning objectives are to review prediction rule (model) methods, learn how prediction models may be applied to feature analysis, and understand the challenges of developing, testing, and implementing prediction models.

Decision Support Techniques↗

Infection control and quality health care in the new millennium.

Health care-associated infection remains a major issue of patient safety. It complicates a significant proportion of patient care deliveries, adds to the burden of resource use, and contributes to unexpected deaths. Early infection control pioneers showed that surveillance and prevention programs can be successful and have set the scene for today's infection control activities. Parameters for success include those to recognize and explain health care-associated infections and implement interventions to decrease infection rates and limit antimicrobial resistance spread. Current major challenges facing infection control programs are reviewed with an emphasis on recent trends in health care delivery systems, together with some vision on future activities and interactions toward such changes. Benchmarking of infection rates is considered inevitable, and, thus, surveillance strategies, adapted to changing health care systems, should improve and emphasize intervention and standardization. Major challenges for the future include antimicrobial use and control of resistances, new materials, emerging pathogens, infection control issues related to transgenic therapy, massive and complete immunosuppression and xenotransplantation, prion diseases, use of fully computerized patient record and data-mining-derived epidemiology, development of evidence-based recommendations for infection control and prevention, addressing cost constraints and newly apparent health care system trends, and health care worker behavior modification.

Cost-Benefit Analysis↗

A landmark extraction method for protein 2DE gel images based on multi-dimensional clustering.

OBJECTIVE: Two-dimensional electrophoresis (2DE) is a separation technique that can identify target proteins existing in a tissue. Its result is represented by a gel image that displays an individual protein in a tissue as a spot. However, because the technique suffers from low reproducibility, a user should manually annotate landmark spots on each gel image to analyze the spots of different images together. This operation is an error-prone and tedious job. For this reason, this paper proposes a method of extracting landmark spots automatically by using a data mining technique. METHOD AND MATERIAL: A landmark profile which summarizes the characteristics of landmark spots in a set of training gel images of the same tissue is generated by extracting the common properties of the landmark spots. On the basis of the landmark profile, candidate landmark spots in a new gel image of the same tissue are identified, and final landmark spots are determined by the well-known A* search algorithm. RESULT AND CONCLUSIONS: The performance of the proposed method is analyzed through a series of experiments in order to identify its various characteristics.

Algorithms↗

Temporal representation and reasoning in medicine: Research directions and challenges.

OBJECTIVE: The main aim of this paper is to propose and discuss promising directions of research in the field of temporal representation and reasoning in medicine, taking into account the recent scientific literature and challenging issues of current interest as viewed from the different research perspectives of the authors of the paper. BACKGROUND: Temporal representation and reasoning in medicine is a well-known field of research in the medical as well as computer science community. It encompasses several topics, such as summarizing data from temporal clinical databases, reasoning on temporal clinical data for therapeutic assessments, and modeling uncertainty in clinical knowledge and data. It is also related to several medical tasks, such as monitoring intensive care patients, providing treatments for chronic patients, as well as planning and scheduling clinical routine activities within complex healthcare organizations. METHODOLOGY: The authors jointly identified significant research areas based on their importance as for temporal representation and reasoning issues; the subjects were considered to be promising topics of future activity. Every subject was addressed in detail by one or two authors and then discussed with the entire team to achieve a consensus about future fields of research. RESULTS: We identified and focused on four research areas, namely (i) fuzzy logic, time, and medicine, (ii) temporal reasoning and data mining, (iii) health information systems, business processes, and time, and (iv) temporal clinical databases. For every area, we first highlighted a few basic notions that would permit any reader--including those who are unfamiliar with the topic--to understand the main goals. We then discuss interesting and promising directions of research, taking into account the recent literature and underlining the yet unresolved medical/clinical issues that deserve further scientific investigation. The considered research areas are by no means disjointed, because they share common theoretical and methodological features. Moreover, subjects of imminent interest in medicine are represented in many of the fields considered. CONCLUSIONS: We propose and discuss promising subjects of future research that deserve investigation to develop software systems that will properly manage the multifaceted temporal aspects of information and knowledge encountered by physicians during their clinical work. As the subjects of research have resulted from merging the different perspectives of the authors involved in this study, we hope the paper will succeed in stimulating discussion and multidisciplinary work in the described fields of research.

Artificial Intelligence↗

Regulatory studies of murine methylenetetrahydrofolate reductase reveal two major promoters and NF-kappaB sensitivity.

Two promoters of the murine methylenetetrahydrofolate reductase gene (Mthfr), a key enzyme in folate metabolism, were characterized in Neuro-2a, NIH/3T3 and RAW 264.7 cells. Sequences of 189 bp and 273 bp were sufficient to achieve maximal activity of the upstream and downstream promoter, respectively. However, subtle differences in minimal promoter lengths and in promoter activities were observed between the cell lines. Both promoters demonstrated comparable activity in NIH/3T3 and RAW 264.7 cells, while in Neuro-2a cells, the upstream promoter was 15-fold more active than the downstream promoter. Alignment and data mining tools identified a candidate nuclear factor kappa B (NF-kappaB) binding site at the 3'end of the downstream promoter that is conserved throughout several species. NF-kappaB activation experiments in cultured cells were associated with increased Mthfr mRNA. Co-transfection of NF-kappaB and promoter constructs demonstrated Mthfr up-regulation by at least 2-fold through its downstream promoter in Neuro-2a cells; this increase was significantly reduced when the putative binding site was mutated. EMSA analysis demonstrated direct binding of NF-kappaB to this non-mutated site. This study, a first step into the elucidation of Mthfr regulation, demonstrates that two TATA-less, GC-rich promoters differentially drive transcription of Mthfr in a cell-specific manner, and provides a novel link of Mthfr to possible roles in the immune response and cell survival.

3T3 Cells↗

Expression and genomic profiling of colorectal cancer.

Colorectal cancer still represents a paradigm for the elucidation of the cellular, genetic and molecular mechanisms that underly solid tumor initiation, progression to malignancy, and metastasis to distal organ sites. The relative ease with which pathological specimens can be obtained by either surgery or endoscopy from different stages of tumor progression has facilitated the application of omics technologies to allow the genome-wide analysis both at the RNA (gene expression) and DNA (aneuploidy) levels. Here, we have reviewed the multiplicity of studies appeared to date in the scientific literature on the expression and genomic analysis of colorectal cancer, and attempted an integration of the profiling data generated and made available in the public domain. This approach is likely to pinpoint specific chromosomal loci and the corresponding genes which (i) play rate-limiting roles in colorectal cancer, (ii) represent putative diagnostic and prognostic markers for the accurate prediction of clinical outcome and response to treatment, and (iii) encompass potential therapeutic targets. Moreover, cross-species data mining and integration of the human colorectal cancer profiles with those obtained from mouse models of intestinal tumorigenesis will even more contribute to the elucidation of highly conserved pathways and cellular functions underlying malignancy in the GI tract. Notwithstanding the above promises, tumor heterogeneity, limited cohort sizes, and methodological differences among experimental and bioinformatic approaches still poses main obstacles towards the optimal utilization and integration of omics profiles.

Adenoma↗

The HBSP gene is expressed during HBV replication, and its coded BH3-containing spliced viral protein induces apoptosis in HepG2 cells.

The mechanisms of liver injury in hepatitis B virus (HBV) infection are defined to be due not to the direct cytopathic effects of viruses, but to the host immune response to viral proteins expressed by infected hepatocytes. We showed here that transfection of mammalian cells with a replicative HBV genome causes extensive cytopathic effects, leading to the death of infected cells. While either necrosis or apoptosis or both may contribute to the death of infected cells, results from flow cytometry suggest that apoptosis plays a major role in HBV-induced cell death. Data mining of the four HBV protein sequences reveals the presence of a Bcl-2 homology domain 3 (BH3) in HBSP, a spliced viral protein previously shown to be able to induce apoptosis and associated with HBV pathogenesis. HBSP is expressed at early stage of our cell-based HBV replication. When transfected into HepG2 cells, HBSP causes apoptosis in a caspase dependent manner. Taken together, our results suggested a direct involvement of HBV viral proteins in cellular apoptosis, which may contribute to liver pathogenesis.

Apoptosis↗

Manifestation, mechanisms and mysteries of gene amplifications.

Gene amplifications are essential features of advanced cancers and have prognostic as well as therapeutic significance in clinical cancer treatment. Models explaining the amplification process, such as breakage-fusion-bridge cycle and excision and unequal segregation of extrachromosomal DNA fragments, predict that independent DNA double-stranded breaks must occur to induce amplification formation. Many cellular, tissue and environmental factors induce DNA damage and amplifications. Also labile DNA sequence features like fragile sites facilitate amplifications. Although, databases and data mining tools of various genomic attributes are already available, extra-large scale systems biology endeavors to decipher dynamics, interactions and dependencies between different factors contributing to amplification process fail, because current databases of DNA copy number aberrations and fragile sites comprise conventional cytogenetics results obtained at far too coarse chromosome band resolution. Array comparative genomic hybridization (aCGH) enables genome-wide gene copy number measurements and amplification detection at molecular genetic resolution. Similarly, cloning and sequencing of fragile sites produce mapping information of vastly improved resolution. In conclusion, databases of aCGH and sequenced fragile sites are needed to resolve the mechanisms of gene amplifications in systems biology configuration.

Chromosome Aberrations↗

Genomic approaches to drug discovery.

Considerable progress has been made in exploiting the enormous amount of genomic and genetic information for the identification of potential targets for drug discovery and development. New tools that incorporate pathway information have been developed for gene expression data mining to reflect differences in pathways in normal and disease states. In addition, forward and reverse genetics used in a high-throughput mode with full-length cDNA and RNAi libraries enable the direct identification of components of signaling pathways. The discovery of the regulatory function of microRNAs highlights the importance of continuing the investigation of the genome with sophisticated tools. Furthermore, epigenetic information including DNA methylation and histone modifications that mediate important biological processes add to the possibilities to identify novel drug targets and patient populations that will benefit from new therapies.

Animals↗

Differential expression and molecular characterisation of Lmo7, Myo1e, Sash1, and Mcoln2 genes in Btk-defective B-cells.

PURPOSE: Bruton's tyrosine kinase is crucial for B-lymphocyte development. By the use of gene expression profiling, we have identified four expressed sequence tags among 38 potential Btk target genes, which have now been characterised. METHODS: Bioinformatics tools including data mining of additional unpublished gene expression profiles, sequence verification of PCR products and qualitative RT-PCR were used. Stimulations targeting the B-cell receptor and the protein kinase C were used to activate whole B-cell splenocytes. RESULTS: Target genes were characterised as Lim domain only 7 (Lmo7); Myosin1e (Myo1e); SAM and SH3 domain containing 1 (Sash1); and Mucolipin2 (Mcoln2). Expression was found in cell lines of different origin and developmental stages as well as in whole B-cell splenocytes and Transitional type 1 (T1) splenic B-cells from wild type and Btk-defective mice, respectively. By the use of semi-quantitative RT-PCR we found Sash1 not to be expressed in the investigated haematopoietic cell lines, while transcripts were found in whole splenic B-cells from both wild type and Btk-defective mice, whereas Lmo7, Myo1e, and Mcoln2 were expressed in both B-cell lines and primary B-lymphocytes. Except for Lmo7, the transcript level was similarly affected by stimulation in control and Btk-defective cells.

Agammaglobulinaemia Tyrosine Kinase↗

Data quality aspects of a database for abdominal septic shock patients.

Since many years, medical researchers have investigated the mechanisms that may cause a septic shock. Despite many approaches that analyzed smaller parts of the relevant data or single variables, respectively, no larger database with all the possible relevant data existed. Our work was to bridge this gap. We built a large database for abdominal septic shock patients. While building it, we were confronted with many problems concerning the database realization and the data quality. Thus, we will demonstrate how we built our database and how we assured data quality. This is of interest for all medical or computer scientists who are concerned with building medical databases with retrospective data, e.g. for data mining purposes.

Abdomen↗

Identification of mouse mslp2 gene from EST databases by repeated searching, comparison, and assembling.

The NPHS2 gene is expressed in podocytes and encodes the integral membrane protein called podocin, which is believed to play an important role in the renal function of glomerular filtration. Mutations in this gene can cause serious renal function disorders. In this study, we used data-mining techniques and bioinformatic tools to search for the mouse ortholog of the NPHS2-related gene. It might be valuable for future studies of renal diseases. We employed repeated cycles of searching, comparison, and assembling to extend the assembled EST sequences. The discovered gene sequence mslp2, an ortholog of the human SLP2 gene, was found to have a total length of 1253 bp with the amino acid coding region located in 32-1093 nt. It was further verified using the RT-PCR and RACE techniques to ensure its biological accuracy and then registered with the GenBank. When ClustalW was used for comparing the mslp2 and human SLP2 genes for similarities, the similarities were as high as 88% for nucleotide and 92% for amino acid sequences. In conclusion, we propose a method for rapid identification of the mouse ortholog gene from the human genome.

Animals↗