Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Identification and functional analysis of six mycolyltransferase genes of Corynebacterium glutamicum ATCC 13032: the genes cop1, cmt1, and cmt2 can replace each other in the synthesis of trehalose dicorynomycolate, a component of the mycolic acid layer of the cell envelope.

By data mining in the sequence of the Corynebacterium glutamicum ATCC 13032 genome, six putative mycolyltransferase genes were identified that code for proteins with similarity to the N-terminal domain of the mycolic acid transferase PS1 of the related C. glutamicum strain ATCC 17965. The genes identified were designated cop1, cmt1, cmt2, cmt3, cmt4, and cmt5 ( cmt from corynebacterium mycolyl transferases). cop1 encodes a protein of 657 amino acids, which is larger than the proteins encoded by the cmt genes with 365, 341, 483, 483, and 411 amino acids. Using bioinformatics tools, it was shown that all six gene products are equipped with signal peptides and esterase domains. Proteome analyses of the cell envelope of C. glutamicum ATCC 13032 resulted in identification of the proteins Cop1, Cmt1, Cmt2, and Cmt4. All six mycolyltransferase genes were used for mutational analysis. cmt4 could not be mutated and is considered to be essential. cop1 was found to play an additional role in cell shape formation. A triple mutant carrying mutations in cop1, cmt1, and cmt2 aggregated when cultivated in MM1 liquid medium. This mutant was also no longer able to synthesize trehalose di coryno mycolate (TDCM). Since single and double mutants of the genes cop1, cmt1, and cmt2 could form TDCM, it is concluded that the three genes, cop1, cmt1, and cmt2, are involved in TDCM biosynthesis. The presence of the putative esterase domain makes it highly possible that cop1, cmt1, and cmt2 encode enzymes synthesizing TDCM from trehalose monocorynomycolate.

Arabidopsis Proteins↗

Development of an on-line automated sample clean-up method and liquid chromatography-tandem mass spectrometry analysis: application in an in vitro proteolytic assay.

Fluorescence detection has been a method of choice in industry for screening assays, including identification of enzyme inhibitors, owing to its high-throughput capabilities, excellent reproducibility, and sensitivity. Occasionally, inhibitors are identified that challenge the fluorescence assay limit, necessitating the development of more sensitive detection methods to assess these compounds. For data mining purposes, however, original assay conditions may be required. A direct method transfer to highly sensitive and specific LC-MS-based methods has not always been possible due to the presence of MS-incompatible neutral detergents and non-volatile salts in the assay matrix. Utilizing an in vitro proteolytic screening assay for the serine protease hepatitis C virus (HCV) nonstructural (NS) 3 protease as a test case, we report the development of an automated sample clean-up procedure implemented on-line with liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis to complement fluorescence detection. Ion exchange and peptide microtraps were employed to remove MS-incompatible assay matrix components. Three protease inhibitors were used to validate the MS/MS method. Comparable potencies were achieved for these compounds when assessed by fluorescence and MS/MS detection. Furthermore, four-fold less enzyme could be utilized when employing the MS/MS method compared to fluorescence detection. The longer analysis time, however, resulted in reduced sample capacity. The potency of our designed HCV NS3 protease inhibitors are thus routinely evaluated using a continuous fluorescence-based assay. Only pertinent inhibitors approaching the fluorescence assay sensitivity limit are subsequently analyzed further by LC-MS/MS. This methodology allows us to maintain a database and to compare results independent of the detection method. Despite the relatively slow sample turnaround time of this LC-MS approach, the versatility of the automated on-line clean-up procedure and sample analysis can be applied to assays containing reagents which were historically considered to be MS incompatible.

Chromatography, Liquid↗

The Tübingen approach: identification, selection, and validation of tumor-associated HLA peptides for cancer therapy.

There is substantial need for molecularly defined tumor antigens to prime cytotoxic T cells in vivo for cancer immunotherapy, especially in the case of tumor entities for which only a few tumor antigens have been defined so far. In this review, we present the "Tübingen approach" to identify, select, and validate large numbers of MHC/HLA class I-associated peptides derived from tumor-associated antigens. Step 1 is the identification of naturally presented HLA-associated peptides directly from primary tumor cells. Step 2 is selection of tumor-associated peptides from step 1 by differential gene expression analysis and data mining. Step 3 is validation of selected candidates by monitoring in vivo T-cell responses in the context of patient-individualized immunizations. Our approach combines methods from genomics, proteomics, bioinformatics, and T-cell immunology. The aim is to develop effective immunotherapeutics consisting of multiple tumor-associated epitopes in order to induce a broad and specific immune response against cancer cells.

Antigens, Neoplasm↗

A two-step strategy for detecting differential gene expression in cDNA microarray data.

A mixed-model approach is proposed for identifying differential gene expression in cDNA microarray experiments. This approach is implemented by two interconnected steps. In the first step, we choose a subset of genes that are potentially expressed differentially among treatments with a loose criterion. In the second step, these potential genes are used for further analyses and data-mining with a stringent criterion, in which differentially expressed genes (DEGs) are confirmed and some quantities of interest (such as gene x treatment interaction) are estimated. By simulating datasets with DEGs, we compare our statistical method with a widely used method, the t-statistic, for single genes. Simulation results show that our approach produces a high power and a low false discovery rate for DEG identification. We also investigate the impacts of various source variations resulting from microarray experiments on the efficiency of DEG identification. Analysis of a published experiment studying unstable transcripts in Arabidopsis illustrates the utility of our method. Our method identifies more novel and biologically interesting unstable transcripts than those reported in the original literature.

Arabidopsis↗

A comprehensive mouse IBD database for the efficient localization of quantitative trait loci.

Traditional fine-mapping approaches in mouse genetics that go from a linkage region to a candidate gene are very costly and time consuming. Shared ancestry regions, along with the combination of genetics and genomics approaches, provide a powerful tool to shorten the time and effort required to identify a causative gene. In this article we present a novel methodology that predicts IBD (identical by descent) regions between pairs of inbred strains using single nucleotide polymorphism (SNP) maps. We have validated this approach by comparing the IBD regions, estimated using different algorithms, to the results derived using the sequence information in the strains present in the Celera Mouse Database. We showed that based on the current publicly available SNP genotypes, large IBD regions (>1 Mb) can be identified successfully. By assembling a list of 21,514 SNPs in 61 common inbred strains, we inferred IBD regions between all pairs of strains and confirmed, for the first time, that existing quantitative trait genes (QTG) and susceptibility genes all lie outside of IBD regions. We also illustrated how knowledge of IBD structures can be applied to strain selection for future crosses. We have made our results available for data mining and download through a public website ( http://www.mouseibd.florida.scripps.edu ).

Algorithms↗

Identification of candidate maternal-effect genes through comparison of multiple microarray data sets.

Transcriptional profiling by microarray hybridization has become a standard method to analyze global gene expression and has resulted in the availability of enormous amounts of experimental data. Given the number of different microarray platforms currently in use, it is critical to determine how reproducible results are from one platform to another. Additional variability may also arise from tissue collection and protocol differences among laboratories. In an effort to identify genes whose maternal mRNA pools are critical during preimplantation development, we compared published results of three independent studies of the mouse preimplantation embryo transcriptome, each performed in a different laboratory using different microarray platforms. We searched the combined data set for genes whose expression patterns were consistent among the three experiments. Querying for presence or absence at single developmental windows indicates that between 52% and 60% of genes are in agreement among the three experiments. Searching for expression patterns across three developmental windows (oocyte + 1-cell, 2- through 8-cell, and blastocyst stage) revealed approximately 33% agreement among the three experiments, although the majority of these genes were either always present or always absent. Using this approach, we identified 51 genes with a predicted expression pattern of maternal RNA only (not present during 2-cell through 8-cell or at the blastocyst stage). RT-PCR validation indicates 37 (72%) of these candidates have the microarray-predicted expression pattern and represent candidate maternal-effect genes. Based on our analysis, we conclude that data mining microarray experiments in this way greatly enhances candidate gene expression pattern accuracy.

Animals↗

Gene expression profiling of plant responses to abiotic stress.

Expression profiling has become an important tool to investigate how an organism responds to environmental changes. Plants, being sessile, have the ability to dramatically alter their gene expression patterns in response to environmental changes such as temperature, water availability or the presence of deleterious levels of ions. Sometimes these transcriptional changes are successful adaptations leading to tolerance while in other instances the plant ultimately fails to adapt to the new environment and is labeled as sensitive to that condition. Expression profiling can define both tolerant and sensitive responses. These profiles of plant response to environmental extremes (abiotic stresses) are expected to lead to regulators that will be useful in biotechnological approaches to improve stress tolerance as well as to new tools for studying regulatory genetic circuitry. Finally, data mining of the alterations in the plant transcriptome will lead to further insights into how abiotic stress affects plant physiology.

Adaptation, Physiological↗

MED12-STAT1-TAP2 axis regulates CD8 + T cell cytotoxicity and mediates immunotherapy outcome in non-small cell lung cancer.

Although immunotherapy for late-stage non-small cell lung carcinoma (NSCLC) has been clinically utilized, its prognosis remains highly heterogeneous, prompting us to investigate novel predictive immunotherapy biomarkers for NSCLC. We analyzed the correlations between MED12 nonsynonymous mutations and survival, clinical, genomic, transcriptomic information, and immune infiltration information through data mining across multiple datasets. We also investigated the mechanism of MED12 using luciferase assay, Western blot, ChIP-PCR, and siRNA. MED12 is significantly associated with survival in completely independent immunotherapy datasets, including MSKCC (N = 350), Naiyer2015 (N = 34), our own (N = 295) and the pan-cancer dataset, but not in the TCGA dataset, where patients received non-immunotherapy regimens. Mutations in MED12 showed no significant correlation with known metrics (TMB, IPS/CTLA4/PD1 status, PD-1/PD-L1 expression, and TCR/BCR status) or DNA Damage Repair (DDR) pathway mutations, yet they carried independent prognostic information according to the Cox multivariate regression. On the other hand, MED12 mutation is significantly associated with multiple immune-related pathways and immune infiltration of CD8 + T cells and activated NK cells. Lactate dehydrogenase assay revealed that knockdown of TAP2 restored the upregulation of CD8 + T cell cytotoxicity triggered by MED12 knockdown. ChIP-PCR, luciferase assay and siRNA knock down assay indicate that MED12 binds to the promoter region of STAT1 to suppress its transcription, while the transcription factor STAT1 promotes the transcription of TAP2, thus inhibiting the antigen processing and presentation. Collectively, MED12 mutation is an independent and valuable biomarker for predicting the response to immune checkpoint inhibitor (ICI)therapy in NSCLC by modulating CD8 + T cell cytotoxicity via the STAT1/TAP2 axis.

Humans↗

OpenRIMS: an open architecture radiology informatics management system.

The benefits of an integrated picture archiving and communication system/radiology information system (PACS/RIS) archive built with open source tools and methods are 2-fold. Open source permits an inexpensive development model where interfaces can be updated as needed, and the code is peer reviewed by many eyes (analogous to the scientific model). Integration of PACS/RIS functionality reduces the risk of inconsistent data by reducing interfaces among databases that contain largely redundant information. Also, wide adoption would promote standard data mining tools--reducing user needs to learn multiple methods to perform the same task. A model has been constructed capable of accepting HL7 orders, performing examination and resource scheduling, providing digital imaging and communications in medicine (DICOM) worklist information to modalities, archiving studies, and supporting DICOM query/retrieve from third party viewing software. The multitiered architecture uses a single database communicating via an ODBC bridge to a Linux server with HL7, DICOM, and HTTP connections. Human interaction is supported via a web browser, whereas automated informatics services communicate over the HL7 and DICOM links. The system is still under development, but the primary database schema is complete as well as key pieces of the web user interface. Additional work is needed on the DICOM/HL7 interface broker and completion of the base DICOM service classes.

Databases, Factual↗

OpenRIMS: an open architecture radiology informatics Management system.

The following are benefits of an integrated picture archiving and communication system/radiology information system archive built with open-source tools and methods: open source, inexpensive interfaces can be updated as needed, and reduced risk of redundant and inconsistent data. Also, wide adoption would promote standard data mining tools, reducing user needs to learn multiple methods to perform the same task. A model has been constructed capable of accepting orders, performing exam resource scheduling, providing Digital communications in Medicine (DICOM) work list information to modalities, archiving studies, and supporting DICOM query/retrieve from third-party viewing software. The multitiered architecture uses a single database communicating via an open database connectivity bridge to a Linux server with Health Level 7 (HL7), DICOM, and HTTP connections. Human interaction is supported via a browser, whereas other informatics systems communicate over the HL7 and DICOM links. The system is still under development, but the primary database schema is complete, as are key pieces of the Web user interface. Additional work is needed on the DICOM/HL7 interface broker and completion of the base DICOM service classes.

Databases as Topic↗

Six characteristics of effective structured reporting and the inevitable integration with speech recognition.

The reporting of radiological images is undergoing dramatic changes due to the introduction of two new technologies: structured reporting and speech recognition. Each technology has its own unique advantages. The highly organized content of structured reporting facilitates data mining and billing, whereas speech recognition offers a natural succession from the traditional dictation-transcription process. This article clarifies the distinction between the process and outcome of structured reporting, describes fundamental requirements for any effective structured reporting system, and describes the potential development of a novel, easy-to-use, customizable structured reporting system that incorporates speech recognition. This system should have all the advantages derived from structured reporting, accommodate a wide variety of user needs, and incorporate speech recognition as a natural component and extension of the overall reporting process.

Humans↗

Creating an IHE ATNA-based audit repository.

Compliance with the Health Insurance Portability and Accountability Act (HIPAA) requires gathering audit information from picture archiving and communications systems (PACS) regarding evidence trails of human interactions. Until recently, most PACS users have had limited access to auditing information. Access required resources to handle manual inspection of audit logs, and access to proprietary databases was not always available. Some vendors now produce eXtensible Markup Language (XML) audit logs based on certain events occurring in PACS. However, it is up to the user to convert this information into an easily mined data repository supporting compliance and quality control. This process can be handled in multiple ways, which could mean different audit mechanisms depending on the PACS (or other hospital system) used. It is apparent that an organized method of dealing with audit information is needed. This help may be provided within the Integrating the Healthcare Environment (IHE) framework. The IHE initiative defines a set of profiles, actors, and transactions that create common scenarios for particular workflow processes. The Integration Profiles depict security as a fundamental requirement of the framework. Specifically, the Audit Trail and Node Authentication (ATNA) profile defines standards based mechanisms for securely transmitting and storing audit records in a central repository. The data structure defined by the profile provides a number of record types that capture different audit events. A general feasibility study for storing currently available PACS audit information following the profile is defined, and steps to an automated solution are discussed.

Feasibility Studies↗

Differential effects of omega-3 and omega-6 Fatty acids on gene expression in breast cancer cells.

Essential fatty acids have long been identified as possible oncogenic factors. Existing reports suggest omega-6 (omega-6) essential fatty acids (EFA) as pro-oncogenic and omega-3 (omega-3) EFA as anti-oncogenic factors. The omega-3 fatty acids, eicosapentaenoic acid (EPA) and docosahexaenoic acid (DHA), inhibit the growth of human breast cancer cells while the omega-6 fatty acids induces growth of these cells in animal models and cell lines. In order to explore likely mechanisms for the modulation of breast cancer cell growth by omega-3 and omega-6 fatty acids, we examined the effects of arachidonic acid (AA), linoleic acid (LA), EPA and DHA on human breast cancer cell lines using cDNA microarrays and quantitative polymerase chain reaction. MDA-MB-231, MDA-MB-435s, MCF-7 and HCC2218 cell lines were treated with the selected fatty acids for 6 and 24 h. Microarray analysis of gene expression profiles in the breast cancer cells treated with both classes of fatty acids discerned essential differences among the two classes at the earlier time point. The differential effects of omega-3 and omega-6 fatty acids on the breast cancer cells were lessened at the late time point. Data mining and statistical analyses identified genes that were differentially expressed between breast cancer cells treated with omega-3 and omega-6 fatty acids. Ontological investigations have associated those genes to a broad spectrum of biological functions, including cellular nutrition, cell division, cell proliferation, metastasis and transcription factors etc., and thus presented an important pool of biomarkers for the differential effect of omega-3 and omega-6EFAs.

Breast Neoplasms↗

Potential impact of advanced clinical information technology on cancer care in 2015.

New clinical information technologies now sporadically available will soon be in routine clinical use, bringing many changes to all phases of the cancer care continuum. For example, new technologies such as: (1) The next generation Internet; (2) Real-time clinical decision support systems; (3) Off-line, population-based systems; (4) Large, integrated, individual patient-level phenotypic and genotypic databases with intelligent data mining capabilities; (5) Wireless, invasive and non-invasive physiologic monitoring devices; (6) Natural Language Processing (NLP) systems; and (7) Mathematical models of complex biological systems all have the potential to impact significantly the provision of cancer care throughout its continuum. While new information management and communication techniques and technologies will reduce many of the inefficiencies and inaccuracies of our present systems, there will be an equal, and potentially far more dangerous, set of unintended consequences. Informatics investigators, cancer specialists, and health system administrators must focus on the study of what is working and what is not, as well as, on development and testing of the new clinical information management and communication technologies, if we are to be ready for the future.

Cancer Care Facilities↗

Understanding metastatic SCCHN cells from unique genotypes to phenotypes with the aid of an animal model and DNA microarray analysis.

Metastasis of squamous cell carcinoma of the head and neck (SCCHN) is a significant health-care problem worldwide. The 5-year survival rate is less than 50% for patients with lymph node metastases. Understanding the molecular basis of SCCHN metastasis would facilitate the development of new therapeutic approaches to the disease. To identify proteins that mediate SCCHN metastasis, we established a SCCHN xenograft mouse model and performed in vivo selection from a SCCHN cell line using the model. In the fourth round of in vivo selection, significant incidences of metastases in lymph nodes (7/10) and lungs (6/10) were achieved from a derived SCCHN cell line as compared with its parental cells, 1/5 in lymph nodes and 0/5 in lungs. Metastatic cell lines from lymph node metastases and parental cell lines from non-metastatic xenograft tumors were subjected to DNA microarray analysis using an Affymetrix gene chip HG-U133A, followed by data mining studies. The identified metastasis-related genes were further evaluated for their encoding protein products and the metastatic cells were examined by biological analyses. DNA microarray analysis highlighted molecular features of the metastatic SCCHN cells, including alteration of expression of cell-cell adhesion proteins, epithelial cell markers, apoptosis and cell cycle regulatory molecules. Further biological analyses of phenotypic alterations revealed that the metastatic cells gained epithelial-mesenchymal transition (EMT) features and were more resistant to anoikis, which are two of the important phenotypes for metastatic SCCHN.

Animals↗

Reverse-engineering gene-regulatory networks using evolutionary algorithms and grid computing.

OBJECTIVE: Living organisms regulate the expression of genes using complex interactions of transcription factors, messenger RNA and active protein products. Due to their complexity, gene-regulatory networks are not fully understood.However, by building computational models it is possible to gain insight into their function and operation. METHODS: Evolutionary algorithms are used to create computational models of gene-regulatory networks based on observed microarray data. These algorithms can be computationally intensive. They will be implemented within an existing grid computing infrastructure, that has been developed for data mining purposes, and which is able to deliver the required compute power. RESULTS: We discuss how models can built achieved using distributed and grid computing technology. In particular we investigate how Condor and JavaSpaces technology is suited to the requirements of our modeling approach. CONCLUSIONS: Determining network models of gene-regulatory networks using evolutionary algorithms not only requires considerable computational power, but also a modeling formalism that can explain the underlying dynamics.

Algorithms↗

Healthcare informatics research: from data to evidence-based management.

Healthcare informatics research is a scientific endeavor that applies information science, computer technology, and statistical modeling techniques to develop decision support systems for improving both health service organizations' performance and patient care outcomes. The analytical strategies include (1) the formulation of a data warehouse for exploration, (2) data mining, (3) the application of confirmatory statistical analysis, (4) simulation via an interface with computer and information system technologies, and (5) translational research. Healthcare informatics research will help to direct evidence-based strategic management.

Decision Support Systems, Clinical↗

A global view of CK2 function and regulation.

The wealth of biochemical, molecular, genetic, genomic, and bioinformatic resources available in S. cerevisiae make it an excellent system to explore the global role of CK2 in a model organism. Traditional biochemical and genetic studies have revealed that CK2 is required for cell viability, cell cycle progression, cell polarity, ion homeostasis, and other functions, and have identified a number of potential physiological substrates of the enzyme. Data mining of available bioinformatic resources indicates that (1) there are likely to be hundreds of CK2 targets in this organism, (2) the majority of predicted CK2 substrates are involved in various aspects of global gene expression, (3) CK2 is present in several nuclear protein complexes predicted to have a role in chromatin structure and remodeling, transcription, or RNA metabolism, and (4) CK2 is localized predominantly in the nucleus. These bioinformatic results suggest that the observed phenotypic consequences of CK2 depletion may lie downstream of primary defects in chromatin organization and/or global gene expression. Further progress in defining the physiological role of CK2 will almost certainly require a better understanding of the mechanism of regulation of the enzyme. Beginning with the crystal structure of the human CK2 holoenzyme, we present a molecular model of filamentous CK2 that is consistent with earlier proposals that filamentous CK2 represents an inactive form of the enzyme. The potential role of filamentous CK2 in regulation in vivo is discussed.

Casein Kinase II↗