Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Monitoring genotoxic exposure in uranium mines.

Recent data from deep uranium mines in Czechoslovakia indicated that mines are exposed to other mutagenic factors in addition to radon daughter products. Mycotoxins were identified as a possible source of mutagens in these mines. Mycotoxins were examined in 38 samples from mines and in throat swabs taken from 116 miners and 78 controls. The following mycotoxins were identified from mines samples: aflatoxins B1 and G1, citrinin, citreoviridin, mycophenolic acid, and sterigmatocystin. Some mold strains isolated from mines and throat swabs were investigated for mutagenic activity by the SOS chromotest and Salmonella assay with strains TA100 and TA98. Mutagenicity was observed, especially with metabolic activation in vitro. These data suggest that mycotoxins produced by molds in uranium mines are a new genotoxic factor for uranium miners.

Adult↗

An integrated proteome database for two-dimensional electrophoresis data analysis and laboratory information management system.

We describe an integrated proteome database, termed Yonsei Proteome Research Center Proteome Database (YPRC-PDB) which can store, retrieve and analyze various information including two-dimensional electrophoresis (2-DE) images and associated spot information that were obtained during studies of hepatocellular carcinoma (HCC). YPRC-PDB is also designed to perform as a laboratory information management system that manages sample information, clinical background, conditions of both sample preparation and 2-DE, and entire sets of experimental results. It also features query system and data-mining applications, which are amenable to automatically analyze expression level changes of a specific protein and directly link to clinical information. The user interface is web-based, so that the results from other laboratories can be shared effectively. In particular, the master gel image query is equipped with a graphic tool that can easily identify the relationship between the specific pathological stage of HCC and expression levels of a potential marker protein on the master gel image. Thus, YPRC-PDB is a versatile integrated database suitable for subsequent analyses. The information in YPRC-PDB is updated easily and it is available to authorized users on the World Wide Web (http://yprcpdb.proteomix.org/ approximately damduck/).

Carcinoma, Hepatocellular↗

Separation of human erythrocyte membrane associated proteins with one-dimensional and two-dimensional gel electrophoresis followed by identification with matrix-assisted laser desorption/ionization-time of flight mass spectrometry.

A classical proteomic analysis was used to establish a reference map of proteins associated with healthy human erythrocyte ghosts. Following osmotic lysis and differential centrifugation, ghost proteins were separated by either one-dimensional gel electrophoresis (1-DE) or two-dimensional gel electrophoresis (2-DE). Selected protein bands or spots were excised and trypsinized before mass spectrometric analyses and data mining was performed using the SWISS-PROT and NCBI nonredundant databases. A total of 102 protein spots from a 2-D gel were successfully identified. These corresponded to 59 distinct polypeptides with the remaining 43 being isoforms. As for the 1-D gel, 44 polypeptides were identified, of which 19 were also found on the 2-D gel. Most of the 19 common polypeptides were membrane cytoskeletal proteins that are often referred to as the "band" proteins. The remaining 25 polypeptides that were found exclusively on 1-D gels were proteins with high hydrophobicity (e.g., sorbitol dehydrogenase and glucose transporter) and high molecular mass (e.g., Kell blood group glycoprotein and Janus-kinase 2). A higher number of signaling proteins was also identified on 1-D gels compared to 2-D gels. These included Ras, cAMP dependent protein kinase and TGF-beta receptor type 1 precursor.

Centrifugation↗

DDX21 Enhances Radiosensitivity in Head and Neck Squamous Cell Carcinoma by Suppressing MK2-Mediated DNA Damage Response.

Radioresistance remains a significant challenge in the radiotherapy (RT) of head and neck squamous cell carcinoma (HNSCC). However, the biological factors that govern sensitivity to this therapy are not well-understood. The DEAD-box family is known for its role in genome stability, and inextricably linked to the radiotherapy resistance of tumors. This study found the role of the RNA helicase DDX21 in regulating radiosensitivity through extensive data mining. High DDX21 expression predicted improved survival after postoperative radiotherapy. Overexpression of DDX21 increased radiosensitivity in vitro and in vivo, whereas depletion promoted radioresistance. In vitro, DDX21 enhanced radiation-induced DNA damage, genomic instability, and apoptosis by binding MK2 and suppressing MK2 phosphorylation independently of p38 activity. Meanwhile MK2 inhibition restored and further augmented radiosensitivity in DDX21-deficient cells and xenografts by increasing DNA damage and apoptosis. Overall, DDX21 regulates radiosensitivity in HNSCC by suppressing MK2 signaling and modulating the radiation-induced DNA damage response. Its expression may serve as a potential biomarker associated with radiosensitivity, and MK2 inhibition offers a promising approach to overcome radioresistance in tumors with low DDX21 expression.

DDX21↗

Linkage and association studies of single-nucleotide polymorphism-tagged tumor necrosis factor haplotypes in juvenile oligoarthritis.

OBJECTIVE: The presence of increased levels of tumor necrosis factor (TNF) in serum and synovial fluid of patients and the encouraging outcome of anti-TNF therapy have implicated TNFalpha in the etiopathogenesis of juvenile oligoarthritis. Although the locus is polymorphic, no study has investigated all TNF single-nucleotide polymorphisms (SNPs) with respect to disease. The aim of this study was to examine the association of multiple TNF SNPs with juvenile oligoarthritis and to construct and analyze SNP-tagged TNF haplotypes. METHODS: A total of 144 simplex families consisting of parent and affected child, as well as 88 healthy, unrelated control subjects were available for study. In these individuals, 9 polymorphic positions of TNF were typed by a high-throughput genotyping method based on the SNaPshot assay. The chi-square and extended transmission disequilibrium tests were used to test for association and linkage, respectively. Odds ratios (ORs) with 95% confidence intervals (95% CIs) were also calculated. Haplotype-tagging SNPs (htSNPs) for the locus were identified by ordering the haplotypes according to their frequencies. RESULTS: The study detected association of several TNF SNPs and established linkage of the locus to juvenile oligoarthritis. The most significant association observed was between the intronic +851 TNF SNP and the persistent oligoarthritis subgroup (OR 3.86, 95% CI 1.6-9.2). Haplotype data mining showed that only 4 of the 9 SNPs need to be typed in order to capture the most frequent TNF haplotypes. CONCLUSION: The TNF locus is linked and associated with juvenile oligoarthritis. Information on the htSNPs can be useful in genetic studies of diseases in which TNF may be of relevance.

Adult↗

A structure-information approach to the prediction of biological activities and properties.

The structure-information approach to quantitative biological modeling and prediction is presented in contrast to the mechanism-based approach. Basic structure information is developed from the chemical graph (connection table). The development, beginning with information explicit in the connection table (element identity and skeletal connections), leads to significant structure information useful for establishing sound models of a wide range of properties of interest in drug design. Skeletal branching patterns and valence state definition lead to relationships for valence-state electronegativity and atom or group molar volumes. Based on these important aspects of molecules, both the electrotopological state (E-State) and molecular-connectivity structure descriptors (chi indices) are developed. A summary of QSAR models indicates the wide range of applicability of these structure descriptors and the predictive quality of QSAR models for protein binding, HIV-1 protease inhibition, blood-brain-barrier partitioning, fish toxicity, carcinogenicity risk, structure space for similarity searching, and data mining. These models are independent of three-dimensional structure information and are directly interpretable in terms of structure information useful to the drug-design process.

Animals↗

ERp29, a general endoplasmic reticulum marker, is highly expressed throughout the brain.

ERp29 is a recently discovered resident of the endoplasmic reticulum (ER) that is abundant in brain and most other mammalian tissues. Investigations of nonneural secretory tissues have implicated ERp29 in a major role producing export proteins, but a molecular activity remains wanting for this functional orphan. Intriguingly, ERp29 appears to be heavily utilized in the cerebellum, a brain region not conventionally regarded as neurosecretory. To elucidate this functional quandary, we used immunochemical approaches to characterize the regional, cellular, and subcellular distributions of ERp29 in rat brain. Immunohistochemistry revealed ubiquitous expression in neuronal and nonneuronal cells, with a distinctive variation in somatic ERp29 levels. Highly expressing cells were found in diverse locations, implying that ERp29 is not biased towards the cerebellum functionally. Using immunolocalization data mined from the literature, a proteomic profile was developed to assess the functional significance of ERp29's characteristic expression pattern. Surprisingly, ERp29 correlated poorly with classical markers of neurosecretion, but strongly with a variety of major membrane proteins. Together with immunogold localization of ERp29 to somatic ER, these observations led to a novel hypothesis that ERp29 is involved primarily in production of endomembrane proteins rather than proteins destined for export. This study establishes ERp29 as a general ER marker for brain cells and provides a stimulating clue about ERp29's enigmatic function. ERp29 appears to have broad significance for neural pathophysiology, given its ubiquitous distribution and prominence in brain over classical ER residents like BiP and protein disulfide isomerase.

Animals↗

New multivariate test for linkage, with application to pleiotropy: fuzzy Haseman-Elston.

We propose a new method of linkage analysis based on using the grade of membership scores resulting from fuzzy clustering procedures to define new dependent variables for the various Haseman-Elston approaches. For a single continuous trait with low heritability, the aim was to identify subgroups such that the grade of membership scores to these subgroups would provide more information for linkage than the original trait. For a multivariate trait, the goal was to provide a means of data reduction and data mining. Simulation studies using continuous traits with relatively low heritability (H=0.1, 0.2, and 0.3) showed that the new approach does not enhance power for a single trait. However, for a multivariate continuous trait (with three components), it is more powerful than the principal component method and more powerful than the joint linkage test proposed by Mangin et al. ([1998] Biometrics 54:88-99) when there is pleiotropy.

Cluster Analysis↗

Prediction of survival in surgical unresectable lung cancer by artificial neural networks including genetic polymorphisms and clinical parameters.

Lung cancer, a common malignancy in Taiwan, involves multiple factors, including genetics and environmental factors. The survival time is very short once cancer is diagnosed as being in advanced stage and surgically unresectable. Therefore, a good model of prediction of disease outcome is important for a treatment plan. We investigated the survival time in advanced lung cancer by using computer science from the genetic polymorphism of the p21 and p53 genes in conjunction with patients' general data. We studied 75 advanced and surgical unresectable lung cancer patients. The prediction of survival time was made by comparing real data obtained from follow-up periods with data generated by an artificial neural network (ANN). The most important input variable was the clinical staging of lung cancer patients. The second and third most important variables were pathological type and responsiveness to treatment, respectively. There were 25 neurons in the input layer, four neurons in the hidden layer-1, and one neuron in the output layer. The predicted accuracy was 86.2%. The average survival time was 12.44 +/- 7.95 months according to real data and 13.16 +/- 1.77 months based on the ANN results. ANN provides good prediction results when clinical parameters and genetic polymorphisms are considered in the model. It is possible to use computer science to integrate the genetic polymorphisms and clinical parameters in the prediction of disease outcome. Data mining provides a promising approach to the study of genetic markers for advanced lung cancer.

Adult↗

RepairNET: a bioinformatics toolbox for functional exploration of DNA damage response.

DNA damage response is one of the essential cellular mechanisms to maintaining the genomic integrity of the cell. Aberrations in the mechanism of DNA damage response often result in cancer. We describe here RepairNET, a protein-protein interaction network associated with the DNA damage response. RepairNET is assembled from the published literature by using a protocol that involved computational data mining of the MEDLINE and manual curation. This network represents the current knowledge on the intrinsic signaling pathways related to the DNA damage response process. RepairNET currently contains more than 1,200 proteins with over 2,300 functional interactions. A number of web-interface tools have been implemented to facilitate a user-friendly environment. The users can navigate through the cellular network associated with the DNA damage response via a Java-based interactive graphical interface. In order to help users explore the functional relationships between the interacting proteins, we have assigned functional domains to the proteins in RepairNET based on their sequences. A total of 365 unique functional domains are assigned. RepairNET is available online at http://guanyin.chem.temple.edu/RepairNET.html. It could become an essential resource center for cancer research, providing clues to understanding the functional relationship between proteins in the network, and to building scientific models for the mechanism of DNA damage response and cancer proliferation.

Adaptor Proteins, Signal Transducing↗

Deciphering the human nucleolar proteome.

Nucleoli are plurifunctional nuclear domains involved in the regulation of several major cellular processes such as ribosome biogenesis, the biogenesis of non-ribosomal ribonucleoprotein complexes, cell cycle, and cellular aging. Until recently, the protein content of nucleoli was poorly described. Several proteomic analyses have been undertaken to discover the molecular bases of the biological roles fulfilled by nucleoli. These studies have led to the identification of more than 700 proteins. Extensive bibliographic and bioinformatic analyses allowed the classification of the identified proteins into functional groups and suggested potential functions of 150 human proteins previously uncharacterized. The combination of improvements in mass spectrometry technologies, the characterization of protein complexes, and data mining will assist in furthering our understanding of the role of nucleoli in different physiological and pathological cell states.

Cell Nucleolus↗

SPLASH: systematic proteomics laboratory analysis and storage hub.

In the field of proteomics, the increasing difficulty to unify the data format, due to the different platforms/instrumentation and laboratory documentation systems, greatly hinders experimental data verification, exchange, and comparison. Therefore, it is essential to establish standard formats for every necessary aspect of proteomics data. One of the recently published data models is the proteomics experiment data repository [Taylor, C. F., Paton, N. W., Garwood, K. L., Kirby, P. D. et al., Nat. Biotechnol. 2003, 21, 247-254]. Compliant with this format, we developed the systematic proteomics laboratory analysis and storage hub (SPLASH) database system as an informatics infrastructure to support proteomics studies. It consists of three modules and provides proteomics researchers a common platform to store, manage, search, analyze, and exchange their data. (i) Data maintenance includes experimental data entry and update, uploading of experimental results in batch mode, and data exchange in the original PEDRo format. (ii) The data search module provides several means to search the database, to view either the protein information or the differential expression display by clicking on a gel image. (iii) The data mining module contains tools that perform biochemical pathway, statistics-associated gene ontology, and other comparative analyses for all the sample sets to interpret its biological meaning. These features make SPLASH a practical and powerful tool for the proteomics community.

Database Management Systems↗

Differential gene expression in the rat caudate putamen after "binge" cocaine administration: advantage of triplicate microarray analysis.

Rat genome U34A (Affymetrix) oligonucleotide microarrays were used to analyze changes in gene expression in the caudate putamen (CPu) of Fischer rats induced by 1 and 3 days of "binge" cocaine (or saline) administration. A triplicate array assay of pooled RNA of each treatment group was used to evaluate the technical variability and sensitivity of microarrays. Cocaine-regulated genes were identified using the Affymetrix MAS 5.0 and Data Mining Tool v. 3. Eighty-nine upregulated and eight downregulated genes/ESTs were found after 1 day of "binge" cocaine. Following 3 days of cocaine treatment we identified 21 upregulated and 17 downregulated genes/ESTs. RNase protection assays of selected genes confirmed reliability of changes identified by the microarrays at the level of > or =1.40-fold increase. Many genes upregulated in the CPu by cocaine were immediate early genes for transcription factors and for "effector" proteins (e.g., vesl/Homer1a, Arc, synaptotagmin IV). Acute "binge" cocaine also increased mRNA levels for glutamate receptor GluR2, dopamine receptor D1, and a number of phosphatases. Genes downregulated by cocaine include several genes associated with energy metabolism in mitochondria, as well as the phosphatydylinositol-4 kinase and the regulator of G-protein signaling protein 4 (RGS4). A differential expression of somatostatin receptor SSTR2, not known to be a cocaine-responsive gene, as well as the clock gene Per2, were found by microarrays and confirmed by RNase protection assay. These results demonstrate the potential of microarrays in profiling gene expression with > or =40% increase or > or =14% decrease in mRNA levels for discovery of novel cocaine-responsive genes.

Animals↗

Gene expression profile of human bone marrow stromal cells: high-throughput expressed sequence tag sequencing analysis.

Human bone marrow stromal cells (HBMSC) are pluripotent cells with the potential to differentiate into osteoblasts, chondrocytes, myelosupportive stroma, and marrow adipocytes. We used high-throughput DNA sequencing analysis to generate 4258 single-pass sequencing reactions (known as expressed sequence tags, or ESTs) obtained from the 5' (97) and 3' (4161) ends of human cDNA clones from a HBMSC cDNA library. Our goal was to obtain tag sequences from the maximum number of possible genes and to deposit them in the publicly accessible database for ESTs (dbEST of the National Center for Biotechnology Information). Comparisons of our EST sequencing data with nonredundant human mRNA and protein databases showed that the ESTs represent 1860 gene clusters. The EST sequencing data analysis showed 60 novel genes found only in this cDNA library after BLAST analysis against 3.0 million ESTs in NCBI's dbEST database. The BLAST search also showed the identified ESTs that have close homology to known genes, which suggests that these may be newly recognized members of known gene families. The gene expression profile of this cell type is revealed by analyzing both the frequency with which a message is encountered and the functional categorization of expressed sequences. Comparing an EST sequence with the human genomic sequence database enables assignment of an EST to a specific chromosomal region (a process called digital gene localization) and often enables immediate partial determination of intron/exon boundaries within the genomic structure. It is expected that high-throughput EST sequencing and data mining analysis will greatly promote our understanding of gene expression in these cells and of growth and development of the skeleton.

Bone Marrow Cells↗

Biomass quantification by image analysis.

Microbiologists have always rely on microscopy to examine microorganisms. When microscopy, either optical or electron-based, is coupled to quantitative image analysis, the spectrum of potential applications is widened: counting, sizing, shape characterization, physiology assessment, analysis of visual texture, motility studies are now easily available for obtaining information on biomass. In this chapter the main tools used for cell visualization as well as the basic steps of image treatment are presented. General shape descriptors can be used to characterize the cell morphology, but special descriptors have been defined for filamentous microorganisms. Physiology assessment is often based on the use of fluorescent dyes. The quantitative analysis of visual texture is still limited in bioengineering but the characterization of the surface of microbial colonies may open new prospects, especially for cultures on solid substrates. In many occasions, the number of parameters extracted from images is so large that data-mining tools, such as Principal Components Analysis, are useful for summarizing the key pieces of information.

Algorithms↗

Genome function--a virus-world view.

By studying viruses one may begin to understand how static genomes can define dynamic processes of development. This talk will describe some of the approaches we are taking, using computer simulations and laboratory experiments, to account for the many molecular-level processes and interactions that occur when a common bacterium, E. coli, is infected by one of its viruses, phage T7. We accounted for processes of phage genome entry, transcription, translation, and DNA replication, including protein-DNA and protein-protein regulatory interactions, and we predicted the dynamics of phage progeny formation. The simulations have enabled us to identify limiting host-cell resources in phage growth, discover novel anti-viral strategies, and suggest frameworks for mining data from global mRNA and protein studies.

Bacteriophage T7↗

Assessment of cytochrome P450 sequences offers a useful tool for determining genetic diversity in higher plant species.

To investigate and develop new genetic tools for assessing genome-wide diversity in higher plant-species, polymorphisms of gene analogues of mammalian cytochrome P450 mono-oxygenases were studied. Data mining on Arabidopsis thaliana indicated that a small number of primer-sets derived from P450 genes could provide universal tools for the assessment of genome-wide genetic diversity in diverse plant species that do not have relevant genetic markers, or for which, there is no prior inheritance knowledge of inheritance traits. Results from PCR amplification of 51 plant species from 28 taxonomic families using P450 gene-primer sets suggested that there were at least several mammalian P450 gene mammalian-analogues in plants. Intra- and inter- specific variations were demonstrated following PCR amplifications of P450 analogue fragments, and this suggested that these would be effective genetic markers for the assessment of genetic diversity in plants. In addition, BLAST search analysis revealed that these amplified fragments possessed homologies to other genes and proteins in different plant varieties. We conclude that the sequence diversity of P450 gene-analogues in different plant species reflects the diversity of functional regions in the plant genome and is therefore an effective tool in functional genomic studies of plants.

Arabidopsis↗