Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Oxidative stress is involved in the development of experimental abdominal aortic aneurysm: a study of the transcription profile with complementary DNA microarray.

BACKGROUND: The role of oxidative stress in the formation of aneurysms is not fully understood. We used the complementary DNA (cDNA) microarray technique to determine the transcription profile in the development of elastase-induced abdominal aortic aneurysm in rat models, with an emphasis on the oxidative stress-related genes. MATERIALS AND METHODS: In the experimental group, rat abdominal aortas were perfused with elastase to induce AAA. In the control group, a sham operation was performed with perfusion of the aortas with saline solution. Four or five animals were used for each time point for each of the elastase-treated or saline-treated groups. At day 2, day 7, and day 10 after surgery, the external aortic diameter was measured and AAA formation was estimated. Total RNA was isolated from aortas and subjected to cDNA microarray analysis with the use of the rat genome U34A high-density oligonucleotide DNA chip (Affymetrix, Santa Clara, Calif), which contains a total number of 8799 genes of which 2017 are expressed sequence tag (EST) genes. The data were analyzed with the GENECHIP Data Mining Tool software (Affymetrix). For genes of interest, reverse-transcription polymerase chain reaction was performed to confirm their expression level. RESULTS: Comparison ranking analysis revealed that during AAA development, the expression of 212 genes, including 46 of EST genes, increased by more than two-fold and 229 genes, including 95 of EST genes, decreased by more than two-fold in at least one of the three time points. The regulated genes included those encoding heme oxygenase, inducible nitric oxide synthase, some extracellular matrix proteins, members of the matrix metalloproteinase family, and those associated with prooxidant/antioxidant and inflammatory responses. Reverse-transcription polymerase chain reaction analysis confirmed the upregulation of genes involved in oxidative stress, such as heme oxygenase, inducible nitric oxide synthase, 12-lipoxygenase, and heart cytochrome c oxydase subunit VIa, and the downregulation of antioxidant genes, such as superoxide dismutase, reduced nicotinamide adenine dinucleotide-cytochrome b-5 reductase, and glutathion S-transferase. CONCLUSION: The cDNA microarray technique was useful for investigation of the transcription profiles during the development of AAA. Our results indicate that oxidative stress may play a pivotal role in the pathologic progression of AAA.

Animals↗

Adult mouse brain gene expression patterns bear an embryologic imprint.

The current model to explain the organization of the mammalian nervous system is based on studies of anatomy, embryology, and evolution. To further investigate the molecular organization of the adult mammalian brain, we have built a gene expression-based brain map. We measured gene expression patterns for 24 neural tissues covering the mouse central nervous system and found, surprisingly, that the adult brain bears a transcriptional "imprint" consistent with both embryological origins and classic evolutionary relationships. Embryonic cellular position along the anterior-posterior axis of the neural tube was shown to be closely associated with, and possibly a determinant of, the gene expression patterns in adult structures. We also observed a significant number of embryonic patterning and homeobox genes with region-specific expression in the adult nervous system. The relationships between global expression patterns for different anatomical regions and the nature of the observed region-specific genes suggest that the adult brain retains a degree of overall gene expression established during embryogenesis that is important for regional specificity and the functional relationships between regions in the adult. The complete collection of extensively annotated gene expression data along with data mining and visualization tools have been made available on a publicly accessible web site (www.barlow-lockhart-brainmapnimhgrant.org).

Algorithms↗

Expression profiling of a human cell line model of prostatic cancer reveals a direct involvement of interferon signaling in prostate tumor progression.

Cancer-associated fibroblasts induce malignant behavior in genetically initiated but nontumorigenic human prostatic epithelium. The genetic basis for such transformation is still unknown. By using Affymetrix GeneChip technology, we profiled genomewide gene expression of transformed [tumorigenic benign prostatic hyperplasia (BPH1)(CAFTD)] and parental (nontumorigenic BPH1) cells. We identified differentially expressed genes, which are associated with tumorigenesis or tumor progression. One striking finding is that a significant portion of the down-regulated genes belongs to interferon (IFN)-inducible molecules. We show that IFN inhibited the tumorigenic BPH1(CAFTD) cell proliferation and colony formation in vitro and inhibited tumor growth in xenografts in vivo. Expression of the IFN-inducible molecules correlates with the growth-inhibiting effects of IFN. In addition, these genes are reported to be mapped mainly to two chromosomal regions, 10q23-26 and 17q21, which are frequently deleted in human prostate cancers. Furthermore, in silico data-mining with the GeneLogic database revealed that expression of the IFN-inducible genes was down-regulated in approximately 30% of the 49 clinically characterized samples of prostatic adenocarcinomas. Collectively, we show that there seems to be a direct link between IFN-inducible molecules and prostatic tumor progression. These findings suggest IFN-inducible molecules as potential therapeutic targets for the treatment of prostate cancer.

Cell Division↗

Functionally diverging molecular quasi-species evolve by crossing two enzymes.

Molecular evolution is frequently portrayed by structural relationships, but delineation of separate functional species is more elusive. We have generated enzyme variants by stochastic recombinations of DNA encoding two homologous detoxication enzymes, human glutathione transferases M1-1 and M2-2, and explored their catalytic versatilities. Sampled mutants were screened for activities with eight alternative substrates, and the activity fingerprints were subjected to principal component analysis. This phenotype characterization clearly identified at least three distributions of substrate selectivity, where one was orthogonal to those of the parent-like distributions. This approach to evolutionary data mining serves to identify emerging molecular quasi-species and indicates potential trajectories available for further protein evolution.

Evolution, Molecular↗

Changes in global gene expression patterns during development and maturation of the rat kidney.

We set out to define patterns of gene expression during kidney organogenesis by using high-density DNA array technology. Expression analysis of 8,740 rat genes revealed five discrete patterns or groups of gene expression during nephrogenesis. Group 1 consisted of genes with very high expression in the early embryonic kidney, many with roles in protein translation and DNA replication. Group 2 consisted of genes that peaked in midembryogenesis and contained many transcripts specifying proteins of the extracellular matrix. Many additional transcripts allied with groups 1 and 2 had known or proposed roles in kidney development and included LIM1, POD1, GFRA1, WT1, BCL2, Homeobox protein A11, timeless, pleiotrophin, HGF, HNF3, BMP4, TGF-alpha, TGF-beta2, IGF-II, met, FGF7, BMP4, and ganglioside-GD3. Group 3 consisted of transcripts that peaked in the neonatal period and contained a number of retrotransposon RNAs. Group 4 contained genes that steadily increased in relative expression levels throughout development, including many genes involved in energy metabolism and transport. Group 5 consisted of genes with relatively low levels of expression throughout embryogenesis but with markedly higher levels in the adult kidney; this group included a heterogeneous mix of transporters, detoxification enzymes, and oxidative stress genes. The data suggest that the embryonic kidney is committed to cellular proliferation and morphogenesis early on, followed sequentially by extracellular matrix deposition and acquisition of markers of terminal differentiation. The neonatal burst of retrotransposon mRNA was unexpected and may play a role in a stress response associated with birth. Custom analytical tools were developed including "The Equalizer" and "eBlot," which contain improved methods for data normalization, significance testing, and data mining.

Animals↗

Studying the protein organization of the postsynaptic density by a novel solid phase- and chemical cross-linking-based technology.

Agarose beads carrying a cleavable, fluorescent, and photoreactive cross-linking reagent on the surface were synthesized and used to selectively pull out the proteins lining the surface of supramolecules. A quantitative comparison of the abundances of various proteins in the sample pulled out by the beads from supramolecules with their original abundances could provide information on the spatial arrangement of these proteins in the supramolecule. The usefulness of these synthetic beads was successfully verified by trials using a synthetic protein complex consisting of three layers of different proteins on glass coverslips. By using these beads, we determined the interior or superficial locations of five major and 19 minor constituent proteins in the postsynaptic density (PSD), a large protein complex and the landmark structure of asymmetric synapses in the mammalian central nervous system. The results indicate that alpha,beta-tubulins, dynein heavy chain, microtubule-associated protein 2, spectrin, neurofilament H and M subunits, an hsp70 protein, alpha-internexin, dynamin, and PSD-95 protein reside in the interior of the PSD. Dynein intermediate chain, alpha-amino-3-hydroxy-5-methyl-4-isoxazole propionate receptors, kainate receptors, N-cadherin, beta-catenin, N-ethylmaleimide-sensitive factor, an hsc70 protein, and actin reside on the surface of the PSD. The results further suggest that the N-methyl-d-aspartate receptors and the alpha-subunits of calcium/calmodulin-dependent protein kinase II are likely to reside on the surface of the PSD although with unique local protein organizations. Based on our results and the known interactions between various PSD proteins from data mining, a model for the molecular organization of the PSD is proposed.

Animals↗

Drug exposure and psoriasis vulgaris: case-control and case-crossover studies.

Intake of drugs is considered a risk factor for psoriasis. The aim of this study was to investigate the association between drugs and psoriasis. A case-control study including 110 patients who were hospitalized for extensive psoriasis was performed. A control group (n = 515) was defined as patients who had undergone elective surgery. A case-crossover study included 98 patients with psoriasis. Exposure to drugs was assessed during a hazard period (3 months before hospitalization) and compared to a control period in the patient's past. Data on drug sales were extracted by data mining techniques. Multivariate analyses were performed by logistic regression and conditional logistic regression. In the case-control study, psoriasis was associated with benzodiazepines (OR 6.9), organic nitrates (OR 5.0), angiotensin-converting enzyme (ACE) inhibitors (OR 4.0) and non-steroidal anti-inflammatory drugs (NSAIDs) (OR 3.7). In the case-crossover study, psoriasis was associated with ACE inhibitors (OR 9.9), beta-blockers (OR 9.9), dipyrone (OR 4.9) and NSAIDs (OR 2.1). Extensive psoriasis may be associated with intake of ACE inhibitors, NSAIDs or beta-blockers.

Adrenergic beta-Antagonists↗

Pattern recognition for road traffic accident severity in Korea.

An increasing number of road traffic accidents (RTA) in Korea has emerged as being harmful both for the economy and for safety. An accurately estimated classification model for several severity types of RTA as a function of related factors provides crucial information for the prevention of potential accidents. Here, three data-mining techniques (neural network, logistic regression, decision tree) are used to select a set of influential factors and to build up classification models for accident severity. The three approaches are then compared in terms of classification accuracy. The finding is that accuracy does not differ significantly for each model and that the protective device is the most important factor in the accident severity variation.

Accidents, Traffic↗

Digital microscopy imaging and new approaches in toxicologic pathology.

Digital microscopy, a comprehensive integration of digital imaging and light microscopy, can assist the pathologist to observe, acquire, record, share, analyze, and manage pathology image data. To lead the activity for establishing new generation digital microscopy capacity, novel concepts and strategies of digital pathology information flow and digital pathology platform were designed to integrate personal digital pathology microscopy workstations and other pathology imaging modalities with centralized data storage/management. In addition, a strategy for Web-enabled interactive telepathology that would permit global capacity was designed. A novel concept of high content pathology was also created to develop an automated tissue microscopy imaging and screening approach. These new concepts, strategies, and approaches guided the development and implementation of a digital pathology platform, a telepathology platform, and automated tissue slide imaging capacity. Digital microscopy photography is now able to replace photographic film in toxicologic pathology. Digital pathology and telepathology platforms can provide a networked environment for multisite, global team participation. Our practice also ascertained the central value of digital microscopy which can provide innovative quantitative pathology information and data mining capability with various imaging biomarkers via advanced digital image processing and pathology informatics; these are now the focus of ongoing development.

Image Enhancement↗

Food allergy--towards predictive testing for novel foods.

The risks associated with IgE-mediated food allergy highlight the need for methods to screen for potential food allergens. Clinical and immunological tests are available for the diagnosis of food allergy to known food allergens, but this does not extend to the evaluation, or prediction of allergenicity in novel foods. This category, includes foods produced using novel processes genetically modified (GM) foods, and foods that might be used as alternatives to traditional foods. Through the collation and analysis of the protein sequences of known allergens and their epitopes, it is possible to identify related groups which correlate with observed clinical cross-reactivities. 3-D modelling extends the use of sequence data and can be used to display eptiopes on the surface of a molecule. Experimental models support sequence analysis and 3-D modelling. Observed cross-reactivities can be examined by Western blots prepared from native 2-D gels of a whole food preparation (e.g. hazelnut, peanut), and common proteins identified. IgEs to novel proteins can be raised in Brown Norway rat (a high IgE responder strain) and the proteins tested in simulated digest to determine epitope stability. Using the CSL serum bank, epitope binding can be examined through the ability of an allergen to cross-link the high affinity IgE receptor and thereby release mediators using in vitro cell-based models. This range of methods, in combination with data mining, provides a variety of screening options for testing the potential of a novel food to be allergenic, which does not involve prior exposure to the consumer.

Allergens↗

Predicting employment outcomes of rehabilitation clients with orthopedic disabilities: a CHAID analysis.

PURPOSE: To examine demographic and service factors affecting employment outcomes of people with orthopedic disabilities in public vocational rehabilitation programs in the United States. METHOD: The sample included 74,861 persons (55% men and 45% women) with disabilities involving the limbs or spinal column who were closed either as rehabilitated or not rehabilitated by their state-run vocational rehabilitation agencies in the fiscal year 2001. Mean age of participants was 41.4 years (SD = 11.2). The dependent variable is employment outcomes. The predictor variables include a set of personal history variables and rehabilitation service variables. RESULTS: The chi-squared automatic interaction detector (CHAID) analysis indicated that job placement services significantly enhanced competitive employment outcomes but were significantly underutilized (only 25% of the clients received this service). Physical restoration and assistive technology services along with support services such as counseling also contributed to positive employment outcomes. Importantly, clients who received general assistance, supplementary security income, and/or social security disability insurance benefits had a significant lower competitive employment rates (45%) than clients without such work disincentives (60%). CONCLUSION: The data mining approach (i.e., CHAID analysis) provided detailed information and insight about interactions among demographic variables, service patterns, and competitive employment rates through the segmentation of the sample into mutually exclusive homogeneous subgroups.

Adult↗

Neural networks predict protein folding and structure: artificial intelligence faces biomolecular complexity.

In the genomic era DNA sequencing is increasing our knowledge of the molecular structure of genetic codes from bacteria to man at a hyperbolic rate. Billions of nucleotides and millions of aminoacids are already filling the electronic files of the data bases presently available, which contain a tremendous amount of information on the most biologically relevant macromolecules, such as DNA, RNA and proteins. The most urgent problem originates from the need to single out the relevant information amidst a wealth of general features. Intelligent tools are therefore needed to optimise the search. Data mining for sequence analysis in biotechnology has been substantially aided by the development of new powerful methods borrowed from the machine learning approach. In this paper we discuss the application of artificial feedforward neural networks to deal with some fundamental problems tied with the folding process and the structure-function relationship in proteins.

Databases, Factual↗

Genomics, morphogenesis and biophysics: triangulation of Purkinje cell development.

The cerebellar Purkinje cells (P-cells) comprise an organelle that is suitable for combined analysis by morphology and genomics, using biophysical tools. In some unknown way, genomic information specifies the development of P-cells. One of us (AJP) has previously proposed that fractal processes associated with DNA are in a causal relation to the fractal properties of organelles such as P-cells (FractoGene, 2002, patent pending). This fractal postulate predicts that the dendritic arborization of P-cells will be less complex in lower order vertebrates. The prediction can be tested by systematic comparative neuroanatomy of the P-cell in species for which genome sequences permit inter-species comparison. The Fugu rubripes (Fugu), Danio rerio (Danio) and other species are lower order vertebrates for which genome sequences are available and tests could be conducted. Consistent with the fractal prediction, P-cell dendritic arbor is primitive in Fugu, being much less complex than in Mus musculus and in Homo sapiens. Genomic analysis readily identified PEP19/Pcp4, Calbindin-D28k, and GAD67 genes in Fugu and in Danio that are closely associated with P-cells in Canis familiaris, Rattus norvegicus, Mus musculus and Homo sapiens. Gene L7/Pcp2 exhibits strongest association with P-cells in higher vertebrates. L7/Pcp2 shows strong protein residue homology with genes greater than 600 residues and including 2-3 GoLoco domains, designated as having G protein signaling modulator function (AGS3-like proteins). Fugu has a short gene with a single GoLoco domain, but it has greatest homology with the AGS3-like proteins. No similar short gene is present in Danio or in Xenopus. Classical L7/Pcp2 is only detected in higher vertebrates, suggesting that it may be a marker of more recent evolutionary development of cerebellar P-cells. We expect that a new generation of data mining tools will be required to support recursive fractal geometrical, combinatorial, and neural network models of the genomic basis of morphogenesis.

Animals↗

Eucalyptus ESTs related to genes for oxidative stress.

Oxidative stress generating active oxygen species has been proved to be one of the underlying agents causing tissue injury after the exposure of Eucalyptus (Eucalyptus spp.) plants to a wide variety of stress conditions. The objective of this study was to perform data mining to identify favorable genes and alleles associated with the enzyme systems superoxide dismutase, catalase, peroxidases, and glutathione S-transferase that are related to tolerance for environmental stresses and damage caused by pests, diseases, herbicides, and by weeds themselves. This was undertaken by using the eucalyptus expressed-sequence database (https//forests.esalq.usp.br). The alignment results between amino acid and nucleotide sequences indicated that the studied enzymes were adequately represented in the ESTs database of the FORESTs project.

Catalase↗

GAMOLA: a new local solution for sequence annotation and analyzing draft and finished prokaryotic genomes.

Laboratories working with draft phase genomes have specific software needs, such as the unattended processing of hundreds of single scaffolds and subsequent sequence annotation. In addition, it is critical to follow the "movement" and the manual annotation of single open reading frames (ORFs) within the successive sequence updates. Even with finished genomes, regular database updates can lead to significant changes in the annotation of single ORFs. In functional genomics it is important to mine data and identify new genetic targets rapidly and easily. Often there is no need for sophisticated relational databases (RDB) that greatly reduce the system-independent access of the results. Another aspect is the internet dependency of most software packages. If users are working with confidential data, this dependency poses a security issue. GAMOLA was designed to handle the numerous scaffolds and changing contents of draft phase genomes in an automated process and stores the results for each predicted ORF in flatfile databases. In addition, annotation transfers, ORF designation tracking, Blast comparisons, and primer design for whole genome microarrays have been implemented. The software is available under the license of North Carolina State University. A website and a downloadable example are accessible under (http://fsweb2.schaub. ncsu.edu/TRKwebsite/index.htm).

Algorithms↗

Binary state pattern clustering: a digital paradigm for class and biomarker discovery in gene microarray studies of cancer.

Class and biomarker discovery continue to be among the preeminent goals in gene microarray studies of cancer. We have developed a new data mining technique, which we call Binary State Pattern Clustering (BSPC) that is specifically adapted for these purposes, with cancer and other categorical datasets. BSPC is capable of uncovering statistically significant sample subclasses and associated marker genes in a completely unsupervised manner. This is accomplished through the application of a digital paradigm, where the expression level of each potential marker gene is treated as being representative of its discrete functional state. Multiple genes that divide samples into states along the same boundaries form a kind of gene-cluster that has an associated sample-cluster. BSPC is an extremely fast deterministic algorithm that scales well to large datasets. Here we describe results of its application to three publicly available oligonucleotide microarray datasets. Using an alpha-level of 0.05, clusters reproducing many of the known sample classifications were identified along with associated biomarkers. In addition, a number of simulations were conducted using shuffled versions of each of the original datasets, noise-added datasets, as well as completely artificial datasets. The robustness of BSPC was compared to that of three other publicly available clustering methods: ISIS, CTWC and SAMBA. The simulations demonstrate BSPC's substantially greater noise tolerance and confirm the accuracy of our calculations of statistical significance.

Algorithms↗

LARaLink 2.0: a comprehensive aid to basic and clinical cytogenetic research.

LARaLink 2.0 (Loci Analysis for Rearrangement Link) is an enabling web technology that permits the rapid retrieval of clinical cytogenetic and molecular data. New data mining capabilities have been incorporated into version 2.0, building upon LARaLink 1.0, to extend the utility of the system for applications in both the clinical and basic sciences. These include access to the Chromosomal Variation in Man database and the GEO database. Together these new resources enhance the user's ability to associate genotype with phenotype to identify potential gene candidates. Unlimited access for researchers exploring disease-gene relationships and for clinicians extending practice in patient care is available at LARaLink.bioinformatics.wayne.edu:8080/ unigene.

Chromosome Aberrations↗

Shopping in the genome market with EnsMart.

Life scientists who work with the supermarket of genome data will find the EnsMart database and software package offers a valuable door to a wealth of genes and genome features. Not only available to lab biologists on the web, this popular multi-organism genome database can be installed and used on your own Unix computer with relative ease. It offers a flexible, fast and practical data-mining framework for computer-savvy biologists and bioinformaticians.

Animals↗