Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.

We present the Saccharomyces cerevisiae PeptideAtlas composed from 47 diverse experiments and 4.9 million tandem mass spectra. The observed peptides align to 61% of Saccharomyces Genome Database (SGD) open reading frames (ORFs), 49% of the uncharacterized SGD ORFs, 54% of S. cerevisiae ORFs with a Gene Ontology annotation of 'molecular function unknown', and 76% of ORFs with Gene names. We highlight the use of this resource for data mining, construction of high quality lists for targeted proteomics, validation of proteins, and software development.

Codon↗

Human fetal neuroblast and neuroblastoma transcriptome analysis confirms neuroblast origin and highlights neuroblastoma candidate genes.

BACKGROUND: Neuroblastoma tumor cells are assumed to originate from primitive neuroblasts giving rise to the sympathetic nervous system. Because these precursor cells are not detectable in postnatal life, their transcription profile has remained inaccessible for comparative data mining strategies in neuroblastoma. This study provides the first genome-wide mRNA expression profile of these human fetal sympathetic neuroblasts. To this purpose, small islets of normal neuroblasts were isolated by laser microdissection from human fetal adrenal glands. RESULTS: Expression of catecholamine metabolism genes, and neuronal and neuroendocrine markers in the neuroblasts indicated that the proper cells were microdissected. The similarities in expression profile between normal neuroblasts and malignant neuroblastomas provided strong evidence for the neuroblast origin hypothesis of neuroblastoma. Next, supervised feature selection was used to identify the genes that are differentially expressed in normal neuroblasts versus neuroblastoma tumors. This approach efficiently sifted out genes previously reported in neuroblastoma expression profiling studies; most importantly, it also highlighted a series of genes and pathways previously not mentioned in neuroblastoma biology but that were assumed to be involved in neuroblastoma pathogenesis. CONCLUSION: This unique dataset adds power to ongoing and future gene expression studies in neuroblastoma and will facilitate the identification of molecular targets for novel therapies. In addition, this neuroblast transcriptome resource could prove useful for the further study of human sympathoadrenal biogenesis.

Databases, Genetic↗

Contribution of regulatory and structural variations in APOE to predicting dyslipidemia.

The objective of this study was to evaluate 1) whether non single nucleotide polymorphisms-coding (non-cSNP) in the apolipoprotein E gene (APOE) identified by resequencing studies contribute to statistically explaining dyslipidemia if variations in the two cSNPs in exon 4 that define the 2, 3, and 4 alleles are ignored, and 2) whether the contribution of these additional SNPs persists when variations in the cSNPs are considered. We used an ecological, multiple-population, data-mining strategy to identify single-SNP and two-SNP genotypes that distinguish between high and low levels of plasma lipids in three training samples, European-Americans from Rochester, MN, African-Americans from Jackson, MS, and Europeans from North Karelia, Finland. We found that a pair of SNPs located in the 5' region define genotypes A560T832/A560T832, A560T832/A560G832, and A560T832/T560T832, which distinguish between high and low levels of HDL-cholesterol (HDL-C), triglycerides (TG), and/or total cholesterol (T-C). The A560T832/- genotypes predicted high TG and high T-C in both genders in a large independent test sample from Copenhagen, Denmark. Prediction of high T-C in the Danish females was dependent on genotypes defined by the cSNPs. Our study suggests that both regulatory and structural variations should be considered when evaluating the utility of APOE for predicting dyslipidemia in the population at large.

Black or African American↗

Gene expression profiling of benign and malignant pheochromocytoma.

There are currently no reliable diagnostic and prognostic markers or effective treatments for malignant pheochromocytoma. This study used oligonucleotide microarrays to examine gene expression profiles in pheochromocytomas from 90 patients, including 20 with malignant tumors, the latter including metastases and primary tumors from which metastases developed. Other subgroups of tumors included those defined by tissue norepinephrine compared to epinephrine contents (i.e., noradrenergic versus adrenergic phenotypes), adrenal versus extra-adrenal locations, and presence of germline mutations of genes predisposing to the tumor. Correcting for the confounding influence of noradrenergic versus adrenergic catecholamine phenotype by the analysis of variance revealed a larger and more accurate number of genes that discriminated benign from malignant pheochromocytomas than when the confounding influence of catecholamine phenotype was not considered. Seventy percent of these genes were underexpressed in malignant compared to benign tumors. Similarly, 89% of genes were underexpressed in malignant primary tumors compared to benign tumors, suggesting that malignant potential is largely characterized by a less-differentiated pattern of gene expression. The present database of differentially expressed genes provides a unique resource for mapping the pathways leading to malignancy and for establishing new targets for treatment and diagnostic and prognostic markers of malignant disease. The database may also be useful for examining mechanisms of tumorigenesis and genotype-phenotype relationships. Further progress on the basis of this database can be made from follow-up confirmatory studies, application of bioinformatics approaches for data mining and pathway analyses, testing in pheochromocytoma cell culture and animal model systems, and retrospective and prospective studies of diagnostic markers.

Adrenal Gland Neoplasms↗

Unequivocal delineation of clinicogenetic subgroups and development of a new model for improved outcome prediction in neuroblastoma.

PURPOSE: Neuroblastoma is a genetically heterogeneous pediatric tumor with a remarkably variable clinical behavior ranging from widely disseminated disease to spontaneous regression. In this study, we aimed for comprehensive genetic subgroup discovery and assessment of independent prognostic markers based on genome-wide aberrations detected by comparative genomic hybridization (CGH). MATERIALS AND METHODS: Published CGH data from 231 primary untreated neuroblastomas were converted to a digitized format suitable for global data mining, subgroup discovery, and multivariate survival analyses. RESULTS: In contrast to previous reports, which included only a few genetic parameters, we present here for the first time a strategy that allows unbiased evaluation of all genetic imbalances detected by CGH. The presented approach firmly established the existence of three different clinicogenetic subgroups and indicated that chromosome 17 status and tumor stage were the only independent significant predictors for patient outcome. Important new findings were: (1) a normal chromosome 17 status as a delineator of a subgroup of presumed favorable-stage tumors with highly increased risk; (2) the recognition of a survivor signature conferring 100% 5-year survival for stage 1, 2, and 4S tumors presenting with whole chromosome 17 gain; and (3) the identification of 3p deletion as a hallmark of older age at diagnosis. CONCLUSION: We propose a new regression model for improved patient outcome prediction, incorporating tumor stage, chromosome 17, and amplification/deletion status. These findings may prove highly valuable with respect to more reliable risk assessment, evaluation of clinical results, and optimization of current treatment protocols.

Adolescent↗

Differential gene expression profile in omental adipose tissue in women with polycystic ovary syndrome.

CONTEXT: The polycystic ovary syndrome (PCOS) is frequently associated with visceral obesity, suggesting that omental adipose tissue might play an important role in the pathogenesis of the syndrome. OBJECTIVE: The objective was to study the expression profiles of omental fat biopsy samples obtained from morbidly obese women with or without PCOS at the time of bariatric surgery. DESIGN: This was a case-control study. SETTINGS: We conducted the study in an academic hospital. PATIENTS: Eight PCOS patients and seven nonhyperandrogenic women submitted to bariatric surgery because of morbid obesity. INTERVENTIONS: Biopsy samples of omental fat were obtained during bariatric surgery. MAIN OUTCOME MEASURE: The main outcome measure was high-density oligonucleotide arrays. RESULTS: After statistical analysis, we identified changes in the expression patterns of 63 genes between PCOS and control samples. Gene classification was assessed through data mining of Gene Ontology annotations and cluster analysis of dysregulated genes between both groups. These methods highlighted abnormal expression of genes encoding certain components of several biological pathways related to insulin signaling and Wnt signaling, oxidative stress, inflammation, immune function, and lipid metabolism, as well as other genes previously related to PCOS or to the metabolic syndrome. CONCLUSION: The differences in the gene expression profiles in visceral adipose tissue of PCOS patients compared with nonhyperandrogenic women involve multiple genes related to several biological pathways, suggesting that the involvement of abdominal obesity in the pathogenesis of PCOS is more ample than previously thought and is not restricted to the induction of insulin resistance.

Adipose Tissue↗

A dynamic expression survey identifies transcription factors relevant in mouse digestive tract development.

Tissue-restricted transcription factors (TFs), which confer specialized cellular properties, are usually identified through sequence homology or cis-element analysis of lineage-specific genes; conventional modes of mRNA profiling often fail to report non-abundant TF transcripts. We evaluated the dynamic expression during mouse gut organogenesis of 1381 transcripts, covering nearly every known and predicted TF, and documented the expression of approximately 1000 TF genes in gastrointestinal development. Despite distinctive structures and functions, the stomach and intestine exhibit limited differences in TF genes. Among differentially expressed transcripts, a few are virtually restricted to the digestive tract, including Nr2e3, previously regarded as a photoreceptor-specific product. TFs that are enriched in digestive organs commonly serve essential tissue-specific functions, hence justifying a search for other tissue-restricted TFs. Computational data mining and experimental investigation focused interest on a novel homeobox TF, Isx, which appears selectively in gut epithelium and mirrors expression of the intestinal TF Cdx2. Isx-deficient mice carry a specific defect in intestinal gene expression: dysregulation of the high density lipoprotein (HDL) receptor and cholesterol transporter scavenger receptor class B, type I (Scarb1). Thus, integration of developmental gene expression with biological assessment, as described here for TFs, represents a powerful tool to investigate control of tissue differentiation.

Animals↗

MASSIS: a mass spectrum simulation system 1. Principle and method.

A mass spectrum simulation system was developed. The simulated spectrum for a given target structure is computed based on the cleavage knowledge and statistical rules established and stocked in pivot databases: cleavage rule knowledge, function groups, small fragments and fragment-intensity relationships. These databases were constructed from correlation charts and statistical analysis of large population of organic mass spectra using data mining techniques. Since 1980, several systems were proposed for mass spectrum simulation, but in present there is no any commercial software available. This shows the complexity and difficulties in the development of a such system. The reported mass spectral simulation system in this paper could be the first general software for organic chemistry use

Journal Article↗

Management of medicines information for patient safety.

Ensuring patient safety in the management of medicines requires the efficient collection, administration, analysis and distribution of huge amounts of data from different sources. This requires the development of generic methods and tools for tasks such as: (1) data mining and semantics-based inference; (2) integration of heterogeneous scientific information databases relating to drugs and diseases; (3) terminology and coding issues; (4) adverse drug events and signal detection; and (5) networking and security. Collaborative research and development will be required to develop coordinated health care information services. Easy access to high-quality information will provide better health care and increased cost efficiency.

Consumer Product Safety↗

Following in the footnotes of giants: citation analysis and its discontents.

Reflecting on his own personal history with bibliometrics, the author places it in the broader context of research with available information and data-mining. In so doing, he considers the utility of bibliometrics for raising new questions and its limitations for guiding decision-making.

Anecdotes as Topic↗

Thomas A. Neff lecture. Application of expression profiling to the developing lung: identification of putative regulatory networks controlling matrix production.

If we hope to repair damaged lung tissue associated with a variety of acquired and developmental diseases, we must first gain a full appreciation of normal lung development. As an approach, we have utilized Affymetrix (Santa Clara, CA) high-density, oligonucleotide-based microarrays to generate an expression profile of the entire process of rodent lung development, which will be made publicly available. Our initial results were internally consistent and correlated closely with those generated with standard expression techniques such as Northern hybridization. We have verified known expression of genes, found other genes with previously unsuspected expression during lung development, as well as uncovered many expressed sequence tags whose role in lung development awaits further study. Data mining reveals close relationships of expression profiles between specific genes, suggesting novel regulatory relationships. In the future, application of these methods to the study of gene-targeted mice with abnormal lung development should uncover pathways of airway and alveolar development. Ultimately, expression profiling of diseased lungs might allow us to understand why the lung fails to repair, and strategies to influence repair might become apparent.

Animals↗

Bioinformatics approaches to cancer gene discovery.

The Cancer Gene Anatomy Project (CGAP) database of the National Cancer Institute has thousands of known and novel expressed sequence tags (ESTs). These ESTs, derived from diverse normal and tumor cDNA libraries, offer an attractive starting point for cancer gene discovery. Data-mining the CGAP database led to the identification of ESTs that were predicted to be specific to select solid tumors. Two genes from these efforts were taken to proof of concept for diagnostic and therapeutics indications of cancer. Microarray technology was used in conjunction with bioinformatics to understand the mechanism of one of the targets discovered. These efforts provide an example of gene discovery by using bioinformatics approaches. The strengths and weaknesses of this approach are discussed in this review.

Basic Helix-Loop-Helix Proteins↗

Pathway mapping tools for analysis of high content data.

The complexity of human biology requires a systems approach that uses computational approaches to integrate different data types. Systems biology encompasses the complete biological system of metabolic and signaling pathways, which can be assessed by measuring global gene expression, protein content, metabolic profiles, and individual genetic, clinical, and phenotypic data. High content screening assays can also be used to generate systems biology knowledge. In this review, we will summarize the pathway databases and describe biological network tools used predominantly with this genomics, proteomics, and metabolomics data but which are equally as applicable for high content screening data analysis. We describe in detail the integrated data-mining tools applicable to building biological networks developed by GeneGo, namely, MetaCore and MetaDrug.

Computational Biology↗

Classification of HIV-1-mediated neuronal dendritic and synaptic damage using multiple criteria linear programming.

The ability to identify neuronal damage in the dendritic arbor during HIV-1-associated dementia (HAD) is crucial for designing specific therapies for the treatment of HAD. To study this process, we utilized a computer-based image analysis method to quantitatively assess HIV-1 viral protein gp120 and glutamate-mediated individual neuronal damage in cultured cortical neurons. Changes in the number of neurites, arbors, branch nodes, cell body area, and average arbor lengths were determined and a database was formed (http://dm.ist.unomaha. edu/database.htm). We further proposed a two-class model of multiple criteria linear programming (MCLP) to classify such HIV-1-mediated neuronal dendritic and synaptic damages. Given certain classes, including treatments with brain-derived neurotrophic factor (BDNF), glutamate, gp120 or non-treatment controls from our in vitro experimental systems, we used the two-class MCLP model to determine the data patterns between classes in order to gain insight about neuronal dendritic damages. This knowledge can be applied in principle to the design and study of specific therapies for the prevention or reversal of neuronal damage associated with HAD. Finally, the MCLP method was compared with a well-known artificial neural network algorithm to test for the relative potential of different data mining applications in HAD research.

Animals↗

Neural architectures for robot intelligence.

We argue that direct experimental approaches to elucidate the architecture of higher brains may benefit from insights gained from exploring the possibilities and limits of artificial control architectures for robot systems. We present some of our recent work that has been motivated by that view and that is centered around the study of various aspects of hand actions since these are intimately linked with many higher cognitive abilities. As examples, we report on the development of a modular system for the recognition of continuous hand postures based on neural nets, the use of vision and tactile sensing for guiding prehensile movements of a multifingered hand, and the recognition and use of hand gestures for robot teaching. Regarding the issue of learning, we propose to view real-world learning from the perspective of data-mining and to focus more strongly on the imitation of observed actions instead of purely reinforcement-based exploration. As a concrete example of such an effort we report on the status of an ongoing project in our laboratory in which a robot equipped with an attention system with a neurally inspired architecture is taught actions by using hand gestures in conjunction with speech commands. We point out some of the lessons learnt from this system, and discuss how systems of this kind can contribute to the study of issues at the junction between natural and artificial cognitive systems.

Action Potentials↗

Information management systems for pharmacogenomics.

The value of high-throughput genomic research is dramatically enhanced by association with key patient data. These data are generally available but of disparate quality and not typically directly associated. A system that could bring these disparate data sources into a common resource connected with functional genomic data would be tremendously advantageous. However, the integration of clinical and accurate interpretation of the generated functional genomic data requires the development of information management systems capable of effectively capturing the data as well as tools to make that data accessible to the laboratory scientist or to the clinician. In this review these challenges and current information technology solutions associated with the management, storage and analysis of high-throughput data are highlighted. It is suggested that the development of a pharmacogenomic data management system which integrates public and proprietary databases, clinical datasets, and data mining tools embedded in a high-performance computing environment should include the following components: parallel processing systems, storage technologies, network technologies, databases and database management systems (DBMS), and application services.

Animals↗

Dysregulation of gene expression in the 1-methyl-4-phenyl-1,2,3,6-tetrahydropyridine-lesioned mouse substantia nigra.

Parkinson's disease pathogenesis proceeds through several phases, culminating in the loss of dopaminergic neurons of the substantia nigra (SN). Although the 1-methyl-4-phenyl-1,2,3,6-tetrahydropyridine (MPTP) model of oxidative SN injury is frequently used to study degeneration of dopaminergic neurons in mice and non-human primates, an understanding of the temporal sequence of molecular events from inhibition of mitochondrial complex 1 to neuronal cell death is limited. Here, microarray analysis and integrative data mining were used to uncover pathways implicated in the progression of changes in dopaminergic neurons after MPTP administration. This approach enabled the identification of small, yet consistently significant, changes in gene expression within the SN of MPTP-treated animals. Such an analysis disclosed dysregulation of genes in three main areas related to neuronal function: cytoskeletal stability and maintenance, synaptic integrity, and cell cycle and apoptosis. The discovery and validation of these alterations provide molecular evidence for an evolving cascade of injury, dysfunction, and cell death.

1-Methyl-4-phenyl-1,2,3,6-tetrahydropyridine↗

Identification of wheat chromosomal regions containing expressed resistance genes.

The objectives of this study were to isolate and physically localize expressed resistance (R) genes on wheat chromosomes. Irrespective of the host or pest type, most of the 46 cloned R genes from 12 plant species share a strong sequence similarity, especially for protein domains and motifs. By utilizing this structural similarity to perform modified RNA fingerprinting and data mining, we identified 184 putative expressed R genes of wheat. These include 87 NB/LRR types, 16 receptor-like kinases, and 13 Pto-like kinases. The remaining were seven Hm1 and two Hs1(pro-1) homologs, 17 pathogenicity related, and 42 unique NB/kinases. About 76% of the expressed R-gene candidates were rare transcripts, including 42 novel sequences. Physical mapping of 121 candidate R-gene sequences using 339 deletion lines localized 310 loci to 26 chromosomal regions encompassing approximately 16% of the wheat genome. Five major R-gene clusters that spanned only approximately 3% of the wheat genome but contained approximately 47% of the candidate R genes were observed. Comparative mapping localized 91% (82 of 90) of the phenotypically characterized R genes to 18 regions where 118 of the R-gene sequences mapped.

Base Sequence↗