Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

In silico characterization of the INO80 subfamily of SWI2/SNF2 chromatin remodeling proteins.

Proteins belonging to SNF2 family of DNA dependent ATPases are important members of the chromatin remodeling complexes that are implicated in epigenetic control of gene expression. The yeast Ino80, the catalytic ATPase subunit of the INO80 complex, is the most recently described member of the SNF2 family. Outside the conserved ATPase domain, it has very little similarity with other well-characterized SNF2 proteins hence it is believed to represent a new subfamily. We have identified new members of this subfamily in different organisms and have detected characteristic features of this subfamily. Using various data mining tools we have identified a new, previously undetected domain in all members of this subfamily. This domain designated DBINO is characteristic of the INO80 subfamily and is predicted to have DNA-binding function. The presence of this domain in all the INO80 subfamily proteins from different organisms suggests its conserved function in evolution.

Amino Acid Sequence↗

Temporal profiling of the transcriptional basis for the development of corticosteroid-induced insulin resistance in rat muscle.

Elevated systemic levels of glucocorticoids are causally related to peripheral insulin resistance. The pharmacological use of synthetic glucocorticoids (corticosteroids) often results in insulin resistance/type II diabetes. Skeletal muscle is responsible for close to 80% of the insulin-induced systemic disposal of glucose and is a major target for glucocorticoid-induced insulin resistance. We used Affymetrix gene chips to profile the dynamic changes in mRNA expression in rat skeletal muscle in response to a single bolus dose of the synthetic glucocorticoid methyl-prednisolone. Temporal expression profiles (analyzed on individual chips) were obtained from tissues of 48 drug-treated animals encompassing 16 time points over 72 h following drug administration along with four vehicle-treated controls. Data mining identified 653 regulated probe sets out of 8799 present on the chip. Of these 653 probe sets we identified 29, which represented 22 gene transcripts, that were associated with the development of insulin resistance. These 29 probe sets were regulated in three fundamental temporal patterns. 16 probe sets coding for 12 different genes had a profile of enhanced expression. 10 probe sets coding for eight different genes showed decreased expression and three probe sets coding for two genes showed biphasic temporal signatures. These transcripts were grouped into four general functional categories: signal transduction, transcription regulation, carbohydrate/fat metabolism, and regulation of blood flow to the muscle. The results demonstrate the polygenic nature of transcriptional changes associated with insulin resistance that can provide a temporal scaffolding for translational and post-translational data as they become available.

Animals↗

The histamine H4 receptor as a new therapeutic target for inflammation.

Following the sequencing of the human genome, data-mining efforts have revealed the existence of a new histamine receptor that is expressed at high levels in mast cells and leukocytes. The histamine H(4) receptor has a distinct pharmacological profile and the first compounds that act selectively on the H(4) receptor have been developed. Initial experiments in vivo with H(4) receptor antagonists indicate a role for the H(4) receptor in inflammatory conditions.

Amino Acid Sequence↗

Combined histomorphometric and gene-expression profiling applied to toxicology.

We have developed a unique methodology for the combined analysis of histomorphometric and gene-expression profiles amenable to intensive data mining and multisample comparison for a comprehensive approach to toxicology. This hybrid technology, termed extensible morphometric relational gene-expression analysis (EMeRGE), is applied in a toxicological study of time-varied vehicle- and carbon-tetrachloride (CCl4)-treated rats, and demonstrates correlations between specific genes and tissue structures that can augment interpretation of biological observations and diagnosis.

Animals↗

Application of quantitative signal detection in the Dutch spontaneous reporting system for adverse drug reactions.

The primary aim of spontaneous reporting systems (SRSs) is the timely detection of unknown adverse drug reactions (ADRs), or signal detection. Generally this is carried out by a systematic manual review of every report sent to an SRS. Statistical analysis of the data sets of an SRS, or quantitative signal detection, can provide additional information concerning a possible relationship between a drug and an ADR. We describe the role of quantitative signal detection and the way it is applied at the Netherlands Pharmacovigilance Centre Lareb. Results of the statistical analysis are implemented in the traditional case-by-case analysis. In addition, for data-mining purposes, a list of associations of ADRs and suspected drugs that are disproportionally present in the database is periodically generated. Finally, quantitative signal generation can be used to study more complex relationships, such as drug-drug interactions and syndromes. The results of quantitative signal detection should be considered as an additional source of information, complementary to the traditional analysis. Techniques for the detection of drug interactions and syndromes offer a new challenge for pharmacovigilance in the near future.

Adverse Drug Reaction Reporting Systems↗

Melanin-concentrating hormone and its receptors: state of the art.

Melanin-concentrating hormone (MCH) is a cyclic neuropeptide of nineteen amino acids in mammals. Its involvement in the feeding behaviour has been well established during the last few years. A first receptor subtype, now termed MCHIR, was discovered in 1999, following the desorphanisation of the SLCI orphan receptor, using either reverse pharmacology or systematic screening of agonist candidates. A second MCH receptor, MCH2R, has been discovered recently, by several groups working on data mining of genomic banks. The molecular pharmacology of these two receptors is only described on the basis of the action of peptides derived from MCH. The present review tentatively summarizes the knowledge on these two receptors and presents the first attempts to discover new classes of antagonists that might have major roles in the control of obesity and feeding behaviour.

Amino Acid Sequence↗

The influence of organ donor factors on early allograft function.

PURPOSE OF REVIEW: Postischaemic acute renal allograft failure is among the main risk factors for reduced transplant survival. Although new immunosuppressive protocols have reduced the number of acute rejections, the incidence of acute renal failure remained unchanged. On the basis of histomorphology it is not possible to predict donor kidneys at risk of subsequent failure. Some factors are associated with failure, but even combinations of these risk factors can not precisely predict the development of acute renal failure. Studies have therefore evaluated the influence of demographic donor and recipient factors on acute renal failure. New biotechnology and data mining tools are currently being used to study and identify the molecular predictors of acute renal failure. RECENT FINDINGS: Recent studies showed that donor factors contributed to approximately 40% of the variability in early allograft function. Deductive approaches identified some isolated molecular targets, such as adhesion molecules, as risk factors. Explorative analysis of the entire human genome, however, identified several predictive clusters of genes, which can be functionally grouped into categories such as cell death, stress response, cell adhesion, transcription factors, inflammatory response or cell cycle-related genes. Based on this information, preventative strategies using antisense oligonucleotides or antibodies were adopted. Clinical studies identified the use of catecholamines in the organ donor as beneficial. All these efforts aim to reduce renal tubular damage. SUMMARY: A detailed analysis of the molecular events and pathways of renal gene expression in the donor and after reperfusion, together with sophisticated data analysis tools, will provide new insights into the pathophysiology of acute renal failure.

Acute Kidney Injury↗

Prediction of maximum exposure in poor metabolizers following inhibition of nonpolymorphic pathways.

Marked increases in exposure of some substrates have been noted in poor metabolizers given inhibitors of nonpolymorphic enzymes. Among the small number of clinical trials conducted to investigate this problem, a wide variation in the degree of maximum exposure ratios (area under the curve in poor metabolizers in the presence of inhibitor/area under the curve in extensive metabolizers) among the different substrates has been reported, with some trials reporting profound increases (> tenfold), and others demonstrating less remarkable changes (< twofold). The conduct of such trials raises safety concerns for the trial participants, in addition to other ethical and logistic concerns; therefore, the possibility was investigated that maximum exposure (area under the curve in poor metabolizers in the presence of an inhibitor) could be predicted, and that substrates susceptible to large increases in exposure could be identified. Existing clinical trials were identified by data mining the literature. A theoretical approach was developed to predict maximum exposure in poor metabolizers from studies in extensive metabolizers treated with an inhibitor of the nonpolymorphic pathway. Maximum exposure was predicted in eleven instances and the mean percentage difference between predicted and observed was 11.9%. Substrates with a fraction of substrate dose metabolized by the polymorphic enzyme (fm(POLY)) higher than 75% are at greater risk of exhibiting maximum exposure ratios of more than tenfold.

Algorithms↗

Components of the antigen processing and presentation pathway revealed by gene expression microarray analysis following B cell antigen receptor (BCR) stimulation.

BACKGROUND: Activation of naïve B lymphocytes by extracellular ligands, e.g. antigen, lipopolysaccharide (LPS) and CD40 ligand, induces a combination of common and ligand-specific phenotypic changes through complex signal transduction pathways. For example, although all three of these ligands induce proliferation, only stimulation through the B cell antigen receptor (BCR) induces apoptosis in resting splenic B cells. In order to define the common and unique biological responses to ligand stimulation, we compared the gene expression changes induced in normal primary B cells by a panel of ligands using cDNA microarrays and a statistical approach, CLASSIFI (Cluster Assignment for Biological Inference), which identifies significant co-clustering of genes with similar Gene Ontology annotation. RESULTS: CLASSIFI analysis revealed an overrepresentation of genes involved in ion and vesicle transport, including multiple components of the proton pump, in the BCR-specific gene cluster, suggesting that activation of antigen processing and presentation pathways is a major biological response to antigen receptor stimulation. Proton pump components that were not included in the initial microarray data set were also upregulated in response to BCR stimulation in follow up experiments. MHC Class II expression was found to be maintained specifically in response to BCR stimulation. Furthermore, ligand-specific internalization of the BCR, a first step in B cell antigen processing and presentation, was demonstrated. CONCLUSION: These observations provide experimental validation of the computational approach implemented in CLASSIFI, demonstrating that CLASSIFI-based gene expression cluster analysis is an effective data mining tool to identify biological processes that correlate with the experimental conditional variables. Furthermore, this analysis has identified at least thirty-eight candidate components of the B cell antigen processing and presentation pathway and sets the stage for future studies focused on a better understanding of the components involved in and unique to B cell antigen processing and presentation.

Algorithms↗

A knowledge based approach for automated signal generation in pharmacovigilance.

BACKGROUND: Pharmacovigilance experts detect new adverse drug reactions (ADR) by manually reviewing spontaneous reporting systems. Automated signal generation aims to focus the attention of experts on drug-adverse event associations which are disproportionally present in the database. Although adverse events are coded by means of controlled vocabularies such as the MedDRA dictionary, this semantic information is not taken into account for signal generation. OBJECTIVE: To improve the performance of current signal detection algorithms using knowledge based approach. METHOD: We developed a formal ontology of ADRs and built a data mining tool that uses description logic representations of MedDRA terms to group medically related case reports. RESULTS: This knowledge based approach increased the sensitivity of signal detection with no decrease in specificity. DISCUSSION: A knowledge based approach improved the performance of signal detection tools. However, the huge work-load involved in the knowledge engineering step limits the use of this approach for machine learning.

Adverse Drug Reaction Reporting Systems↗

An interactive tool for extracting exons and SNP from genomic sequence: isolation of HCN1 and HCN3 ion channel genes.

Genome Analyzer (GenoA) with a relational database back-end, was developed to extract information from mammalian genomic sequences. This data mining and visualization tool-set enables laboratory bench scientists to identify and assemble virtual cDNA from genomic exon sequences, and provides a starting point to identify potential alternative splice variants and polymorphisms in silico. The study described in this paper demonstrates the use of GenoA to study human brain hyperpolarization-activated cation channel genes HCN1 and HCN3.

Amino Acid Sequence↗

Irreversible UV inactivation of Cryptosporidium spp. despite the presence of UV repair genes.

Ultraviolet light is being considered as a disinfectant by the water industry because it appears to be very effective for inactivating pathogens, including Cryptosporidium parvum. However, many organisms have mechanisms for repairing ultraviolet light-induced DNA damage, which may limit the utility of this disinfection technology. Inactivation of C. parvum was assessed by measuring infectivity in cells of the human ileocecal adenocarcinoma HCT-8 cell line, with an assay targeting a heat shock protein gene and using a reverse transcriptase polymerase chain reaction to detect infections. Oocysts of five different isolates displayed similar sensitivity to ultraviolet light. An average dosage of 7.6 mJ/cm2 resulted in 99.9% inactivation, providing the first evidence that multiple isolates of C. parvum are equally sensitive to ultraviolet disinfection. Irradiated oocysts were unable to regain pre-irradiation levels of infectivity, following exposure to a broad array of potential repair conditions, such as prolonged incubation, pre-infection excystation triggers, and post-ultraviolet holding periods. A combination of data-mining and sequencing was used to identify genes for all of the major components of a nucleotide excision repair complex in C. parvum and Cryptosporidium hominis. The average similarity between the two organisms for the various genes was 96.4% (range, 92-98%). Thus, while Cryptosporidum spp. may have the potential to repair ultraviolet light-induced damage, oocyst reactivation will not occur under the standard conditions used for storage and distribution of treated drinking water.

Animals↗

Temporally precise cortical firing patterns are associated with distinct action segments.

Despite many reports indicating the existence of precise firing sequences in cortical activity, serious objections have been raised regarding the statistics used to detect them and the relations of these sequences to behavior. We show that in behaving monkeys, pairs of spikes from different neurons tend to prefer certain time delays when measured in relation to a specific behavior. Single-unit activity was recorded from eight microelectrodes inserted into the motor and premotor cortices of two monkeys while they were performing continuous drawinglike hand movements. Repeated scribbling paths, termed drawing components, were extracted by data-mining techniques. The set of the least predictable relations between drawing components and pairs of neurons was determined and represented by one statistic termed the relations score. The chance probability of the relations score was evaluated by teetering the spike times: 1,000 surrogates were generated by randomly teetering the original time of each spike in a small window. In nine of 13 experimental days the precision was better than 12 ms and, in the best case, spike precision reached 0.5 ms.

Animals↗

Bioinformatics in microbial biotechnology--a mini review.

The revolutionary growth in the computation speed and memory storage capability has fueled a new era in the analysis of biological data. Hundreds of microbial genomes and many eukaryotic genomes including a cleaner draft of human genome have been sequenced raising the expectation of better control of microorganisms. The goals are as lofty as the development of rational drugs and antimicrobial agents, development of new enhanced bacterial strains for bioremediation and pollution control, development of better and easy to administer vaccines, the development of protein biomarkers for various bacterial diseases, and better understanding of host-bacteria interaction to prevent bacterial infections. In the last decade the development of many new bioinformatics techniques and integrated databases has facilitated the realization of these goals. Current research in bioinformatics can be classified into: (i) genomics--sequencing and comparative study of genomes to identify gene and genome functionality, (ii) proteomics--identification and characterization of protein related properties and reconstruction of metabolic and regulatory pathways, (iii) cell visualization and simulation to study and model cell behavior, and (iv) application to the development of drugs and anti-microbial agents. In this article, we will focus on the techniques and their limitations in genomics and proteomics. Bioinformatics research can be classified under three major approaches: (1) analysis based upon the available experimental wet-lab data, (2) the use of mathematical modeling to derive new information, and (3) an integrated approach that integrates search techniques with mathematical modeling. The major impact of bioinformatics research has been to automate the genome sequencing, automated development of integrated genomics and proteomics databases, automated genome comparisons to identify the genome function, automated derivation of metabolic pathways, gene expression analysis to derive regulatory pathways, the development of statistical techniques, clustering techniques and data mining techniques to derive protein-protein and protein-DNA interactions, and modeling of 3D structure of proteins and 3D docking between proteins and biochemicals for rational drug design, difference analysis between pathogenic and non-pathogenic strains to identify candidate genes for vaccines and anti-microbial agents, and the whole genome comparison to understand the microbial evolution. The development of bioinformatics techniques has enhanced the pace of biological discovery by automated analysis of large number of microbial genomes. We are on the verge of using all this knowledge to understand cellular mechanisms at the systemic level. The developed bioinformatics techniques have potential to facilitate (i) the discovery of causes of diseases, (ii) vaccine and rational drug design, and (iii) improved cost effective agents for bioremediation by pruning out the dead ends. Despite the fast paced global effort, the current analysis is limited by the lack of available gene-functionality from the wet-lab data, the lack of computer algorithms to explore vast amount of data with unknown functionality, limited availability of protein-protein and protein-DNA interactions, and the lack of knowledge of temporal and transient behavior of genes and pathways.

Journal Article↗

Identification and characterization of Bombyx mori eIF5A gene through bioinformatics approaches.

As the genome of B. mori is available in GenBank and the EST database of B. mori is expanding, identification of novel genes of B. mori was conceivable by data-mining techniques and bioinformatics tools. In this study, we used the in silico cloning method to identify eukaryotic initiation factor 5A (eIF5A) gene in B. mori. With the hypusine formation, eIF5A is involved in the regulation of cell proliferation and apoptosis. Using the computer program MEGA3, we conducted a search for homologs of eIF5A among many eukaryotic species and confirmed that the eIF5A was conserved in all organisms investigated. This gene has been registered in GenBank under the accession number DQ104412.

Amino Acid Sequence↗

Identification of interactive gene networks: a novel approach in gene array profiling of myometrial events during guinea pig pregnancy.

OBJECTIVE: The transition from myometrial quiescence to activation is poorly understood, and the analysis of array data is limited by the available data mining tools. We applied functional analysis and logical operations along regulatory gene networks to identify molecular processes and pathways underlying quiescence and activation. STUDY DESIGN: We analyzed some 18,400 transcripts and variants in guinea pig myometrium at stages corresponding to quiescence and activation, and compared them to the nonpregnant (control) counterpart using a functional mapping tool, MetaCore (GeneGo, St Joseph, MI) to identify novel gene networks composed of biological pathways during mid (MP) and late (LP) pregnancy. RESULTS: Genes altered during quiescence and or activation were identified following gene specific comparisons with myometrium from nonpregnant animals, and then linked to curated pathways and formulated networks. The MP and LP networks were subtracted from each other to identify unique genomic events during those periods. For example, changes 2-fold or greater in genes mediating protein biosynthesis, programmed cell death, microtubule polymerization, and microtubule based movement were noted during the transition to LP. CONCLUSION: We describe a novel approach combining microarrays and genetic data to identify networks associated with normal myometrial events. The resulting insights help identify potential biomarkers and permit future targeted investigations of these pathways or networks to confirm or refute their importance.

Animals↗

Bioinformatic screening of human ESTs for differentially expressed genes in normal and tumor tissues.

BACKGROUND: Owing to the explosion of information generated by human genomics, analysis of publicly available databases can help identify potential candidate genes relevant to the cancerous phenotype. The aim of this study was to scan for such genes by whole-genome in silico subtraction using Expressed Sequence Tag (EST) data. METHODS: Genes differentially expressed in normal versus tumor tissues were identified using a computer-based differential display strategy. Bcl-xL, an anti-apoptotic member of the Bcl-2 family, was selected for confirmation by western blot analysis. RESULTS: Our genome-wide expression analysis identified a set of genes whose differential expression may be attributed to the genetic alterations associated with tumor formation and malignant growth. We propose complete lists of genes that may serve as targets for projects seeking novel candidates for cancer diagnosis and therapy. Our validation result showed increased protein levels of Bcl-xL in two different liver cancer specimens compared to normal liver. Notably, our EST-based data mining procedure indicated that most of the changes in gene expression observed in cancer cells corresponded to gene inactivation patterns. Chromosomes and chromosomal regions most frequently associated with aberrant expression changes in cancer libraries were also determined. CONCLUSION: Through the description of several candidates (including genes encoding extracellular matrix and ribosomal components, cytoskeletal proteins, apoptotic regulators, and novel tissue-specific biomarkers), our study illustrates the utility of in silico transcriptomics to identify tumor cell signatures, tumor-related genes and chromosomal regions frequently associated with aberrant expression in cancer.

Algorithms↗

tRNA maturation in Aquifex aeolicus.

Several tRNAs in the hyperthermophilic bacterium Aquifex aeolicus are encoded in clusters and as part of ribosomal RNA operons, implying the requirement for tRNA processing by ribonuclease P (RNase P). Intriguingly, neither a gene for the RNA nor the protein component of this ubiquitous ribonucleoprotein enzyme has been hitherto identified in the sequenced genome of A. aeolicus, despite extensive data mining. As a result of the present study, primer extension analysis revealed that tRNAs in A. aeolicus possess canonical mature 5' ends; yet we were unable to demonstrate RNase P holoenzyme or RNase P RNA alone activity in A. aeolicus extracts under a variety of reaction conditions utilizing mono- and dimeric ptRNA substrates. Processing of dimeric ptRNA transcripts in extracts of A. aeolicus disclosed at least one endoribonuclease which cleaves in the A/U-rich spacer of the tandem ptRNA, reminiscent of bacterial RNase E-like enzymes.

Bacteria↗