Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

A novel class of inhibitors of peptide deformylase discovered through high-throughput screening and virtual ligand screening.

Peptide deformylase (PDF) has been identified as a promising antibacterial and herbicide target. A structurally novel class of inhibitors containing a 2-thioxo-thiazolidin-4-one heterocycle substituted by an arylidene group at the 5-position and a hexanoic acid side chain at the 3-position was discovered independently via high-throughput screening and virtual ligand screening. Data mining and analogue synthesis established a structure--activity relationship for the side chain region that is consistent with the docked structure.

Amidohydrolases↗

A 3D similarity method for scaffold hopping from known drugs or natural ligands to new chemotypes.

A primary goal of 3D similarity searching is to find compounds with similar bioactivity to a reference ligand but with different chemotypes, i.e., "scaffold hopping". However, an adequate description of chemical structures in 3D conformational space is difficult due to the high-dimensionality of the problem. We present an automated method that simplifies flexible 3D chemical descriptions in which clustering techniques traditionally used in data mining are exploited to create "fuzzy" molecular representations called FEPOPS (feature point pharmacophores). The representations can be used for flexible 3D similarity searching given one or more active compounds without a priori knowledge of bioactive conformations or pharmacophores. We demonstrate that similarity searching with FEPOPS significantly enriches for actives taken from in-house high-throughput screening datasets and from MDDR activity classes COX-2, 5-HT3A, and HIV-RT, while also scaffold or ring-system hopping to new chemical frameworks. Further, inhibitors of target proteins (dopamine 2 and retinoic acid receptor) are recalled by FEPOPS by scaffold hopping from their associated endogenous ligands (dopamine and retinoic acid). Importantly, the method excels in comparison to commonly used 2D similarity methods (DAYLIGHT, MACCS, Pipeline Pilot fingerprints) and a commercial 3D method (Pharmacophore Distance Triplets) at finding novel scaffold classes given a single query molecule.

Cyclooxygenase 2↗

Phylomat: an automated protein motif analysis tool for phylogenomics.

Recent progress in genomics, proteomics, and bioinformatics enables unprecedented opportunities to examine the evolutionary history of molecular, cellular, and developmental pathways through phylogenomics. Accordingly, we have developed a motif analysis tool for phylogenomics (Phylomat, http://alg.ncsa.uiuc.edu/pmat) that scans predicted proteome sets for proteins containing highly conserved amino acid motifs or domains for in silico analysis of the evolutionary history of these motifs/domains. Phylomat enables the user to download results as full protein or extracted motif/domain sequences from each protein. Tables containing the percent distribution of a motif/domain in organisms normalized to proteome size are displayed. Phylomat can also align the set of full protein or extracted motif/domain sequences and predict a neighbor-joining tree from relative sequence similarity. Together, Phylomat serves as a user-friendly data-mining tool for the phylogenomic analysis of conserved sequence motifs/domains in annotated proteomes from the three domains of life.

Algorithms↗

Correcting common errors in identifying cancer-specific serum peptide signatures.

"Molecular signatures" are the qualitative and quantitative patterns of groups of biomolecules (e.g., mRNA, proteins, peptides, or metabolites) in a cell, tissue, biological fluid, or an entire organism. To apply this concept to biomarker discovery, the measurements should ideally be noninvasive and performed in a single read-out. We have therefore developed a peptidomics platform that couples magnetics-based, automated solid-phase extraction of small peptides with a high-resolution MALDI-TOF mass spectrometric readout (Villanueva, J.; Philip, J.; Entenberg, D.; Chaparro, C. A.; Tanwar, M. K.; Holland, E. C.; Tempst, P. Anal. Chem. 2004, 76, 1560-1570). Since hundreds of peptides can be detected in microliter volumes of serum, it allows to search for disease signatures, for instance in the presence of cancer. We have now evaluated, optimized, and standardized a number of clinical and analytical chemistry variables that are major sources of bias; ranging from blood collection and clotting, to serum storage and handling, automated peptide extraction, crystallization, spectral acquisition, and signal processing. In addition, proper alignment of spectra and user-friendly visualization tools are essential for meaningful, certifiable data mining. We introduce a minimal entropy algorithm, "Entropycal", that simplifies alignment and subsequent statistical analysis and increases the percentage of the highly distinguishing spectral information being retained after feature selection of the datasets. Using the improved analytical platform and tools, and a commercial statistics program, we found that sera from thyroid cancer patients can be distinguished from healthy controls based on an array of 98 discriminant peptides. With adequate technological and computational methods in place, and using rigorously standardized conditions, potential sources of patient related bias (e.g., gender, age, genetics, environmental, dietary, and other factors) may now be addressed.

Algorithms↗

Functional proteomics approach to investigate the biological activities of cDNAs implicated in breast cancer.

Functional proteomics approaches that comprehensively evaluate the biological activities of human cDNAs may provide novel insights into disease pathogenesis. To systematically investigate the functional activity of cDNAs that have been implicated in breast carcinogenesis, we generated a collection of cDNAs relevant to breast cancer, the Breast Cancer 1000 (BC1000), and conducted screens to identify proteins that induce phenotypic changes that resemble events which occur during tumor initiation and progression. Genes were selected for this set using bioinformatics and data mining tools that identify genes associated with breast cancer. Greater than 1000 cDNAs were assembled and sequence verified with high-throughput recombination-based cloning. To our knowledge, the BC1000 represents the first publicly available sequence-validated human disease gene collection. The functional activity of a subset of the BC1000 collection was evaluated in cell-based assays that monitor changes in cell proliferation, migration, and morphogenesis in MCF-10A mammary epithelial cells expressing a variant of ErbB2 that can be inducibly activated through dimerization. Using this approach, we identified many cDNAs, encoding diverse classes of cellular proteins, that displayed activity in one or more of the assays, thus providing insights into a large set of cellular proteins capable of inducing functional alterations associated with breast cancer development.

Breast Neoplasms↗

Genomic analysis of rodent pulmonary tissue following bis-(2-chloroethyl) sulfide exposure.

Bis-(2-chloroethyl) sulfide (sulfur mustard, SM) is a carcinogenic alkylating agent that has been utilized as a chemical warfare agent. To understand the mechanism of SM-induced lung injury, we analyzed global changes in gene expression in a rat lung SM exposure model. Rats were injected in the femoral vein with liquid SM, which circulates directly to the pulmonary vein and then to the lung. Rats were exposed to 1, 3, or 6 mg/kg of SM, and lungs were harvested at 0.5, 1, 3, 6, and 24 h postinjection. Three biological replicates were used for each time point and dose tested. RNA was extracted from the lungs and used as the starting material for the probing of replicate oligonucleotide microarrays. The gene expression data were analyzed using principal component analysis and two-way analysis of variance to identify the genes most significantly changed across time and dose. These genes were ranked by p value and categorized based on molecular function and biological process. Computer-based data mining algorithms revealed several biological processes affected by SM exposure, including protein catabolism, apoptosis, and glycolysis. Several genes that are significantly upregulated in a dose-dependent fashion have been reported as p53 responsive genes, suggesting that cell cycle regulation and p53 activation are involved in the response to SM exposure in the lung. Thus, SM exposure induces transcriptional changes that reveal the cellular response to this potent alkylating agent.

Animals↗

Cuticular hydrocarbons of Tetramorium ants from central Europe: analysis of GC-MS data with self-organizing maps (SOM) and implications for systematics.

Cuticular hydrocarbons were extracted from workers of 63 different nests of five species of Tetramorium ants (Hymenoptera: Formicidae) from Austria, Hungary, and Spain. The GC-MS data were classified (data mining) by self-organizing maps (SOM). SOM neurons derived from primary neuron separation were subjected to hierarchical SOM (HSOM) and were grouped to neuron areas on the basis of vicinity in the hexagonal output grid. While primary neuron separation and HSOM resulted in classifications on a level more sensitive than species differences, neuron areas resulted in chemical phenotypes apparently of the order of species. These chemical phenotypes have implications for systematics: while the chemical phenotypes for T. ferox and T. moravicum correspond to morphological determination, in T. caespitum and T. impurum a total of six chemical phenotypes is found. Three hypotheses are discussed to explain this disparity between morphological and chemical classifications, including in particular the possibility of hybridization and the existence of cryptic species. Overall, the GC-MS profiles classified by SOM prove to be a practical alternative to morphological determination (T. ferox, T. moravicum) and indicate the need to revisit systematics (T. caespitum, T. impurum).

Animals↗

E-medicine and health care consumers: recognizing current problems and possible resolutions for a safer environment.

Millions of Americans access the Internet for health information, which is changing the way patients seek information about, and often treat, certain medical conditions. It is estimated that there may be as many as 100,000 health-related Web sites. The availability of so much health information permits consumers to assume more responsibility for their own health care. At the same time, it raises a number of issues that need to be addressed. The health information available to Internet users may be inaccurate or out-of-date. Potential conflicts of interest result from the blurring of the distinction between advertising and professional health information. Also, potential threats to privacy may result from data mining. Health care consumers need to be able to evaluate the quality of the information provided on the Internet. Various evaluative mechanisms such as codes of ethics, rating systems, and seals of approval have been developed to aid in this process. The effectiveness of these solutions is evaluated in this paper. Finally, the paper addresses the importance of including patients in developing standardized quality assurance systems for online health information.

Conflict of Interest↗

In silico models for the prediction of dose-dependent human hepatotoxicity.

The liver is extremely vulnerable to the effects of xenobiotics due to its critical role in metabolism. Drug-induced hepatotoxicity may involve any number of different liver injuries, some of which lead to organ failure and, ultimately, patient death. Understandably, liver toxicity is one of the most important dose-limiting considerations in the drug development cycle, yet there remains a serious shortage of methods to predict hepatotoxicity from chemical structure. We discuss our latest findings in this area and present a new, fully general in silico model which is able to predict the occurrence of dose-dependent human hepatotoxicity with greater than 80% accuracy. Utilizing an ensemble recursive partitioning approach, the model classifies compounds as toxic or non-toxic and provides a confidence level to indicate which predictions are most likely to be correct. Only 2D structural information is required and predictions can be made quite rapidly, so this approach is entirely appropriate for data mining applications and for profiling large synthetic and/or virtual libraries.

Computer Simulation↗

Computational approaches to protein-protein interaction.

The interactions between proteins allow the cell's life. A number of experimental, genome-wide, high-throughput studies have been devoted to the determination of protein-protein interactions and the consequent interaction networks. Here, the bioinformatics methods dealing with protein-protein interactions and interaction network are overviewed. 1. Interaction databases developed to collect and annotate this immense amount of data; 2. Automated data mining techniques developed to extract information about interactions from the published literature; 3. Computational methods to assess the experimental results developed as a consequence of the finding that the results of high-throughput methods are rather inaccurate; 4. Exploitation of the information provided by protein interaction networks in order to predict functional features of the proteins; and 5. Prediction of protein-protein interactions.

Algorithms↗

Towards patient-specific tumor antigen selection for vaccination.

In this review, we discuss the possibilities for combining the power of molecular analysis of the antigens expressed in a given individual tumor with the design of a tailored vaccine containing defined antigens. Step 1 is a differential gene expression analysis of tumor and corresponding normal tissue. Step 2 is the analysis of human leukocyte antigen (HLA) ligands on tumor cells. Step 3 is data mining with the aim to select those antigens that might be suitable for tumor attack by the adaptive immune system. Step 4 is the on-the-spot clinical grade production of the constituents of the patient tailored vaccine, e.g. peptides. Step 5 is then vaccination and monitoring. Although it will not be possible to cover all relevant antigens expressed in a tumor, the antigens that can be identified with our present technical possibilities might be enough for improved immunotherapy. The scope of the present review is to explore the possibilities and the formidable technical and logistical challenge for such individual patient-oriented antigen definition to be used for therapeutic immunization.

Algorithms↗

Combinatorial informatics in the post-genomics ERA.

The multitude of potential drug targets emerging from genome sequencing demands new approaches to drug discovery. A chemogenomics strategy, which involves the generation of small-molecule compounds that can be used both as tools to probe biological mechanisms and as leads for drug-property optimization, provides a highly parallel, industrialized solution. Key to the success of this strategy is an integrated suite of chemi-informatics applications that can allow the rapid and directed optimization of chemical compounds with drug-like properties using 'just-in-time' combinatorial chemical synthesis. An effective embodiment of this process requires new computational and data-mining tools that cover all aspects of library generation, compound selection and experimental design, and work effectively on a massive scale.

Combinatorial Chemistry Techniques↗

Testing the hypothesis of recent population expansions in nematode parasites of human-associated hosts.

It has been predicted that parasites of human-associated organisms (eg humans, domestic pets, farm animals, agricultural and silvicultural plants) are more likely to show rapid recent population expansions than are parasites of other hosts. Here, we directly test the generality of this demographic prediction for species of parasitic nematodes that currently have mitochondrial sequence data available in the literature or the public-access genetic databases. Of the 23 host/parasite combinations analysed, there are seven human-associated parasite species with expanding populations and three without, and there are three non-human-associated parasite species with expanding populations and 10 without. This statistically significant pattern confirms the prediction. However, it is likely that the situation is more complicated than the simple hypothesis test suggests, and those species that do not fit the predicted general pattern provide interesting insights into other evolutionary processes that influence the historical population genetics of host-parasite relationships. These processes include the effects of postglacial migrations, evolutionary relationships and possibly life-history characteristics. Furthermore, the analysis highlights the limitations of this form of bioinformatic data-mining, in comparison to controlled experimental hypothesis tests.

Animals↗

Identification of genomic organisation, sequence variants and analysis of the role of the human dishevelled 1 gene in late onset Alzheimer's disease.

Alzheimer's disease (AD) is a disorder characterised by a progressive deterioration in memory and other cognitive functions. Neurofibrillary tangles (NFT) are a major pathological hallmark of AD, these are aggregations of paired helical filaments (PHF) comprised of the hyperphosphorylated microtubule associated protein tau. Several kinases, such as glycogen synthase kinase 3 beta (GSK3beta) and c-Jun N-terminal kinase (JNK), phosphorylate tau at sites that are phosphorylated in PHF. Dishevelled 1 (DVL1) is thought to act as a positive regulator of the wnt signalling pathway, and inhibits GSK3beta activity preventing beta-catenin degradation and thus allowing wnt target gene expression. JNK activation is also regulated by DVL1, however it is unclear if this is via the wnt signalling pathway. These observations suggest a central role for DVL1 in tau phosphorylation and AD and led us to investigate DVL1 as a candidate gene for this disorder. We determined the genomic structure of the DVL1 gene by sequencing and data mining and searched for sequence variations in the coding sequences and flanking introns. The DVL1 gene spans a region of approximately 13.8 kb (not including the 5' untranslated region) and is encoded by 15 exons. Analysis of over 4.3 kb of sequence, including 98% of exonic sequences and introns 2, 3, 6, 7, 9, 10, 11 and 12, revealed there to be six rare (< or =6%) sequence variations. None of these had any association with late onset AD. This would suggest that polymorphic variations in the coding sequences of DVL1 are not important in AD. However further analysis of regulatory regions may lead to the identification of other sequence variations which may be implicated in AD.

Adaptor Proteins, Signal Transducing↗

Pharmacogenomic analysis: correlating molecular substructure classes with microarray gene expression data.

Genomic studies are producing large databases of molecular information on cancers and other cell and tissue types. Hence, we have the opportunity to link these accumulating data to the drug discovery processes. Our previous efforts at 'information-intensive' molecular pharmacology have focused on the relationship between patterns of gene expression and patterns of drug activity. In the present study, we take the process a step further-relating gene expression patterns, not just to the drugs as entities, but to approximately 27,000 substructures and other chemical features within the drugs. This coupling of genomic information with structure-based data mining can be used to identify classes of compounds for which detailed experimental structure-activity studies may be fruitful. Using a systematic substructure analysis coupled with statistical correlations of compound activity with differential gene expression, we have identified two subclasses of quinones whose patterns of activity in the National Cancer Institute's 60-cell line screening panel (NCI-60) correlate strongly with the expression patterns of particular genes: (i) The growth inhibitory patterns of an electron-withdrawing subclass of benzodithiophenedione-containing compounds over the NCI-60 are highly correlated with the expression patterns of Rab7 and other melanoma-specific genes; (ii) the inhibitory patterns of indolonaphthoquinone-containing compounds are highly correlated with the expression patterns of the hematopoietic lineage-specific gene HS1 and other leukemia genes. As illustrated by these proof-of-principle examples, we introduce here a set of conceptual tools and fluent computational methods for projecting directly from gene expression patterns to drug substructures and vice versa. The analysis is presented in terms of the NCI-60 cell lines and microarray-based gene expression patterns, but the concept and methods are broadly applicable to other large-scale pharmacogenomic database sets as well. The approach (SAT for Structure-Activity-Target) provides a systematic way to mine databases for the design of further structure-activity studies, particularly to aid in target and lead identification.

Algorithms↗

Optimization of fertirrigation efficiency in strawberry crops by application of fuzzy logic techniques.

A high level of price support has favoured intensive agriculture and an increasing use of fertilisers and pesticides. This has resulted in the pollution of water and soils and damage to certain eco-systems. The target relationship that must be established between agriculture and environment can be called "sustainable agriculture". In this work we aim at relating strawberry total yield with nitrate concentration in water at different soil depths. To achieve this objective, we have used the Predictive Fuzzy Rules Generator (PreFuRGe) tool, based on fuzzy logic and data mining, by means of which the dose that allows a balance between yield and environmental damage minimization can be determined. This determination is quite simple and is done directly from the obtained charts. This technique can be used in other types of crops permitting one to determine in a precise way at which depth the appropriate dose of nitrate fertilizer must be correctly applied, on the one hand providing the maximum yield but, on the other hand, with the minimum loss of nitrates that leachate through the saturated zone polluting aquifers.

Agriculture↗

Advances in two-dimensional gel matching technology.

For many years, two-dimensional gel electrophoresis has been the method of choice for the investigation of complex mixtures of proteins. Although there are a number of emerging technologies that can be applied to proteomics, none can yet yield routinely the breadth of information available from two-dimensional gels. To be able to obtain instant information regarding molecular mass and pI, as well as to highlight quickly the expression changes or unique proteins across a gel series requires sophisticated and powerful image analysis software. The range of software products offered by Nonlinear Dynamics covers all levels of user application and throughput, from the user-guided Phoretix two-dimensional approach, when working with a small number of gels, to the automatic processing of large numbers of gels with minimal user intervention with Progenesis. Integration of the analysis software with powerful database components allows advanced gel comparisons and data mining to be performed with statistical verification of the results. Spot pick lists can be quickly created and automatically linked to a number of commercially available spot picking robots further increasing the support for proteomics research. The importance of image analysis for accurate, reliable and meaningful results will be discussed. Recent advances in development, with particular attention placed on the impact of noise contamination within gels, are illustrated and how the Progenesis product from Nonlinear Dynamics can be utilized to get the most from two-dimensional gel electrophoresis is shown.

Electrophoresis, Gel, Two-Dimensional↗

Loss of detoxification in inflammatory bowel disease: dysregulation of pregnane X receptor target genes.

BACKGROUND & AIMS: Phase 1, phase 2, and cellular efflux transporters are critical components in intestinal barrier function against xenobiotics and bacteria. We therefore performed global gene expression profiling in patients with ulcerative colitis (UC) and Crohn's disease as well as control specimens, with a special emphasis on genes involved in detoxification and epithelial membrane integrity. METHODS: Mucosal biopsy specimens from nonaffected regions of the colon and the terminal ileum were subjected to DNA microarray analysis and pathway-related data mining. Real-time reverse-transcription polymerase chain reaction was used for verification of selected regulated candidate genes in larger inflammatory bowel disease sample numbers and intestinal cell lines. RESULTS: Several dysregulated genes were identified in both disease groups and tissues. A set of genes coordinately down-regulated in the colon of patients with UC was composed of cellular detoxification and defense genes, which are target genes for the transcription factor pregnane X receptor (PXR). Messenger RNA expression of ABCB1 (MDR1) and PXR was significantly reduced in the colon of patients with UC but was unaffected in patients with Crohn's disease. In contrast to some of its target genes, the expression of PXR was not sensitive to tumor necrosis factor alpha stimulation of intestinal cell lines. CONCLUSIONS: A disease- and tissue-specific decrease in the expression of detoxification enzymes and ABC transporters was observed, which may be explained by a loss of PXR expression. Thus, dysregulation of xenobiotic metabolism and PXR activity in the gut is likely to contribute to the pathophysiology of UC.

ATP-Binding Cassette Transporters↗