Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

The multiple outcomes bias in antidepressants research.

Despite the widespread use of antidepressant medication, there are no signs that the burden of depression and suicide is decreasing in the industrialised world. This is generating mounting scepticism on the effectiveness of this class of drugs as an approach for the treatment of mood disorders. These doubts are also fuelled by the increasing awareness that the literature on antidepressants is fundamentally flawed and under the control of the pharmaceutical companies. This article describes systematically for the first time what is probably the most insidious and misleading of the biases that affect this area of research: the "multiple outcomes bias". Most trials on the effectiveness of antidepressants, instead of first establishing a hypothesis and then trying to demonstrate it, following the scientific method, start instead "data mining", without a clear hypothesis, and then select for publication, amongst a multitude of outcomes, only the ones that favour the antidepressant drug, ignoring the others. This method has obviously no scientific validity and is very misleading, allowing the manipulation of the data without any overt fraudulent action. There is the need to generate new research, independently funded and with clear hypotheses established "a priori ". What is at stake is not only the appraisal of the balance between benefits and potential damage to the patients when using this class of medications, after the realisation that they are not as harmless as believed. It is also to establish whether the research on antidepressant medication has gone on a "wild goose chase" over the last half century, concentrating almost exclusively on molecules that modify the monoaminergic transmission at synaptic level and virtually ignoring any other avenue.

Antidepressive Agents↗

Genomics, evolution and biological functions of the pacifastin peptide family: a conserved serine protease inhibitor family in arthropods.

The last decade, a new serine protease inhibitor family has been described in arthropods. Eight members were purified from the locusts Locusta migratoria (LMPI-1-2 and HI) and Schistocerca gregaria (SGPI-1-5). The light chain of the heterodimeric protease inhibitor pacifastin, from the freshwater crayfish Pacifastacus leniusculus, was found to be composed of nine consecutive inhibitory domains (PLDs). These domains share a pattern of six conserved cysteine residues (Cys-Xaa(9-12)-Cys-Asn-Xaa-Cys-Xaa-Cys-Xaa(2-3)-Gly-Xaa(3-6)-Cys-Thr-Xaa(3)-Cys) with the locust inhibitors. Via cDNA cloning, eight pacifastin-related precursors have been identified in locusts. Interestingly, additional pacifastin-related precursors have been identified in Diptera, Lepidoptera and Coleoptera utilising an in silico data mining approach.

Amino Acid Sequence↗

The complete human olfactory subgenome.

Olfactory receptors likely constitute the largest gene superfamily in the vertebrate genome. Here we present the nearly complete human olfactory subgenome elucidated by mining the genome draft with gene discovery algorithms. Over 900 olfactory receptor genes and pseudogenes (ORs) were identified, two-thirds of which were not annotated previously. The number of extrapolated ORs is in good agreement with previous theoretical predictions. The sequence of at least 63% of the ORs is disrupted by what appears to be a random process of pseudogene formation. ORs constitute 17 gene families, 4 of which contain more than 100 members each. "Fish-like" Class I ORs, previously considered a relic in higher tetrapods, constitute as much as 10% of the human repertoire, all in one large cluster on chromosome 11. Their lower pseudogene fraction suggests a functional significance. ORs are disposed on all human chromosomes except 20 and Y, and nearly 80% are found in clusters of 6-138 genes. A novel comparative cluster analysis was used to trace the evolutionary path that may have led to OR proliferation and diversification throughout the genome. The results of this analysis suggest the following genome expansion history: first, the generation of a "tetrapod-specific" Class II OR cluster on chromosome 11 by local duplication, then a single-step duplication of this cluster to chromosome 1, and finally an avalanche of duplication events out of chromosome 1 to most other chromosomes. The results of the data mining and characterization of ORs can be accessed at the Human Olfactory Receptor Data Exploratorium Web site (http://bioinfo.weizmann.ac.il/HORDE).

Chromosome Mapping↗

Whole-proteome interaction mining.

MOTIVATION: A major post-genomic scientific and technological pursuit is to describe the functions performed by the proteins encoded by the genome. One strategy is to first identify the protein-protein interactions in a proteome, then determine pathways and overall structure relating these interactions, and finally to statistically infer functional roles of individual proteins. Although huge amounts of genomic data are at hand, current experimental protein interaction assays must overcome technical problems to scale-up for high-throughput analysis. In the meantime, bioinformatics approaches may help bridge the information gap required for inference of protein function. In this paper, a previously described data mining approach to prediction of protein-protein interactions (Bock and Gough, 2001, Bioinformatics, 17, 455-460) is extended to interaction mining on a proteome-wide scale. An algorithm (the phylogenetic bootstrap) is introduced, which suggests traversal of a phenogram, interleaving rounds of computation and experiment, to develop a knowledge base of protein interactions in genetically-similar organisms. RESULTS: The interaction mining approach was demonstrated by building a learning system based on 1,039 experimentally validated protein-protein interactions in the human gastric bacterium Helicobacter pylori. An estimate of the generalization performance of the classifier was derived from 10-fold cross-validation, which indicated expected upper bounds on precision of 80% and sensitivity of 69% when applied to related organisms. One such organism is the enteric pathogen Campylobacter jejuni, in which comprehensive machine learning prediction of all possible pairwise protein-protein interactions was performed. The resulting network of interactions shares an average protein connectivity characteristic in common with previous investigations reported in the literature, offering strong evidence supporting the biological feasibility of the hypothesized map. For inferences about complete proteomes in which the number of pairwise non-interactions is expected to be much larger than the number of actual interactions, we anticipate that the sensitivity will remain the same but precision may decrease. We present specific biological examples of two subnetworks of protein-protein interactions in C. jejuni resulting from the application of this approach, including elements of a two-component signal transduction systems for thermoregulation, and a ferritin uptake network.

Algorithms↗

Isolation, expression pattern of a novel human RAB gene RAB41 and characterization of its intronless homolog RAB41P.

Small GTPases form a big family including Ras, Rho, Rac, Rab, Sar1/Arf subfamilies and Ran homologs, playing important roles in diverse cellular processes. Through data mining, a novel human RAB41 gene was predicted and subsequently isolated from human testis. The open reading frame of RAB41 is 636bp in length. RAB41 was composed of three exons, and it was mapped to chromosome 3q21.3 by comparing with the human genomic data. RAB41 protein contains a RAB domain. Similarity analysis indicated that RAB41 is closely similar to RAB19. The results of PCR amplification indicated that human RAB41 is widely expressed in brain, testis, lung, heart, ovary, colon, kidney, uterus and spleen but not in liver. By genomic searching, an intronless pseudogene homologous to RAB41 at chromosome 16--RAB41P was identified at human chromosome 16q11.2. Flanking the pseudogene RAB41P in human genome, there exist some transposable elements LINEs--L1, whose contribution in the generation of this intronless homolog is discussed.

Amino Acid Sequence↗

Neuropeptides and their precursors in the fruitfly, Drosophila melanogaster.

Neuropeptides form the most diverse class of chemical messenger molecules in metazoan nervous systems. They are usually generated from biosynthetic precursor polypeptides by enzymatic processing and modification. Many different peptides belonging to a number of distinct neuropeptide families have already been characterized from various insect species. The Drosophila Genome Sequencing Project has important implications for the future of neurobiological research. This paper describes the discovery of several new fruitfly neuropeptides by an in silico data mining approach. In addition, the state-of-the-art of Drosophila peptide research is reviewed.

Amino Acid Sequence↗

InfoEvolve: moving from data to knowledge using information theory and genetic algorithms.

InfoEvolve is a unified suite of data mining and empirical modeling tools capable of discovering low-bias and low-variance solutions to complex processes. The method is based on a common set of principles involving information theory and genetic algorithms. InfoEvolve can also discover multiple strategies embedded in complex data sets for achieving a desired target or goal. This latter aspect may prove to be very useful in drug design. The paper analyzes the following: InfoEvolve from a theoretical standpoint; a conceptual overview of InfoEvolve with a short description of the modeling method; the method using the example of homogeneous identification of DNA from an analysis of its melting curve behavior; and key learnings and additional applications of the technology for both drug design and genome analysis.

Algorithms↗

G2D: a tool for mining genes associated with disease.

BACKGROUND: Human inherited diseases can be associated by genetic linkage with one or more genomic regions. The availability of the complete sequence of the human genome allows examining those locations for an associated gene. We previously developed an algorithm to prioritize genes on a chromosomal region according to their possible relation to an inherited disease using a combination of data mining on biomedical databases and gene sequence analysis. RESULTS: We have implemented this method as a web application in our site G2D (Genes to Diseases). It allows users to inspect any region of the human genome to find candidate genes related to a genetic disease of their interest. In addition, the G2D server includes pre-computed analyses of candidate genes for 552 linked monogenic diseases without an associated gene, and the analysis of 18 asthma loci. CONCLUSION: G2D can be publicly accessed at http://www.ogic.ca/projects/g2d_2/.

Algorithms↗

An open source protein gel documentation system for proteome analyses.

Data organization and data mining represents one of the main challenges for modern high throughput technologies in pharmaceutical chemistry and medical chemistry. The presented open source documentation and analysis system provides an integrated solution (tutorial, setup protocol, sources, executables) aimed at substituting the traditionally used lab-book. The data management solution provided incorporates detailed information about the processing of the gels and the experimental conditions used and includes basic data analysis facilities which can be easily extended. The sample database and User-Interface are available free of charge under the GNU license from http://webber.physik.uni-freiburg.de/~fallerd/tutorial.htm.

Documentation↗

Statistical methods for analyzing tissue microarray data.

Tissue microarrays (TMAs) are a new high-throughput tool for the study of protein expression patterns in tissues and are increasingly used to evaluate the diagnostic and prognostic importance of biomarkers. TMA data are rather challenging to analyze. Covariates are highly skewed, non-normal, and may be highly correlated. We present statistical methods for relating TMA data to censored time-to-event data. We review methods for evaluating the predictive power of Cox regression models and show how to test whether biomarker data contain predictive information above and beyond standard pathology covariates. We use nonparametric bootstrap methods to validate model fitting indices such as the concordance index. We also present data mining methods for characterizing high risk patients with simple biomarker rules. Since researchers in the TMA community routinely dichotomize biomarker expression values, survival trees are a natural choice. We also use bump hunting (patient rule induction method), which we adapt to the use with survival data. The proposed methods are applied to a kidney cancer tissue microarray data set.

Algorithms↗

A strategy for primary high throughput cytotoxicity screening in pharmaceutical toxicology.

PURPOSE: Recent advances in combinatorial chemistry and high throughput screens for pharmacologic activity have created an increasing demand for in vitro high throughput screens for toxicological evaluation in the early phases of drug discovery. METHODS: To develop a strategy for such a screen, we have conducted a data mining study of the National Cancer Institute's Developmental Therapeutics Program (DTP) cytotoxicity database. RESULTS: Using hierarchical cluster analysis, we confirmed that the different tissues of origin and individual cell lines showed differential sensitivity to compounds in the DTP Standard Agents database. Surprisingly, however, approaching the data globally, linear regression analysis showed that the differences were relatively minor. Comparison with the literature on acute toxicity in mice showed that the predictive power of growth inhibition was marginally superior to that of cell death. CONCLUSIONS: This datamining study suggests that in designing a strategy for high throughput cytotoxicity screening: a single cell line, the choice of which may not be critical, can be used as a primary screen; a single end point may be an adequate measure and a cut off value for 50% growth inhibition between 10(-6) and 10(-8) M may be a reasonable starting point for accepting a cytotoxic compound for scale up and further study.

Animals↗

Linking the growth inhibition response from the National Cancer Institute's anticancer screen to gene expression levels and other molecular target data.

MOTIVATION: Data mining tools are proposed to establish mechanistic connections between chemotypes and specific cellular functions. Drawing on a previous study that classified the cellular response patterns of growth inhibition measurements log( GI(50)) from the National Cancer Institute's (NCI's) anticancer screen, we have examined additional data for mRNA expression, sets of known molecular targets and mutational status against these same tumor cell lines to relate chemosensitivity more precisely to biochemical pathways. RESULTS: Our analysis finds that gene expression levels do not, in general, correlate with log(GI(50)) measurements, instead they reflect a generic toxic condition. Within the remaining set of non-generic conditions, examples were found where a correlation suggesting a biochemical basis for cellular cytotoxicity could be supported. These included reconfirmation of previously observed associations between mutant and wild-type status of p53, and chemosensitivity to alkylating agents, while extending these results to reveal associations with gamma-induced expressions of MDM2, WAF1 and GADD45, signals that were not apparent in measurements of basal mRNA expression levels for any of these genes. Additional examinations revealed that mRNA expression levels directly correlated with paclitaxel chemosensitivity to mitosis, while also identifying additional chemotypes as P-glycoprotein substrates. Our analysis revealed well-known direct associations between p16 mutant status and chemotypes implicated in cell cycle control, and extended these results to include expression levels for three additional tyrosine kinase proteins (TEK, transgelin and hCdc4). Links were also found that suggested associations between chemosensitivity and the endocrine, paracrine ligand-receptor loops, via expression of the adrenergic receptor, calcium second messenger pathways via expression levels of carbonic anhydrase and cellular communication pathways via fibrillin.

Antineoplastic Agents↗

VirGen: a comprehensive viral genome resource.

VirGen is a comprehensive viral genome resource that organizes the 'sequence space' of viral genomes in a structured fashion. It has been developed with the objective of serving as an annotated and curated database comprising complete genome sequences of viruses, value-added derived data and data mining tools. The current release (v1.1) contains 559 complete genomes in addition to 287 putative genomes of viruses belonging to eight viral families for which the host range includes animals and plants. Viral genomes in VirGen are annotated using sequence-based Bioinformatics approaches. The genomic data is also curated to identify 'alternate names' of viral proteins, where available. VirGen archives the results of comparisons of genomes, proteomes and individual proteins within and between viral species. It is the first resource to provide phylogenetic trees of viral species computed using whole-genome sequence data. The module of predicted B-cell antigenic determinants in VirGen is an attempt to link the genome to its vaccinome. Comparative genome analysis data facilitate the study of genome organization and evolution of viruses, which would have implications in applied research to identify candidates for the design of vaccines and antiviral drugs. VirGen is a relational database and is available at http://bioinfo. ernet.in/virgen/virgen.html.

Antigens, Viral↗

Computational cluster validation in post-genomic data analysis.

MOTIVATION: The discovery of novel biological knowledge from the ab initio analysis of post-genomic data relies upon the use of unsupervised processing methods, in particular clustering techniques. Much recent research in bioinformatics has therefore been focused on the transfer of clustering methods introduced in other scientific fields and on the development of novel algorithms specifically designed to tackle the challenges posed by post-genomic data. The partitions returned by a clustering algorithm are commonly validated using visual inspection and concordance with prior biological knowledge--whether the clusters actually correspond to the real structure in the data is somewhat less frequently considered. Suitable computational cluster validation techniques are available in the general data-mining literature, but have been given only a fraction of the same attention in bioinformatics. RESULTS: This review paper aims to familiarize the reader with the battery of techniques available for the validation of clustering results, with a particular focus on their application to post-genomic data analysis. Synthetic and real biological datasets are used to demonstrate the benefits, and also some of the perils, of analytical clustervalidation. AVAILABILITY: The software used in the experiments is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/. SUPPLEMENTARY INFORMATION: Enlarged colour plots are provided in the Supplementary Material, which is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/.

Algorithms↗

Automated detection of informative combined effects in genetic association studies of complex traits.

There is a growing body of evidence suggesting that the relationships between gene variability and common disease are more complex than initially thought and require the exploration of the whole polymorphism of candidate genes as well as several genes belonging to biological pathways. When the number of polymorphisms is relatively large and the structure of the relationships among them complex, the use of data mining tools to extract the relevant information is a necessity. Here, we propose an automated method for the detection of informative combined effects (DICE) among several polymorphisms (and nongenetic covariates) within the framework of association studies. The algorithm combines the advantages of the regressive approaches with those of data exploration tools. Importantly, DICE considers the problem of interaction between polymorphisms as an effect of interest and not as a nuisance effect. We illustrate the method with three applications on the relationship between (1). the P-selectin gene and myocardial infarction, (2). the cholesteryl ester transfer protein gene and plasma high-density-lipoprotein cholesterol concentration, and (3). genes of the renin-angiotensin-aldosterone system and myocardial infarction. The applications demonstrated that the method was able to recover results already found using other approaches, but in addition detected biologically sensible effects not previously described.

Algorithms↗

Geovisualization to support the exploration of large health and demographic survey data.

BACKGROUND: Survey data are increasingly abundant from many international projects and national statistics. They are generally comprehensive and cover local, regional as well as national levels census in many domains including health, demography, human development, and economy. These surveys result in several hundred indicators. Geographical analysis of such large amount of data is often a difficult task and searching for patterns is particularly a difficult challenge. Geovisualization research is increasingly dealing with the exploration of patterns and relationships in such large datasets for understanding underlying geographical processes. One of the attempts has been to use Artificial Neural Networks as a technology especially useful in situations where the numbers are vast and the relationships are often unclear or even hidden. RESULTS: We investigate ways to integrate computational analysis based on a Self-Organizing Map neural network, with visual representations of derived structures and patterns in a framework for exploratory visualization to support visual data mining and knowledge discovery. The framework suggests ways to explore the general structure of the dataset in its multidimensional space in order to provide clues for further exploration of correlations and relationships. CONCLUSION: In this paper, the proposed framework is used to explore a demographic and health survey data. Several graphical representations (information spaces) are used to depict the general structure and clustering of the data and get insight about the relationships among the different variables. Detail exploration of correlations and relationships among the attributes is provided. Results of the analysis are also presented in maps and other graphics.

Journal Article↗

Metrics in the science of surge.

Metrics are the driver to positive change toward better patient care. However, the research into the metrics of the science of surge is incomplete, research funding is inadequate, and we lack a criterion standard metric for identifying and quantifying surge capacity. Therefore, a consensus working group was formed through a "viral invitation" process. With a combination of online discussion through a group e-mail list and in-person discussion at a breakout session of the Academic Emergency Medicine 2006 Consensus Conference, "The Science of Surge," seven consensus statements were generated. These statements emphasize the importance of funded research in the area of surge capacity metrics; the utility of an emergency medicine research registry; the need to make the data available to clinicians, administrators, public health officials, and internal and external systems; the importance of real-time data, data standards, and electronic transmission; seamless integration of data capture into the care process; the value of having data available from a single point of access through which data mining, forecasting, and modeling can be performed; and the basic necessity of a criterion standard metric for quantifying surge capacity. Further consensus work is needed to select a criterion standard metric for quantifying surge capacity. These consensus statements cover the future research needs, the infrastructure needs, and the data that are needed for a state-of-the-art approach to surge and surge capacity.

Consensus↗

MS1, MS2, and SQT-three unified, compact, and easily parsed file formats for the storage of shotgun proteomic spectra and identifications.

As the speed with which proteomic labs generate data increases along with the scale of projects they are undertaking, the resulting data storage and data processing problems will continue to challenge computational resources. This is especially true for shotgun proteomic techniques that can generate tens of thousands of spectra per instrument each day. One design factor leading to many of these problems is caused by storing spectra and the database identifications for a given spectrum as individual files. While these problems can be addressed by storing all of the spectra and search results in large relational databases, the infrastructure to implement such a strategy can be beyond the means of academic labs. We report here a series of unified text file formats for storing spectral data (MS1 and MS2) and search results (SQT) that are compact, easily parsed by both machine and humans, and yet flexible enough to be coupled with new algorithms and data-mining strategies.

Database Management Systems↗