Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Proteomic fingerprinting for the diagnosis of human African trypanosomiasis.

Papadopoulos et al. recently reported the discovery of a diagnostic serum proteomic signature for human African trypanosomiasis (HAT), using a combination of surface-enhanced laser desorption-ionization time-of-flight (SELDI-TOF) mass spectrometry and data-mining algorithms. This novel approach, coupled with biochemical characterization of the proteins that contribute to the signature, provides powerful new tools for the development of improved diagnostic tests, disease staging and identification of potential novel drug targets in HAT.

Algorithms↗

Bioinformatics: harvesting information for plant and crop science.

Bioinformatics is an integral aspect of plant and crop science research. Developments in data management and analytical software are reviewed with an emphasis on applications in functional genomics. This includes information resources for Arabidopsis and crop species, and tools available for analysis and visualisation of comparative genomic data. Approaches used to explore relationships between plant genes and expressed sequences are compared, including use of ontologies. The impact of bioinformatics in forward and reverse genetics is described, together with the potential from data mining. The role of bioinformatics is explored in the wider context of plant and crop science.

Algorithms↗

The latest phospholipase C, PLCeta, is implicated in neuronal function.

Members of the phosphoinositide-specific phospholipase C (PLC) family have key roles in cell signalling. In response to many extracellular stimuli, such as hormones, neurotransmitters, antigens and growth factors, PLCs catalyse the hydrolysis of phosphatidylinositol (4,5)-bisphosphate [PtdIns(4,5)P(2)], thereby generating two well-established second messengers, inositol (1,4,5)-trisphosphate and diacylglycerol. Eleven PLC isozymes encoded by different genes have been identified in mammals and, on the basis of their structure and sequence relationships, have been classified into five families designated PLCbeta (1-4), PLCgamma (1 and 2), PLCdelta (1, 3 and 4), PLCepsilon (1) and PLCzeta (1). All PLCs contain the catalytic X and Y domain, in addition to other regulatory domains including the C2 domain and the EF-hand domain. In 2005, four groups independently identified an entirely new family of PLCs--eta1 and eta2--using data mining of mammalian genomes. The properties of the PLCeta enzyme suggest that it might act as a Ca(2+) sensor, in particular, functioning during formation and maintenance of the neuronal network in the postnatal brain.

Animals↗

Techniques: Bioprospecting historical herbal texts by hunting for new leads in old tomes.

Ethnobotany has led to the identification of novel pharmacological agents but many challenges to using ethnobotany as a research tool remain. In particular, the loss of traditional knowledge together with the advent of high-throughput screening has made ethnobotanical techniques laborious and potentially unnecessary. However, historical herbal texts provide a preexisting resource that documents the traditional uses of various species as medicines. As generational losses of traditional knowledge accrue, these herbal texts become increasingly valuable. The methodology for extracting useful information contained within these resources had been cumbersome and consuming. However, the application of new bioinformatics data-mining systems to herbal texts holds great promise for identifying novel pharmacotherapeutic leads for bioactive compounds.

Ethnobotany↗

Construction, database integration, and application of an Oenothera EST library.

Coevolution of cellular genetic compartments is a fundamental aspect in eukaryotic genome evolution that becomes apparent in serious developmental disturbances after interspecific organelle exchanges. The genus Oenothera represents a unique, at present the only available, resource to study the role of the compartmentalized plant genome in diversification of populations and speciation processes. An integrated approach involving cDNA cloning, EST sequencing, and bioinformatic data mining was chosen using Oenothera elata with the genetic constitution nuclear genome AA with plastome type I. The Gene Ontology system grouped 1621 unique gene products into 17 different functional categories. Application of arrays generated from a selected fraction of ESTs revealed significantly differing expression profiles among closely related Oenothera species possessing the potential to generate fertile and incompatible plastid/nuclear hybrids (hybrid bleaching). Furthermore, the EST library provides a valuable source of PCR-based polymorphic molecular markers that are instrumental for genotyping and molecular mapping approaches.

Cell Nucleus↗

Intron length and accelerated 3' gene evolution.

Genetic evolution depends in part upon a balance between negative selection and environmentally driven mutation. To explore whether this balance is affected by gene structure, we have used phylogenetic data mining to compare gene compositions across a range of species. Here we show that genomes of higher species exhibit a greater frequency of 5' CpG islands and of CpG-->TpG/CpA transitions. This latter mutational pattern exhibits a 5'-to-3' trend in higher species, consistent with a length-dependent effect on methylation-dependent CpG suppression. Associated strand asymmetry (TpG>CpA) declines with gene length, implying attenuation of transcription-coupled repair 3' to introns. A sharp 3' rise in coding region single-nucleotide polymorphism frequency further supports a mechanistic role for intron length in promoting genetic variation by reducing repair and/or weakening negative selection. Consistent with this, the Ka/Ks ratio of 3' exons exceeds that of centrally located exons in intron-containing, but not in intronless, genes (p<0.0003). We conclude that the efficiency of transcription-coupled repair decreases with gene length, suggesting in turn that 3' gene evolution is accelerated both by introns and by gene methylation.

3' Untranslated Regions↗

Production of custom microarrays for neuroscience research.

Microarray chips produced by commercial vendors and academic laboratories are mostly generic in nature to facilitate wide applicability. With the sequencing of the human, mouse, and rat genomes, the thrust is to expand clone and oligonucleotide sets and increase the number of genes represented on a particular array. This is appropriate for discovery based investigations where microarray technology has been successfully utilized. However, array technology can also be employed to perform hypothesis based studies if optimized chips can be produced with relevant content. Existing array technology available at core facilities can be effectively utilized to produce a custom microarrays with genes that are most relevant to the research interests of individual investigators or research groups for use as a standard molecular tool. The power of this technology can be harnessed to further our understanding of specific biological problems without involvement in extensive data mining and analysis. The custom microarray approach is presented with procedural details for design and production in the context of neurobiological investigations.

Animals↗

Automated extraction of information in molecular biology.

We review data mining techniques in molecular biology, specifically those that extract information from the scientific literature itself. As more of the biological literature is published electronically, there is an opportunity, and even a need, to automatically summarize the literature in a customized way, for example by associating keywords to a topic. These keywords can be extracted from relevant publications. The process of keyword extraction can be automated and optimized to keep literature pointers automatically up-to-date or to filter relevant information from the literature. To illustrate these points, OMIM (Online Mendelian Inheritance in Man), a database of human inherited diseases, was linked to the literature and keywords were derived that covered distinct aspects such as genetic information on the one hand and disease-specific protein and phenotypic information on the other. They were used to extract information that is helpful for keeping entries about disease up-to-date.

Databases, Factual↗

Identification and mapping of small-molecule binding sites in proteins: computational tools for structure-based drug design.

The number of protein structures is currently increasing at an impressive rate. The growing wealth of data calls for methods to efficiently exploit structural information for medicinal and pharmaceutical purposes. Given the three-dimensional (3D) structure of a validated protein target, the identification of functionally relevant binding sites and the analysis ('mapping') of these sites with respect to molecular recognition properties are important initial tasks in structure-based drug design. To address these tasks, a variety of computational tools have been developed. Approaches to identify binding pockets include geometric analyses of protein surfaces, comparisons of protein structures, similarity searches in databases of protein cavities, and docking scans to reveal areas of high ligand complementarity. In the context of binding-site analysis, powerful data mining tools help to retrieve experimental information about related protein-ligand complexes. To identify interaction hot spots, various potential functions and knowledge-based approaches are available for mapping binding regions. The results may subsequently be used to guide virtual screenings for new ligands via pharmacophore searches or docking simulations.

Artificial Intelligence↗

Promising directions for the diagnosis and management of gynecological cancers.

Diagnosis and management of cancer requires tools with both high sensitivity and specificity. The minimally invasive cervical smear has demonstrated how a test, even one with low specificity, can change the public health profile of a cancer from a late stage deadly disease to early diagnosis with rare tumor-related deaths. The benefit of such a test is best demonstrated by the low frequency of cervix cancer and its good outcome in countries where this test is readily available and used with appropriate secondary follow up. Early and specific symptoms, and identification and prevention for high risk groups has had similar impact for endometrial cancer. Neither a robust test, nor reliable or specific early symptoms are available for ovarian cancer, making clinical and scientific advances in this area a critical world-wide need. Current approaches testing one protein or gene marker at a time will not address this crisis expeditiously. New sensitive, specific, accurate, and reliable technologies that can be implemented using high throughput mechanisms are needed at as low a cost as possible. Ideally, these technologies should be focused on readily available patient resources, such as blood or urine, or as in the case of cervix cancer, minimally invasive informative approaches such as cervical smears. Techniques that allow data mining from a large input database overcome the slow advances of one protein-one gene investigation, and further address the multi-faceted carcinogenesis process occurring even in germ line mutation-associated malignancy. Proteomics, the study of the cellular proteins and their activation states, has led the progress in biomarker development for ovarian and other cancers and is being applied to management assessment. Amenable to high throughput, internet interface, and representative of the proteome spectrum, proteomic technology is the newest and most promising direction for translational developments in gynecologic cancers.

Biomarkers, Tumor↗

Parasite genome initiatives.

During 1993-1994, scientists from developing and developed countries planned and initiated a number of parasite genome projects and several consortiums for the mapping and sequencing of these medium-sized genomes were established, often based on already ongoing scientific collaborations. Financial and other support came from WHO/TDR, Wellcome Trust and other funding agencies. Thus, the genomes of Plasmodium falciparum, Schistosoma mansoni, Trypanosoma cruzi, Leishmania major, Trypanosoma brucei, Brugia malayi and other pathogenic nematodes are now under study. From an initial phase of network formation, mapping efforts and resource building (EST, GSS, phage, cosmid, BAC and YAC library constructions), sequencing was initiated in gene discovery projects but soon also on a small chromosome, and now on a fully fledged genome scale. Proteomics, functional analysis, genetic manipulation and microarray analysis are ongoing to different degrees in the respective genome initiatives, and as the funding for the whole genome sequencing becomes secured, most of the participating laboratories, apart from larger sequencing centres, become oriented to post-genomics. Bioinformatics networks are being expanded, including in developing countries, for data mining, annotation and in-depth analysis.

Animals↗

Relibase: design and development of a database for comprehensive analysis of protein-ligand interactions.

Knowledge discovery from the exponentially growing body of structurally characterised protein-ligand complexes as a source of information in structure-based drug design is a major challenge in contemporary drug research. Given the need for powerful data retrieval, integration and analysis tools, Relibase was developed as a database system particularly designed to handle protein-ligand related problems and tasks. Here, we describe the design and functionality of the Relibase core database system. Features of Relibase include, e.g. the detailed analysis of superimposed ligand binding sites, ligand similarity and substructure searches, and 3D searches for protein-ligand and protein-protein interaction patterns. The broad range of functions provided in Relibase and its high level of data integration, along with its flexible and intuitive interface, makes Relibase an invaluable data mining tool which can significantly enhance the drug development process. An example, illustrating a 3D query for quarternary ligand nitrogen atoms interacting with aromatic ring systems in proteins, a pattern found in pharmaceutically relevant target proteins such as, e.g. acetylcholine-esterase, is discussed.

Animals↗

Microarray analysis of bacterial pathogenicity.

The DNA microarray, a surface that contains an ordered arrangement of each identified open reading frame of a sequenced genome, is the engine of functional genomics. Its output, the expression profile, provides a genome wide snap-shot of the transcriptome. Refined by array-specific statistical instruments and data-mined by clustering algorithms and metabolic pathway databases, the expression profile discloses, at the transcriptional level, how the microbe adapts to new conditions of growth--the regulatory networks that govern the adaptive response and the metabolic and biosynthetic pathways that effect the new phenotype. Adaptation to host microenvironments underlies the capacity of infectious agents to persist in and damage host tissues. While monitoring the whole genome transcriptional response of bacterial pathogens within infected tissues has not been achieved, it is likely that the complex, tissue-specific response is but the sum of individual responses of the bacteria to specific physicochemical features that characterize the host milieu. These are amenable to experimentation in vitro and whole-genome expression studies of this kind have defined the transcriptional response to iron starvation, low oxygen, acid pH, quorum-sensing pheromones and reactive oxygen intermediates. These have disclosed new information about even well-studied processes and provide a portrait of the adapting bacterium as a 'system', rather than the product of a few genes or even a few regulons. Amongst the regulated genes that compose this adaptive system are transcription factors. Expression profiling experiments of transcription factor mutants delineate the corresponding regulatory cascade. The genetic basis for pathogenicity can also be studied by using microarray-based comparative genomics to characterize and quantify the extent of genetic variability within natural populations at the gene level of resolution. Also identified are differences between pathogen and commensal that point to possible virulence determinants or disclose evolutionary history. The host vigorously engages the pathogen; expression studies using host genome microarrays and bacterially infected cell cultures show that the initial host reaction is dominated by the innate immune response. However, within the complex expression profile of the host cell are components mediated by pathogen-specific determinants. In the future, the combined use of bacterial and host microarrays to study the same infected tissue will reveal the dialogue between pathogen and host in a gene-by-gene and site- and time-specific manner. Translating this conversation will not be easy and will probably require a combination of powerful bioinformatic tools and traditional experimental approaches--and considerable effort and time.

Animals↗

TM4 microarray software suite.

Powerful specialized software is essential for managing, quantifying, and ultimately deriving scientific insight from results of a microarray experiment. We have developed a suite of software applications, known as TM4, to support such gene expression studies. The suite consists of open-source tools for data management and reporting, image analysis, normalization and pipeline control, and data mining and visualization. An integrated MIAME-compliant MySQL database is included. This chapter describes each component of the suite and includes a sample analysis walk-through.

Algorithms↗

Non-equilibrium proteins.

There exist no methodical studies concerning non-equilibrium systems in cellular biology. This paper is an attempt to partially fill this shortcoming. We have undertaken an extensive data-mining operation in the existing scientific literature to find scattered information about non-equilibrium subcellular systems, in particular concerning fast proteins, i.e. those with short turnover half-time. We have advanced the hypothesis that functionality in fast proteins emerges as a consequence of their intrinsic physical instability that arises due to conformational strains resulting from co-translational folding (the interdependence between chain elongation and chain folding during biosynthesis on ribosomes). Such intrinsic physical instability, a kind of conformon (Klonowski-Klonowska conformon, according to Ji, (Molecular Theories of Cell Life and Death, Rutgers University Press, New Brunswick, 1991)) is probably the most important feature determining functionality and timing in these proteins. If our hypothesis is true, the turnover half-time of fast proteins should be positively correlated with their molecular weight, and some experimental results (Ames et al., J. Neurochem. 35 (1980) 131) indeed demonstrated such a correlation. Once the native structure (and function) of a fast protein macromolecule is lost, it may not be recovered--denaturation of such proteins will always be irreversible; therefore, we searched for information on irreversible denaturation. Only simulation and modeling of protein co-translational folding may answer the questions concerning fast proteins (Ruggiero and Sacile, Med. Biol. Eng. Comp. 37 (Suppl. 1) (1999) 363). Non-equilibrium structures may also be built up of protein subunits, even if each one taken by itself is in thermodynamic equilibrium (oligomeric proteins; sub-cellular sol-gel dissipative network structures).

Amino Acid Sequence↗

Medical target prediction from genome sequence: combining different sequence analysis algorithms with expert knowledge and input from artificial intelligence approaches.

By exploiting the rapid increase in available sequence data, the definition of medically relevant protein targets has been improved by a combination of: (i) differential genome analysis (target list): and (ii) analysis of individual proteins (target analysis). Fast sequence comparisons, data mining, and genetic algorithms further promote these procedures. Mycobacterium tuberculosis proteins were chosen as applied examples.

Algorithms↗

Selective serotonin reuptake inhibitors in pregnant women and neonatal withdrawal syndrome: a database analysis.

BACKGROUND: Selective serotonin reuptake inhibitors (SSRIs) have been associated with withdrawal symptoms. We investigated whether use of these drugs in pregnant women might cause neonatal withdrawal syndrome. METHODS: An association between paroxetine and neonatal convulsions was identified in December, 2001, by the data mining method routinely used to screen the WHO database of adverse drug reactions. An information component (IC) measure was used to screen for unexpected adverse reactions relative to the information in the database. We then assessed cases of neonatal convulsions and neonatal withdrawal syndrome associated with drugs included in the anatomical therapeutic chemical groups N06AB and N06AX. FINDINGS: By November, 2003, a total of 93 suspected cases of SSRI-induced neonatal withdrawal syndrome had been reported, and were regarded as enough information to confirm a possible causal relation. 64 of the cases were associated with paroxetine, 14 with fluoxetine, nine with sertraline, and seven with citalopram. The IC-2 SD for the group became greater than 0 in the first quarter of 1991, and the IC increased to 2.68 (IC-2 SD 0.32) by the second quarter of 2003. For each individual compound, the IC-2 SD was greater than 0. INTERPRETATION: SSRIs, especially paroxetine, should be cautiously managed in the treatment of pregnant women with a psychiatric disorder.

Antidepressive Agents, Second-Generation↗

Association of asthma therapy and Churg-Strauss syndrome: an analysis of postmarketing surveillance data.

BACKGROUND: Churg-Strauss syndrome (CSS), also known as allergic granulomatous angiitis (AGA), is a rare vasculitis that occurs in patients with bronchial asthma. The nature of the association of CSS with various asthma therapies is unclear. OBJECTIVE: This study investigated the associations of different multidrug asthma therapy regimens and the reporting of AGA (the preferred code for CSS in the coding dictionary for the Adverse Event Reporting System [AERS]) by applying an iterative method of disproportionally analysis to th AERS database maintained by the US Food and Drug Administration. METHODS: The public-release version of the AERS database was used to identify reports of AGA in patients receiving asthma therapy. Reporting of AGA was examined using iterative disproportionality methods in patients receiving > or =1 of the following drug classes: inhaled corticosteroid (ICS), leukotriene receptor antagonist (LTRA), short-acting beta(2)-agonist (SABA), or long-acting beta(2)-agonist (LABA). The Bayesian data-mining algorithm known as the multi-item gamma poisson shrinker was used to determine the relative reporting rates by calculation of the empirical Bayes geometric mean (EBGM) and its 90% CI (EB05 = lower limit and EB95 = upper limit) for each drug. Subset analyses were performed for each drug with different medication combinations to differentiate the relative reporting of AGA for each. RESULTS: A strong association was found between LTRA use and AGA (EBGM = 104.0, EB05 = 95.0, EB95 = 113.8) that persisted with all combinations of therapy studied. AGA was also associated with the ICS, SABA and LABA classes (EBGM values of 27.8, 14.6 and 40.4, respectively). However, the latter associations were mostly dependent on the presence of concurrent LTRA and, to a lesser extemt, oral corticosteroid therapy and became negligible (ie, EB05 < 2) for patients who were not receiving these concurrent treatments. CONCLUSIONS: Differences based on relative reporting were observed in the patterns of association of AGA with LTRA, ICS, and beta(2)-agonist therapies. A strong association between LTRA use and AGA was present regardless of the use of other asthma drugs.

Administration, Inhalation↗