Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

Quantitative methods in pharmacovigilance: focus on signal detection.

Pharmacovigilance serves to detect previously unrecognised adverse events associated with the use of medicines. The simplest method for detecting signals of such events is crude inspection of lists of spontaneously reported drug-event combinations. Quantitative and automated numerator-based methods such as Bayesian data mining can supplement or supplant these methods. The theoretical basis and limitations of these methods should be understood by drug safety professionals, and automated methods should not be automatically accepted. Published evaluations of these techniques are mainly limited to large regulatory databases, and performance characteristics may differ in smaller safety databases of drug developers. Head-to-head comparisons of the major techniques have not been published. Regardless of previous statistical training, pharmacovigilance practitioners should understand how these methods work. The mathematical basis of these techniques should not obscure the numerous confounders and biases inherent in the data. This article seeks to make automated signal detection methods transparent to drug safety professionals of various backgrounds. This is accomplished by first providing a brief overview of the evolution of signal detection followed by a series of sections devoted to the methods with the greatest utilisation and evidentiary support: proportional reporting rations, the Bayesian Confidence Propagation Neural Network and empirical Bayes screening. Sophisticated yet intuitive explanations are provided for each method, supported by figures in which the underlying statistical concepts are explored. Finally the strengths, limitations, pitfalls and outstanding unresolved issues are discussed. Pharmacovigilance specialists should not be intimidated by the mathematics. Understanding the theoretical basis of these methods should enhance the effective assessment and possible implementation of these techniques by drug safety professionals.

Adverse Drug Reaction Reporting Systems↗

Computing motif correlations in proteins.

Protein motifs, which are specific regions and conserved regions, are found by comparing multiple protein sequences. These conserved regions in general play an important role in protein functions and protein folds, for example, for their binding properties or enzymatic activities. The aim here is to find the existence correlations of protein motifs. The knowledge of protein motif/domain sharing should be important in shedding new light on the biologic functions of proteins and offering a basis in analyzing the evolution in the human genome or other genomes. The protein sequences used here are obtained from the PIR-NREF database and the protein motifs are retrieved from the PROSITE database. We apply data mining approach to discover the occurrence correlations of motif in protein sequences. The correlation of motifs mined can be used in evolution analyses and protein structure prediction. We discuss the latter, i.e., protein structure prediction in this study. The correlations mined are stored and maintained in a database system. The database is now available at http://bioinfo.csie.ncu.edu.tw/ProMotif/.

Algorithms↗

Molecular characterization and quantitative analysis of superoxide dismutases in virulent and avirulent strains of Aeromonas salmonicida subsp. salmonicida.

Aeromonas salmonicida subsp. salmonicida is a facultatively intracellular gram-negative bacterium that is the etiological agent of furunculosis, a bacterial septicemia of salmonids that causes significant economic loss to the salmon farming industry. The mechanisms by which A. salmonicida evades intracellular killing may be relevant in understanding virulence and the eventual design of appropriate treatment strategies for furunculosis. We have identified two open reading frames (ORFs) and related upstream sequences that code for two putative superoxide dismutases (SODs), sodA and sodB. The sodA gene encoded a protein of 204 amino acids with a molecular mass of approximately 23.0 kDa (SodA) that had high similarity to other prokaryotic Mn-SODs. The sodB gene encoded a protein of 194 amino acids with a molecular mass of approximately 22.3 kDa that had high similarity to other prokaryotic Fe-SODs. Two enzymes with activities consistent with both these ORFs were identified by inhibition of O(2)(-)-catalyzed tetrazolium salt reduction in both gels and microtiter plate assays. The two enzymes differed in their expression patterns in in vivo- and in vitro-cultured bacteria. The regulatory sequences upstream of putative sodA were consistent with these differences. We could not identify other SOD isozymes such as sodC either functionally or through data mining. Levels of SOD were significantly higher in virulent than in avirulent strains of A. salmonicida subsp. salmonicida strain A449 when cultured in vitro and in vivo.

Aeromonas↗

Affinity capillary electrophoresis for the screening of novel antimicrobial targets.

The increasing number of multiantibiotic-resistant organisms, including methicillin-resistant Staphylococcus aureus (MRSA), requires the development of novel chemotherapies that are structurally distinct and exempt from current resistance mechanisms. Bioinformatics data mining of microbial genomes has revealed numerous previously unexploited essential open reading frames (ORFs) of unknown biochemical function. The potential of these proteins as screening targets is not readily apparent because most screening technologies rely on knowledge of biological function. To address this problem, the authors employed affinity capillary electrophoresis (ACE) to identify antimicrobial compounds that bound the novel target YihA. Screening a small-molecule library of 44,000 compounds initially identified 115 binders, of which 76% were confirmed. Furthermore, the ACE assay distinguished diverse compounds that possessed drug-like properties and antimicrobial activity against drug-resistant clinical isolates. These data validate ACE as a valuable tool for the fast, efficient detection of specific binding molecules that possess biological activity.

Anti-Bacterial Agents↗

In vivo epinephrine-mediated regulation of gene expression in human skeletal muscle.

The stress hormone epinephrine produces major physiological effects on skeletal muscle. Here we determined skeletal muscle mRNA expression profiles before and during a 6-h epinephrine infusion performed in nine young men. Stringent statistical analysis of data obtained using 43000 cDNA element microarrays showed that 1206 and 474 genes were up- and down-regulated, respectively. Microarray data were validated using reverse transcription quantitative PCR. Gene classification was performed through data mining of Gene Ontology annotations, cluster analysis of regulated genes among 14 human tissues, and correlation analysis of mRNA and clinical parameter variations. Evidence of an autoregulatory control was provided by the regulation of key genes of the cAMP-dependent transcription pathway. Genes with known functional cAMP response elements were regulated by the hormone. The impact on metabolism was illustrated by coordinated regulations of genes involved in carbohydrate and protein metabolisms. Epinephrine had a profound effect on genes involved in immunity and inflammatory response, a previously unappreciated aspect of catecholamine action. Information on 526 mRNAs corresponded to genes of unknown function. These data define the molecular signatures of epinephrine action in human skeletal muscle. They may contribute to the understanding of skeletal muscle alterations observed in pathological conditions characterized by sympathetic nervous system overdrive.

Adrenergic Agonists↗

An updated physiologically based pharmacokinetic model for hexachlorobenzene: incorporation of pathophysiological states following partial hepatectomy and hexachlorobenzene treatment.

Physiologically based pharmacokinetic (PBPK) modeling is generally used for describing xenobiotic disposition in animals and humans with normal physiological conditions. We describe here an updated PBPK model for hexachlorobenzene (HCB) in male F344 rats with the incorporation of pathophysiological conditions. Two more features contribute to the distinctness of this model from the earlier published versions. This model took erythrocyte binding into account, and a particular elimination process of HCB, the plasma-to-gastrointestinal (GI) lumen passive diffusion (i.e., exsorption), was incorporated. Our PBPK model was developed using data mined from multiple pharmacokinetic studies in the literature, and then modified to simulate HCB disposition under the conditions of our integrated pharmacokinetics/liver foci bioassay. This model included plasma, erythrocytes, liver, fat, rapidly and slowly perfused compartments, and GI lumen. To account for the distinct characteristics of HCB absorption, the GI lumen was split into an upper and a lower part. HCB was eliminated through liver metabolism and the exsorption process. The pathophysiological changes after partial hepatectomy, such as alterations in the liver and body weights and fat volume, were incorporated in our model. With adjustment of the transluminal diffusion-related parameters, the model adequately described the data from the literature and our bioassay. Our PBPK model simulation suggests that HCB absorption and exsorption processes depend on exposure conditions; different exposure conditions dictate different absorption and exsorption rates. This model forms a foundation for our further exploration of the quantitative relationship between HCB exposure and development of preneoplastic liver foci.

Animals↗

Improved validation of peptide MS/MS assignments using spectral intensity prediction.

A major limitation in identifying peptides from complex mixtures by shotgun proteomics is the ability of search programs to accurately assign peptide sequences using mass spectrometric fragmentation spectra (MS/MS spectra). Manual analysis is used to assess borderline identifications; however, it is error-prone and time-consuming, and criteria for acceptance or rejection are not well defined. Here we report a Manual Analysis Emulator (MAE) program that evaluates results from search programs by implementing two commonly used criteria: 1) consistency of fragment ion intensities with predicted gas phase chemistry and 2) whether a high proportion of the ion intensity (proportion of ion current (PIC)) in the MS/MS spectra can be derived from the peptide sequence. To evaluate chemical plausibility, MAE utilizes similarity (Sim) scoring against theoretical spectra simulated by MassAnalyzer software (Zhang, Z. (2004) Prediction of low-energy collision-induced dissociation spectra of peptides. Anal. Chem. 76, 3908-3922) using known gas phase chemical mechanisms. The results show that Sim scores provide significantly greater discrimination between correct and incorrect search results than achieved by Sequest XCorr scoring or Mascot Mowse scoring, allowing reliable automated validation of borderline cases. To evaluate PIC, MAE simplifies the DTA text files summarizing the MS/MS spectra and applies heuristic rules to classify the fragment ions. MAE output also provides data mining functions, which are illustrated by using PIC to identify spectral chimeras, where two or more peptide ions were sequenced together, as well as cases where fragmentation chemistry is not well predicted.

Amino Acid Sequence↗

Developmental gene expression profiling of mammalian, fetal orofacial tissue.

BACKGROUND: The embryonic orofacial region is an excellent developmental paradigm that has revealed the centrality of numerous genes encoding proteins with diverse and important biological functions in embryonic growth and morphogenesis. DNA microarray technology presents an efficient means of acquiring novel and valuable information regarding the expression, regulation, and function of a panoply of genes involved in mammalian orofacial development. METHODS: To identify differentially expressed genes during mammalian orofacial ontogenesis, the transcript profiles of GD-12, GD-13, and GD-14 murine orofacial tissue were compared utilizing GeneChip arrays from Affymetrix. Changes in gene expression were verified by TaqMan quantitative real-time PCR. Cluster analysis of the microarray data was done with the GeneCluster 2.0 Data Mining Tool and the GeneSpring software. RESULTS: Expression of >50% of the approximately 12,000 genes and expressed sequence tags examined in this study was detected in GD-12, GD-13, and GD-14 murine orofacial tissues and the expression of several hundred genes was up- and downregulated in the developing orofacial tissue from GD-12 to GD-13, as well as from GD-13 to GD-14. Such differential gene expression represents changes in the expression of genes encoding growth factors and signaling molecules; transcription factors; and proteins involved in epithelial-mesenchymal interactions, extracellular matrix synthesis, cell adhesion, proliferation, differentiation, and apoptosis. Following cluster analysis of the microarray data, eight distinct patterns of gene expression during murine orofacial ontogenesis were selected for graphic presentation of gene expression patterns. CONCLUSIONS: This gene expression profiling study identifies a number of potentially unique developmental participants and serves as a valuable aid in deciphering the complex molecular mechanisms crucial for mammalian orofacial development.

Animals↗

Maximizing the entropy of histogram bar heights to explore neural activity: a simulation study on auditory and tactile fibers.

Neurophysiologists often use histograms to explore patterns of activity in neural spike trains. The bin size selected to construct a histogram is crucial: too large bin widths result in coarse histograms, too small bin widths expand unimportant detail. Peri-stimulus time (PST) histograms of simulated nerve fibers were studied in the current article. This class of histograms gives information about neural activity in the temporal domain and is a density estimate for the spike rate. Scott's rule based on modem statistical theory suggests that the optimal bin size is inversely proportional to the cube root of sample size. However, this estimate requires a priori knowledge about the density function. Moreover, there are no good algorithms for adaptive-mesh histograms, which have variable bin sizes to minimize estimation errors. Therefore, an unconventional technique is proposed here to help experimenters in practice. This novel method maximizes the entropy of histogram-bar heights to find the unique bin size, which generates the highest disorder in a histogram (i.e., the most complex histogram), and is useful as a starting point for neural data mining. Although the proposed method is ad hoc from a density-estimation point of view, it is simple, efficient and more helpful in the experimental setting where no prior statistical information on neural activity is available. The results of simulations based on the entropy method are also discussed in relation to Ellaway's cumulative-sum technique, which can detect subtle changes in neural activity in certain conditions.

Algorithms↗

Analysis of respiratory pressure-volume curves in intensive care medicine using inductive machine learning.

We present a case study of machine learning and data mining in intensive care medicine. In the study, we compared different methods of measuring pressure-volume curves in artificially ventilated patients suffering from the adult respiratory distress syndrome (ARDS). Our aim was to show that inductive machine learning can be used to gain insights into differences and similarities among these methods. We defined two tasks: the first one was to recognize the measurement method producing a given pressure-volume curve. This was defined as the task of classifying pressure-volume curves (the classes being the measurement methods). The second was to model the curves themselves, that is, to predict the volume given the pressure, the measurement method and the patient data. Clearly, this can be defined as a regression task. For these two tasks, we applied C5.0 and CUBIST, two inductive machine learning tools, respectively. Apart from medical findings regarding the characteristics of the measurement methods, we found some evidence showing the value of an abstract representation for classifying curves: normalization and high-level descriptors from curve fitting played a crucial role in obtaining reasonably accurate models. Another useful feature of algorithms for inductive machine learning is the possibility of incorporating background knowledge. In our study, the incorporation of patient data helped to improve regression results dramatically, which might open the door for the individual respiratory treatment of patients in the future.

Adult↗

Ant-based clustering and topographic mapping.

Ant-based clustering and sorting is a nature-inspired heuristic first introduced as a model for explaining two types of emergent behavior observed in real ant colonies. More recently, it has been applied in a data-mining context to perform both clustering and topographic mapping. Early work demonstrated some promising characteristics of the heuristic but did not extend to a rigorous investigation of its capabilities. We describe an improved version, called ATTA, incorporating adaptive, heterogeneous ants, a time-dependent transporting activity, and a method (for clustering applications) that transforms the spatial embedding produced by the algorithm into an explicit partitioning. ATTA is then subjected to the most rigorous experimental evaluation of an ant-based clustering and sorting algorithm undertaken to date: we compare its performance with standard techniques for clustering and topographic mapping using a set of analytical evaluation functions and a range of synthetic and real data collections. Our results demonstrate the ability of ant-based clustering and sorting to automatically identify the number of clusters inherent in a data collection, and to produce high quality solutions; indeed, we show that it is particularly robust for clusters of differing sizes and for overlapping clusters. The results obtained for topographic mapping are, however, disappointing. We provide evidence that the solutions generated by the ant algorithm are barely topology-preserving, and we explain in detail why results have--in spite of this--been misinterpreted (much more positively) in previous research.

Algorithms↗

Active learning with support vector machines in the drug discovery process.

We investigate the following data mining problem from computer-aided drug design: From a large collection of compounds, find those that bind to a target molecule in as few iterations of biochemical testing as possible. In each iteration a comparatively small batch of compounds is screened for binding activity toward this target. We employed the so-called "active learning paradigm" from Machine Learning for selecting the successive batches. Our main selection strategy is based on the maximum margin hyperplane-generated by "Support Vector Machines". This hyperplane separates the current set of active from the inactive compounds and has the largest possible distance from any labeled compound. We perform a thorough comparative study of various other selection strategies on data sets provided by DuPont Pharmaceuticals and show that the strategies based on the maximum margin hyperplane clearly outperform the simpler ones.

Computer-Aided Design↗

Nuclear heat shock response and novel nuclear domain 10 reorganization in respiratory syncytial virus-infected a549 cells identified by high-resolution two-dimensional gel electrophoresis.

The pneumovirus respiratory syncytial virus (RSV) is a leading cause of epidemic respiratory tract infection. Upon entry, RSV replicates in the epithelial cytoplasm, initiating compensatory changes in cellular gene expression. In this study, we have investigated RSV-induced changes in the nuclear proteome of A549 alveolar type II-like epithelial cells by high-resolution two-dimensional gel electrophoresis (2DE). Replicate 2D gels from uninfected and RSV-infected nuclei were compared for changes in protein expression. We identified 24 different proteins by peptide mass fingerprinting after matrix-assisted laser desorption ionization-time of flight mass spectrometry (MS), whose average normalized spot intensity was statistically significant and differed by +/-2-fold. Notable among the proteins identified were the cytoskeletal cytokeratins, RNA helicases, oxidant-antioxidant enzymes, the TAR DNA binding protein (a protein that associates with nuclear domain 10 [ND10] structures), and heat shock protein 70- and 60-kDa isoforms (Hsp70 and Hsp60, respectively). The identification of Hsp70 was also validated by liquid chromatography quadropole-TOF tandem MS (LC-MS/MS). Separate experiments using immunofluorescence microscopy revealed that RSV induced cytoplasmic Hsp70 aggregation and nuclear accumulation. Data mining of a genomic database showed that RSV replication induced coordinate changes in Hsp family proteins, including the 70, 70-2, 90, 40, and 40-3 isoforms. Because the TAR DNA binding protein associates with ND10s, we examined the effect of RSV infection on ND10 organization. RSV induced a striking dissolution of ND10 structures with redistribution of the component promyelocytic leukemia (PML) and speckled 100-kDa (Sp100) proteins into the cytoplasm, as well as inducing their synthesis. Our findings suggest that cytoplasmic RSV replication induces a nuclear heat shock response, causes ND10 disruption, and redistributes PML and Sp100 to the cytoplasm. Thus, a high-resolution proteomics approach, combined with immunofluorescence localization and coupled with genomic response data, yielded unexpected novel insights into compensatory nuclear responses to RSV infection.

Cell Nucleus↗

Computational identification of residues that modulate voltage sensitivity of voltage-gated potassium channels.

BACKGROUND: Studies of the structure-function relationship in proteins for which no 3D structure is available are often based on inspection of multiple sequence alignments. Many functionally important residues of proteins can be identified because they are conserved during evolution. However, residues that vary can also be critically important if their variation is responsible for diversity of protein function and improved phenotypes. If too few sequences are studied, the support for hypotheses on the role of a given residue will be weak, but analysis of large multiple alignments is too complex for simple inspection. When a large body of sequence and functional data are available for a protein family, mature data mining tools, such as machine learning, can be applied to extract information more easily, sensitively and reliably. We have undertaken such an analysis of voltage-gated potassium channels, a transmembrane protein family whose members play indispensable roles in electrically excitable cells. RESULTS: We applied different learning algorithms, combined in various implementations, to obtain a model that predicts the half activation voltage of a voltage-gated potassium channel based on its amino acid sequence. The best result was obtained with a k-nearest neighbor classifier combined with a wrapper algorithm for feature selection, producing a mean absolute error of prediction of 7.0 mV. The predictor was validated by permutation test and evaluation of independent experimental data. Feature selection identified a number of residues that are predicted to be involved in the voltage sensitive conformation changes; these residues are good target candidates for mutagenesis analysis. CONCLUSION: Machine learning analysis can identify new testable hypotheses about the structure/function relationship in the voltage-gated potassium channel family. This approach should be applicable to any protein family if the number of training examples and the sequence diversity of the training set that are necessary for robust prediction are empirically validated. The predictor and datasets can be found at the VKCDB web site.

Algorithms↗

Trypanosoma cruzi: analysis of the complete PUF RNA-binding protein family.

The members of the PUF family of RNA-binding proteins regulate the fate of mRNAs by binding to their 3'UTR sequence elements in eukaryotes. In trypanosomes, for which gene expression is polycistronic and controlled almost exclusively by post-transcriptional processes, PUF proteins could play a crucial role. We report here the complete analysis of the PUF protein family of Trypanosoma cruzi composed of 10 members. In silico analysis predicts the existence of at least three major groups within the T. cruzi family, based on their putative binding specificity. Using yeast three hybrid assays, we tested some of these predictions for TcPUF1, TcPUF3, TcPUF5, and TcPUF8 as representatives of these groups. Data mining of the T. cruzi genome led us to describe putative binding targets for the TcPUFs of the most conserved group, TcPUF1 and TcPUF2. The targets include genes for mitochondrial proteins and protein kinases. Finally, immunolocalization experiments showed that TcPUF1 is localized in multiple discrete foci in the cytoplasm supporting its proposed function.

3' Untranslated Regions↗

Imitating manual curation of text-mined facts in biomedicine.

Text-mining algorithms make mistakes in extracting facts from natural-language texts. In biomedical applications, which rely on use of text-mined data, it is critical to assess the quality (the probability that the message is correctly extracted) of individual facts--to resolve data conflicts and inconsistencies. Using a large set of almost 100,000 manually produced evaluations (most facts were independently reviewed more than once, producing independent evaluations), we implemented and tested a collection of algorithms that mimic human evaluation of facts provided by an automated information-extraction system. The performance of our best automated classifiers closely approached that of our human evaluators (ROC score close to 0.95). Our hypothesis is that, were we to use a larger number of human experts to evaluate any given sentence, we could implement an artificial-intelligence curator that would perform the classification job at least as accurately as an average individual human evaluator. We illustrated our analysis by visualizing the predicted accuracy of the text-mined relations involving the term cocaine.

Abstracting and Indexing↗

[Improved medical care for miners with ischemic heart disease].

The article is devoted to coronary disease in miners of deep Donbass mines. Data of its prevalence, chemical and functional features are given. Rapid progress of the disease was found to correlate with unfavourable factors of occupational environment. Mechanisms of dangerous heart rythm disorders formation during the work are shown. The main points of the programme improving the health care of miners suffering from coronary heart disease are described.

Adult↗

Purification of an eight subunit RNA polymerase I complex in Trypanosoma brucei.

Trypanosoma brucei harbors a unique multifunctional RNA polymerase (pol) I which transcribes, in addition to ribosomal RNA genes, the gene units encoding the major cell surface antigens variant surface glycoprotein and procyclin. In consequence, this RNA pol I is recruited to three structurally different types of promoters and sequestered to two distinct nuclear locations, namely the nucleolus and the expression site body. This versatility may require parasite-specific protein-protein interactions, subunits or subunit domains. Thus far, data mining of trypanosomatid genomes have revealed 13 potential RNA pol I subunits which include two paralogous sets of RPB5, RPB6, and RPB10. Here, we analyzed a cDNA library prepared from procyclic insect form T. brucei and found that all 13 candidate subunits are co-expressed. Moreover, we PTP-tagged the largest subunit TbRPA1, tandem affinity-purified the enzyme complex to homogeneity, and determined its subunit composition. In addition to the already known subunits RPA1, RPA2, RPC40, 1RPB5, and RPA12, the complex contained RPC19, RPB8, and 1RPB10. Finally, to evaluate the absence of RPB6 in our purifications, we used a combination of epitope-tagging and reciprocal coimmunoprecipitation to demonstrate that 1RPB6 but not 2RPB6 binds to RNA pol I albeit in an unstable manner. Collectively, our data strongly suggest that T. brucei RNA pol I binds a distinct set of the RPB5, RPB6, and RPB10 paralogs.

Amino Acid Sequence↗