Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Multiple terminal uridylyltransferases of trypanosomes.

The transferase activities that add uridylyl residues to RNA have been reported in several unicellular and metazoan organisms. Thus far, the two terminal uridylyltransferases (TUTases) involved in uridine insertion/deletion mRNA editing in mitochondria of trypanosomes were the only known enzymes with confirmed UTP specificity. Here, we demonstrate that protein sequences of editing TUTases may be used to predict novel UTP-specific enzymes by data mining. The highest-scoring open reading frame from Trypanosoma brucei was expressed and recombinant protein purified. This enzyme catalyzes a processive UMP incorporation and is not localized to the mitochondria suggesting a non-editing biological function.

Amino Acid Sequence↗

Gene expression profiling of 12633 genes in Alzheimer hippocampal CA1: transcription and neurotrophic factor down-regulation and up-regulation of apoptotic and pro-inflammatory signaling.

Alterations in transcription, RNA editing, translation, protein processing, and clearance are a consistent feature of Alzheimer's disease (AD) brain. To extend our initial study (Alzheimer Reports [2000] 3:161-167), RNA samples isolated from control and AD hippocampal cornu ammonis 1 (CA1) were analyzed for 12633 gene and expressed sequence tag (EST) expression levels using DNA microarrays (HG-U95Av2 Genechips; Affymetrix, Santa Clara, CA). Hippocampal CA1 tissues were carefully selected from several hundred potential specimens obtained from domestic and international brain banks. To minimize the effects of individual differences in gene expression, RNA of high spectral quality (A(260/280) > or= 1.9) was pooled from CA1 of six control or six AD subjects. Results were compared as a group; individual gene expression patterns for the most-changed RNA message levels were also profiled. There were no significant differences in age, postmortem interval (mean < or = 2.1 hr) nor tissue pH (range 6.6-6.9) between the two brain groups. AD tissues were derived from subjects clinically classified as CDR 2-3 (CERAD/NIA). Expression data were analyzed using GeneSpring (Silicon Genetics, Redwood City, CA) and Microarray Data Mining Tool (Affymetrix) software. Compared to controls and 354 background/alignment markers, AD brain showed a generalized depression in brain gene transcription, including decreases in RNA encoding transcription factors (TFs), neurotrophic factors, signaling elements involved in synaptic plasticity such as synaptophysin, metallothionein III, and metal regulatory factor-1. Three- or morefold increases in RNAs encoding DAXX, cPLA(2), CDP5, NF-kappaBp52/p100, FAS, betaAPP, DPP1, NFIL6, IL precursor, B94, HB15, COX-2, and CEX-1 signals were strikingly apparent. These data support the hypothesis of widespread transcriptional alterations, misregulation of RNAs involved in metal ion homeostasis, TF signaling deficits, decreases in neurotrophic support and activated apoptotic and neuroinflammatory signaling in moderately affected AD hippocampal CA1.

Aged↗

LitMiner and WikiGene: identifying problem-related key players of gene regulation using publication abstracts.

The LitMiner software is a literature data-mining tool that facilitates the identification of major gene regulation key players related to a user-defined field of interest in PubMed abstracts. The prediction of gene-regulatory relationships is based on co-occurrence analysis of key terms within the abstracts. LitMiner predicts relationships between key terms from the biomedical domain in four categories (genes, chemical compounds, diseases and tissues). Owing to the limitations (no direction, unverified automatic prediction) of the co-occurrence approach, the primary data in the LitMiner database represent postulated basic gene-gene relationships. The usefulness of the LitMiner system has been demonstrated recently in a study that reconstructed disease-related regulatory networks by promoter modelling that was initiated by a LitMiner generated primary gene list. To overcome the limitations and to verify and improve the data, we developed WikiGene, a Wiki-based curation tool that allows revision of the data by expert users over the Internet. LitMiner (http://andromeda.gsf.de/litminer) and WikiGene (http://andromeda.gsf.de/wiki) can be used unrestricted with any Internet browser.

Abstracting and Indexing↗

Characterisation and application of a bovine U6 promoter for expression of short hairpin RNAs.

BACKGROUND: The use of small interfering RNA (siRNA) molecules in animals to achieve double-stranded RNA-mediated interference (RNAi) has recently emerged as a powerful method of sequence-specific gene knockdown. As DNA-based expression of short hairpin RNA (shRNA) for RNAi may offer some advantages over chemical and in vitro synthesised siRNA, a number of vectors for expression of shRNA have been developed. These often feature polymerase III (pol. III) promoters of either mouse or human origin. RESULTS: To develop a shRNA expression vector specifically for bovine RNAi applications, we identified and characterised a novel bovine U6 small nuclear RNA (snRNA) promoter from bovine sequence data. This promoter is the putative bovine homologue of the human U6-8 snRNA promoter, and features a number of functional sequence elements that are characteristic of these types of pol. III promoters. A PCR based cloning strategy was used to incorporate this promoter sequence into plasmid vectors along with shRNA sequences for RNAi. The promoter was then used to express shRNAs, which resulted in the efficient knockdown of an exogenous reporter gene and an endogenous bovine gene. CONCLUSION: We have mined data from the bovine genome sequencing project to identify a functional bovine U6 promoter and used the promoter sequence to construct a shRNA expression vector. The use of this native bovine promoter in shRNA expression is an important component of our future development of RNAi therapeutic and transgenic applications in bovine species.

Animals↗

Predicting outcomes of hospitalization for heart failure using logistic regression and knowledge discovery methods.

The purpose of this study is to determine the best prediction of heart failure outcomes, resulting from two methods -- standard epidemiologic analysis with logistic regression and knowledge discovery with supervised learning/data mining. Heart failure was chosen for this study as it exhibits higher prevalence and cost of treatment than most other hospitalized diseases. The prevalence of heart failure has exceeded 4 million cases in the U.S.. Findings of this study should be useful for the design of quality improvement initiatives, as particular aspects of patient comorbidity and treatment are found to be associated with mortality. This is also a proof of concept study, considering the feasibility of emerging health informatics methods of data mining in conjunction with or in lieu of traditional logistic regression methods of prediction. Findings may also support the design of decision support systems and quality improvement programming for other diseases.

Databases as Topic↗

Strategies for efficient lead structure discovery from natural products.

This investigation aims to evaluate strategies for an efficient selection of bioactive compounds from the multitude and biodiversity of the plant kingdom. Statistics prove natural products (NPs) as a source leading most consistently to successful development of new drugs. However, there are several reasons why the interest in finding bioactive NPs has generally declined at several major pharmaceutical companies. Their substantial argument is that the research in this field is time-consuming, highly complex and ineffective. A more rational and economic search for new lead structures from nature must therefore be a priority in order to overcome these problems. In this paper, different strategies are described to exploit the molecular diversity of bioactive secondary metabolites, namely classical pharmacognostic approaches and computational methods. The latter include various data mining tools, like virtual screening filtering experiments using pharmacophore models, docking studies, and neural networks, which help to establish a relationship between chemical structure and biological activity. The strengths and weaknesses of these methods will be shown in this review. Focusing on selected targets within the arachidonic acid cascade (phospholipase A(2), 5-lipoxygenase, cyclooxygenase-1 and -2), several studies of successful discoveries in the field of anti-inflammatory NPs were scrutinized for the applied strategies. Both the compilation of relevant published data and recent studies supported by our own research clearly demonstrate the benefits of the synergistic effect of a hybridization of these strategies for an effective drug discovery from natural ingredients.

Anti-Inflammatory Agents↗

Gene expression profiling in schizophrenia and related mental disorders.

The etiology and pathophysiology of schizophrenia and related mental disorders such as bipolar disorder and major depression remain largely unclear. Recent advances in mRNA profiling techniques made it possible to perform genome-wide gene expression analysis in a hypothesis-free manner. It was thought that this large-scale data mining approach would reveal unknown molecular cascades involved in mental disorders. Contrary to this initial expectation, however, DNA microarray results in psychiatric fields have been notoriously discordant. Here the authors review the findings of DNA microarray analysis, focusing on systematic gene expression changes in schizophrenia, as well as alterations in the expression of specific genes, that have been reported and replicated. The authors also address the probable causes for the discordance among studies, possible ways to solve the problem, and their preferred approach for data interpretation.

Bipolar Disorder↗

Robust regression with asymmetric heavy-tail noise distributions.

In the presence of a heavy-tail noise distribution, regression becomes much more difficult. Traditional robust regression methods assume that the noise distribution is symmetric, and they downweight the influence of so-called outliers. When the noise distribution is asymmetric, these methods yield biased regression estimators. Motivated by data-mining problems for the insurance industry, we propose a new approach to robust regression tailored to deal with asymmetric noise distribution. The main idea is to learn most of the parameters of the model using conditional quantile estimators (which are biased but robust estimators of the regression) and to learn a few remaining parameters to combine and correct these estimators, to minimize the average squared error in an unbiased way. Theoretical analysis and experiments show the clear advantages of the approach. Results are on artificial data as well as insurance data, using both linear and neural network predictors.

Algorithms↗

Enhanced care of hypertensive patients using the Internet.

BACKGROUND: Medical guidelines provide recommendations that ought to help physicians in medical decision-making under different circumstances. To develop electronic medical guidelines and to make them available for physicians using the Internet can further enhance the quality and efficiency of health care, especially with the simultaneous use of electronic health records. OBJECTIVES: This study aims to outline the needs for web-based systems that support the use of medical guidelines in practice; and to focus on the development of web-based electronic medical guidelines for treatment of arterial hypertension. METHODS: The importance of electronic health record in cardiology for data acquisition, data storage and data mining is considered. The anonymized database of approximately 1800 hypertensive patients has been created to compare medical practice with guidelines and discover features of diseases that can help with their management. Using this database we evaluated several web-based electronic guidelines systems. RESULTS: The 1999 WHO/ISH Guidelines for the Management of Hypertension were formalized and interpreted using the Guide-X methodology, using the Apollo system and web-based electronic guidelines. The web-based electronic guidelines were tested on the smaller anonymous recent data set of 840 hypertensive patients. CONCLUSIONS: An easy transfer of knowledge from medical guidelines to structured electronic guidelines opens new possibilities for easy reusability of medical knowledge by general practitioners and clinicians.

Cardiology↗

Using dependency/association rules to find indications for computed tomography in a head trauma dataset.

Analysis of a clinical head trauma dataset was aided by the use of a new, binary-based data mining technique, termed Boolean analyzer (BA), which finds dependency/association rules. With initial guidance from a domain user or domain expert, the BA algorithm is given one or more metrics to partition the entire dataset. The weighted rules are in the form of Boolean expressions. To augment the analysis of the rules produced, we applied a probabilistic interestingness measure (PIM) to order the generated rules based on event dependency, where events are combinations of primed and unprimed variables. Interpretation of the dependency rules generated on the clinical head trauma data resulted in a set of criteria that identified minor head trauma patients needing computed tomography (CT) scans. The BA criteria contained fewer variables than were found using recursive partitioning of Chi-square values (five variables versus seven variables, respectively). The BA five-variable criteria set was more sensitive but less specific than the seven-variable Chi-square criteria set. We believe that the BA method has broad applicability in the medical domain, and hope that this paper will stimulate other creative applications of the technique.

Algorithms↗

PhosphoregDB: the tissue and sub-cellular distribution of mammalian protein kinases and phosphatases.

BACKGROUND: Protein kinases and protein phosphatases are the fundamental components of phosphorylation dependent protein regulatory systems. We have created a database for the protein kinase-like and phosphatase-like loci of mouse http://phosphoreg.imb.uq.edu.au that integrates protein sequence, interaction, classification and pathway information with the results of a systematic screen of their sub-cellular localization and tissue specific expression data mined from the GNF tissue atlas of mouse. RESULTS: The database lets users query where a specific kinase or phosphatase is expressed at both the tissue and sub-cellular levels. Similarly the interface allows the user to query by tissue, pathway or sub-cellular localization, to reveal which components are co-expressed or co-localized. A review of their expression reveals 30% of these components are detected in all tissues tested while 70% show some level of tissue restriction. Hierarchical clustering of the expression data reveals that expression of these genes can be used to separate the samples into tissues of related lineage, including 3 larger clusters of nervous tissue, developing embryo and cells of the immune system. By overlaying the expression, sub-cellular localization and classification data we examine correlations between class, specificity and tissue restriction and show that tyrosine kinases are more generally expressed in fewer tissues than serine/threonine kinases. CONCLUSION: Together these data demonstrate that cell type specific systems exist to regulate protein phosphorylation and that for accurate modelling and for determination of enzyme substrate relationships the co-location of components needs to be considered.

Amino Acid Sequence↗

Advanced visualization of self-organizing maps with vector fields.

Self-Organizing Maps have been applied in various industrial applications and have proven to be a valuable data mining tool. In order to fully benefit from their potential, advanced visualization techniques assist the user in analyzing and interpreting the maps. We propose two new methods for depicting the SOM based on vector fields, namely the Gradient Field and Borderline visualization techniques, to show the clustering structure at various levels of detail. We explain how this method can be used on aggregated parts of the SOM that show which factors contribute to the clustering structure, and show how to use it for finding correlations and dependencies in the underlying data. We provide examples on several artificial and real-world data sets to point out the strengths of our technique, specifically as a means to combine different types of visualizations offering effective multidimensional information visualization of SOMs.

Algorithms↗

The effect of diabetes on sensorineural hearing loss.

OBJECTIVE: To identify whether patients with diabetes have a higher incidence of sensorineural hearing loss than the general population and examine whether control of diabetes is related to severity of hearing loss. STUDY DESIGN: Retrospective database review; complete data mining of electronic medical record from 1989 to present. SETTING: Tertiary referral center. PATIENTS: Electronic medical records from 53461 nondiabetic age-matched patients and 12575 diabetic patients were reviewed. MAIN OUTCOME MEASURES: Presence or absence of diabetes and/or sensorineural hearing loss, serum creatinine, pure tone hearing (dB), speech discrimination (%), serum cholesterol, and triglycerides. RESULTS: Sensorineural hearing loss was more common in the diabetic patients than in age0matched nondiabetic patients from the same institutions. Poor control of diabetes, as measured by increasing serum creatinine, but not apparent in hemoglobin A1C laboratory data, correlated with worsening hearing in patients with diabetes who had sensorineural hearing loss. CONCLUSIONS: Sensorineural hearing loss was more common in patients with diabetes than in the control nondiabetic patients, and severity of hearing loss seemed to correlate with progression of disease as reflected in serum creatinine. This may have been due to microangiopathic disease in the inner ear.

Audiometry, Pure-Tone↗

Domain-specific language models and lexicons for tagging.

Accurate and reliable part-of-speech tagging is useful for many Natural Language Processing (NLP) tasks that form the foundation of NLP-based approaches to information retrieval and data mining. In general, large annotated corpora are necessary to achieve desired part-of-speech tagger accuracy. We show that a large annotated general-English corpus is not sufficient for building a part-of-speech tagger model adequate for tagging documents from the medical domain. However, adding a quite small domain-specific corpus to a large general-English one boosts performance to over 92% accuracy from 87% in our studies. We also suggest a number of characteristics to quantify the similarities between a training corpus and the test data. These results give guidance for creating an appropriate corpus for building a part-of-speech tagger model that gives satisfactory accuracy results on a new domain at a relatively small cost.

Humans↗

Rapid and accurate calculation of protein 1H, 13C and 15N chemical shifts.

A computer program (SHIFTX) is described which rapidly and accurately calculates the diamagnetic 1H, 13C and 15N chemical shifts of both backbone and sidechain atoms in proteins. The program uses a hybrid predictive approach that employs pre-calculated, empirically derived chemical shift hypersurfaces in combination with classical or semi-classical equations (for ring current, electric field, hydrogen bond and solvent effects) to calculate 1H, 13C and 15N chemical shifts from atomic coordinates. The chemical shift hypersurfaces capture dihedral angle, sidechain orientation, secondary structure and nearest neighbor effects that cannot easily be translated to analytical formulae or predicted via classical means. The chemical shift hypersurfaces were generated using a database of IUPAC-referenced protein chemical shifts--RefDB (Zhang et al., 2003), and a corresponding set of high resolution (<2.1 A) X-ray structures. Data mining techniques were used to extract the largest pairwise contributors (from a list of approximately 20 derived geometric, sequential and structural parameters) to generate the necessary hypersurfaces. SHIFTX is rapid (<1 CPU second for a complete shift calculation of 100 residues) and accurate. Overall, the program was able to attain a correlation coefficient (r) between observed and calculated shifts of 0.911 (1Halpha), 0.980 (13Calpha), 0.996 (13Cbeta), 0.863 (13CO), 0.909 (15N), 0.741 (1HN), and 0.907 (sidechain 1H) with RMS errors of 0.23, 0.98, 1.10, 1.16, 2.43, 0.49, and 0.30 ppm, respectively on test data sets. We further show that the agreement between observed and SHIFTX calculated chemical shifts can be an extremely sensitive measure of the quality of protein structures. Our results suggest that if NMR-derived structures could be refined using heteronuclear chemical shifts calculated by SHIFTX, their precision could approach that of the highest resolution X-ray structures. SHIFTX is freely available as a web server at http://redpoll.pharmacy.ualberta.ca.

Animals↗

cDNA cloning of two different serine protease inhibitor precursors in the migratory locust, Locusta migratoria.

Recently, a novel serine protease-inhibiting peptide family, designated as the 'pacifastin family', has been described in locusts and crayfish. All members of this family possess a characteristic cysteine-rich domain. The present study describes the cDNA cloning, sequencing and transcript distribution of two novel pacifastin-related peptide precursors in the migratory locust, Locusta migratoria. Only one of the encoded peptides (HI) was identified previously, whereas six others represent new members of the pacifastin family. Northern blot analysis showed that both precursor transcripts are present in adult locust fat body. These could not be detected in the midgut. Interestingly, an in silico data mining approach of the expressed sequence tags (EST) database revealed the existence of Manduca sexta and Bombyx mori cDNAs that display pronounced sequence similarities with these locust pacifastin-related transcripts.

Amino Acid Sequence↗

CKLFSF2 is highly expressed in testis and can be secreted into the seminiferous tubules.

CKLFSF2 is a member of the chemokine-like factor superfamily (CKLFSF), a novel gene family containing CKLF and CKLFSF1-8. Using a combination of data mining and polymerase chain reactions, we determined the full cDNA sequence and genomic structure of human CKLFSF2, a 4-exon gene encoding 248 amino acids and spanning approximately 8.8 kb on chromosome 16q22.1. Expression profile analyses indicated that CKLFSF2 is expressed in a limited number of tissues. Specifically, immunohistochemistry indicated that CKLFSF2 is highly expressed in testis, mainly in spermatogonia and the seminiferous tubular fluid. Subcellular localization experiments suggested that CKLFSF2 is equally distributed in the cytoplasm, and Western blot analysis revealed that overexpressed CKLFSF2 is secreted into the supernatant of cultured cells. The data therefore strongly suggest that CKLFSF2 is a secreted protein that may be functionally relevant during spermatogenesis.

Amino Acid Sequence↗

Supervised machine learning techniques for the classification of metabolic disorders in newborns.

MOTIVATION: During the Bavarian newborn screening programme all newborns have been tested for about 20 inherited metabolic disorders. Owing to the amount and complexity of the generated experimental data, machine learning techniques provide a promising approach to investigate novel patterns in high-dimensional metabolic data which form the source for constructing classification rules with high discriminatory power. RESULTS: Six machine learning techniques have been investigated for their classification accuracy focusing on two metabolic disorders, phenylketo nuria (PKU) and medium-chain acyl-CoA dehydrogenase deficiency (MCADD). Logistic regression analysis led to superior classification rules (sensitivity >96.8%, specificity >99.98%) compared to all investigated algorithms. Including novel constellations of metabolites into the models, the positive predictive value could be strongly increased (PKU 71.9% versus 16.2%, MCADD 88.4% versus 54.6% compared to the established diagnostic markers). Our results clearly prove that the mined data confirm the known and indicate some novel metabolic patterns which may contribute to a better understanding of newborn metabolism.

Algorithms↗