Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Identification of a binding site for the anti-inflammatory tripeptide feG.

The mechanism of action of feG, an anti-inflammatory peptide, was explored using data mining, molecular modeling, and enzymatic techniques. The molecular coordinates of protein kinase A (PKA) were used to create six virtual isoforms of protein kinase C (PKCalpha, betaI, betaII, delta, iota, and zeta). With in silico techniques a binding site for feG was identified on PKCbetaI that correlated significantly with a biological activity, the inhibition of intestinal anaphylaxis. Since feG selectively increased the binding of a PKCbetaI antibody, it is proposed that this peptide inhibits the reassociation of the hydrophobic tail of PKCbetaI with its binding site and prevents the enzyme from assuming an inactive conformation.

Amino Acid Sequence↗

Cell organisation, sulphur metabolism and ion transport-related genes are differentially expressed in Paracoccidioides brasiliensis mycelium and yeast cells.

BACKGROUND: Mycelium-to-yeast transition in the human host is essential for pathogenicity by the fungus Paracoccidioides brasiliensis and both cell types are therefore critical to the establishment of paracoccidioidomycosis (PCM), a systemic mycosis endemic to Latin America. The infected population is of about 10 million individuals, 2% of whom will eventually develop the disease. Previously, transcriptome analysis of mycelium and yeast cells resulted in the assembly of 6,022 sequence groups. Gene expression analysis, using both in silico EST subtraction and cDNA microarray, revealed genes that were differential to yeast or mycelium, and we discussed those involved in sugar metabolism. To advance our understanding of molecular mechanisms of dimorphic transition, we performed an extended analysis of gene expression profiles using the methods mentioned above. RESULTS: In this work, continuous data mining revealed 66 new differentially expressed sequences that were MIPS(Munich Information Center for Protein Sequences)-categorised according to the cellular process in which they are presumably involved. Two well represented classes were chosen for further analysis: (i) control of cell organisation - cell wall, membrane and cytoskeleton, whose representatives were hex (encoding for a hexagonal peroxisome protein), bgl (encoding for a 1,3-beta-glucosidase) in mycelium cells; and ags (an alpha-1,3-glucan synthase), cda (a chitin deacetylase) and vrp (a verprolin) in yeast cells; (ii) ion metabolism and transport - two genes putatively implicated in ion transport were confirmed to be highly expressed in mycelium cells - isc and ktp, respectively an iron-sulphur cluster-like protein and a cation transporter; and a putative P-type cation pump (pct) in yeast. Also, several enzymes from the cysteine de novo biosynthesis pathway were shown to be up regulated in the yeast form, including ATP sulphurylase, APS kinase and also PAPS reductase. CONCLUSION: Taken together, these data show that several genes involved in cell organisation and ion metabolism/transport are expressed differentially along dimorphic transition. Hyper expression in yeast of the enzymes of sulphur metabolism reinforced that this metabolic pathway could be important for this process. Understanding these changes by functional analysis of such genes may lead to a better understanding of the infective process, thus providing new targets and strategies to control PCM.

Biological Transport↗

Gene functional annotation by statistical analysis of biomedical articles.

BACKGROUND: Functional annotation of genes is an important task in biology since it facilitates the characterization of genes relationships and the understanding of biochemical pathways. The various gene functions can be described by standardized and structured vocabularies, called bio-ontologies. The assignment of bio-ontology terms to genes is carried out by means of applying certain methods to datasets extracted from biomedical articles. These methods originate from data mining and machine learning and include maximum entropy or support vector machines (SVM). PURPOSE: The aim of this paper is to propose an alternative to the existing methods for functionally annotating genes. The methodology involves building of classification models, validation and graphical representations of the results and reduction of the dimensions of the dataset. METHODS: Classification models are constructed by Linear discriminant analysis (LDA). The validation of the models is based on statistical analysis and interpretation of the results involving techniques like hold-out samples, test datasets and metrics like confusion matrix, accuracy, recall, precision and F-measure. Graphical representations, such as boxplots, Andrew's curves and scatterplots of the variables resulting from the classification models are also used for validating and interpreting the results. RESULTS: The proposed methodology was applied to a dataset extracted from biomedical articles for 12 Gene Ontology terms. The validation of the LDA models and the comparison with the SVM show that LDA (mean F-measure 75.4%) outperforms the SVM (mean F-measure 68.7%) for the specific data. CONCLUSION: The application of certain statistical methods can be beneficial for functional gene annotation from biomedical articles. Apart from the good performance the results can be interpreted and give insight of the bio-text data structure.

Abstracting and Indexing↗

Gender difference in HIV-1 RNA viral loads.

OBJECTIVES: To test and characterize the dependence of viral load on gender in different countries and racial groups as a function of CD4 T-cell count. METHODS: Plasma viral load data were analysed for > 30,000 HIV-infected patients attending clinics in the USA [HIV Insight (Cerner Corporation, Vienna, VA, USA) and Plum Data Mining LLC (East Meadow, NY, USA) databases] and the Netherlands (Athena database; HIV Monitoring Foundation, Amsterdam, Netherlands). Log-normal regression models were used to test for an effect of gender on viral load while adjusting for covariates and allowing the effect to depend on CD4 T-cell count. Sensitivity analyses were performed to test the robustness of conclusions to assumptions regarding viral loads below the lower limit of quantification (LLOQ). RESULTS: After adjusting for covariates, women had (nonsignificantly) lower viral loads than men (HIV Insight: -0.053 log(10) HIV-1 RNA copies/mL, P = 0.202; Athena: -0.005 log(10) copies/mL, P = 0.667; Plum: -0.072 log(10) copies/mL, P = 0.273). However, further investigation revealed that the gender effect depended on CD4 T-cell count. Women had consistently higher viral loads than men when CD4 T-cell counts were at most 50 cells/microL, and consistently lower viral loads than men when CD4 T-cell counts were greater than 350 cells/microL. These effects were remarkably consistent when estimated independently for the racial groups with sufficient data available in the HIV Insight and Plum databases. CONCLUSIONS: The consistent relationship between gender-related differences in viral load and CD4 T-cell count demonstrated here explains the diverse findings previously published.

Adult↗

GeneFarm, structural and functional annotation of Arabidopsis gene and protein families by a network of experts.

Genomic projects heavily depend on genome annotations and are limited by the current deficiencies in the published predictions of gene structure and function. It follows that, improved annotation will allow better data mining of genomes, and more secure planning and design of experiments. The purpose of the GeneFarm project is to obtain homogeneous, reliable, documented and traceable annotations for Arabidopsis nuclear genes and gene products, and to enter them into an added-value database. This re-annotation project is being performed exhaustively on every member of each gene family. Performing a family-wide annotation makes the task easier and more efficient than a gene-by-gene approach since many features obtained for one gene can be extrapolated to some or all the other genes of a family. A complete annotation procedure based on the most efficient prediction tools available is being used by 16 partner laboratories, each contributing annotated families from its field of expertise. A database, named GeneFarm, and an associated user-friendly interface to query the annotations have been developed. More than 3000 genes distributed over 300 families have been annotated and are available at http://genoplante-info.infobiogen.fr/Genefarm/. Furthermore, collaboration with the Swiss Institute of Bioinformatics is underway to integrate the GeneFarm data into the protein knowledgebase Swiss-Prot.

Arabidopsis↗

Techniques in plant telomere biology.

The role model systems have played in understanding telomere biology has been enormous, and understanding has rapidly transferred to human telomere research. Most work using model organisms to study telomerase and nontelomerase-based telomere-maintenance systems has centered on yeasts, ciliates, and insects. But it is now timely to put considerably more effort into plant models for a number of reasons: (i) the rice and Arabidopsis genome sequencing projects make data mining possible; (ii) extensive collections of insertion mutants of Arabidopsis thaliana enable phenotypic effects of protein gene knockouts to be analyzed, including for those genes involved in telomere structure, function (including, for example, in meiosis), and maintenance; and (iii) the variability of plant telomeres is considerable and ranges from the telomerase-mediated synthesis of the Arabidopsis-type (TTTAGGG) and vertebrate-type (TTAGGG) repeats to sequences synthesized by telomerase-independent mechanism(s) that are still to be discovered. Here we describe how the understanding of telomere biology has been advanced by methods used to isolate telomeric sequences and prove that the putative sequences isolated are indeed telomeric. We show how assays designed to prove the activity of telomerase [e.g., telomeric repeat amplification protocol (TRAP)] lead not only to an understanding of telomere structure and function, but also to the understanding of cell activity in development and in the cell cycle. We review how assays designed to reveal protein/protein and protein/nucleic acid interactions promote understanding of the structure and activities of plant telomeres. Together, the data are making significant contributions to telomere biology in general and could have medical implications.

Cell Nucleus↗

Statistical and visual morph movie analysis of crystallographic mutant selection bias in protein mutation resource data.

The relationship between protein mutations and conformational change can potentially decipher the language relating sequence to structure. Elsewhere, we presented the Protein Mutant Resource (PMR), an online tool that systematically identified related mutants in the Protein DataBank (PDB), inferred mutant Gene Ontology classifications using data-mining, and allowed intuitive exploration of relationships between mutant structures. Here, we perform a comprehensive statistical analysis of PMR mutants. Although the PMR contains spectacular conformational changes, generally there is a counter-intuitive inverse relationship between conformational change and the number of mutations. That is, PDB mutations contrast naturally evolved mutations. We compare the frequencies of mutations in the PMR/PDB datasets against the PAM250 natural mutation frequencies to confirm this. We make available morph movies from PMR structure pairs, allowing visual analysis of conformational change and the ability to distinguish visually between conformational change due to motions (e.g., ligand binding)and mutations. The PMR is at http://pmr.sdsc.edu.

Bias↗

Adaptive user displays for intelligent tutoring software.

Intelligent tutoring software (ITS) holds great promise for K-12 instruction. Yet it is difficult to obtain rich information about users that can be used in realistic educational delivery settings--public school classrooms--in which eye tracking and other user sensing technologies are not suitable. We are pursuing three "cheap and cheerful" strategies to meet this challenge in the context of an ITS for high school math instruction. First, we use detailed representations of student cognitive skills, including tasks to assess individual users' proficiency with abstract reasoning, proficiency with simple math facts and computational skill, and spatial ability. Second, we are using data mining and machine learning algorithms to identify instructional sequences that have been effective with previous students, and to use these patterns to make decisions about current students. Third, we are integrating a simple focus-of-attention tracking system into the software, using inexpensive, web cameras. This coarse-grained information can be used to time the display of multimedia hints, explanations, and examples when the user is actually looking at the screen, and to diagnose causes of problem-solving errors. The ultimate goal is to create non-intrusive software that can adapt the display of instructional information in real time to the user's cognitive strengths, motivation, and attention.

Algorithms↗

Identification and expression analysis of hepcidin-like antimicrobial peptides in bony fish.

Antimicrobial peptides play a crucial role as the first line of defense against invading pathogens. Several types of antimicrobial peptides have been isolated from fish, mostly of the cationic alpha-helical variety. Here, we present the cDNA sequences of five highly disulphide-bonded hepcidin-like peptides from winter flounder, Pseudopleuronectes americanus (Walbaum) and two from Atlantic salmon, Salmo salar (L.). These hepcidin-like molecules consist of a 24 amino acid signal peptide and an acidic propiece of 38-40 amino acids in addition to the mature processed peptide of 19-27 amino acids. Exhaustive data mining of GenBank with these sequences revealed that similar peptides are encoded in the genomes of Japanese flounder, rainbow trout, hybrid striped bass and medaka, indicating that they are widespread among fish. Southern hybridization analysis suggests that closely related hepcidin-like genes are present in other flatfish species, and that they exist as a multigene family clustered on the winter flounder genome. Hepcidin variants are differentially expressed during bacterial challenge, during larval development of P. americanus and in different tissues of adult fish.

Amino Acid Sequence↗

Identification of three isoforms for mitochondrial adenine nucleotide translocator in the pufferfish Takifugu rubripes.

Three adenine nucleotide translocator (ANT) genes were identified through in silico data mining of the Fugu genome database along with isolation of their corresponding cDNAs in vivo from the pufferfish (Takifugu rubripes). As a result of phylogenetic analysis, the ANT gene on scaffold_254 corresponded to mammalian ANT1, whereas both of those on scaffold_6 and scaffold_598 to mammalian ANT3. The ANT gene encoded by scaffold_6 was expressed ubiquitously in various tissues, whereas the ANT genes encoded by scaffold_254 and scaffold_598 were predominantly expressed in skeletal muscle and heart, respectively.

Amino Acid Sequence↗

Plant metabolomics: from holistic hope, to hype, to hot topic.

In a short time, plant metabolomics has gone from being just an ambitious concept to being a rapidly growing, valuable technology applied in the stride to gain a more global picture of the molecular organization of multicellular organisms. The combination of improved analytical capabilities with newly designed, dedicated statistical, bioinformatics and data mining strategies, is beginning to broaden the horizons of our understanding of how plants are organized and how metabolism is both controlled but highly flexible. Metabolomics is predicted to play a significant, if not indispensable role in bridging the phenotype-genotype gap and thus in assisting us in our desire for full genome sequence annotation as part of the quest to link gene to function. Plants are a fabulously rich source of diverse functional biochemicals and metabolomics is also already proving valuable in an applied context. By creating unique opportunities for us to interrogate plant systems and characterize their biochemical composition, metabolomics will greatly assist in identifying and defining much of the still unexploited biodiversity available today.

Biotechnology↗

Development and characterization of microsatellite markers for the canine hookworm, Ancylostoma caninum.

Microsatellites are repetitive genomic elements that show high levels of variation and therefore provide excellent tools to study the genetics of eukaryotic organisms. Hookworms are extremely common and important nematode parasites of humans and animals, causing potentially serious disease morbidity. Control of hookworms in dogs is achieved by frequent treatment with anthelmintics, and in humans, anthelmintics are frequently administered in a mass-treatment community-wide approach. Understanding the population genetics of hookworms has important implications for studies on the development and spread of drug resistance. We investigated the genome of Ancylostoma caninum for microsatellites by developing and then screening an enriched genomic library as well as by data mining published sequences of a whole genome shotgun library. Investigations revealed a high abundance of trinucleotide repeats. Dinucleotide repeats were characterized by a high number of AT, GA, and GT repeats. After testing and optimization of 68 markers, a panel of 34 polymorphic microsatellite markers were selected. Microsatellite analysis of hookworm isolates revealed a high degree of polymorphism, which was not influenced by the length of the repeats. This panel of microsatellite markers makes it possible to pursue investigations on the population genetics of A. caninum. Furthermore, a number of the markers demonstrated suitability for analysis of the human hookworm species Necator americanus and A. duodenale.

Ancylostoma↗

The human olfactory subgenome: from sequence to structure and evolution.

Olfactory receptors (ORs) constitute the largest multigene family in multicellular organisms. Their evolutionary proliferation has been driven by the need to provide recognition capacity for millions of potential odorants with arbitrary chemical configurations. Human genome sequencing has provided a highly informative picture of the "olfactory subgenome", the repertoire of OR genes. We describe here an analysis of 224 human OR genes, a much larger number than hitherto systematically analyzed. These are derived by literature survey, data mining at 14 genomic clusters, and by an OR-targeted experimental sequencing strategy. The presented set contains at least 53% pseudogenes and is minimally divided into 11 gene families. One of these (no. 7) has undergone a particularly extensive expansion in primates. The analysis of this collection leads to insight into the origin of OR genes, suggesting a graded expansion through mammalian evolution. It also allows us to delineate a structural map of the respective proteins. A sequence database and analysis package is provided (http://bioinformatics.weizmann.ac.il/HORDE), which will be useful for analyzing human OR sequences genome-wide.

Amino Acid Sequence↗

Tissue distribution and functional expression of a cDNA encoding a novel mixed lineage kinase.

Hypertrophy is an adaptive response of the heart to myocardial injury or hemodynamic overload that may progress and contribute to cardiac decompensation and eventually to heart failure. The signaling pathways controlling this response in the cardiac myocyte are poorly understood. A data mining effort of a human failed heart cDNA library was undertaken in an effort to identify novel signaling molecules involved in cardiac hypertrophy. This effort identified a novel kinase (MLK7) homologous to the mixed lineage kinase family of proteins. The mixed lineage kinases are mitogen-activated protein kinase kinase kinases (MAPKKKs) which activate stress activated protein kinase/c-Jun N-terminal kinase (SAPK/JNK) and p38 kinase pathways. They contain a catalytic domain with homology to both serine/threonine and tyrosine-specific kinases and a dual leucine zipper. MLK7 is identical to leucine zipper and sterile-alpha motif protein kinase (ZAK) through the leucine zipper domain but has a completely divergent COOH-terminus and shares approximately 40% homology with the other MLKs overall. Expression of MLK7 mRNA is most abundant in skeletal muscle and heart, with expression restricted to the cardiac myocyte. The recombinant histidine tagged MLK7 expressed and purified from insect cells exhibited serine/threonine kinase activity in vitro with myelin basic protein as substrate. When expressed in cardiac myocytes, MLK7 activated SAPK/JNK1, and ERK and p38 to a lesser extent. Additionally, MLK7 altered fetal gene expression and increased protein synthesis in cardiac myocytes. These data suggest that MLK7 is a new member of the mixed lineage kinase family that modulates cardiac SAPK/JNK pathway and may play a role in cardiac hypertrophy and progression to heart failure.

Adult↗

Improving end of life care: an information systems approach to reducing medical errors.

Chronic and terminally ill patients are disproportionately affected by medical errors. In addition, the elderly suffer more preventable adverse events than younger patients. Targeting system wide "error-reducing" reforms to vulnerable populations can significantly reduce the incidence and prevalence of human error in medical practice. Recent developments in health informatics, particularly the application of artificial intelligence (AI) techniques such as data mining, neural networks, and case-based reasoning (CBR), presents tremendous opportunities for mitigating error in disease diagnosis and patient management. Additionally, the ubiquity of the Internet creates the possibility of an almost ideal network for the dissemination of medical information. We explore the capacity and limitations of web-based palliative information systems (IS) to transform the delivery of care, streamline processes and improve the efficiency and appropriateness of medical treatment. As a result, medical error(s) that occur with patients dealing with severe, chronic illness and the frail elderly can be reduced.The palliative model grew out of the need for pain relief and comfort measures for patients diagnosed with cancer. Applied definitions of palliative care extend this convention, but there is no widely accepted definition. This research will discuss the development life cycle of two palliative information systems: the CONFER QOLP management information system (MIS), currently used by a community-based palliative care program in Brooklyn, New York, and the CAREN case-based reasoning prototype. CONFER is a web platform based on the idea of "eCare". CONFER uses XML (extensible mark-up language), a W3C-endorced standard mark up to define systems data. The second system, CAREN, is a CBR prototype designed for palliative care patients in the cancer trajectory. CBR is a technique, which tries to exploit the similarities of two situations and match decision-making to the best-known precedent cases. The prototype uses the opensource CASPIAN shell developed by the University of Aberystwyth, Wales and is available by anonymous FTP. We will discuss and analyze the preliminary results we have obtained using this CBR tool. Our research suggests that automated information systems can be used to improve the quality of care at the end of life and disseminate expert level 'know how' to palliative care clinicians. We will present how our CBR prototype can be successfully deployed, capable of securely transferring information using a Secure File Transfer Protocol (SFTP) and using a JAVA CBR engine.

Artificial Intelligence↗

Resources for comparing the speed and performance of medical autocoders.

BACKGROUND: Concept indexing is a popular method for characterizing medical text, and is one of the most important early steps in many data mining efforts. Concept indexing differs from simple word or phrase indexing because concepts are typically represented by a nomenclature code that binds a medical concept to all equivalent representations. A concept search on the term renal cell carcinoma would be expected to find occurrences of hypernephroma, and renal carcinoma (concept equivalents). The purpose of this study is to provide freely available resources to compare speed and performance among different autocoders. These tools consist of: 1) a public domain autocoder written in Perl (a free and open source programming language that installs on any operating system); 2) a nomenclature database derived from the unencumbered subset of the publicly available Unified Medical Language System; 3) a large corpus of autocoded output derived from a publicly available medical text. METHODS: A simple lexical autocoder was written that parses plain-text into a listing of all 1,2,3, and 4-word strings contained in text, assigning a nomenclature code for text strings that match terms in the nomenclature. The nomenclature used is the unencumbered subset of the 2003 Unified Medical Language System (UMLS). The unencumbered subset of UMLS was reduced to exclude homonymous one-word terms and proper names, resulting in a term/code data dictionary containing about a half million medical terms. The Online Mendelian Inheritance in Man (OMIM), a 92+ Megabyte publicly available medical opus, was used as sample medical text for the autocoder. RESULTS: The autocoding Perl script is remarkably short, consisting of just 38 command lines. The 92+ Megabyte OMIM file was completely autocoded in 869 seconds on a 2.4 GHz processor (less than 10 seconds per Megabyte of text). The autocoded output file (9,540,442 bytes) contains 367,963 coded terms from OMIM and is distributed with this manuscript. CONCLUSIONS: A public domain Perl script is provided that can parse through plain-text files of any length, matching concepts against an external nomenclature. The script and associated files can be used freely to compare the speed and performance of autocoding software.

Abstracting and Indexing↗

Comparing social work's role in renal dialysis in Israel and the United States: the practice-based research potential of available clinical information.

This paper demonstrates the use of clinical data-mining in a study of social work interventions with dialysis patients in two countries, the US and Israel. We aimed to examine the role of social workers in improving kidney patient outcomes and to determine the potential of readily available patient information for studying this process. The findings showed considerable differences between the patient samples in both countries, as far as the socio-demographic background was considered. In spite of this, there were numerous similarities in the type of psycho-social problems and reactions, as well as the social workers' interventions. Differences which arose in various patient states and outcomes were examined in light of variations in the health care systems and socio-cultural contexts of renal dialysis in both sites.

Adult↗

The olfactory receptor gene superfamily of the mouse.

Olfactory receptor (OR) genes are the largest gene superfamily in vertebrates. We have identified the mouse OR genes from the nearly complete Celera mouse genome by a comprehensive data mining strategy. We found 1,296 mouse OR genes (including 20% pseudogenes), which can be classified into 228 families. OR genes are distributed in 27 clusters on all mouse chromosomes except 12 and Y. One OR gene cluster matches a known locus mediating a specific anosmia, indicating the anosmia may be due directly to the loss of receptors. A large number of apparently functional 'fish-like' Class I OR genes in the mouse genome may have important roles in mammalian olfaction. Human ORs cover a similar 'receptor space' as the mouse ORs, suggesting that the human olfactory system has retained the ability to recognize a broad spectrum of chemicals even though humans have lost nearly two-thirds of the OR genes as compared to mice.

Animals↗