Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Serum proteomic features for detection of endometrial cancer.

To find new potential biomarkers for detection of endometrial cancer (EC), 70 serum samples including 40 from EC patients and 30 from normal healthy females were detected by surface-enhanced laser desorption-ionization time-of-flight mass spectrometry (SELDI-TOF-MS) using WCX2 (weak cation exchange) protein chip. Mass spectra were then assessed with three powerful data-mining tools: a tree classifier, Biomarker Wizard software, and Biomarker Patterns System. The diagnostic pattern combined with 13 potential biomarkers could differentiate EC patients from healthy persons, with a specificity of 100%, sensitivity of 92.5%, and total coincidence of 95.7%. The combination of surface-enhanced laser desorption-ionization with bioinformatics tools could help find new biomarkers and establish with high sensitivity and specificity for the detection of EC.

Adult↗

Modulatory effects of plant phenols on human multidrug-resistance proteins 1, 4 and 5 (ABCC1, 4 and 5).

Plant flavonoids are polyphenolic compounds, commonly found in vegetables, fruits and many food sources that form a significant portion of our diet. These compounds have been shown to interact with several ATP-binding cassette transporters that are linked with anticancer and antiviral drug resistance and, as such, may be beneficial in modulating drug resistance. This study investigates the interactions of six common polyphenols; quercetin, silymarin, resveratrol, naringenin, daidzein and hesperetin with the multidrug-resistance-associated proteins, MRP1, MRP4 and MRP5. At nontoxic concentrations, several of the polyphenols were able to modulate MRP1-, MRP4- and MRP5-mediated drug resistance, though to varying extents. The polyphenols also reversed resistance to NSC251820, a compound that appears to be a good substrate for MRP4, as predicted by data-mining studies. Furthermore, most of the polyphenols showed direct inhibition of MRP1-mediated [3H]dinitrophenyl S-glutathione and MRP4-mediated [3H]cGMP transport in inside-out vesicles prepared from human erythrocytes. Also, both quercetin and silymarin were found to inhibit MRP1-, MRP4- and MRP5-mediated transport from intact cells with high affinity. They also had significant effects on the ATPase activity of MRP1 and MRP4 without having any effect on [32P]8-azidoATP[alphaP] binding to these proteins. This suggests that these flavonoids most likely interact at the transporter's substrate-binding sites. Collectively, these results suggest that dietary flavonoids such as quercetin and silymarin can modulate transport activities of MRP1, -4 and -5. Such interactions could influence bioavailability of anticancer and antiviral drugs in vivo and thus, should be considered for increasing efficacy in drug therapies.

ATP Binding Cassette Transporter, Subfamily B↗

An informatics infrastructure for patient safety and evidence-based practice in home healthcare.

The informatics infrastructure for patient safety and evidence-based practice (EBP) in home healthcare comprises data acquisition methods, healthcare standards including standardized terminologies, data repositories and clinical event monitors, data-mining techniques, digital sources of evidence, and communication technologies. Although the components of an informatics infrastructure are available and applications that bring these components together to promote patient safety and enable EBP have demonstrated positive or promising results in the acute care setting, a number of challenges hinder implementation in home healthcare. Resolution of these challenges requires commitment and collaboration among key stakeholders.

Benchmarking↗

Genetic characterization of the resorcinol catabolic pathway in Corynebacterium glutamicum.

Corynebacterium glutamicum grew on resorcinol as a sole source of carbon and energy. By genome-wide data mining, two gene clusters, designated NCgl1110-NCgl1113 and NCgl2950-NCgl2953, were proposed to encode putative proteins involved in resorcinol catabolism. Deletion of the NCgl2950-NCgl2953 gene cluster did not result in any observable phenotype changes. Disruption and complementation of each gene at NCgl1110-NCgl1113, NCgl2951, and NCgl2952 indicated that these genes were involved in resorcinol degradation. Expression of NCgl1112, NCgl1113, and NCgl2951 in Escherichia coli revealed that NCgl1113 and NCgl2951 both coded for hydroxyquinol 1,2-dioxygenases and NCgl1112 coded for maleylacetate reductases. NCgl1111 encoded a putative monooxygenase, but this putative hydroxylase was very different from previously functionally identified hydroxylases. Cloning and expression of NCgl1111 in E. coli revealed that NCgl1111 encoded a resorcinol hydroxylase that needs NADPH as a cofactor. E. coli cells containing Ncgl1111 and Ncgl1113 sequentially converted resorcinol into maleylacetate. NCgl1110 and NCgl2950 both encoded putative TetR family repressors, but only NCgl1110 was transcribed and functional. NCgl2953 encoded a putative transporter, but disruption of this gene did not affect resorcinol degradation by C. glutamicum. The function of NCgl2953 remains unclear.

Bacterial Proteins↗

Functional identification of novel genes involved in the glutathione-independent gentisate pathway in Corynebacterium glutamicum.

Corynebacterium glutamicum used gentisate and 3-hydroxybenzoate as its sole carbon and energy source for growth. By genome-wide data mining, a gene cluster designated ncg12918-ncg12923 was proposed to encode putative proteins involved in gentisate/3-hydroxybenzoate pathway. Genes encoding gentisate 1,2-dioxygenase (ncg12920) and fumarylpyruvate hydrolase (ncg12919) were identified by cloning and expression of each gene in Escherichia coli. The gene of ncg12918 encoding a hypothetical protein (Ncg12918) was proved to be essential for gentisate-3-hydroxybenzoate assimilation. Mutant strain RES167Deltancg12918 lost the ability to grow on gentisate or 3-hydroxybenzoate, but this ability could be restored in C. glutamicum upon the complementation with pXMJ19-ncg12918. Cloning and expression of this ncg12918 gene in E. coli showed that Ncg12918 is a glutathione-independent maleylpyruvate isomerase. Upstream of ncg12920, the genes ncg12921-ncg12923 were located, which were essential for gentisate and/or 3-hydroxybenzoate catabolism. The Ncg12921 was able to up-regulate gentisate 1,2-dioxygenase, maleylpyruvate isomerase, and fumarylpyruvate hydrolase activities. The genes ncg12922 and ncg12923 were deduced to encode a gentisate transporter protein and a 3-hydroxybenzoate hydroxylase, respectively, and were essential for gentisate or 3-hydroxybenzoate assimilation. Based on the results obtained in this study, a GSH-independent gentisate pathway was proposed, and genes involved in this pathway were identified.

Bacterial Proteins↗

Genetically determined susceptibility to tuberculosis in mice causally involves accelerated and enhanced recruitment of granulocytes.

Classical twin studies and recent linkage analyses of African populations have revealed a potential involvement of host genetic factors in susceptibility or resistance to Mycobacterium tuberculosis infection. In order to identify the candidate genes involved and test their causal implication, we capitalized on the mouse model of tuberculosis, since inbred mouse strains also differ substantially in their susceptibility to infection. Two susceptible and two resistant mouse strains were aerogenically infected with 1,000 CFU of M. tuberculosis, and the regulation of gene expression was examined by Affymetrix GeneChip U74A array with total lung RNA 2 and 4 weeks postinfection. Four weeks after infection, 96 genes, many of which are involved in inflammatory cell recruitment and activation, were regulated in common. One hundred seven genes were differentially regulated in susceptible mouse strains, whereas 43 genes were differentially expressed only in resistant mice. Data mining revealed a bias towards the expression of genes involved in granulocyte pathophysiology in susceptible mice, such as an upregulation of those for the neutrophil chemoattractant LIX (CXCL5), interleukin 17 receptor, phosphoinositide kinase 3 delta, or gamma interferon-inducible protein 10. Following M. tuberculosis challenge in both airways or peritoneum, granulocytes were recruited significantly faster and at higher numbers in susceptible than in resistant mice. When granulocytes were efficiently depleted by either of two regimens at the onset of infection, only susceptible mice survived aerosol challenge with M. tuberculosis significantly longer than control mice. We conclude that initially enhanced recruitment of granulocytes contributes to susceptibility to tuberculosis.

Animals↗

Molecular characteristics and global spread of Mycobacterium tuberculosis with a western cape F11 genotype.

In order to fully understand the global tuberculosis (TB) epidemic it is important to investigate the population structure and dissemination of the causative agent that drives the epidemic. Mycobacterium tuberculosis strain family 11 (F11) genotype isolates (found in 21.4% of all infected patients) are at least as successful as the Beijing genotype family isolates (16.5%) in contributing to the TB problem in some Western Cape communities of South Africa. This study describes key molecular characteristics that define the F11 genotype. A data-mining approach coupled with additional molecular analysis showed that members of F11 can easily and uniquely be identified by PCR-based techniques such as spoligotyping and dot blot screening for a specific rrs491 polymorphism. Isolates of F11 not only are a major contributor to the TB epidemic in South Africa but also are present in four different continents and at least 25 other countries in the world. Careful study of dominant compared to rare strains should provide clues to their success and possibly provide new ideas for combating TB.

Genotype↗

Discovery and characterization of complete genomes of 38 head-tailed proviruses in four predominant phyla of archaea.

Archaea play a significant role in natural ecosystems and the human body. Archaeal viruses exert a considerable influence on the structure and composition of archaeal communities and their associated ecological environments. The present study revealed the complete genomes of 38 archaeal head-tailed proviruses through comprehensive data mining. The hosts of these proviruses were identified as belonging to the following four dominant phyla: Halobacteriota, Thermoplasmatota, Thermoproteota, and Nanoarchaeota. In addition to the 14 proviruses of halophilic archaea related to the Graaviviridae family, the remaining proviruses exhibited limited genetic similarities to known (pro)viruses, suggesting the existence of 14 potential novel families. Of the 38 archaeal proviruses, 30 have the potential to lyse host cells. Eleven proviruses contain genes linked to antiviral defense mechanisms, including those involved in restriction modification (RM), clustered regularly interspaced short palindromic repeat (CRISPR)-associated (CRISPR-Cas) nucleases, defense island system associated with restriction-modification (DISARM), and DNA degradation (Dnd). Moreover, auxiliary metabolic genes were identified in the proviruses of Bathyarchaeia and Halobacteriota archaea, including those involved in carbohydrate and amino acid metabolism. Our findings indicate the diversity of archaeal viruses, their interactions with archaeal hosts, and their roles in the adaptation of the host.IMPORTANCEThe field of archaeal virology has seen a rapid expansion through the use of metagenomics, yet the diversity of these viruses remains largely uncharted. In this study, the complete genomes of 38 novel archaeal proviruses were identified for the following four dominant phyla: Halobacteriota, Thermoplasmatota, Thermoproteota, and Nanoarchaeota. Two families and six genera of Archaea were the first to be identified as hosts for viruses. The proviruses were found to contain diverse genes that were involved in distinct adaptation strategies of viruses to hosts. Our findings contribute to the expansion of the lineages of archaeal viruses and highlight their intricate interactions and essential roles in enabling host survival and adaptation to diverse environmental conditions.

Archaea↗

Online practice guidelines: issues, obstacles, and future prospects.

The "guidelines movement" was formed to reduce variability in practice, control costs, and improve patient care outcomes. Yet the overall impact on practice and outcomes has been disappointing. Evidence demonstrates that the most effective method of stimulating awareness of and compliance with best practices is computer-generated reminders provided at the point of care. This paper reviews five steps along the path from the development of a guideline to its integration into practice and the subsequent evaluation of its impact on practice and outcomes. Issues arising at each step and obstacles to moving from one step to the next are described. Last, developments that could help overcome the obstacles are highlighted. These include 1) more rapid knowledge acquisition using data mining, 2) better accommodation to imprecise knowledge in clinical algorithms using fuzzy logic, 3) development of a shareable model for guideline representation and execution, and 4) more widespread availability of clinically robust information systems that support decision-making at the point of care.

Algorithms↗

Opportunities at the intersection of bioinformatics and health informatics: a case study.

This paper provides a "viewpoint discussion" based on a presentation made to the 2000 Symposium of the American College of Medical Informatics. It discusses potential opportunities for researchers in health informatics to become involved in the rapidly growing field of bioinformatics, using the activities of the Yale Center for Medical Informatics as a case study. One set of opportunities occurs where bioinformatics research itself intersects with the clinical world. Examples include the correlations between individual genetic variation with clinical risk factors, disease presentation, and differential response to treatment; and the implications of including genetic test results in the patient record, which raises clinical decision support issues as well as legal and ethical issues. A second set of opportunities occurs where bioinformatics research can benefit from the technologic expertise and approaches that informaticians have used extensively in the clinical arena. Examples include database organization and knowledge representation, data mining, and modeling and simulation. Microarray technology is discussed as a specific potential area for collaboration. Related questions concern how best to establish collaborations with bioscientists so that the interests and needs of both sets of researchers can be met in a synergistic fashion, and the most appropriate home for bioinformatics in an academic medical center.

Computational Biology↗

Simulating the growth of viruses.

To explore how the genome of an organism defines its growth, we have developed a computer simulation for the intracellular growth of phage T7 on its E. coli host. Our simulation, which incorporates 30 years of genetic, biochemical, physiological, and biophysical data, is used here to study how the intracellular resources of the host, determined by the specific growth rate of the host, contribute toward phage development. It is also used to probe how changes in the linear organization of genetic elements on the T7 genome can affect T7 development. Further, we show how time-series trajectories of T7 mRNA and protein levels generated by the simulation may be used as raw data to test data-mining strategies, specifically, to identify partners in protein-protein interactions. Finally, we suggest how generalization of this work can lead to a knowledge-driven simulation for the growth of any virus.

Algorithms↗

Exploring alternative transcript structure in the human genome using blocks and InterPro.

Understanding how alternative splicing affects gene function is an important challenge facing modern-day molecular biology. Using homology-based, protein sequence analysis methods, it should be possible to investigate how transcript diversity impacts protein function. To test this, high-quality exon-intron structures were deduced for over 8000 human genes, including over 1300 (17 percent) that produce multiple transcript variants. A data mining technique (DiffMotif) was developed to identify genes in which transcript variation coincides with changes in conserved motifs between variants. Applying this method, we found that 30 percent of the multi-variant genes in our test set exhibited a differential profile of conserved InterPro and/or BLOCKS motifs across different mRNA variants. To investigate these, a visualization tool (ProtAnnot) that displays amino acid motifs in the context of genomic sequence was developed. Using this tool, genes revealed by the DiffMotif method were analyzed, and when possible, hypotheses regarding the potential role of alternative transcript structure in modulating gene function were developed. Examples of these, including: MEOX1, a homeobox-containing protein; AIRE, involved in auto-immune disease; PLAT, tissue type plasminogen activator; and CD79b, a component of the B-cell receptor complex, are presented. These results demonstrate that amino acid motif databases like BLOCKS and InterPro are useful tools for investigating how alternative transcript structure affects gene function.

Algorithms↗

Predicting risk of coronary artery disease from DNA microarray-based genotyping using neural networks and other statistical analysis tool.

This paper presents a novel approach for complex disease prediction that we have developed, exemplified by a study on risk of coronary artery disease (CAD). This multi-disciplinary approach straddles fields of microarray technology and genetics, neural networks (NN), data mining and machine learning, as well as traditional statistical analysis techniques, namely principal components analysis (PCA) and factor analysis (FA). A description of the biological background of the study is given, followed by a detailed description of how the problem has been modeled for analyses by neural networks and FA. A committee learning approach for NN has been used to improve generalization rates. We show that our NN approach is able to yield promising prediction results despite using only the most fundamental network structures. More interestingly, through the statistical analysis process, genes of similar biological functions have been clustered. In addition, a gene marker involved in breaking down lipids has been found to be the most correlated to CAD.

Algorithms↗

Microarray analysis of the rat lacrimal gland following the loss of parasympathetic control of secretion.

Previous studies showed that loss of muscarinic parasympathetic input to the lacrimal gland (LG) leads to a dramatic reduction in tear secretion and profound changes to LG structure. In this study, we used DNA microarrays to examine the regulation of the gene expression of the genes for secretory function and organization of the LG. Long-Evans rats anesthetized with a mixture of ketamine/xylazine (80:10 mg/kg) underwent unilateral sectioning of the greater superficial petrosal nerve, the input to the pterygopalatine ganglion. After 7 days, tear secretion was measured, the animals were killed, and structural changes in the LG were examined by light microscopy. Total RNA from control and experimental LGs (n = 5) was used for DNA microarray analysis employing the U34A GeneChip. Three statistical algorithms (detection, change call, and signal log ratio) were used to determine differential gene expression using the Microarray Suite (5.0) and Data Mining Tools (3.0). Tear secretion was significantly reduced and corneal ulcers developed in all experimental eyes. Light microscopy showed breakdown of the acinar structure of the LG. DNA microarray analysis showed downregulation of genes associated with the endoplasmic reticulum and Golgi, including genes involved in protein folding and processing. Conversely, transcripts for cytoskeleton and extracellular matrix components, inflammation, and apoptosis were upregulated. The number of significantly upregulated genes (116) was substantially greater than the number of downregulated genes (49). Removal of the main secretory input to the rat LG resulted in clinical symptoms associated with severe dry eye. Components of the secretory pathway were negatively affected, and the increase in cell proliferation and inflammation may lead to loss of organization in the parasympathectomized lacrimal gland.

Algorithms↗

Bioinformatics Resources for In Silico Proteome Analysis.

In the growing field of proteomics, tools for the in silico analysis of proteins and even of whole proteomes are of crucial importance to make best use of the accumulating amount of data. To utilise this data for healthcare and drug development, first the characteristics of proteomes of entire species-mainly the human-have to be understood, before secondly differentiation between individuals can be surveyed. Specialised databases about nucleic acid sequences, protein sequences, protein tertiary structure, genome analysis, and proteome analysis represent useful resources for analysis, characterisation, and classification of protein sequences. Different from most proteomics tools focusing on similarity searches, structure analysis and prediction, detection of specific regions, alignments, data mining, 2D PAGE analysis, or protein modelling, respectively, comprehensive databases like the proteome analysis database benefit from the information stored in different databases and make use of different protein analysis tools to provide computational analysis of whole proteomes.

Journal Article↗

Identification of a PAX-FKHR gene expression signature that defines molecular classes and determines the prognosis of alveolar rhabdomyosarcomas.

Alveolar rhabdomyosarcomas (ARMS) are aggressive soft-tissue sarcomas affecting children and young adults. Most ARMS tumors express the PAX3-FKHR or PAX7-FKHR (PAX-FKHR) fusion genes resulting from the t(2;13) or t(1;13) chromosomal translocations, respectively. However, up to 25% of ARMS tumors are fusion negative, making it unclear whether ARMS represent a single disease or multiple clinical and biological entities with a common phenotype. To test to what extent PAX-FKHR determine class and behavior of ARMS, we used oligonucleotide microarray expression profiling on 139 primary rhabdomyosarcoma tumors and an in vitro model. We found that ARMS tumors expressing either PAX-FKHR gene share a common expression profile distinct from fusion-negative ARMS and from the other rhabdomyosarcoma variants. We also observed that PAX-FKHR expression above a minimum level is necessary for the detection of this expression profile. Using an ectopic PAX3-FKHR and PAX7-FKHR expression model, we identified an expression signature regulated by PAX-FKHR that is specific to PAX-FKHR-positive ARMS tumors. Data mining for functional annotations of signature genes suggested a role for PAX-FKHR in regulating ARMS proliferation and differentiation. Cox regression modeling identified a subset of genes within the PAX-FKHR expression signature that segregated ARMS patients into three risk groups with 5-year overall survival estimates of 7%, 48%, and 93%. These prognostic classes were independent of conventional clinical risk factors. Our results show that PAX-FKHR dictate a specific expression signature that helps define the molecular phenotype of PAX-FKHR-positive ARMS tumors and, because it is linked with disease outcome in ARMS patients, determine tumor behavior.

Biomarkers, Tumor↗

The interaction of four genes in the inflammation pathway significantly predicts prostate cancer risk.

It is widely hypothesized that the interactions of multiple genes influence individual risk to prostate cancer. However, current efforts at identifying prostate cancer risk genes primarily rely on single-gene approaches. In an attempt to fill this gap, we carried out a study to explore the joint effect of multiple genes in the inflammation pathway on prostate cancer risk. We studied 20 genes in the Toll-like receptor signaling pathway as well as several cytokines. For each of these genes, we selected and genotyped haplotype-tagging single nucleotide polymorphisms (SNP) among 1,383 cases and 780 controls from the CAPS (CAncer Prostate in Sweden) study population. A total of 57 SNPs were included in the final analysis. A data mining method, multifactor dimensionality reduction, was used to explore the interaction effects of SNPs on prostate cancer risk. Interaction effects were assessed for all possible n SNP combinations, where n = 2, 3, or 4. For each n SNP combination, the model providing lowest prediction error among 100 cross-validations was chosen. The statistical significance levels of the best models in each n SNP combination were determined using permutation tests. A four-SNP interaction (one SNP each from IL-10, IL-1RN, TIRAP, and TLR5) had the lowest prediction error (43.28%, P = 0.019). Our ability to analyze a large number of SNPs in a large sample size is one of the first efforts in exploring the effect of high-order gene-gene interactions on prostate cancer risk, and this is an important contribution to this new and quickly evolving field.

Case-Control Studies↗

Applications for microarrays in renal biology and medicine.

Groundbreaking recent developments, such as the near completion of human and mouse genome sequencing efforts and the emergence of robust microarray (gene chip) technologies, enabling comprehensive analysis of transcriptomes, provide new opportunities of unprecedented scale for researchers of kidney biology and disease. Combined with advanced computational and mathematical approaches for microarray data analysis, microarray applications promise to revolutionize our understanding of molecular mechanisms of kidney development and renal pathogenesis. New knowledge in this field will facilitate new approaches for molecular diagnostics, drug discovery, and eventually "personalized" renal medicine. In this review, we outline current and future research applications of microarray and computational approaches in renal biology and disease. We describe basic steps in microarray data analysis and introduce advanced computational approaches to optimize data mining of vast microarray datasets.

Genomics↗