Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Data mining the Arabidopsis genome reveals fifteen 14-3-3 genes. Expression is demonstrated for two out of five novel genes.

In plants, 14-3-3 proteins are key regulators of primary metabolism and membrane transport. Although the current dogma states that 14-3-3 isoforms are not very specific with regard to target proteins, recent data suggest that the specificity may be high. Therefore, identification and characterization of all 14-3-3 (GF14) isoforms in the model plant Arabidopsis are important. Using the information now available from The Arabidopsis Information Resource, we found three new GF14 genes. The potential expression of these three genes, and of two additional novel GF14 genes (Rosenquist et al., 2000), in leaves, roots, and flowers was examined using reverse transcriptase-polymerase chain reaction and cDNA library polymerase chain reaction screening. Under normal growth conditions, two of these genes were found to be transcribed. These genes were named grf11and grf12, and the corresponding new 14-3-3 isoforms were named GF14omicron and GF14iota, respectively. The gene coding for GF14omicron was expressed in leaves, roots, and flowers, whereas the gene coding for GF14iota was only expressed in flowers. Gene structures and relationships between all members of the GF14 gene family were deduced from data available through The Arabidopsis Information Resource. The data clearly support the theory that two 14-3-3 genes were present when eudicotyledons diverged from monocotyledons. In total, there are 15 14-3-3 genes (grfs 1-15) in Arabidopsis, of which 12 (grfs 1-12) now have been shown to be expressed.

14-3-3 Proteins↗

Knowledge discovery in gene-expression-microarray data: mining the information output of the genome.

A key aspect of the genomics revolution is the transformation of large amounts of biological information into an electronic format, leading to an information-based approach to biomedical problems. Large-scale RNA assays and gene-expression-microarray studies, in particular, represent the second wave of the genomics revolution, providing gene-expression data that complement gene-sequence data and help our understanding of the molecular basis of health and disease. They are being applied at several stages in the drug-development process and could ultimately have broad applications in disease diagnosis and patient prognosis.

Animals↗

Inside the data mine: showcasing UR/QA.

The latest look inside San Ramon, CA-based health care consulting company GE Medical Systems Health Care Solutions (HCS--formerly MECON)Ddata Mine shows that health care facilities spend an average of about $29 per adjusted discharge for services rendered by their utilization review and quality assurance (UR/QA) departments. How do you compare?

California↗

TreeGeneBrowser: phylogenetic data mining of gene sequences from public databases.

MOTIVATION: Sequence databases represent an enormous resource of phylogenetic information, but there is a lack of tools for accessing that information in order to assess the amount of evolutionary information in these databases that may be suitable for phylogenetic reconstruction and for identifying areas of the taxonomy that are under-represented for specific gene sequences. RESULTS: We have developed TreeGeneBrowser which allows inspection and evaluation of gene sequence data for phylogenetic reconstruction. This program improves the efficiency of identification of genes that may be useful for particular phylogenetic studies and identifies taxa and taxonomic branches that are under-represented in sequence databases.

Algorithms↗

SwinePan for pig graph-based pangenome and multiomics data mining.

Pigs are one of the most important livestock species worldwide. Although multiple high-quality reference genomes exist, reliance on a single linear reference limits the detection of structural variants (SVs) and the characterization of population-specific genetic diversity. To address this limitation, we developed SwinePan, a comprehensive and integrated multiomics database for pigs built on a graph-based pangenome framework. SwinePan incorporates a variome derived from the graph-based pangenome, covering 2,598 individuals across 35 breeds, including 185,759 SVs, 117 million SNPs, and 6.8 million indels. The database also integrates transcriptomic data from liver, loin muscle, abdominal fat, and backfat, along with over 150,000 phenotypic records. The online toolkit deployed in SwinePan enables genome-wide association studies (GWAS), expression quantitative trait locus (eQTL) mapping, and colocalization, while interactive modules visualize population structure and multiomics associations, streamlining candidate gene and variant exploration. Additionally, two proof-of-concept analyses demonstrate how SwinePan pinpoints trait-associated loci and deciphers their potential regulatory mechanisms.

Journal Article↗

Methods for data mining from large multinational surveillance studies.

Traditionally, large surveillance studies have been analyzed by the use of the MICs at which 90% of isolates tested are inhibited (MIC(90)s), MIC(50)s, frequency distributions, and percent susceptibility. In the past, these approaches have proved satisfactory for the monitoring of resistance. From these traditional uses, one can readily detect an increase in MICs for organism and drug combinations. Now that large surveillance studies have been conducted for a number of years and databases have grown to include a large number of datum points, new approaches to the extraction of useful information from these studies are needed. The present study proposes approaches, including the use of antibiotypes, principal components analysis, phylogenetics, and population genetic analysis, to the evaluation of data from large multinational surveillance studies. Application of these types of analyses can be used to describe genetic diversity, analyze changes in susceptibility patterns over time, and possibly, shed light on the origins and evolution of antimicrobial resistance. As global surveillance studies become more common and new questions concerning the evolution of resistance are raised, innovative approaches to analysis of the data will increase in importance.

Algorithms↗

Data-mining approaches reveal hidden families of proteases in the genome of malaria parasite.

The search for novel antimalarial drug targets is urgent due to the growing resistance of Plasmodium falciparum parasites to available drugs. Proteases are attractive antimalarial targets because of their indispensable roles in parasite infection and development, especially in the processes of host erythrocyte rupture/invasion and hemoglobin degradation. However, to date, only a small number of proteases have been identified and characterized in Plasmodium species. Using an extensive sequence similarity search, we have identified 92 putative proteases in the P. falciparum genome. A set of putative proteases including calpain, metacaspase, and signal peptidase I have been implicated to be central mediators for essential parasitic activity and distantly related to the vertebrate host. Moreover, of the 92, at least 88 have been demonstrated to code for gene products at the transcriptional levels, based upon the microarray and RT-PCR results, and the publicly available microarray and proteomics data. The present study represents an initial effort to identify a set of expressed, active, and essential proteases as targets for inhibitor-based drug design.

Amino Acid Sequence↗

More than 1,000 putative new human signalling proteins revealed by EST data mining.

Cloning procedures aided by homology searches of EST databases have accelerated the pace of discovery of new genes, but EST database searching remains an involved and onerous task. More than 1.6 million human EST sequences have been deposited in public databases, making it difficult to identify ESTs that represent new genes. Compounding the problems of scale are difficulties in detection associated with a high sequencing error rate and low sequence similarity between distant homologues. We have developed a new method, coupling BLAST-based searches with a domain identification protocol, that filters candidate homologues. Application of this method in a large-scale analysis of 100 signalling domain families has led to the identification of ESTs representing more than 1,000 novel human signalling genes. The 4,206 publicly available ESTs representing these genes are a valuable resource for rapid cloning of novel human signalling proteins. For example, we were able to identify ESTs of at least 106 new small GTPases, of which 6 are likely to belong to new subfamilies. In some cases, further analyses of genomic DNA led to the discovery of previously unidentified full-length protein sequences. This is exemplified by the in silico cloning (prediction of a gene product sequence using only genomic and EST sequence data) of a new type of GTPase with two catalytic domains.

Amino Acid Sequence↗

The olfactory receptor gene superfamily: data mining, classification, and nomenclature.

The vertebrate olfactory receptor (OR) subgenome harbors the largest known gene family, which has been expanded by the need to provide recognition capacity for millions of potential odorants. We implemented an automated procedure to identify all OR coding regions from published sequences. This led us to the identification of 831 OR coding regions (including pseudogenes) from 24 vertebrate species. The resulting dataset was subjected to neighbor-joining phylogenetic analysis and classified into 32 distinct families, 14 of which include only genes from tetrapodan species (Class II ORs). We also report here the first identification of OR sequences from a marsupial (koala) and a monotreme (platypus). Analysis of these OR sequences suggests that the ancestral mammal had a small OR repertoire, which expanded independently in all three mammalian subclasses. Classification of "fish-like" (Class I) ORs indicates that some of these ancient ORs were maintained and even expanded in mammals. A nomenclature system for the OR gene superfamily is proposed, based on a divergence evolutionary model. The nomenclature consists of the root symbol 'OR', followed by a family numeral, subfamily letter(s), and a numeral representing the individual gene within the subfamily. For example, OR3A1 is an OR gene of family 3, subfamily A, and OR7E12P is an OR pseudogene of family 7, subfamily E. The symbol is to be preceded by a species indicator. We have assigned the proposed nomenclature symbols for all 330 human OR genes in the database. A WWW tool for automated name assignment is provided.

Animals↗

Data mining: Efficiency of using sequence databases for polymorphism discovery.

An open question in research on Single Nucleotide Polymorphisms (SNPs) is, what is the percentage of true SNPs found by in silico pre-screening? To this end, we selected 13 genes, and determined the complete collection of "true" polymorphisms, or polymorphisms experimentally detected, existing in these genes in our laboratory using Denaturing High Performance Liquid Chromatography (DHPLC) and fluorescent sequencing, or in other laboratories using comparable methods. The genes studied by our group were PTGS2, IGFBP1, IGFBP3, and CYP19. GenBank sequence information was then aligned using two methods, and sequence differences termed "candidate" polymorphisms. We then compared the series of SNPs obtained experimentally and in silico and we have found that in silico methods are relatively specific (up to 55% of candidate SNPs found by SNPFinder have been discovered by experimental procedure) but have low sensitivity (not more than 27% of true SNPs are found by in silico methods).

Aromatase↗

Data mining of public SNP databases for the selection of intragenic SNPs.

Different strategies to search public single nucleotide polymorphism (SNP) databases for intragenic SNPs were evaluated. First, we assembled a strategy to annotate SNPs onto candidate genes based on a BLAST search of public SNP databases (Intragenic SNP Annotation by BLAST, ISAB). Only BLAST hits that complied with stringent criteria according to 1) percentage identity (minimum 98%), 2) BLAST hit length (the hit covers at least 98% of the length of the SNP entry in the database, or the hit is longer than 250 base pairs), and 3) location in non-repetitive DNA, were considered as valid SNPs. We assessed the intragenic context and redundancy of these SNPs, and demonstrated that the SNP content of the dbSNP and HGBASE/HGVbase databases are highly complementary but also overlap significantly. Second, we assessed the validity of intragenic SNP annotation available on the dbSNP and HGVbase websites by comparison with the results of the ISAB strategy. Only a minority of all annotated SNPs was found in common between the respective public SNP database websites and the ISAB annotation strategy. A detailed analysis was performed aiming to explain this discrepancy. As a conclusion, we recommend the application of an independent strategy (such as ISAB) to annotate intragenic SNPs, complementary to the annotation provided at the dbSNP and HGVbase websites. Such an approach might be useful in the selection process of intragenic SNPs for genotyping in genetic studies. Hum Mutat 20:162-173, 2002.

DNA↗

Oligonucleotide microarray data mining: search for age-dependent gene expression.

Information on gene expression in colon tumors versus normal human colon was recently generated by an oligonucleotide microarray study. We used the associated database to search for genes that display age-dependent variations in expression. Statistically significant evidence was obtained that such genes are present in both the tumor and normal tissue databases. Besides the analysis of all genes included in the database, three subsets of genes were analyzed separately: genes controlled by p53, and genes coding for ribosomal proteins and for nuclear-encoded mitochondrial proteins. Among the genes controlled by p53 some show an age-dependent change in expression in tumor tissues, in the sense compatible with an activation of p53 at higher age. A decreased expression of some ribosomal genes at advanced age was detected both in tumor and normal tissues. No significant age-dependent expression could be detected for genes encoding mitochondrial proteins.

Adult↗

A data mining approach for the elucidation of the action of putative etiological agents: application to the non-genotoxic carcinogenicity of genistein.

A procedure designated "the virtual similarity index" (VSI) is described to determine the probability that two or more toxicants are related mechanistically. The approach is structure-activity relationship (SAR) based and generates the virtual toxicological profiles of the chemicals under investigation. It also determines the similarities between them. That commonality is compared to the frequency with which it is found among a population of 10,000 chemicals representing the "universe of chemicals". The similarities between the candidate chemicals and chemicals known to act by other recognized mechanisms are also determined. If the similarities between the candidate chemicals are significantly greater than for the non-related ones, the chemicals are assumed to act by a common mechanism. In that context, the putative non-genotoxic mechanism responsible for the carcinogenicity of genistein (GEN) and its relationship to the action of diethylstilbestrol is examined.

Animals↗

Data mining for indicators of early mortality in a database of clinical records.

This paper describes the analysis of a database of diabetic patients' clinical records and death certificates. The objective of the study was to find rules that describe associations between observations made of patients at their first visit to the hospital and early mortality.Pre-processing was carried out and a knowledge discovery in databases (KDD) package, developed by the Lanner Group and the University of East Anglia, was used for rule induction using simulated annealing.The most significant discovered rules describe an association that was not generally known or accepted by the medical community, however, recent independent studies confirm their validity.

Aged↗