Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Medical data mining using evolutionary computation.

In this paper, we introduce a system for discovering medical knowledge by learning Bayesian networks and rules. Evolutionary computation is used as the search algorithm. The Bayesian networks can provide an overall structure of the relationships among the attributes. The rules can capture detailed and interesting patterns in the database. The system is applied to real-life medical databases for limb fracture and scoliosis. The knowledge discovered provides insights to and allows better understanding of these two medical domains.

Algorithms↗

Data mining for simple sequence repeats in expressed sequence tags from barley, maize, rice, sorghum and wheat.

Plant genomics projects involving model species and many agriculturally important crops are resulting in a rapidly increasing database of genomic and expressed DNA sequences. The publicly available collection of expressed sequence tags (ESTs) from several grass species can be used in the analysis of both structural and functional relationships in these genomes. We analyzed over 260000 EST sequences from five different cereals for their potential use in developing simple sequence repeat (SSR) markers. The frequency of SSR-containing ESTs (SSR-ESTs) in this collection varied from 1.5% for maize to 4.7% for rice. In addition, we identified several ESTs that are related to the SSR-ESTs by BLAST analysis. The SSR-ESTs and the related sequences were clustered within each species in order to reduce the redundancy and to produce a longer consensus sequence. The consensus and singleton sequences from each species were pooled and clustered to identify cross-species matches. Overall a reduction in the redundancy by 85% was observed when the resulting consensus and singleton sequences (3569) were compared to the total number of SSR-EST and related sequences analyzed (24 606). This information can be useful for the development of SSR markers that can amplify across the grass genera for comparative mapping and genetics. Functional analysis may reveal their role in plant metabolism and gene evolution.

Computational Biology↗

Association of genes to genetically inherited diseases using data mining.

Although approximately one-quarter of the roughly 4,000 genetically inherited diseases currently recorded in respective databases (LocusLink, OMIM) are already linked to a region of the human genome, about 450 have no known associated gene. Finding disease-related genes requires laborious examination of hundreds of possible candidate genes (sometimes, these are not even annotated; see, for example, refs 3,4). The public availability of the human genome draft sequence has fostered new strategies to map molecular functional features of gene products to complex phenotypic descriptions, such as those of genetically inherited diseases. Owing to recent progress in the systematic annotation of genes using controlled vocabularies, we have developed a scoring system for the possible functional relationships of human genes to 455 genetically inherited diseases that have been mapped to chromosomal regions without assignment of a particular gene. In a benchmark of the system with 100 known disease-associated genes, the disease-associated gene was among the 8 best-scoring genes with a 25% chance, and among the best 30 genes with a 50% chance, showing that there is a relationship between the score of a gene and its likelihood of being associated with a particular disease. The scoring also indicates that for some diseases, the chance of identifying the underlying gene is higher.

Chromosome Mapping↗

Data mining the p53 pathway in the Fugu genome: evidence for strong conservation of the apoptotic pathway.

The p53 tumour suppressor gene belongs to a small family of related proteins that includes two other members, p63 and p73. Phylogenetic and functional studies suggest that p63 and p73 are ancient genes that have essential roles in normal development, whereas p53 seems to have evolved more recently to prevent cell transformation. In mammalian cells, a plethora of proteins have been found to specifically regulate p53 activity. The genome of the fish Fugu rubripes has been recently published. It is the second vertebrate genome for which the entire sequence is now available. Phylogenetic studies are essential in order to analyse and define signalling pathways important for cell cycle regulation. The presence or absence of a critical member in any pathway can shed light about the evolution of these pathways. The Fugu genome databank has been analysed for several members of the p53 network, including p53, p63 and p73. A good conservation of the network that regulates p53 stability and apoptosis has been found. We also discovered that some cofactors that cooperate with p53 for apoptosis are also well conserved and belong to multigene families not detected in the human genome.

Animals↗

G-language Genome Analysis Environment: a workbench for nucleotide sequence data mining.

SUMMARY: G-language Genome Analysis Environment (G-language GAE) is an open source generic software package aimed for higher efficiency in bioinformatics analysis. G-language GAE has an interface as a set of Perl libraries for software development, and a graphical user interface for easy manipulation. Both Windows and Linux versions are available. AVAILABILITY: From http://www.g-language.org/ under GNU General Public License. CD-ROMs are distributed freely in major conferences.

Database Management Systems↗

Data mining techniques to study the disulfide-bonding state in proteins: signal peptide is a strong descriptor.

In the eucaryotic cell, the formation of disulfide bonds takes place in general inside the endoplasmic reticulum which provides a unique folding environment. The DisulfideDB database gathers information about this biological process with structural, evolutionary and neighborhood information on cysteines in proteins. Mining this information with an association rule discovery program permits to extract some strong rules for the prediction of the disulfide-bonding state of cysteines.

Binding Sites↗

Predicting osteoarthritic knee rehabilitation outcome by using a prediction model developed by data mining techniques.

Artificial neural networks (ANN) have been applied to assist in clinical decision-making and prediction. While we consider possible effective treatments for patients with osteoarthritic knee such as Transcutaneous Electrical Nerve Stimulation (TENS), exercise, and TENS with exercise respectively, we have to select a treatment protocol for patients such that they would gain the best improvements according to their clinical conditions. To facilitate this functionality with the existing patient assessment, we hope to apply the ANN programming techniques to develop a computerized prediction system. A preliminary validation was performed to test the validity of the newly developed prediction protocol on knee rehabilitation. We input the key clinical attributes of 62 patients who have undergone the three above-mentioned knee treatments to the protocol. The expected pain improvement of each patient as predicted by the protocol was obtained. Spearman rank-order correlation was used to identify whether there was a significant correlation between the rankings of the observed and expected pain improvement. We found that the Spearman's rho was 0.424, which is statistically significant at p < 0.001. From this preliminary analysis, we are confident that this newly developed prediction protocol will be useful when deciding which treatment regime best suits a patient.

Combined Modality Therapy↗

Experimental validation of data mined single nucleotide polymorphisms from several databases and consecutive dbSNP builds.

Rapid development in the annotation of human genetic variation has increased the numbers of single nucleotide polymorphisms (SNPs) in candidate genes by several orders of magnitude. The selection of both useful target SNPs for disease-gene association studies and SNPs associated with the treatment response is therefore an increasingly challenging task. We describe a workflow for selecting SNPs based on their putative function and frequency in candidate genes extracted from PubMed resources. The annotation of each SNP and its frequency in a Caucasian population was assessed in several databases. Approximately 4000 SNPs were identified from an initial 233 candidate genes. In a case study, we performed actual genotyping of 1030 of these SNPs in 213 genes and obtained 710 successfully genotyped SNPs. Using the flow-chart outlined here, only 87 SNPs were monomorphic (approximately 12%). This study reports the frequency of SNPs in a Caucasian population, selected in silico, using a candidate gene approach and validated by actually genotyping 193 individuals. The selected genotypes represent a valuable set of verified candidate SNPs for pharmacogenetic studies in Caucasian populations.

Breast Neoplasms↗

Exploring relationships and mining data with the UCSC Gene Sorter.

In parallel with the human genome sequencing and assembly effort, many tools have been developed to examine the structure and function of the human gene set. The University of California Santa Cruz (UCSC) Gene Sorter has been created as a gene-based counterpart to the chromosome-oriented UCSC Genome Browser to facilitate the study of gene function and evolution. This simple, but powerful tool provides a graphical display of related genes that can be sorted and filtered based on a variety of criteria. Genes may be ordered based on such characteristics as expression profiles, proximity in genome, shared Gene Ontology (GO) terms, and protein similarity. The display can be restricted to a gene set meeting a specific set of constraints by filtering on expression levels, gene name or ID, chromosomal position, and so on. The default set of information for each gene entry-gene name, selected expression data, a BLASTP E-value, genomic position, and a description-can be configured to include many other types of data, including expanded expression data, related accession numbers and IDs, orthologs in other species, GO terms, and much more. The Gene Sorter, a CGI-based Web application written in C with a MySQL database, is tightly integrated with the other applications in the UCSC Genome Browser suite. Available on a selected subset of the genome assemblies found in the Genome Browser, it further enhances the usefulness of the UCSC tool set in interactive genomic exploration and analysis.

Computational Biology↗

Predicting crystal structures with data mining of quantum calculations.

Predicting and characterizing the crystal structure of materials is a key problem in materials research and development. It is typically addressed with highly accurate quantum mechanical computations on a small set of candidate structures, or with empirical rules that have been extracted from a large amount of experimental information, but have limited predictive power. In this Letter, we transfer the concept of heuristic rule extraction to a large library of ab initio calculated information, and we demonstrate that this can be developed into a tool for crystal structure prediction.

Journal Article↗