Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Data mining of VDJ genes reveals interesting clues.

Hypervariability of the complementary determining regions in characteristic structure of Immunoglobulins and the distinct, cell-specific expressions of the genes coding for this important class of proteins pose intriguing problems in experimental and computational/informatics research requiring a special approach different from those for the other proteins. We present here an Average Linkage Hierarchical Clustering of the Homosapien VDJ genes and the Immunoglobulin polypeptides generated by them using special kind of data structures and correlation matrices in place of the microarray data. The results reveal interesting clues on the heterogeneity of exon - intron locations in these gene-families and its possible role in hypervariability of the Immunoglobulins.

Chromosome Mapping↗

Genome data mining of lactic acid bacteria: the impact of bioinformatics.

Lactic acid bacteria (LAB) have been widely used in food fermentations and, more recently, as probiotics in health-promoting food products. Genome sequencing and functional genomics studies of a variety of LAB are now rapidly providing insights into their diversity and evolution and revealing the molecular basis for important traits such as flavor formation, sugar metabolism, stress response, adaptation and interactions. Bioinformatics plays a key role in handling, integrating and analyzing the flood of 'omics' data being generated. Reconstruction of metabolic potential using bioinformatics tools and databases, followed by targeted experimental verification and exploration of the metabolic and regulatory network properties, are the present challenges that should lead to improved exploitation of these versatile food bacteria.

Adaptation, Biological↗

On distance computation in space of mixed-type variables in medical data mining.

In the next study we consider two distance metrics that were presented in the machine leaming literature for mixed-type variables. We show that they are not really metrics, but pseudometrics. The problem arose from missing values. The metrics can be redefined to satisfy the metricity definition. Distance computation can then be performed reliably without a possibility that a distance between exactly similar patient cases would not be zero. We experimented with both procedures and so-called ignore mode using two medical data sets and observed that the redefined procedure was as good as the original approach as measured with the classification ability of one-nearest neighbour searching. It is noticeable that the described problem caused by missing values can be with any distance measure if its missing values are treated in the pessimistic manner as in the original version of these distance measures.

Computer Simulation↗

Developing a web-based data mining application to impact community health improvement initiatives: the Virginia Atlas of Community Health.

This article describes how a team from the Virginia Department of Health (VDH) and the Virginia Center for Healthy Communities (VCHC) attended the UNC Management Academy for Public Health to learn skills to address Virginia's commitment to using technology to improve the public's health. After creating a business plan for a food-safety information Web site, team members used that experience as well as Management Academy training in information technology, the management of data and finances, and strategic partnering to create a comprehensive tool with which to place customizable population data in the hands of anyone interested in pursuing population health improvement. The Virginia Atlas of Community Health, launched through the VCHC in 2003, places clear, compelling data in the hands of those who can influence decisions at the local level and create the most impact for health. Since the program's inception, more than 2,000 individuals have registered as ongoing users of the Virginia Atlas. Initially funded by a Turning Point grant from the Robert Wood Johnson Foundation, the program is sustained through a series of smaller grants and funding from the VDH.

Atlases as Topic↗

SwinePan for pig graph-based pangenome and multiomics data mining.

Pigs are one of the most important livestock species worldwide. Although multiple high-quality reference genomes exist, reliance on a single linear reference limits the detection of structural variants (SVs) and the characterization of population-specific genetic diversity. To address this limitation, we developed SwinePan, a comprehensive and integrated multiomics database for pigs built on a graph-based pangenome framework. SwinePan incorporates a variome derived from the graph-based pangenome, covering 2,598 individuals across 35 breeds, including 185,759 SVs, 117 million SNPs, and 6.8 million indels. The database also integrates transcriptomic data from liver, loin muscle, abdominal fat, and backfat, along with over 150,000 phenotypic records. The online toolkit deployed in SwinePan enables genome-wide association studies (GWAS), expression quantitative trait locus (eQTL) mapping, and colocalization, while interactive modules visualize population structure and multiomics associations, streamlining candidate gene and variant exploration. Additionally, two proof-of-concept analyses demonstrate how SwinePan pinpoints trait-associated loci and deciphers their potential regulatory mechanisms.

Journal Article↗

Microarray data mining using gene ontology.

DNA microarray technology allows scientists to study the expression of thousands of genes--potentially entire genomes--simultaneously. However the large number of genes, variety of statistical methods employed and the complexity of biologic systems complicate analysis of microarray results. We have developed a web based environment that simplifies the presentation of microarray results by combining microarray results processed for statistical significance with probe set annotation by Genbank, NCBI RefSeqs, GeneCards and the Gene Ontology. This allows rapid examination and classification of microarray experiments--annotated by NCIBI tools --by Statistical Significance and Gene Oncology Classes. By providing a simple, easily understood interface to large microarray data sets, this tool has been particularly useful for small research groups focused on a small number of related genes and for researchers who want to ask simple questions without the overhead of complex data management and analysis.

Computational Biology↗

In vivo metabolite detection and identification in drug discovery via LC-MS/MS with data-dependent scanning and postacquisition data mining.

An important aspect in drug discovery is the early structural identification of the metabolites of potential new drugs. This gives information on the metabolically labile points in the molecules under investigation, suggesting structural modifications to improve their metabolic stability, and allowing an early safety assessment via the identification of metabolic activation products. From an analytical point of view, metabolite identification still remains a challenging task, especially for in vivo samples, in which they occur at trace levels together with high amounts of endogenous compounds. Here we describe a method, based on LC-ion trap tandem MS, for the rapid in vivo metabolite identification. It is based on the automatic, data-dependent acquisition of multiple product ion MS/MS scans, followed by a postacquisition search, within the entire MS/MS data set obtained, for specific neutral losses or marker ions in the tandem mass spectra of parent molecule and putative metabolites. One advantage of the method is speed, since it requires minimum sample preparation and all the necessary data can be obtained in one chromatographic run. In addition, it is highly sensitive and selective, allowing detection of trace metabolites even in the presence of a complex matrix. As an example of application, we present the studies of the in vivo metabolism of the compound MEN 15916 (1). The method allowed identification of monohydroxy ([M + H](+) = m/z 655), dihydroxy ([M + H](+) = m/z 671), and trihydroxy ([M + H](+) = m/z 687) metabolites, as well as some unexpected biotransformation products such as a carboxylic acid ([M + H](+) = m/z 669), a N-dealkylated metabolite ([M + H](+) = m/z 541), and its hydroxy-analog ([M + H](+) = m/z 557).

Animals↗

CGH-Profiler: data mining based on genomic aberration profiles.

BACKGROUND: CGH-Profiler is a program that supports the analysis of genomic aberrations measured by Comparative Genomic Hybridisation (CGH). Comparative genomic hybridisation (CGH) is a well-established, molecular cytogenetic method that allows the detection of chromosomal imbalances in entire genomes. This technique is widely used in routine molecular diagnostics. Typically, chromosomal imbalances are described in a complex syntax based on the International Standard for Cytogenetic Nomenclature (ISCN). This semantic description of chromosomal imbalances hinders a large-scale statistical analysis across different experiments, e.g. for finding aberration patterns associated with a particular disease type or state. RESULTS: CGH-Profiler circumvents the semantic ISCN description by importing data from different CGH system vendors and by directly transferring the data into a table format that is readily accessible for subsequent statistical analysis. CGH-profiler comes with different consistency checks, calculates various statistics and automatically assigns a median copy number ratio to each chromosomal band. Import of CGH profiles from different CGH system vendors is already supported; its extension to other systems can be readily achieved through Perl scripts.CGH profiler can also be used to analyse comparative expressed sequence hybridisation (CESH) data. CESH reveals gene expression patterns according to chromosomal locations in a similar manner as CGH detects chromosomal imbalances. CONCLUSION: CGH-Profiler is a useful tool for processing of CGH and CESH data.

Chromosome Aberrations↗

Methods for data mining from large multinational surveillance studies.

Traditionally, large surveillance studies have been analyzed by the use of the MICs at which 90% of isolates tested are inhibited (MIC(90)s), MIC(50)s, frequency distributions, and percent susceptibility. In the past, these approaches have proved satisfactory for the monitoring of resistance. From these traditional uses, one can readily detect an increase in MICs for organism and drug combinations. Now that large surveillance studies have been conducted for a number of years and databases have grown to include a large number of datum points, new approaches to the extraction of useful information from these studies are needed. The present study proposes approaches, including the use of antibiotypes, principal components analysis, phylogenetics, and population genetic analysis, to the evaluation of data from large multinational surveillance studies. Application of these types of analyses can be used to describe genetic diversity, analyze changes in susceptibility patterns over time, and possibly, shed light on the origins and evolution of antimicrobial resistance. As global surveillance studies become more common and new questions concerning the evolution of resistance are raised, innovative approaches to analysis of the data will increase in importance.

Algorithms↗

Data mining for protein-protein interactions in invertebrate model organisms.

Well-annotated genome databases are available for many invertebrate species, notably the fruitfly, Drosophila melanogaster, and the nematode, Caenorhabditis elegans. An adequate interpretation of this information at the biological level requires the exploration of the interactions between the gene products. Knowledge of protein interactions and the components of cell signalling pathways in the fly and worm are particularly valuable as hypotheses can be rapidly tested using the powerful genetic toolkits available. Invertebrates offer additional experimental advantages when attempting to characterise protein-protein interactions (PPIs). Their relatively small genome size compared to mammals helps to reduce missed interactions due to redundancy, and their function can be addressed using forward (mutants) and reverse (RNA interference) genetics. However, the researcher looking for evidence of PPIs for a protein of interest is faced with the challenge of extracting interaction data from sources that are highly varied, such as the results of microarray experiments in the unstructured text of research papers. This challenge is greatly reduced by a range of public databases of curated information, as well as publicly available, enhanced search engines, which can provide either direct experimental evidence for a PPI, or valuable clues for generating new hypotheses.

Animals↗

Data-mining approaches reveal hidden families of proteases in the genome of malaria parasite.

The search for novel antimalarial drug targets is urgent due to the growing resistance of Plasmodium falciparum parasites to available drugs. Proteases are attractive antimalarial targets because of their indispensable roles in parasite infection and development, especially in the processes of host erythrocyte rupture/invasion and hemoglobin degradation. However, to date, only a small number of proteases have been identified and characterized in Plasmodium species. Using an extensive sequence similarity search, we have identified 92 putative proteases in the P. falciparum genome. A set of putative proteases including calpain, metacaspase, and signal peptidase I have been implicated to be central mediators for essential parasitic activity and distantly related to the vertebrate host. Moreover, of the 92, at least 88 have been demonstrated to code for gene products at the transcriptional levels, based upon the microarray and RT-PCR results, and the publicly available microarray and proteomics data. The present study represents an initial effort to identify a set of expressed, active, and essential proteases as targets for inhibitor-based drug design.

Amino Acid Sequence↗

Automated knowledge extraction for decision model construction: a data mining approach.

Combinations of Medical Subject Headings (MeSH) and Subheadings in MEDLINE citations may be used to infer relationships among medical concepts. To facilitate clinical decision model construction, we propose an approach to automatically extract semantic relations among medical terms from MEDLINE citations. We use the Apriori association rule mining algorithm to generate the co-occurrences of medical concepts, which are then filtered through a set of predefined semantic templates to instantiate useful relations. From such semantic relations, decision elements and possible relationships among them may be derived for clinical decision model construction. To evaluate the proposed method, we have conducted a case study in colorectal cancer management; preliminary results have shown that useful causal relations and decision alternatives can be extracted.

Algorithms↗

More than 1,000 putative new human signalling proteins revealed by EST data mining.

Cloning procedures aided by homology searches of EST databases have accelerated the pace of discovery of new genes, but EST database searching remains an involved and onerous task. More than 1.6 million human EST sequences have been deposited in public databases, making it difficult to identify ESTs that represent new genes. Compounding the problems of scale are difficulties in detection associated with a high sequencing error rate and low sequence similarity between distant homologues. We have developed a new method, coupling BLAST-based searches with a domain identification protocol, that filters candidate homologues. Application of this method in a large-scale analysis of 100 signalling domain families has led to the identification of ESTs representing more than 1,000 novel human signalling genes. The 4,206 publicly available ESTs representing these genes are a valuable resource for rapid cloning of novel human signalling proteins. For example, we were able to identify ESTs of at least 106 new small GTPases, of which 6 are likely to belong to new subfamilies. In some cases, further analyses of genomic DNA led to the discovery of previously unidentified full-length protein sequences. This is exemplified by the in silico cloning (prediction of a gene product sequence using only genomic and EST sequence data) of a new type of GTPase with two catalytic domains.

Amino Acid Sequence↗

Reassessment of the TP53 mutation database in human disease by data mining with a library of TP53 missense mutations.

TP53 alteration is the most frequent genetic alteration found in human cancers. To date, more than 15,000 tumors with TP53 mutations have been published, leading to the description of more than 1,500 different TP53 mutants (http://p53.curie.fr). The frequency of these mutants is highly heterogeneous, with 11 hotspot mutants found more than 100 times, whereas 306 mutants have been reported only once. So far, little is known concerning the biological significance of these rare mutants, as the majority of biological studies have focused on classic hotspot mutants. In order to gain a deeper knowledge about the significance of all of these mutants, we have cross-checked each mutant of the TP53 mutation database for its activity, derived from a library of 2,314 TP53 mutants representing all possible amino acid substitutions caused by a point mutation. The transactivation activity of all of these mutant was analyzed with respect to eight transcription promoters [Kato S, et al., Proc Natl Acad Sci USA (2003)100:8424-8429]. Although the most frequent TP53 mutants sustain a clear loss of transactivation activity, more than 50% of the rare TP53 mutants display significant activity. Analysis in specific types of cancer or in normal skin patches demonstrates a similar distribution of TP53 loss of activity, with the exception of melanoma, in which the majority of TP53 mutants display significant activity. Our data indicate that TP53 mutants represent a highly heterogeneous population with a large diversity in terms of loss of transactivation activity that could account for the heterogeneous tumor phenotypes and the difficulty of clinical studies.

Arthritis, Rheumatoid↗

Enantiophore modeling in 3D-QSAR. A data mining application on Whelk-O1 chiral stationary phase.

A combination of the enantiophore concept described in a previous study and a quantitative structure enantioselective relationship (QSER) based on partial least squares (PLS) analysis is presented. In the present study, a comprehensive approach for describing the enantioselective binding properties of the Whelk-O1 chiral HPLC receptor is achieved using molecular descriptors calculated by the GRID program. The GRID descriptors allow us to describe the molecules in terms of their ability to form favorable interactions with independent chemical groups (probes) that can be related to receptor sites. For each molecule, we compute 120 enantiophore descriptors representing the energy contributions from all possible pairwise combinations of probes. The overall procedure was simplified by considering only the most energetically favorable locations and converting selected grid-point energies into alignment-independent descriptors. By using a training set of 143 diverse chiral compounds, an optimal PLS model requiring seven components was chosen by using the cross-validation method resulting in a correlation coefficient R2= 0.88 and a cross validated correlation coefficient Q2= 0.85. An interpretation of the model is proposed based upon a visual inspection of the regression coefficient plots. From these plots, the influence of particular molecular features for selective binding of solutes was estimated and used to outline the chiral recognition sites in the Whelk-O1 receptor. The predictive power of our model has been estimated by means of an external data set emphasizing the suitability of the procedure also for predictive aims.

Journal Article↗

Data mining of the E-pelvis simulator database: a quest for a generalized algorithm for objectively assessing medical skill.

Inherent difficulties in evaluating clinical competence of physicians has lead to the widespread use of subjective skill assessment techniques. Inspired by an analogy between medical procedure and spoken language, proven modeling methods in the field of speech recognition were adapted for use as objective skill assessment techniques. A generalized methodology using Markov Models (MM) was developed. The database under study was collected with the E-Pelvis physical simulator. The simulator incorporates an array of five contact force sensors located in key anatomical landmarks. Two 32-state fully connected MMs are used, one for each skill level. Each state in the model corresponds to one of the possible combinations of the 5 active contact force sensors distributed in the simulator. Statistical distances measured between models representing subjects with different skill levels are sensitive enough to provide an objective measure of medical skill level. The method was tested with 41 expert subjects and 41 novice subjects in addition to the 30 subjects used for training the MM. Of the 82 subjects, 76 were classified correctly (92%). Moreover, unique state transitions as well as force magnitudes for corresponding states (expert/novice) were found to be skill dependent. Given the white box nature of the model, analyzing the MMs provides insight into the examination process performed. This methodology is independent of the modality under study. It was previously used to assess surgical skill in a minimally invasive surgical setup using the Blue DRAGON, and it is currently applied to data collected using the E-Pelvis.

Algorithms↗

Data mining by total ranking methods: a case study on optimisation of the "pulp and bleaching" process in the paper industry.

Total order ranking methods are multicriteria decision making techniques used for the ranking of various alternatives on the basis of more than one criterion. The criteria, which are the standards by which the elements of the system are judged are not always in agreement, they can be conflicting, motivating the need to find an overall optimum that can deviate from the optima of one or more of the single criteria. Total order ranking methods are based on an aggregation of the criteria in a scalar function, i.e. an order or ranking index, which allow to sort elements according to its numerical value. Several evaluation methods which define a ranking parameter generating a total order ranking have been proposed in the literature. Four total order ranking methods are here described: Desirability functions, Utility functions, Dominance functions and Absolute Reference method. These methods have been compared to each other by applying them to a decision making problem in the paper industry. Various bleaching processes have been analysed and compared on the basis of multiple criteria, the aim being to find out best bleaching process among the ones proposed in the last years as alternative to chlorine bleaching process which is of high environmental impact due to the potential for chlorinated dioxin production.

Case-Control Studies↗

The olfactory receptor gene superfamily: data mining, classification, and nomenclature.

The vertebrate olfactory receptor (OR) subgenome harbors the largest known gene family, which has been expanded by the need to provide recognition capacity for millions of potential odorants. We implemented an automated procedure to identify all OR coding regions from published sequences. This led us to the identification of 831 OR coding regions (including pseudogenes) from 24 vertebrate species. The resulting dataset was subjected to neighbor-joining phylogenetic analysis and classified into 32 distinct families, 14 of which include only genes from tetrapodan species (Class II ORs). We also report here the first identification of OR sequences from a marsupial (koala) and a monotreme (platypus). Analysis of these OR sequences suggests that the ancestral mammal had a small OR repertoire, which expanded independently in all three mammalian subclasses. Classification of "fish-like" (Class I) ORs indicates that some of these ancient ORs were maintained and even expanded in mammals. A nomenclature system for the OR gene superfamily is proposed, based on a divergence evolutionary model. The nomenclature consists of the root symbol 'OR', followed by a family numeral, subfamily letter(s), and a numeral representing the individual gene within the subfamily. For example, OR3A1 is an OR gene of family 3, subfamily A, and OR7E12P is an OR pseudogene of family 7, subfamily E. The symbol is to be preceded by a species indicator. We have assigned the proposed nomenclature symbols for all 330 human OR genes in the database. A WWW tool for automated name assignment is provided.

Animals↗