Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Object-oriented development of a concept learning system for time-centered clinical data.

A concept learning system is expected to be a powerful tool for filtering and analyzing a large amount of data in a variety of scientific fields. A simple application of it to clinical data, however, fails to mine medical information and knowledge. One of the major obstacles in mining a clinical database is time, which is a very important concept in clinical medicine. To be successful in data mining in clinical medicine, an efficient model of clinical data with time and a flexible concept learning system augmented to handle the model are both necessary. Herein we modeled clinical data to easily express and manipulate time and extended a concept learning system to utilize a time-centered clinical data model. The modified concept learning system is based on the message-value method rather than the traditional attribute-value method. The object-oriented technology was of great help in modeling time-centered clinical data and in developing a modified concept learning system.

Algorithms↗

NetAffx Gene Ontology Mining Tool: a visual approach for microarray data analysis.

SUMMARY: The NetAffx Gene Ontology (GO) Mining Tool is a web-based, interactive tool that permits traversal of the GO graph in the context of microarray data. It accepts a list of Affymetrix probe sets and renders a GO graph as a heat map colored according to significance measurements. The rendered graph is interactive, with nodes linked to public web sites and to lists of the relevant probe sets. The GO Mining Tool provides visualization combining biological annotation with expression data, encompassing thousands of genes in one interactive view. AVAILABILITY: GO Mining Tool is freely available at http://www.affymetrix.com/analysis/query/go_analysis.affx

Abstracting and Indexing↗

The concept of degraded images applied to hazard recognition training in mining for reduction of lost-time injuries.

INTRODUCTION: This paper discusses the application of a training intervention that uses degraded images for improving the hazard recognition skills of miners. METHOD: NIOSH researchers, in an extensive literature review, identified fundamental psychological principles on perception that may be employed to enhance the ability of miners to recognize and respond to hazards in their dangerous work environment. Three studies were conducted to evaluate the effectiveness of the degraded image training intervention. A model of hazard recognition was developed to guide the study. RESULTS: In the first study, miners from Pennsylvania, West Virginia and Alabama, who were taught with the aid of degraded images, scored significantly better on follow-up hazard recognition performance measures than those trained using traditional instructional methodologies. The second and third studies investigated the effectiveness of the training intervention at two mining companies. Data collected over a 3-year period showed that lost-time injuries at mines in Alabama and Illinois declined soon after the training intervention was instituted. IMPACT ON INDUSTRY: Further exploration of the hazard recognition model and the development of other interventions based on the model could support the validity of the steps in the hazard recognition model.

Accidents, Occupational↗

On the way to understand biological complexity in plants: S-nutrition as a case study for systems biology.

The establishment of technologies for high-throughput DNA sequencing (genomics), gene expression (transcriptomics), metabolite and ion analysis (metabolomics/ionomics) and protein analysis (proteomics) carries with it the challenge of processing and interpreting the accumulating data sets. Publicly accessible databases and newly development and adapted bioinformatic tools are employed to mine this data in order to filter relevant correlations and create models describing physiological states. These data allow the reconstruction of networks of interactions of the various cellular components as enzyme activities and complexes, gene expression, metabolite pools or pathway flux modes. Especially when merging information from transcriptomics, metabolomics and proteomics into consistent models, it will be possible to describe and predict the behaviour of biological systems, for example with respect to endogenous or environmental changes. However, to capture the interactions of network elements requires measurements under a variety of conditions to generate or refine existing models. The ultimate goal of systems biology is to understand the molecular principles governing plant responses and consistently explain plant physiology.

Arabidopsis↗

Epidemiological geomatics in evaluation of mine risk education in Afghanistan: introducing population weighted raster maps.

Evaluation of mine risk education in Afghanistan used population weighted raster maps as an evaluation tool to assess mine education performance, coverage and costs. A stratified last-stage random cluster sample produced representative data on mine risk and exposure to education. Clusters were weighted by the population they represented, rather than the land area. A "friction surface" hooked the population weight into interpolation of cluster-specific indicators. The resulting population weighted raster contours offer a model of the population effects of landmine risks and risk education. Five indicator levels ordered the evidence from simple description of the population-weighted indicators (level 0), through risk analysis (levels 1-3) to modelling programme investment and local variations (level 4). Using graphic overlay techniques, it was possible to metamorphose the map, portraying the prediction of what might happen over time, based on the causality models developed in the epidemiological analysis. Based on a lattice of local site-specific predictions, each cluster being a small universe, the "average" prediction was immediately interpretable without losing the spatial complexity.

Afghanistan↗

Cancer proteomics: Serum diagnostics for tumor marker discovery.

Cancer proteomics is an exciting field that is witnessing many new developments in recent years. It is hoped that these advances will result in decreased cancer death rates, which have not declined dramatically in the last several decades. Some of the problems with current tumor markers include the lack of sensitivity and specificity, factors that prevent their use in population-based screening of disease. Thus, there is an urgent need to identify novel biomarkers that can faithfully detect the disease state. As we are now in the post-genome era, many opportunities have been created. Genomic sequence data are available for human, as well as several other species. We are now poised to mine these data and to determine the functions of the encoded proteins constituting the human genome. Proteomics affords this opportunity by providing enhanced procedures and tools for discovery and also a framework for understanding these components in terms of pathogenesis. New technologies and improvements in existing methodologies will allow for the rapid growth in the identification and characterization of peptides and proteins that are unique to various clinical states. This technology can be successfully applied to clinical specimens for the identification of new tumor markers.

Algorithms↗

MOLE: a data management application based on a protein production data model.

MOLE (mining, organizing, and logging experiments) has been developed to meet the growing data management and target tracking needs of molecular biologists and protein crystallographers. The prototype reported here will become a Laboratory Information Management System (LIMS) to help protein scientists manage the large amounts of laboratory data being generated due to the acceleration in proteome research and will furthermore facilitate collaborations between groups based at different sites. To achieve this, MOLE is based on the data model for protein production devised at the European Bioinformatics Institute (Pajon A, et al., Proteins in press).

Algorithms↗

Information mining over heterogeneous and high-dimensional time-series data in clinical trials databases.

An effective analysis of clinical trials data involves analyzing different types of data such as heterogeneous and high dimensional time series data. The current time series analysis methods generally assume that the series at hand have sufficient length to apply statistical techniques to them. Other ideal case assumptions are that data are collected in equal length intervals, and while comparing time series, the lengths are usually expected to be equal to each other. However, these assumptions are not valid for many real data sets, especially for the clinical trials data sets. An addition, the data sources are different from each other, the data are heterogeneous, and the sensitivity of the experiments varies by the source. Approaches for mining time series data need to be revisited, keeping the wide range of requirements in mind. In this paper, we propose a novel approach for information mining that involves two major steps: applying a data mining algorithm over homogeneous subsets of data, and identifying common or distinct patterns over the information gathered in the first step. Our approach is implemented specifically for heterogeneous and high dimensional time series clinical trials data. Using this framework, we propose a new way of utilizing frequent itemset mining, as well as clustering and declustering techniques with novel distance metrics for measuring similarity between time series data. By clustering the data, we find groups of analytes (substances in blood) that are most strongly correlated. Most of these relationships already known are verified by the clinical panels, and, in addition, we identify novel groups that need further biomedical analysis. A slight modification to our algorithm results an effective declustering of high dimensional time series data, which is then used for "feature selection." Using industry-sponsored clinical trials data sets, we are able to identify a small set of analytes that effectively models the state of normal health.

Algorithms↗

Mining SARS-CoV protease cleavage data using non-orthogonal decision trees: a novel method for decisive template selection.

MOTIVATION: Although the outbreak of the severe acute respiratory syndrome (SARS) is currently over, it is expected that it will return to attack human beings. A critical challenge to scientists from various disciplines worldwide is to study the specificity of cleavage activity of SARS-related coronavirus (SARS-CoV) and use the knowledge obtained from the study for effective inhibitor design to fight the disease. The most commonly used inductive programming methods for knowledge discovery from data assume that the elements of input patterns are orthogonal to each other. Suppose a sub-sequence is denoted as P2-P1-P1'-P2', the conventional inductive programming method may result in a rule like 'if P1 = Q, then the sub-sequence is cleaved, otherwise non-cleaved'. If the site P1 is not orthogonal to the others (for instance, P2, P1' and P2'), the prediction power of these kind of rules may be limited. Therefore this study is aimed at developing a novel method for constructing non-orthogonal decision trees for mining protease data. RESULT: Eighteen sequences of coronavirus polyprotein were downloaded from NCBI (http://www.ncbi.nlm.nih.gov). Among these sequences, 252 cleavage sites were experimentally determined. These sequences were scanned using a sliding window with size k to generate about 50,000 k-mer sub-sequences (for short, k-mers). The value of k varies from 4 to 12 with a gap of two. The bio-basis function proposed by Thomson et al. is used to transform the k-mers to a high-dimensional numerical space on which an inductive programming method is applied for the purpose of deriving a decision tree for decision-making. The process of this transform is referred to as a bio-mapping. The constructed decision trees select about 10 out of 50,000 k-mers. This small set of selected k-mers is regarded as a set of decisive templates. By doing so, non-orthogonal decision trees are constructed using the selected templates and the prediction accuracy is significantly improved.

Algorithms↗

Heavy metal bioavailability and effects: II. Histopathology-bioaccumulation relationships caused by mining activities in the Gulf of Cádiz (SW, Spain).

The relationship between bioaccumulation of heavy metals (Zn, Cd, Pb and Cu) and histological lesions in different tissues of organisms is assessed in three different areas located in the southwest of Spain in the Gulf of Cádiz (Ría of Huelva, Guadalquivir estuary and Bay of Cádiz) affected and non-affected by mining activities. Data included in these relationships were obtained along the years 2000 and 2001 to address the impact of the Aznalcóllar mining spill on the Guadalquivir estuary. The bioaccumulation and the histological lesions measured in this seasonal study in the Gudalquivir estuary were linked to derive tissue quality guidelines (TQGs) by means of a multivariate analysis approach (MAA). Sediments collected in the same areas of study were used to expose organisms during the survey carried out in autumn 2001 and to address the relationship between bioaccumulation and histological lesions under laboratory conditions and related to chemicals bound to sediments. Lesions show that the organisms collected in the ría of Huelva and exposed to their sediments were severe, intermediate in the Guadalquivir estuary and absent in the Bay of Cádiz. Results show that the Guadalquivir estuary trends to recover its initial status quo previous to the mining spill. The link between chemical concentration and the lesions measured in the same tissues using MAA permits to derive tissue quality guidelines for two organisms, oysters (Crassostrea angulata) and clams (Scrobicularia plana) collected in the Guadalquivir estuary and associated with the heavy metals from the mining spill (Zn and Cd). The TQG values expressed as concentrations (mgkg-1--dry weight) not associated with biological effects are for oysters, Zn, 8603, Cd, 3.42; and for clams Zn, 800, Cd, 2.6.

Animals↗

The UCSC genome browser database: update 2007.

The University of California, Santa Cruz Genome Browser Database contains, as of September 2006, sequence and annotation data for the genomes of 13 vertebrate and 19 invertebrate species. The Genome Browser displays a wide variety of annotations at all scales from the single nucleotide level up to a full chromosome and includes assembly data, genes and gene predictions, mRNA and EST alignments, and comparative genomics, regulation, expression and variation data. The database is optimized for fast interactive performance with web tools that provide powerful visualization and querying capabilities for mining the data. In the past year, 22 new assemblies and several new sets of human variation annotation have been released. New features include VisiGene, a fully integrated in situ hybridization image browser; phyloGif, for drawing evolutionary tree diagrams; a redesigned Custom Track feature; an expanded SNP annotation track; and many new display options. The Genome Browser, other tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Animals↗

Occupational injuries in the mining industry and their association with statewide cold ambient temperatures in the USA.

BACKGROUND: Relatively few occupational epidemiological studies have been conducted concerning the association between cold ambient temperatures and cold exposure injuries, and fewer still of traumatic occupational injuries and cold ambient temperatures. METHODS: The association of ambient temperature and wind data from the National Climatic Data Center with injury data from mines reported to the Mine Safety and Health Administration (MSHA) was evaluated over a 6 year period from 1985-1990; 72,716 injuries from the seven states with the most numerous injuries were included. Temperature and wind data from each state's metropolitan weather stations were averaged for each day of the 6 year period. A weighted linear regression tested the relationship of ungrouped daily temperature and injury rate for all injury classes. For cold exposure injuries and fall injuries, relative incidence rates for grouped temperature data were fit with Poisson regression. RESULTS: As temperatures decreased, injury rates increased for both cold exposure injuries and slip and fall injuries. The association of slip and fall injuries with temperature was inverse but not strictly linear. The strongest association appeared with temperatures 29 degrees F and below. The injury rates for other accident categories increased with increasing ambient temperatures. CONCLUSIONS: This study suggests that statewide average ambient temperature reflects the expected association between the thermal environment and cold exposure injuries for workers, but more importantly, documents an association between ambient temperatures and occupational slip and fall injuries.

Accidental Falls↗

Toxicity and occupational health hazards of coal fly ash (CFA). A review of data and comparison to coal mine dust.

Coal fly ashes (CFA) are complex particles of a variable composition, which is mainly dependent on the combustion process, the source of coal and the precipitation technique. Toxic constituents in these particles are considered to be metals, polycyclic aromatic hydrocarbons and silica. The purpose of this review was to study the in vitro and in vivo data on coal fly ash and relate the studied endpoints to the role of (crystalline) silica, considering its recent classification as a human carcinogen. For most of the effects coal mine dust was chosen as a reference, since it contains up to 10% of crystalline silica (alpha-quartz) and is well studied both in vivo and in vitro. Most studies on fly ash toxicity were not designed to elucidate the effect of its silica-content nor did they include coal mine dust as a reference. Taking this into account, both in vitro and in vivo experimental studies show lower toxicity, inflammatory potential and fibrogenicity of CFA compared to silica and coal mine dust. Although in vitro and in vivo studies suggest genotoxic effects of fly ash, the data are limited and do not clarify the role of silica. Epidemiological studies in fly ash exposed working populations have found no evidence for effects commonly seen in coal workers (pneumoconiosis, emphysema) with the exception of airway obstruction at high exposure. In conclusion, the available data suggest that the hazard of coal fly ash is not to be assessed by merely adding the hazards of individual components. A closer investigation of 'matrix' effects on silica's toxicity in general seems an obligatory step in future risk assessment on fly ashes and other particles that incorporate silica as a component.

Animals↗

Identification and quantification of glycerolipids in cotton fibers: reconciliation with metabolic pathway predictions from DNA databases.

The lipid profiles of cotton fiber cells were determined from total lipid extracts of elongating and maturing cotton fiber cells to see whether the membrane lipid composition changed during the phases of rapid cell elongation or secondary cell wall thickening. Total FA content was highest or increased during elongation and was lower or decreased thereafter, likely reflecting the assembly of the expanding cell membranes during elongation and the shift to membrane maintenance (and increase in secondary cell wall content) in maturing fibers. Analysis of lipid extracts by electrospray ionization and tandem MS (ESI-MS/MS) revealed that in elongating fiber cells (7-10 d post-anthesis), the polar lipids-PC, PE, PI, PA, phosphatidylglycerol, monogalactosyldiacylglycerol, digalactosyldiacylglycerol, and phosphatidylglycerol-were most abundant. These same glycerolipids were found in similar proportions in maturing fiber cells (21 dpa). Detailed molecular species profiles were determined by ESI-MS/MS for all glycerolipid classes, and ESI-MS/MS results were consistent with lipid profiles determined by HPLC and ELSD. The predominant molecular species of PC, PE, PI, and PA was 34:3 (16:0, 18:3), but 36:6 (18:3,18:3) also was prevalent. Total FA analysis of cotton lipids confirmed that indeed linolenic (18:3) and palmitic (16:0) acids were the most abundant FA in these cell types. Bioinformatics data were mined from cotton fiber expressed sequence tag databases in an attempt to reconcile expression of lipid metabolic enzymes with lipid metabolite data. Together, these data form a foundation for future studies of the functional contribution of lipid metabolism to the development of this unusual and economically important cell type.

Chromatography, High Pressure Liquid↗

The UCSC Genome Browser Database: update 2006.

The University of California Santa Cruz Genome Browser Database (GBD) contains sequence and annotation data for the genomes of about a dozen vertebrate species and several major model organisms. Genome annotations typically include assembly data, sequence composition, genes and gene predictions, mRNA and expressed sequence tag evidence, comparative genomics, regulation, expression and variation data. The database is optimized to support fast interactive performance with web tools that provide powerful visualization and querying capabilities for mining the data. The Genome Browser displays a wide variety of annotations at all scales from single nucleotide level up to a full chromosome. The Table Browser provides direct access to the database tables and sequence data, enabling complex queries on genome-wide datasets. The Proteome Browser graphically displays protein properties. The Gene Sorter allows filtering and comparison of genes by several metrics including expression data and several gene properties. BLAT and In Silico PCR search for sequences in entire genomes in seconds. These tools are highly integrated and provide many hyperlinks to other databases and websites. The GBD, browsing tools, downloadable data files and links to documentation and other information can be found at http://genome.ucsc.edu/.

Amino Acid Sequence↗

A task exposure database for use in the alumina and primary aluminium industry.

A task exposure database (TED) was developed to facilitate data collation for construction of a task exposure matrix (TEM) for Healthwise, a series of studies on cancer and respiratory morbidity in the alumina and primary aluminium industry. Following the construction of job classifications for the eight study sites, the site hygienists identified all historical air monitoring time-weighted average (TWA) data, from their respective sites. The earliest data were sampled in the late 1970s, and over 17,000 personal samples were recorded over the eight sites over a twenty-year period. TED, a Microsoft Access database, was developed for use by site occupational hygienists to collate these exposure data across the mines, refineries and smelters. All data conforming to strict criteria for use were recorded using TED and provided to the study group. Following the individual data point entry, a calculator program in TED systematically calculated the geometric means, arithmetic means, and maximum and minimum results at the task level. Other features of TED included fields for flagging "significant changes" and "stepwise changes" in exposure. TED established a standardised means of data collation that later formed the basis for the construction of a TEM for the study. A TEM is similar to a job exposure matrix (JEM) except that the basic unit of categorization is at the task level instead of at the job level. Both a TEM and JEM have been constructed independently for Healthwise. The possible reduction of exposure misclassification and improvement in validity of exposure characterization with the use of the TEM, is currently under investigation. The Healthwise TEM consists of annual TWA and peak data results for each site for various airborne contaminants, including fluorides, coal tar pitch volatiles, sulfur dioxide, inspirable dust, alumina dust, bauxite dust, and oil mist. Construction of the TEM for the Healthwise study was completed in late 1998 and consists of over 33,700 TWA years of task exposure data.

Aluminum↗

Analyzing microarray data using cluster analysis.

As pharmacogenetics researchers gather more detailed and complex data on gene polymorphisms that effect drug metabolizing enzymes, drug target receptors and drug transporters, they will need access to advanced statistical tools to mine that data. These tools include approaches from classical biostatistics, such as logistic regression or linear discriminant analysis, and supervised learning methods from computer science, such as support vector machines and artificial neural networks. In this review, we present an overview of another class of models, cluster analysis, which will likely be less familiar to pharmacogenetics researchers. Cluster analysis is used to analyze data that is not a priori known to contain any specific subgroups. The goal is to use the data itself to identify meaningful or informative subgroups. Specifically, we will focus on demonstrating the use of distance-based methods of hierarchical clustering to analyze gene expression data.

Cluster Analysis↗