Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

The design of a web-based decision support system for the sustainable management of an urban river system.

The effects of urbanization on the aquatic environment, and solutions to the deterioration of water quality and stream ecology in the Love River have long been the main focus of environmental management in southern Taiwan. Apart from choosing the regular strategies of installing an intercept and sewer system, coastal wastewater treatment plant, and ocean outfall pipe, a new opportunity for improving the overall managerial efficiency is to design and implement a web-based Decision Support System (DSS). This DSS must be capable of managing storm water impacts when overflow is inevitable, river water quality variations loading to influences on the ecosystem, and changing land use programs along river corridors and adjacent urban areas simultaneously. This paper presents a new design framework for building such a DSS. By using the advanced information technology in the "Internet" environment, the proposed DSS may perform normal queries and statistical analyses in a web-enabled database, spatial analysis via the use of a Geographic Information System (GIS) in the Internet environment, and essential data warehousing/data mining. Possible linkage with various analytical models is anticipated. Such a DSS must be helpful for achieving the rehabilitation of the estuarine ecosystem and satisfying the goals for sustainable management in a regional sense.

Cities↗

Genomic and proteomic profiling for biomarkers and signature profiles of toxicity.

Toxicity profiling measures and compares all gene expression changes among biological samples after toxicant exposure. Toxicity profiling with DNA microarrays to measure all mRNA transcripts (transcriptomics), or by global separation and identification of proteins (proteomics), has led to the discovery of better descriptors of toxicity, toxicant classification and exposure monitoring than current indicators. A shared goal in transcript and proteomic profiling is the development of biomarkers and signatures of chemical toxicity. In this review, biomarkers and signature profiles are described for specific chemical toxicants that affect target organs such as liver, kidney, neural tissues, gastrointestinal tract and skeletal muscle, for specific disease models such as cancer and inflammation, and for unique chemical-protein adducts underlying cell injury. The recent introduction of toxicogenomics databases support researchers in sharing, analyzing, visualizing and mining expression data, assist the integration of transcriptomics, proteomics and toxicology datasets, and eventually will permit in silico biomarker and signature pattern discovery.

Animals↗

The use of KDD to identify eligible patients for case management programs.

UNLABELLED: Patients diagnosed with co-morbidities are clinically compromised, facing social and psychological challenges. The identification of high risk patients is a key question to maximize their quality of life and the Health System's resources. OBJECTIVE: to determine patterns that identify eligible patients for Case Management Programs, with KDD and AI tools. METHODOLOGY: analysis, selection and correction, formatting and mining of data, analysis of results, data submission to an AI tool. RESULTS: solid recommendations on benefits that Health Systems could obtain by using these tools, regarding the acknowledgement of their users profile and the use of predictive modeling.

Brazil↗

[Structure-based drug design].

Structure-based drug design is not only a fundamental tool in the search of new drugs but also an area of intensive scientific research. The present contribution starts with a critical analysis of the potentiality of the structure based design and then reviews its methods. Protein crystallography, NMR, homology modelling and the use of antibodies are either referenced or briefly described as the methods which are able to provide us with the atomic structure of receptors. The search of lead molecules by data base mining and by generating structures with computers is discussed. Various methods and computer programs for calculating the interaction energy between receptors and ligands are reviewed. Several examples of successful structure base design are cited and some of them are briefly discussed.

Computer Simulation↗

Relational patterns of gene expression via non-metric multidimensional scaling analysis.

MOTIVATION: Microarray experiments result in large-scale data sets that require extensive mining and refining to extract useful information. We demonstrate the usefulness of (non-metric) multidimensional scaling (MDS) method in analyzing a large number of genes. Applying MDS to the microarray data is certainly not new, but the existing works are all on small numbers (< 100) of points to be analyzed. We have been developing an efficient novel algorithm for non-metric MDS (nMDS) analysis for very large data sets as a maximally unsupervised data mining device. We wish to demonstrate its usefulness in the context of bioinformatics (unraveling relational patterns among genes from time series data in this paper). RESULTS: The Pearson correlation coefficient with its sign flipped is used to measure the dissimilarity of the gene activities in transcriptional response of cell-cycle-synchronized human fibroblasts to serum. These dissimilarity data have been analyzed with our nMDS algorithm to produce an almost circular relational pattern of the genes. The obtained pattern expresses a temporal order in the data in this example; the temporal expression pattern of the genes rotates along this circular arrangement and is related to the cell cycle. For the data we analyze in this paper we observe the following. If an appropriate preparation procedure is applied to the original data set, linear methods such as the principal component analysis (PCA) could achieve reasonable results, but without data preprocessing linear methods such as PCA cannot achieve a useful picture. Furthermore, even with an appropriate data preprocessing, the outcomes of linear procedures are not as clear-cut as those by nMDS without preprocessing.

Algorithms↗

Pathological findings in mine workers: II. Quality of the PATHAUT data.

To assess the feasibility of using the pathology automation system (PATHAUT) for research, the quality of the data was explored by examining who comes to autopsy, the quality of the autopsy material, interobserver variability, and repeatability of diagnoses. The data indicated that autopsy rates in the gold mining industry, especially for whites, are high and that even among blacks, gold miners are represented in proportions exceeding the relative size of the working population. Because of the perception of the autopsy service as a means of obtaining compensation, miners with occupational diseases fully compensated in life are probably underrepresented. The autopsy material submitted for full autopsy is generally better preserved than cardiorespiratory organs that are sent for examination. The gold mining industry has a high proportion of full autopsies as does the Iron and Steel Corporation of South Africa. Full autopsies are more commonly performed on older deceased miners. This was true for both blacks and whites. The allocation of material to pathologists for full autopsies and examinations of the cardiorespiratory organs were clearly not random, and this may affect comparisons among pathologists. Active tuberculosis, silicosis, and emphysema prevalences appeared fairly comparable across pathologists; however, there was wide variability in the prevalence of bronchiolitis as determined by the pathologists. Agreement between the diagnoses on PATHAUT and reclassifications by single pathologists was very good for the severity of emphysema and the histological type of bronchogenic carcinoma.

Autopsy↗

Noise exposure and hearing conservation in U.S. coal mines--a surveillance report.

This study examines the patterns and trends in noise exposure documented in data collected by Mine Safety and Health Administration inspectors at U.S. coal mines from 1987 through 2004. During this period, MSHA issued a new regulation on occupational noise exposure that changed the regulatory requirements and enforcement policies. The data were examined to identify potential impacts from these changes. The overall annual median noise dose declined 67% for surface coal mining and 24% for underground coal mining, and the reduction in each group accelerated after promulgation of the new noise rule. However, not all mining occupations experienced a decrease. The exposure reduction was accompanied by an increase of shift length as represented by dosimeter sample duration. For coal miners exposed above the permissible exposure level, use of hearing protection devices increased from 61% to 89% during this period. Participation of miners exposed at or above the action level in hearing conservation programs rapidly reached 86% following the effective date of the noise rule. Based on inspection data, the occupational noise regulation appears to be having a strong positive impact on hearing conservation by reducing exposures and increasing the use of hearing protection devices and medical surveillance. However, the increase in shift duration and resulting reduction in recovery time may mitigate the gains somewhat.

Coal Mining↗

Validation of two air quality models for Indian mining conditions.

All major mining activity particularly opencast mining contributes to the problem of suspended particulate matter (SPM) directly or indirectly. Therefore, assessment and prediction are required to prevent and minimize the deterioration of SPM due to various opencast mining operations. Determination of emission rate of SPM for these activities and validation of air quality models are the first and foremost concern. In view of the above, the study was taken up for determination of emission rate for SPM to calculate emission rate of various opencast mining activities and validation of commonly used two air quality models for Indian mining conditions. To achieve the objectives, eight coal and three iron ore mining sites were selected to generate site specific emission data by considering type of mining, method of working, geographical location, accessibility and above all resource availability. The study covers various mining activities and locations including drilling, overburden loading and unloading, coal/mineral loading and unloading, coal handling or screening plant, exposed overburden dump, stock yard, workshop, exposed pit surface, transport road and haul road. Validation of the study was carried out through Fugitive Dust Model (FDM) and Point, Area and Line sources model (PAL2) by assigning the measured emission rate for each mining activity, meteorological data and other details of the respective mine as an input to the models. Both the models were run separately for the same set of input data for each mine to get the predicted SPM concentration at three receptor locations for each mine. The receptor locations were selected such a way that at the same places the actual filed measurement were carried out for SPM concentration. Statistical analysis was carried out to assess the performance of the models based on a set measured and predicted SPM concentration data. The value of coefficient of correlation for PAL2 and FDM was calculated to be 0.990-0.994 and 0.966-0.997, respectively, which shows a fairly good agreement between measured and predicted values of SPM concentration. The average index of agreement values for PAL2 and FDM was found to be 0.665 and 0.752, respectively, which represents that the prediction by PAL2 and FDM models are accurate by 66.5 and 75.2%, respectively. These indicate that FDM model is more suited for Indian mining conditions.

Air Pollution, Indoor↗

Tissues and hair residues and histopathology in wild rats (Rattus rattus L.) and Algerian mice (Mus spretus Lataste) from an abandoned mine area (Southeast Portugal).

Data gathered in this study suggested the exposure of rats and Algerian mice, living in an abandoned mining area, to a mixture of heavy metals. Although similar histopathological features were recorded in the liver and spleen of both species, the Algerian mouse has proved to be the strongest bioaccumulator species. Hair was considered to be a good biological material to monitor environmental contamination of Cr in rats. Significant positive associations were found between the levels of this element in hair/kidney (r=0.826, n=9, p<0.01) and hair/liver (r=0.697, n=9, p=0.037). Although no association was found between the levels of As recorded in the hair and in the organs, the levels of this element recorded in the hair, of both species, were significantly higher in animals captured in the mining area, which met the data from the organs analysed. Nevertheless, more studies will be needed to reduce uncertainty about cause-effect relationships.

Animals↗

The Botany Array Resource: e-Northerns, Expression Angling, and promoter analyses.

The Botany Array Resource provides the means for obtaining and archiving microarray data for Arabidopsis thaliana as well as biologist-friendly tools for viewing and mining both our own and other's data, for example, from the AtGenExpress Consortium. All the data produced are publicly available through the web interface of the database at http://bbc.botany.utoronto.ca. The database has been designed in accordance with the Minimum Information About a Microarray Experiment convention -- all expression data are associated with the corresponding experimental details. The database is searchable and it also provides a set of useful and easy-to-use web-based data-mining tools for researchers with sophisticated yet understandable output graphics. These include Expression Browser for performing 'electronic Northerns', Expression Angler for identifying genes that are co-regulated with a gene of interest, and Promomer for identifying potential cis-elements in the promoters of individual or co-regulated genes.

Arabidopsis↗

Radon update: facts concerning environmental radon: levels, mitigation strategies, dosimetry, effects and guidelines. SNM Committee on Radiobiological Effects of Ionizing Radiation.

The risk from environmental radon levels is not higher now than in the past, when residential exposures were not considered to be a significant health hazard. The majority of the radon dose is not from radon itself, but from short-lived alpha-emitting radon daughters, most notably 218Po(T1/2 3 min) and 214Po (T1/2 0.164 msec) along with beta particles from 214Bi (T1/2 19.7 min). Radon gas can penetrate homes from many sources and in various fashions. Measuring radon in homes is simple and relatively inexpensive and may be accomplished in a variety of ways. Although it is not possible to radon-proof a house, it is possible to reduce the level. In high radon areas, if the average level is higher than 4-8 pCi/liter (NCRP recommended level is 8 pCi/liter; EPA recommended level is 4 pCi/liter), appropriate action is advised. The shape of the dose response curves for miners exposed to alpha-emitting particles in the workplace is consistent with current biologic knowledge. It is linear in the low dose range and saturates in the high dose range. No detectable increase in lung cancer frequency is seen in the lowest exposed miners (those with exposures < 120 WLM, the relevant dose interval for most homes). Evidence for a health effect from radon exposure is based on data from animal studies and epidemiologic studies of mines. Extensive radiobiologic data predict a linear dose-response curve in the low dose region due to poor biological repair mechanisms for the high density of ionizing events that alpha particles create. However, no compelling evidence for increased cancer risks has yet been demonstrated from "acceptable" levels (< 4-8 pCi/liter).

Air Pollutants↗

Health plan finds answers with data warehouse that stores 4,500 data elements.

A gold mine of information, not just for computer geeks: Pardon the stereotype, but managers of health care have also stereotyped the data warehouse--as more of an unwieldy behemoth. But at Blue Cross/Blue Shield of Tennessee, more than 200 users log on regularly. Quality management, HEDIS reporting, managed care contracting, provider profiling, utilization management, product management, and forecasting all have been made easier with this tool that allows operational data to become a gold mine of information.

Benchmarking↗

An information theoretic approach for analyzing temporal patterns of gene expression.

MOTIVATION: Arrays allow measurements of the expression levels of thousands of mRNAs to be made simultaneously. The resulting data sets are information rich but require extensive mining to enhance their usefulness. Information theoretic methods are capable of assessing similarities and dissimilarities between data distributions and may be suited to the analysis of gene expression experiments. The purpose of this study was to investigate information theoretic data mining approaches to discover temporal patterns of gene expression from array-derived gene expression data. RESULTS: The Kullback-Leibler divergence, an information-theoretic distance that measures the relative dissimilarity between two data distribution profiles, was used in conjunction with an unsupervised self-organizing map algorithm. Two published, array-derived gene expression data sets were analyzed. The patterns obtained with the KL clustering method were found to be superior to those obtained with the hierarchical clustering algorithm using the Pearson correlation distance measure. The biological significance of the results was also examined. AVAILABILITY: Software code is available by request from the authors. All programs were written in ANSI C and Matlab (Mathworks Inc., Natick, MA).

Algorithms↗

Predicting missing values in a home care database using an adaptive uncertainty rule method.

OBJECTIVES: Contemporary literature illustrates an abundance of adaptive algorithms for mining association rules. However, most literature is unable to deal with the peculiarities, such as missing values and dynamic data creation, that are frequently encountered in fields like medicine. This paper proposes an uncertainty rule method that uses an adaptive threshold for filling missing values in newly added records. A new approach for mining uncertainty rules and filling missing values is proposed, which is in turn particularly suitable for dynamic databases, like the ones used in home care systems. METHODS: In this study, a new data mining method named FiMV (Filling Missing Values) is illustrated based on the mined uncertainty rules. Uncertainty rules have quite a similar structure to association rules and are extracted by an algorithm proposed in previous work, namely AURG (Adaptive Uncertainty Rule Generation). The main target was to implement an appropriate method for recovering missing values in a dynamic database, where new records are continuously added, without needing to specify any kind of thresholds beforehand. RESULTS: The method was applied to a home care monitoring system database. Randomly, multiple missing values for each record's attributes (rate 5-20% by 5% increments) were introduced in the initial dataset. FiMV demonstrated 100% completion rates with over 90% success in each case, while usual approaches, where all records with missing values are ignored or thresholds are required, experienced significantly reduced completion and success rates. CONCLUSIONS: It is concluded that the proposed method is appropriate for the data-cleaning step of the Knowledge Discovery process in databases. The latter, containing much significance for the output efficiency of any data mining technique, can improve the quality of the mined information.

Algorithms↗

Effects of the intensity and timing of asbestos exposure on lung cancer risk at two mining areas in Quebec.

Mortality data from 9609 workers at two asbestos mining areas in Quebec were analyzed to assess the effects of the intensity and timing of exposure on lung cancer risk. Summary exposure measures based on differing assumption were computed for lung cancer cases and matched controls and were fitted to the data using conditional logistic regression. A non-linear relationship between intensity and risk fit both mining areas, but risk was greater at one area than the other. At the mine with lower risk, exposure occurring more than 30 years prior to death had little effect, while at the other mine risk did not vary with time since exposure and men starting employment before 1924 were at elevated risk. The results point to differences in dust composition at the two areas and illustrate the difficulties in estimating risk.

Adult↗

Advances in high-throughput mass spectrometry.

The evolution of high-throughput drug discovery is readily apparent as the pharmaceutical industry continues to stress the rapid progression of new chemical entities and biological agents through drug discovery and development pipelines. Mass spectrometry and high performance liquid chromatography-mass spectrometry have played an instrumental role in the support and advancement of all facets of high-throughput drug discovery. The introduction of new instrumentation has extended the breadth of mass spectrometric-based capabilities from the characterization of high-throughput organic synthesis products to early adsorption, distribution, metabolism and excretion profiling. Additionally, advances in the capacity and throughput of mass spectrometry systems have concurrently led to the introduction of data management tools to address automated data reduction, archival and mining, as well as analytical data integration to chemical and biological databases.

Chromatography, High Pressure Liquid↗