Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

Expectancy of working life of mine workers in Hunan province.

From demographic mortality and disability of, in total, 424965 y from workers in 29 coal or coloured/metal mines in Human province of China during the period 1980-1984, we calculated the length of expected working life (EWL) as well as the life expectancy (EL) of the workers in the different types of mines and between those working on the surface and those working underground. The average life expectancy in the coal mines for those starting work at 15 y was found to be 58.91 y and 49.23 y for surface and underground workers respectively. In the coloured/metal mines they were 60.24 y and 56.55 y respectively. If only the mortality data was taken into account; for the coal mines, the EWL at the age of 15 y of surface workers and underground workers was 48.46 and 45.10 y respectively; whilst in the coloured/metal mines the figures were 48.25 and 45.99 y respectively. If both mortality and morbidity data is used the EWL at the same age and in the same population is for the coal mines 39.42 (surface workers) 25.54 (underground workers) and for the coloured/metal mines 42.05 (surface workers) and 30.02 (underground workers). The main causes of the lowered EWL of the underground workers were industrial accidents and pneumoconiosis which indicates that safety measures need to be increased and working conditions improved to protect underground workers.

Accidents, Occupational↗

GEMS: a web server for biclustering analysis of expression data.

The advent of microarray technology has revolutionized the search for genes that are differentially expressed across a range of cell types or experimental conditions. Traditional clustering methods, such as hierarchical clustering, are often difficult to deploy effectively since genes rarely exhibit similar expression pattern across a wide range of conditions. Biclustering of gene expression data (also called co-clustering or two-way clustering) is a non-trivial but promising methodology for the identification of gene groups that show a coherent expression profile across a subset of conditions. Thus, biclustering is a natural methodology as a screen for genes that are functionally related, participate in the same pathways, affected by the same drug or pathological condition, or genes that form modules that are potentially co-regulated by a small group of transcription factors. We have developed a web-enabled service called GEMS (Gene Expression Mining Server) for biclustering microarray data. Users may upload expression data and specify a set of criteria. GEMS then performs bicluster mining based on a Gibbs sampling paradigm. The web server provides a flexible and an useful platform for the discovery of co-expressed and potentially co-regulated gene modules. GEMS is an open source software and is available at http://genomics10.bu.edu/terrence/gems/.

Algorithms↗

A systematic review of modafinil: Potential clinical uses and mechanisms of action.

BACKGROUND: Modafinil is a novel wake-promoting agent that has U.S. Food and Drug Administration approval for narcolepsy and shift work sleep disorder and as adjunctive treatment of obstructive sleep apnea/hypopnea syndrome. Modafinil has a novel mechanism and is theorized to work in a localized manner, utilizing hypocretin, histamine, epinephrine, gamma-aminobutyric acid, and glutamate. It is a well-tolerated medication with low propensity for abuse and is frequently used for off-label indications. The objective of this study was to systematically review the available evidence supporting the clinical use of modafinil. DATA SOURCES: The search term modafinil OR Provigil was searched on PubMed. Selected articles were mined for further potential sources of data. Abstracts from major scientific conferences were reviewed. Lastly, the manufacturer of modafinil in the United States was asked to provide all publications, abstracts, and unpublished data regarding studies of modafinil. DATA SYNTHESIS: There have been 33 double-blind, placebo-controlled trials of modafinil. Additionally, numerous smaller studies have been performed, and case reports of modafinil's use abound in the literature. CONCLUSIONS: Modafinil is a promising drug with a large potential for many uses in psychiatry and general medicine. Treating daytime sleepiness is complex, and determining the precise nature of the sleep disorder is vital. Modafinil may be an effective agent in many sleep conditions. To date, the strongest evidence among off-label uses exists for the use of modafinil in attention-deficit disorder, postanesthetic sedation, and cocaine dependence and withdrawal and as an adjunct to antidepressants for depression.

Animals↗

PROPHECY--a database for high-resolution phenomics.

The rapid recent evolution of the field phenomics--the genome-wide study of gene dispensability by quantitative analysis of phenotypes--has resulted in an increasing demand for new data analysis and visualization tools. Following the introduction of a novel approach for precise, genome-wide quantification of gene dispensability in Saccharomyces cerevisiae we here announce a public resource for mining, filtering and visualizing phenotypic data--the PROPHECY database. PROPHECY is designed to allow easy and flexible access to physiologically relevant quantitative data for the growth behaviour of mutant strains in the yeast deletion collection during conditions of environmental challenges. PROPHECY is publicly accessible at http://prophecy.lundberg.gu.se.

Computer Graphics↗

[Analysis on the risk factors in patients with diabetes mellitus from population in mining districts--a population-based case-control study].

According to data from prevalence study on population from Pingdingshan coal mining districts in Henan province, we analysed 174 patients with diabetes mellitus(DM) and 3,066 control subjects with normal blood glucose(NGT) by a population-based case-control study. After the adjustment of other factors and controlled on confounding factors, the results of unconditional logistic multivariate regression analysis demonstrated that age, DM history of mother and sib, highest BMI through one's life, higher concurrent WHR, higher systolic blood pressure, frequently eating Chinese sorghum and legume may serve as independent risk factors of DM, their odds ratios(OR) were 2.04, 6.04, 2.24, 1.85, 2.57, 1.51, 2.22, 1.25 and their population attribution rates (PAR%) were 80.04%, 7.19%, 3.18%, 37.35%, 48.80%, 8.15%, 3.20%, 10.63% respectively. Higher occupational physical activity and frequently eating vegetables of light colour might serve as independent protective factors of DM, with ORs 0.89 and 0.50 and PAR% of -19.20% and -269.5% respectively. Confounding analysis showed that age was both a positive and negative confounding factor to other factors in the logistic regression model.

Adult↗

Historical exposure to inorganic mercury at the smelter works of Abbadia San Salvatore, Italy.

Metallic mercury production from cinnabar ore may result in high exposures to inorganic mercury, that are difficult to assess separately from the exposures originating from underground extraction, and previously have only been scantily described. We retrieved and analysed the air and biological mercury determinations on workers involved in the smelting process of the Abbadia San Salvatore mine (Monte Amiata, Italy). Native mercury was not present in the ore, and the exposure in the underground extraction was low. The smelter operated from 1897 to 1983. Blood and urine (24/h urine collections and concentration samples) had been sampled in 1968 to 1982, and analysed for mercury by atomic absorption spectrophotometry, and relate to all subjects. Exposure to mercury in air had been determined in a small set of personal samples in 1982. The data relate to all jobs in the smelter process, and all jobs entailed substantial exposure to mercury. The overall distribution of breathing zone air, blood and urinary levels is right-skewed and similar to the log-normal distribution (air, median 48 micrograms/m3, n = 49; blood, arithmetic mean AM 49 micrograms/L; geometric mean GM 26 micrograms/L, n = 192; urinary excretion, AM 140 micrograms/24 h, GM 78 micrograms/24 h, n = 839; and urinary concentration, AM 160 micrograms/L, GM 83 micrograms/L, n = 632). Air, blood and urinary values show a high ratio of the between- and within-job variance, indicating differences in exposure by job. Cinnabar pigment production, of which the exposure has not been characterised previously, was the job with the highest air (AM 160 micrograms/m3) and urinary levels (excretion AM 690 micrograms/24 h; concentration AM 1100 micrograms/L). Other jobs with high urinary levels were soot purification, laboratory work, and bottling. Cleaning of condensers showed the highest blood level (AM 280 micrograms/L). There is a downwards time trend in mercury concentration in blood and in urine. The corresponding trend is not seen for urinary excretion levels, the reason for this being unclear. Roasters, which is the most frequently monitored group, show however a decreasing trend in all sets of data (e.g. the mean of urinary excretion decreased from 300 micrograms/24 h in 1968/69 to 50 micrograms/24 h in 1980/81). The mercury exposure experienced by the smelters of Abbadia San Salvatore is in line with the few available data on workers from other mercury mines and smelters, and our data confirm the high exposure levels in this occupational group, in particular at cinnabar pigment production, soot purification, and condenser cleaning.

Aerosols↗

Ethnic differences in the prevalence of nonmalignant respiratory disease among uranium miners.

OBJECTIVES: This study (1) investigates the relationship of nonmalignant respiratory disease to underground uranium mining and to cigarette smoking in Native American, Hispanic, and non-Hispanic White miners in the Southwest and (2) evaluates the criteria for compensation of ethnic minorities. METHODS: Risk for mining-related lung disease was analyzed by stratified analysis, multiple linear regression, and logistic regression with data on 1359 miners. RESULTS: Uranium mining is more strongly associated with obstructive lung disease and radiographic pnuemoconiosis in Native Americans than in Hispanics and non-Hispanic Whites. Obstructive lung disease in Hispanic and non-Hispanic White miners is mostly related to cigarette smoking. Current compensation criteria excluded 24% of Native Americans who, by ethnic-specific standards, had restrictive lung disease and 4.8% who had obstructive lung disease. Native Americans have the highest prevalence of radiographic pneumoconiosis, but are less likely to meet spirometry criteria for compensation. CONCLUSIONS: Native American miners have more nonmalignant respiratory disease from underground uranium mining, and less disease from smoking, than the other groups, but are less likely to receive compensation for mining-related disease.

Colorado↗

Diabetes: a candidate disease for efficient DNA methylation profiling.

Methylation has been implied in a number of biological processes and has been shown to vary under environmental influences as well as in age. Most results on the correlation of methylation patterns with phenotypic characteristics of cells have been obtained by analysis of very few or even single genomic fragments for methylation. However, variation of methylation may more often than not be a phenomenon that affects multiple genomic loci. The role of methylation has been most conclusively demonstrated in complex disease, with cancer being the most prominent example. The influence of aging and environmental influences such as diet seems to be on global methylation patterns, in turn exerting local effects on groups of genes. Hence, methylation seems literally to be orchestrating complex genetic systems. It could, therefore, be considered an archetypal "genomics" parameter. In consequence, technologies used to analyze methylation patterns should be as industrialized as possible to capture the local events across the entire genome. Epigenomics' research team is the first to have achieved the industrialized production of genome sequence-specific wide methylation data. Our microarray and mass-spectrometry-based detection platform currently allow the analysis of up to 50,000 methylation positions per day, for the first time making methylation data amenable to sophisticated information mining. The information content of methylation position has never been analyzed using the high-dimensional statistical methods that are recognized to be required for the analysis of, for example, mRNA expression profiles or proteomic data. As methylation patterns are nothing but a quasi-digital form of expression data, their information content must be evaluated using similar but adapted algorithms. This article presents a broad set of studies that demonstrate that methylation yields information that is comparable or even superior to the current state of the art, namely, mRNA profiling. We argue that the resulting robust, digital and-because of the highly stable nature of DNA as the analyte-more reproducible information could become the "gold standard" for clinical diagnostics and disease gene identification in age-related, environmentally influenced and epigenetic disease in general, substituting for mRNA expression.

DNA Methylation↗

Nonoccupational exposure to chrysotile asbestos and the risk of lung cancer.

BACKGROUND: Heavy industrial exposure to asbestos causes lung cancer and mesothelioma, but it remains unknown whether much lower environmental exposure to asbestos also causes these cancers. Nevertheless, regulatory agencies, including the Environmental Protection Agency (EPA), have assessed the risk of lung cancer by extrapolating known risks from past industrial exposure to asbestos to today's much lower environmental asbestos levels (roughly 100,000 times lower). We also tested the EPA's model for predicting the risk of asbestos-induced lung cancer in a population of women with relatively high levels of nonoccupational exposure to asbestos. METHODS: Mortality among women in 2 chrysotile-asbestos-mining areas of the province of Quebec was compared with mortality among women in 60 control areas, and age-standardized mortality ratios were derived. With the help of an expert panel, we estimated past exposure to asbestos among women in the mining areas and used these data with the EPA's model to predict the relative risk of lung cancer. We then compared this prediction with the observed mortality ratios. RESULTS: On the basis of the estimated exposure in the asbestos-mining areas, a relative risk of death due to lung cancer of 2.1 was predicted by the EPA's model, amounting to about 75 excess deaths from lung cancer in this population. By contrast, we calculated a standardized mortality ratio of 1.0 and a standardized proportionate mortality ratio of 1.1 (P> 0.05), suggesting that there were between 0 and 6.5 excess deaths from lung cancer among the women with nonoccupational exposure to asbestos. Seven deaths from pleural cancer were observed (relative risk=7.63; P<0.05). CONCLUSIONS: We found no measurable excess risk of death due to lung cancer among women in two chrysotile-asbestos-mining regions. The EPA's model overestimated the risk of asbestos-induced lung cancer by at least a factor of 10.

Adult↗

[Evaluation of historical exposure to silica dust for the workers employed in geologic exploration industry].

OBJECTIVE: To evaluate their level of exposure to silica dust in the workers employed in geologic exploration industry using a quantitative method. METHODS: Monitoring data of exposure to silica dust since 1950s were collected for the workers in various jobs from the factories and mines of the bureaux of geological exploration in the nine provinces, with 30,000 pieces of figures in total. Job title and history of employment were determined for 1,627 employees. Levels of exposure to silica dust in various factories and mines were estimated based on the data mentioned-above, including amount of respirable dust, total dust in the lungs, and proportion of free silica dust to the total silica dust. RESULTS: Concentration of total silica dust was 14 mg/m3 in average in the factories and mines, with a range of 29 mg/m3 during the earlier years to 3 mg/m3 in recent years, and that of respirable silica dust was 3 mg/m3, in average, with (28.0 +/- 8.2)% of free silica dust. Different indices of exposure to silica dust were calculated for individuals, based on their history of employment and level of exposure. CONCLUSION: Assessment of exposure to silica dust during the past years, based on the monitoring data, could provide basis for evaluation of dose-response relationship.

Air Pollutants, Occupational↗

Consolidating the set of known human protein-protein interactions in preparation for large-scale mapping of the human interactome.

BACKGROUND: Extensive protein interaction maps are being constructed for yeast, worm, and fly to ask how the proteins organize into pathways and systems, but no such genome-wide interaction map yet exists for the set of human proteins. To prepare for studies in humans, we wished to establish tests for the accuracy of future interaction assays and to consolidate the known interactions among human proteins. RESULTS: We established two tests of the accuracy of human protein interaction datasets and measured the relative accuracy of the available data. We then developed and applied natural language processing and literature-mining algorithms to recover from Medline abstracts 6,580 interactions among 3,737 human proteins. A three-part algorithm was used: first, human protein names were identified in Medline abstracts using a discriminator based on conditional random fields, then interactions were identified by the co-occurrence of protein names across the set of Medline abstracts, filtering the interactions with a Bayesian classifier to enrich for legitimate physical interactions. These mined interactions were combined with existing interaction data to obtain a network of 31,609 interactions among 7,748 human proteins, accurate to the same degree as the existing datasets. CONCLUSION: These interactions and the accuracy benchmarks will aid interpretation of current functional genomics data and provide a basis for determining the quality of future large-scale human protein interaction assays. Projecting from the approximately 15 interactions per protein in the best-sampled interaction set to the estimated 25,000 human genes implies more than 375,000 interactions in the complete human protein interaction network. This set therefore represents no more than 10% of the complete network.

Algorithms↗

[Analysis on pneumoconiosis characteristic and its prediction in one coal mine].

OBJECTIVE: We aimed to investigate the characteristic of pneumoconiosis on coal miners and provide scientific evidences for its prevention. METHODS: To analyze the data of pneumoconiosis from one coal mine with the historical study and to predict its development tendency by the grey model of GM (1,1). RESULTS: (1) The year of work experience of pneumoconiosis (stage I) was 19.9 years and the age of its diagnosis (stage I) was 51.4 years. There was an obvious tendency that they became longer with years' back-shift. (2) During near forty years, the progression rates of pneumoconiosis from stage I to II and II to III were 13.6% and 11.2%, and the mean time was 8.3 and 8.1 years, respectively. It was obvious that the progression rate decreased gradually and the span prolonged with years' back-shift. (3) The complication rate of pneumoconiosis with lung tuberculosis was 12.5% and it increased with the progression of pneumoconiosis. The rate in dead cases was significantly higher than that in live cases (P < 0.01). (4) The sequence of death causes was pneumoconiosis (20.0%), lung tuberculosis (18.3%), chronic cor pulmonale (17.9%), and pulmonary carcinoma (9.0%), et al. (5) It was predicted that there would be 28 new cases every year during 2001-2020 years and the accumulated numbers of pneumoconiosis would be 2854 cases in 2020, but with a downward trends in the prevalence of -1.7%. CONCLUSION: It was suggested that the prevalence of pneumoconiosis should decrease obviously. However, it still remains a challenge about the task of effectively preventing and curing it or its complication.

Adult↗

Making sense of the metabolome using evolutionary computation: seeing the wood with the trees.

One should perhaps start off by asking the question, 'But what wood is it we want to see?' There are so many trees that make up the wood; within a post-genomics context, genes, transcripts, proteins, and metabolites are the more tangible ones. Rather than studying these components in isolation, a more holistic approach is to unravel the interactions between the myriad of subcellular components and this is vital to systems biology. Moreover, this will help define the phenotype of the organism under investigation. Metabolomics is complementary to transcriptomics and proteomics, and despite the immense metabolite diversity observed in plants, metabolomics has been embraced by the plant community and in particular for studying metabolic networks. Whilst post-genomic science is producing vast data torrents, it is well known that data do not equal knowledge and so the extraction of the most meaningful parts of these data is key to the generation of useful new knowledge. A metabolomics experiment is guaranteed to generate thousands of data points (e.g. samples multiplied by the levels of particular metabolites) of which only a handful might be needed to describe the problem adequately. Evolutionary computational-based methods such as genetic algorithms and genetic programming are ideal strategies for mining such high-dimensional data to generate useful relationships, rules, and predictions. This article describes these techniques and highlights their usefulness within metabolomics.

Algorithms↗

Cancer in asbestos-mining and other areas of Quebec.

Employing incidence data from the Quebec Tumor Registry, we examined the relative risks of cancer of all sites for the years 1969-73 in the asbestos-mining, rural, and metropolitan counties of Quebec Province, Canada. Generally, rates for males exceeded those for females, and the relative risks in the asbestos-mining counties for 7-10 different sites of cancer, all of low incidence, were from 1.50 to 8.08 times those of other rural counties of the Province for both sexes. Metropolitan counties exhibited equally high risk for many of these sites. We discovered higher risks among males in asbestos-mining counties for cancer of the pleura, peritoneum, lip, tongue, salivary gland, mouth, and small intestine and higher risks among females for cancer of the pleura, lip, kidney, salivary gland, and for melanoma. Because of the likelihood of a long latent period for asbestos-related cancers, the risks we observed were possibly the product of since-altered occupational and environmental conditions existing 20-30 years ago in the asbestos-mining areas. The similarities in risks for most cancers in asbestos-mining and urban areas were noteworthy.

Asbestos↗

GeneOrder: comparing the order of genes in small genomes.

MOTIVATION: The recent rapid rise in the availability of whole genome DNA sequence data has led to bottlenecks in their complete analysis. Specifically, there is a need for software tools that will allow mining of gene and putative gene data at a whole genome level. These new tools will complement the current set already in use for studying specific aspects of individual genes and putative genes in detail. A key software challenge is to make them user-friendly, without losing their flexibility and capability for use in research. RESULTS: The creation of GeneOrder-a web-based interactive, computational tool-allows researchers to compare the order of genes in two genomes. It has been tested on full genome sequence data for viruses, mitochondria and chloroplasts that were obtained from the NCBI GenBank database. It is accessible at http://www.bif.atcc.org/GENEOrder/index.html. GeneOrder prepares the comparison in table form, listing the order of similar genes. Hyperlinks are provided from this output; these lead to the 'Protein Coding Regions' in the NCBI database.

Animals↗

Machine learning based pattern recognition applied to microarray data.

MOTIVATION: Microarrays have allowed the expression level of thousands of genes or proteins to be measured simultaneously. Data sets generated by these arrays consist of a small number of observations (e.g., 20-100 samples) on a very large number of variables (e.g., 10,000 genes or proteins). The observations in these data sets often have other attributes associated with them such as a class label denoting the pathology of the subject. Finding the genes or proteins that are correlated to these attributes is often a difficult task since most of the variables do not contain information about the pathology and as such can mask the identity of the relevant features. We describe a genetic algorithm (GA) that employs both supervised and unsupervised learning to mine gene expression and proteomic data. The pattern recognition GA selects features that increase clustering, while simultaneously searching for features that optimize the separation of the classes in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the features chosen by the GA contain information primarily about differences between classes in the data set. The principal component analysis routine embedded in the fitness function of the GA acts as an information filter, significantly reducing the size of the search space since it restricts the search to feature sets whose principal component plots show clustering on the basis of class. The algorithm integrates aspects of artificial intelligence and evolutionary computations to yield a smart one pass procedure for feature selection, clustering, classification, and prediction.

Algorithms↗