Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Underground mining, smoking, and lung cancer: a case-control study in the iron ore municipalities in northern Sweden.

A case-control study of lung cancer in males was performed in two municipalities in northern Sweden with large iron ore mines. Previous studies had revealed an increased lung cancer risk for underground workers in these mines, with all probability related to radon daughter exposure. Data concerning underground mining and smoking were obtained from questionnaires. All analyses suggested an interaction of a multiplicative type between underground mining and smoking in the causation of lung cancer in this population. The calculated population etiologic fraction was about 45% for underground mining and about 80% for smoking.

Aged↗

Acid-base accounting to predict post-mining drainage quality on surface mines.

Acid-base accounting (ABA) is an analytical procedure that provides values to help assess the acid-producing and acid-neutralizing potential of overburden rocks prior to coal mining and other large-scale excavations. This procedure was developed by West Virginia University scientists during the 1960s. After the passage of laws requiring an assessment of surface mining on water quality, ABA became a preferred method to predict post-mining water quality, and permitting decisions for surface mines are largely based on the values determined by ABA. To predict the post-mining water quality, the amount of acid-producing rock is compared with the amount of acid-neutralizing rock, and a prediction of the water quality at the site (whether acid or alkaline) is obtained. We gathered geologic and geographic data for 56 mined sites in West Virginia, which allowed us to estimate total overburden amounts, and values were determined for maximum potential acidity (MPA), neutralization potential (NP), net neutralization potential (NNP), and NP to MPA ratios for each site based on ABA. These values were correlated to post-mining water quality from springs or seeps on the mined property. Overburden mass was determined by three methods, with the method used by Pennsylvania researchers showing the most accurate results for overburden mass. A poor relationship existed between MPA and post-mining water quality, NP was intermediate, and NNP and the NP to MPA ratio showed the best prediction accuracy. In this study, NNP and the NP to MPA ratio gave identical water quality prediction results. Therefore, with NP to MPA ratios, values were separated into categories: <1 should produce acid drainage, between 1 and 2 can produce either acid or alkaline water conditions, and >2 should produce alkaline water. On our 56 surface mined sites, NP to MPA ratios varied from 0.1 to 31, and six sites (11%) did not fit the expected pattern using this category approach. Two sites with ratios <1 did not produce acid drainage as predicted (the drainage was neutral), and four sites with a ratio >2 produced acid drainage when they should not have. These latter four sites were either mined very slowly, had nonrepresentative ABA data, received water from an adjacent underground mine, or had a surface mining practice that degraded the water. In general, an NP to MPA ratio of <1 produced mostly acid drainage sites, between 1 and 2 produced mostly alkaline drainage sites, while NP to MPA ratios >2 produced alkaline drainage with a few exceptions. Using these values, ABA is a good tool to assess overburden quality before surface mining and to predict post-mining drainage quality after mining. The interpretation from ABA values was correct in 50 out of 52 cases (96%), excluding the four anomalous sites, which had acid water for reasons other than overburden quality.

Coal↗

[Physiological and hygienic evaluation of controllers' work in coal mines].

The article contains data on the results of the on-the-spot studies of coal mine controllers' labour conditions and relating hygienic factors. It was established that the work at the dispatcher's control point in coal mines was considerably affected by unfavourable factors (low degree of illumination, constructional shortcomings of the control point and seat) with concomitant neuropsychic stress conditions. Revealed were specific functional shifts in CVS and CNS, neuromuscular disorders and the analyzers' malfunctioning. A set of measures was proposed for labour conditions improvement, ergometric perfection of the working place and reduction of neuroemotional tension.

Administrative Personnel↗

Predictive toxicogenomics approaches reveal underlying molecular mechanisms of nongenotoxic carcinogenicity.

Toxicogenomics technology defines toxicity gene expression signatures for early predictions and hypotheses generation for mechanistic studies, which are important approaches for evaluating toxicity of drug candidate compounds. A large gene expression database built using cDNA microarrays and liver samples treated with over one hundred paradigm compounds was mined to determine gene expression signatures for nongenotoxic carcinogens (NGTCs). Data were obtained from male rats treated for 24 h. Training/testing sets of 24 NGTCs and 28 noncarcinogens were used to select genes. A semiexhaustive, nonredundant gene selection algorithm yielded six genes (nuclear transport factor 2, NUTF2; progesterone receptor membrane component 1, Pgrmc1; liver uridine diphosphate glucuronyltransferase, phenobarbital-inducible form, UDPGTr2; metallothionein 1A, MT1A; suppressor of lin-12 homolog, Sel1h; and methionine adenosyltransferase 1, alpha, Mat1a), which identified NGTCs with 88.5% prediction accuracy estimated by cross-validation. This six genes signature set also predicted NGTCs with 84% accuracy when samples were hybridized to commercially available CodeLink oligo-based microarrays. To unveil molecular mechanisms of nongenotoxic carcinogenesis, 125 differentially expressed genes (P<0.01) were selected by Student's t-test. These genes appear biologically relevant, of 71 well-annotated genes from these 125 genes, 62 were overrepresented in five biochemical pathway networks (most linked to cancer), and all of these networks were linked by one gene, c-myc. Gene expression profiling at early time points accurately predicts NGTC potential of compounds, and the same data can be mined effectively for other toxicity signatures. Predictive genes confirm prior work and suggest pathways critical for early stages of carcinogenesis.

Animals↗

Biomedical informatics: development of a comprehensive data warehouse for clinical and genomic breast cancer research.

The Windber Research Institute is an integrated high-throughput research center employing clinical, genomic and proteomic platforms to produce terabyte levels of data. We use biomedical informatics technologies to integrate all of these operations. This report includes information on a multi-year, multi-phase hybrid data warehouse project currently under development in the Institute. The purpose of the warehouse is to host the terabyte-level of internal experimentally generated data as well as data from public sources. We have previously reported on the phase I development, which integrated limited internal data sources and selected public databases. Currently, we are completing phase II development, which integrates our internal automated data sources and develops visualization tools to query across these data types. This paper summarizes our clinical and experimental operations, the data warehouse development, and the challenges we have faced. In phase III we plan to federate additional manual internal and public data sources and then to develop and adapt more data analysis and mining tools. We expect that the final implementation of the data warehouse will greatly facilitate biomedical informatics research.

Breast Neoplasms↗

Coupled two-way clustering analysis of breast cancer and colon cancer gene expression data.

UNLABELLED: We present and review coupled two-way clustering, a method designed to mine gene expression data. The method identifies submatrices of the total expression matrix, whose clustering analysis reveals partitions of samples (and genes) into biologically relevant classes. We demonstrate, on data from colon and breast cancer, that we are able to identify partitions that elude standard clustering analysis. AVAILABILITY: Free, at http://ctwc.weizmann.ac.il.. SUPPLEMENTARY INFORMATION: http://www.weizmann.ac.il/physics/complex/compphys/bioinfo2/

Algorithms↗

Mining single nucleotide polymorphisms from EST data of silkworm, Bombyx mori, inbred strain Dazao.

We made use of 81,635 expressed sequence tags (ESTs) derived from 12 different cDNA libraries of the silkworm, Bombyx mori, inbred strain Dazao (P50), to identify high-quality candidate single nucleotide polymorphisms (SNPs). By PHRAP assembling, 12,980 contigs containing 11,537 contigs assembled by more than one read were obtained, and 101 candidate SNPs and 27 single base insertions/deletions were identified from 117 contigs assembled from 1576 high-quality reads base-called with PHRED and screened on the basis of the neighborhood quality standard (NQS). Simultaneously, we also predicted 40 SNPs in coding regions (cSNPs), of which 26 were predicted to lead to amino acid non-synonymous variations and 14 synonymous substitutions. Also, the 1.66:1 ratio of transition/transversion is different from that of other insects. As the first SNP analysis of a Lepidoptera, B. mori, the single nucleotide polymorphic density is estimated to be 1.3 x 10(-3) by sequence diversity. This analysis shows that expressed sequences from multiple libraries may provide an abundant source of comparative reads to mine for cSNPs from the silkworm genome.

Amino Acid Substitution↗

Baseline levels of selected trace elements in Colorado oil shale region animals.

Baseline levels of boron, fluorine, molybdenum, and copper are described for 18 mule deer (Odocoileus hemionus) and for 45 composite samples of deer mice (Peromyscus maniculatus) from the Piceance Creek Basin, Rio Blanco County, Colorado. These data were collected before oil shale mining took place, and can be used to compare with levels found after mining is initiated. The data can thus be used to monitor changes in levels in animal tissues and as a basis for mitigating possible harmful effects due to the mining. Mean ppm (+/- S.D.) dry basis of each element is presented for selected tissues of each species. Results are also presented by habitat type for deer mice and by age for mule deer. Significant differences (P < 0.05) in molybdenum levels in deer mice were found between habitats. Significant differences (P < 0.05) were found between fawns and adult mule deer for boron levels, but not for the other elements. A need to standardize bone selection for analysis of fluorine was indicated. Kidneys appeared to be the organ of choice for baseline sampling of molybdenum and copper, and livers may be the organ of choice when toxic levels are suspected.

Animals↗

Integration of different data bodies for humanitarian decision support: an example from mine action.

Geographic information systems (GIS) are increasingly used for integrating data from different sources and substantive areas, including in humanitarian action. The challenges of integration are particularly well illustrated by humanitarian mine action. The informational requirements of mine action are expensive, with socio-economic impact surveys costing over US$1.5 million per country, and are feeding a continuous debate on the merits of considering more factors or 'keeping it simple'. National census offices could, in theory, contribute relevant data, but in practice surveys have rarely overcome institutional obstacles to external data acquisition. A positive exception occurred in Lebanon, where the landmine impact survey had access to agricultural census data. The challenges, costs and benefits of this data integration exercise are analysed in a detailed case study. The benefits are considerable, but so are the costs, particularly the hidden ones. The Lebanon experience prompts some wider reflections. In the humanitarian community, data integration has been fostered not only by the diffusion of GIS technology, but also by institutional changes such as the creation of UN-led Humanitarian Information Centres. There is a question whether the analytic capacity is in step with aggressive data acquisition. Humanitarian action may yet have to build the kind of strong analytic tradition that public health and poverty alleviation have accomplished.

Agriculture↗

Discovery informatics: its evolving role in drug discovery.

Drug discovery and development is a highly complex process requiring the generation of very large amounts of data and information. Currently this is a largely unmet informatics challenge. The current approaches to building information and knowledge from large amounts of data has been addressed in cases where the types of data are largely homogeneous or at the very least well-defined. However, we are on the verge of an exciting new era of drug discovery informatics in which methods and approaches dealing with creating knowledge from information and information from data are undergoing a paradigm shift. The needs of this industry are clear: Large amounts of data are generated using a variety of innovative technologies and the limiting step is accessing, searching and integrating this data. Moreover, the tendency is to move crucial development decisions earlier in the discovery process. It is crucial to address these issues with all of the data at hand, not only from current projects but also from previous attempts at drug development. What is the future of drug discovery informatics? Inevitably, the integration of heterogeneous, distributed data are required. Mining and integration of domain specific information such as chemical and genomic data will continue to develop. Management and searching of textual, graphical and undefined data that are currently difficult, will become an integral part of data searching and an essential component of building information- and knowledge-bases.

Artificial Intelligence↗

Comparison of two carbon analysis methods for monitoring diesel particulate levels in mines.

Two carbon analysis methods are currently being applied to the occupational monitoring of diesel particulate matter. Both methods are based on thermal techniques for the determination of organic and elemental carbon. In Germany, method ZH 1/120.44 has been published. This method, or a variation of it, is being used for compliance measurements in several European countries, and a Comité Européen de Normalization Working Group was formed recently to address the establishment of a European measurement standard. In the USA, a 'thermal-optical' method has been published as Method 5040 by the National Institute for Occupational Safety and Health. As with ZH 1/120.44, organic and elemental carbon are determined through temperature and atmosphere control, but different instrumentation and analysis conditions are used. Although the two methods are similar in principle, they gave statistically different results in a previous interlaboratory comparison. Because different instruments and operating conditions are used, between-method differences can be expected in some cases. Reasonable agreement is expected when the sample contains no other (i.e., non-diesel) sources of carbonaceous particulate and the organic fraction is essentially removed below about 500 degrees C. Airborne particulate samples from some mines may meet these criteria. Comparison data on samples from mines are important because the methods are being applied in this workplace for occupational monitoring and epidemiological studies. In this paper, results of a recent comparison on samples collected in a Canadian mine are reported. As seen in a previous comparison, there was good agreement between the total carbon results found by the two methods, with ZH 1/120.44 giving about 6% less carbon than Method 5040. Differences in the organic and elemental carbon results were again seen, but they were much smaller than those obtained in the previous comparison. The relatively small differences in the split between organic and elemental carbon are attributed to the different thermal programs used.

Air Pollution, Indoor↗

[Occupational morbidity of miners engaged in the capital building of iron ore mines].

The article contains data on the 1961-1985 occupational morbidity rates in the workers engaged in the Krivbass iron-ore mines, and narrates on the morbidity rates, structure and dynamics for different professions, age-groups and duration of work. An analysis is given of the concomitant somatic diseases, along with proposals for investigating ways and means of improving technologies in non-blasting iron-ore mines.

Adult↗

Origin of arsenic pollution in Southwest Tuscany: comparison of fluvial sediments.

The Colline Metallifere in SW Tuscany are characterized by strong anomalies in arsenic concentrations and distribution. The area is sparsely populated and largely wild, though it has been subject to human impact due to mining and metal processing since Etruscan and Roman times. In the Middle Ages it was exploited intensively for silver and copper. Until 1995, pyrite (FeS2) was mined and roasted to produce sulphuric acid and iron. Hypotheses based on geological and mineralogical factors formulated in the last 20 years have failed to explain the peculiar distribution of arsenic in the Colline Metallifere. Here we report preliminary results of widespread sampling and analysis of the fluvial sediments of rivers originating in this mining area. The data was analysed in relation to the archaeological features of the area, since the presence of ancient mining and ore processing sites can shed light on the peculiar distribution of arsenic. Comparison of data from two rivers and their respective contaminated and uncontaminated coastal lagoons also clarified the general mechanisms of arsenic mobility, pinpointing the source of arsenic contamination. The study methods also promise to be useful for discovering unknown archaeological sites.

Archaeology↗

Mining of assembled expressed sequence tag (EST) data for protein families: application to the G protein-coupled receptor superfamily.

The availability of large expressed sequence tag (EST) databases has led to a revolution in the way new genes are identified. Mining of these databases using known protein sequences as queries is a powerful technique for discovering orthologous and paralogous genes. The scientist is often confronted, however, by an enormous amount of search output owing to the inherent redundancy of EST data. In addition, high search sensitivity often cannot be achieved using only a single member of a protein superfamily as a query. In this paper a technique for addressing both of these issues is described. Assembled EST databases are queried with every member of a protein superfamily, the results are integrated and false positives are pruned from the set. The result is a set of assemblies enriched in members of the protein superfamily under consideration. The technique is applied to the G protein-coupled receptor (GPCR) superfamily in the construction of a GPCR Resource. A novel full-length human GPCR identified from the GPCR Resource is presented, illustrating the utility of the method.

Amino Acid Sequence↗

Phenotypic data in FlyBase.

Phenotypic analysis combined with molecular genetics is a powerful tool for mapping gene function onto the genome. Phenotypic data are, by their nature, descriptive, and as varied as the range of mutant phenotypes that can be presented by the organism under study. This paper discusses the mechanisms FlyBase has implemented to systematise published phenotypic data about Drosophila, and provides an introduction to the query tools available for the mining of the data. Though FlyBase is specific to Drosophila, the issues faced in devising protocols for capturing, storing and reporting data are the same issues faced by any database with an interest in using phenotypic data to maximise the potential of genomic analysis.

Alleles↗

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus↗