Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

MetaRouter: bioinformatics for bioremediation.

Bioremediation, the exploitation of biological catalysts (mostly microorganisms) for removing pollutants from the environment, requires the integration of huge amounts of data from different sources. We have developed MetaRouter, a system for maintaining heterogeneous information related to bioremediation in a framework that allows its query, administration and mining (application of methods for extracting new knowledge). MetaRouter is an application intended for laboratories working in biodegradation and bioremediation, which need to maintain and consult public and private data, linked internally and with external databases, and to extract new information from it. Among the data-mining features is a program included for locating biodegradative pathways for chemical compounds according to a given set of constraints and requirements. The integration of biodegradation information with the corresponding protein and genome data provides a suitable framework for studying the global properties of the bioremediation network. The system can be accessed and administrated through a web interface. The full-featured system (except administration facilities) is freely available at http://pdg.cnb.uam.es/MetaRouter. Additional material: http://www.pdg.cnb.uam.es/biodeg_net/MetaRouter.

Bacteria↗

The impact of roster changes on absenteeism and incident frequency in an Australian coal mine.

BACKGROUND: The occupational health and safety implications associated with compressed and extended work periods have not been fully explored in the mining sector. AIMS: To examine the impact on employee health and safety of changes to the roster system in an Australian coal mine. METHODS: Absenteeism and incident frequency rate data were collected over a 33 month period that covered three different roster schedules. Period 1 covered the original 8-hour/7-day roster. Period 2 covered a 12-month period under a 12-hour/7-day schedule, and period 3 covered a 12-month period during which a roster that scheduled shifts only on weekdays, with uncapped overtime on weekends and days off (12-hour/5-day) was in place. Data were collected and analysed from the maintenance, mining, and coal preparation plant (CPP) sectors. RESULTS: The only significant change in absenteeism rates was an increase in the maintenance sector in the third data collection period. Absenteeism rates in the mining and CPP sectors were not different between data collection periods. The increase in the maintenance sector may be owing to: (1) a greater requirement for maintenance employees to perform overtime as a result of the roster change compared to other employee groups; or (2) greater monotony associated with extended work periods for maintenance employees compared to others. After the first roster change, accident incident frequency decreased in the CPP sector but not in the other sectors. There was no effect on incident frequency after the second roster change in any sector. CONCLUSIONS: The current study did not find significant negative effects of a 12-hour pattern, when compared to an 8-hour system. However, when unregulated and excessive overtime was introduced as part of the 12-hour/5-day roster, absenteeism rates were increased in the maintenance sector. The combination of excessive work hours and lack of consultation with employees regarding the second change may have contributed to the overall negative effects.

Absenteeism↗

Extended workdays in an underground mine: a work performance analysis.

Many companies in different industrial sectors are exploring alternative work schedules to deal with diverse problems associated with shiftwork. The use of extended workday schedules (regular shift lengths exceeding 8 h with compressed workweeks) is attracting growing interest in many industries that use continuous operations. To address concerns regarding possible fatigue effects on safety and work performance associated with such schedules, the U.S. Bureau of Mines conducted a two-phase study at an underground metal mine in western Canada. Data were collected before and after a group of workers employed at the mine changed from an 8- to a 12-h schedule. Results indicate nearly unanimous acceptance and improved sleep quality associated with the new schedule. In general, fatigue-sensitive behavioral and physiological performance measures show either no change or improvement with 12-h shifts. We conclude that the extended workday schedule should be retained but periodically reevaluated.

Adult↗

Disease and illness in U.S. mining, 1983-2001.

OBJECTIVES: We describe inconsistencies in disease and illness reporting in U.S. mining, identify under-reporting of disease and illness in U.S. mining, and summarize selected disease and illness in U.S. mining from 1983 through 2001. METHODS: We summarized information on mining-related disease and illness data for the years 1983-2001 from the Mining Safety and Health Administration database (MSHA). RESULTS: Discrepancies exist in types of information collected by the Centers for Disease and Control, the National Institute for Occupational Safety and Health, and the Mining Safety and Health Administration database. Several factors, including a worker's fear of losing his or her job, health insurance, or other job-related benefits contribute to under-reporting of disease and illness information in the US mining industry. CONCLUSIONS: Since 1997, both number of workers employed in mining and disease and illness rates have decreased; however, the highest disease and illness rates in mining continue to be coal worker's pneumoconiosis and hearing loss.

Asbestos↗

Origin and evolution of the chloroplast division machinery.

Chloroplasts were originally established in eukaryotes by the endosymbiosis of a cyanobacterium; they then spread through diversification of the eukaryotic hosts and subsequent engulfment of eukaryotic algae by previously nonphotosynthetic eukaryotes. The continuity of chloroplasts is maintained by division of preexisting chloroplasts. Like their ancestors, chloroplasts use a bacterial division system based on the FtsZ ring and some associated factors, all of which are now encoded in the host nuclear genome. The majority of bacterial division factors are absent from chloroplasts and several new factors have been added by the eukaryotic host. For example, the ftsZ gene has been duplicated and modified, plastid-dividing (PD) rings were most likely added by the eukaryotic host, and a member of the dynamin family of proteins evolved to regulate chloroplast division. The identification of several additional proteins involved in the division process, along with data from diverse lineages of organisms, our current knowledge of mitochondrial division, and the mining of genomic sequence data have enabled us to begin to understand the universality and evolution of the division system. The principal features of the chloroplast division system thus far identified are conserved across several lineages, including those with secondary chloroplasts, and may reflect primeval features of mitochondrial division.

Biological Evolution↗

The separation of two hymenopteran parasitoids, Tersilochus obscurator and Tersilochus microgaster (Ichneumonidae), of stem-mining pests of winter oilseed rape using DNA, morphometric and ecological data.

Tersilochus obscurator Aubert and Tersilochus microgaster (Szépligeti) are larval endoparasitoids of economically-important stem-mining pests of winter oilseed rape (Brassica napus L.) in Europe. They are difficult to separate morphologically. Their hosts are Ceutorhynchus pallidactylus (Marsham) and Psylliodes chrysocephala Linnaeus, respectively. The parasitoids' taxonomic status, identification, host range and phenology were studied using genetic, morphometric and ecological data. The study used 527 female parasitoids from the UK and Germany, either field-collected in emergence traps or reared from field-collected host larvae. Two morphometric characters, the ovipositor sheath to first metasomal tergite ratio and the percentage of the mesopleuron spanned by the sternaulus, were measured. A 440 bp section of mitochondrial DNA cytochrome oxidase subunit I (COI) gene was sequenced from 35 parasitoids reared from C. pallidactylus, 20 reared from P. chrysocephala and individuals from two outgroups, Tersilochus heterocerus Thomson and Phradis interstitialis Thomson. Distinct and invariable COI sequences corresponded exclusively to each parasitoid group, confirming that T. obscurator and T. microgaster are discrete species. Measurements of host-reared and COI-sequenced specimens indicated that the ranges of both morphometric characters overlapped between species. Using these ranges as criteria, all but 3.6% of UK specimens and 2% of German specimens were identifiable to species without reference to host or phenology. There were differences in emergence phenology in the UK, adult T. microgaster emerging from winter diapause by 29 March 2000, T. obscurator emerging between 12 April and 24 May 2000. The value of molecular techniques in the identification of closely-related parasitoid species is discussed.

Animals↗

RESEARCH: Assessment of Nonpoint Source Pollution from Inactive Mines Using a Watershed-Based Approach

/ A watershed-based approach for screening-level assessment of nonpoint source pollution from inactive and abandoned metal mines was developed and illustrated. The methodology was designed to use limited stream discharge and chemical data from synoptic surveys to derive key information required for targeting impaired waterbodies and critical source areas for detailed investigation and remediation. The approach was formulated based on the required attributes of an assessment methodology, information goals for targeting, attributes of data that are typical of basins with inactive mines, and data analysis methods that were useful for the case study. The methodology is presented as steps in a framework including evaluation of existing data/information and identification of data gaps; definition of assessment information goals for targeting and monitoring design; data collection, management, and analysis; and information reporting and use for targeting. Information generated includes the type and extent of and critical conditions for water-quality impairment, concentrations in and loadings to streams, differences between concentrations in and loadings to streams, and risks of exceeding target concentrations and loadings. Data from the Cement Creek Basin, located in the San Juan Mountains of southwestern Colorado, USA, were used to help develop and illustrate application of the methodology. The required information was derived for Cement Creek and used for preliminary targeting of locations for detailed investigation and remediation. Application of the approach to Cement Creek was successful in terms of cost-effective generation of information and use for targeting.KEY WORDS: Water quality assessment; Nonpoint source pollution; Inactive mines; Watershed

Journal Article↗

Radiological classification of Polish underground mines and recommendations of surveillance.

The paper presents the most recent data, collected 1987-1989, on concentrations of 222Rn products in the air of all Polish underground non-uranium mines, and data on the exposure of the miners employed there. The concentrations and exposure of miners were evaluated by using 'passive' dosimeters, based on the track-etch solid state nuclear track detector, worn as small individual cassettes on helmets of representative groups in every mine for one month, four times a year (once in each season of the year). The paper contains the average annual exposure of miners in coal-, metal-ore-, and chemical raw materials oremines. The paper, also presents expected 'frequency' distributions of individual miners' exposure in particular types of mines, as well as the computer simulations of 'relative frequency' distributions of expected miner's exposure when the Annual Limit of Exposure would be adopted at the level 17, 12, 8.6, 6.9 and 3.4.10(-3) Jhm-3 (5.0, 3.5, 2.5, 2 and 1 WLM). The concept and criteria of classification of mines according to the radiation hazards are presented and discussed. According to that concept, all mines in Poland have been considered and classified into four classes of mines with a different level of radiation hazard. The appropriate radiological surveillance to the respective class of mine is proposed and discussed.

Air Pollutants, Occupational↗

Strong-association-rule mining for large-scale gene-expression data analysis: a case study on human SAGE data.

BACKGROUND: The association-rules discovery (ARD) technique has yet to be applied to gene-expression data analysis. Even in the absence of previous biological knowledge, it should identify sets of genes whose expression is correlated. The first association-rule miners appeared six years ago and proved efficient at dealing with sparse and weakly correlated data. A huge international research effort has led to new algorithms for tackling difficult contexts and these are particularly suited to analysis of large gene-expression matrices. To validate the ARD technique we have applied it to freely available human serial analysis of gene expression (SAGE) data. RESULTS: The approach described here enables us to designate sets of strong association rules. We normalized the SAGE data before applying our association rule miner. Depending on the discretization algorithm used, different properties of the data were highlighted. Both common and specific interpretations could be made from the extracted rules. In each and every case the extracted collections of rules indicated that a very strong co-regulation of mRNA encoding ribosomal proteins occurs in the dataset. Several rules associating proteins involved in signal transduction were obtained and analyzed, some pointing to yet-unexplored directions. Furthermore, by examining a subset of these rules, we were able both to reassign a wrongly labeled tag, and to propose a function for an expressed sequence tag encoding a protein of unknown function. CONCLUSIONS: We show that ARD is a promising technique that turns out to be complementary to existing gene-expression clustering techniques.

Algorithms↗

Statistical comparison of diesel particulate matter measurement methods.

Four methods are used to quantify diesel particulate matter (DPM) in the mine environment: respirable combustible dust sampling (RCD), size selective sampling with gravimetric analysis (SSG), respirable dust sampling with elemental carbon (EC) analysis, and respirable dust sampling with total carbon (TC) analysis. The authors assembled data from three underground mine studies to statistically compare these methods. The sampling protocol used in each study was similar. For all the four methods, samples were collected in triplicate at three locations-upwind and downwind of the diesel scoop and on the scoop. The methods were compared with respect to their precision, selectivity, sensitivity, and LOD, as well as their limitations in measuring DPM concentrations. This constitutes a meta-analysis of the available data and provides information over a broader range of mining conditions and DPM concentrations than any of the individual studies. The weighing imprecision for the SSG method is almost twice that for the RCD technique. The imprecision of the EC and TC methods are a function of the mass loading, and EC has a lower imprecision than TC. The EC method was used as the reference "gold standard" against which the other methods were evaluated. The RCD, SSG, and TC methods exhibited substantial levels of interference, leading to much higher minimum concentrations that can be measured by these methods. Of the three, the SSG method has the highest level of interference, primarily from nondiesel material that is collected in the <0.8 microm size range.

Air Pollutants, Occupational↗

Discover protein sequence signatures from protein-protein interaction data.

BACKGROUND: The development of high-throughput technologies such as yeast two-hybrid systems and mass spectrometry technologies has made it possible to generate large protein-protein interaction (PPI) datasets. Mining these datasets for underlying biological knowledge has, however, remained a challenge. RESULTS: A total of 3108 sequence signatures were found, each of which was shared by a set of guest proteins interacting with one of 944 host proteins in Saccharomyces cerevisiae genome. Approximately 94% of these sequence signatures matched entries in InterPro member databases. We identified 84 distinct sequence signatures from the remaining 172 unknown signatures. The signature sharing information was then applied in predicting sub-cellular localization of yeast proteins and the novel signatures were used in identifying possible interacting sites. CONCLUSION: We reported a method of PPI data mining that facilitated the discovery of novel sequence signatures using a large PPI dataset from S. cerevisiae genome as input. The fact that 94% of discovered signatures were known validated the ability of the approach to identify large numbers of signatures from PPI data. The significance of these discovered signatures was demonstrated by their application in predicting sub-cellular localizations and identifying potential interaction binding sites of yeast proteins.

Binding Sites↗

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis↗

A re-evaluation of radiological evidence from a study of U.S. strip coal miners.

In 1972, the U.S. Public Health Service examined 1438 workers employed at seven bituminous and one anthracite U.S. strip coal mines. One conclusion from the study was that workers without previous dust exposures were not at risk of category 2 or higher pneumoconiosis from their strip coal mining environment. Because of recent concerns for silicosis among strip coal miners, the radiographs were reinterpreted and the data re-evaluated. In addition, data from respirable coal mine dust samples collected from 1972 to 1979 in all surface coal mines were analyzed. The results showed that category 2 or higher pneumoconiosis was prevalent among strip coal miners with experience in an underground coal mine. Among those without underground coal mine experience, category 2 or higher was prevalent among anthracite strip miners, but not among bituminous strip miners. Average respirable coal mine dust exposures in the anthracite mine were less than 1 mg/m3 prior to 1975 and, coupled with the radiographic findings, suggest further study of the efficacy of the 2 mg/m3 U.S. Federal surface coal mine dust standard in anthracite coal mines.

Coal Mining↗

Evaluating remedial alternatives for an acid mine drainage stream: application of a reactive transport model.

A reactive transport model based on one-dimensional transport and equilibrium chemistry is applied to synoptic data from an acid mine drainage stream. Model inputs include streamflow estimates based on tracer dilution, inflow chemistry based on synoptic sampling, and equilibrium constants describing acid/base, complexation, precipitation/dissolution, and sorption reactions. The dominant features of observed spatial profiles in pH and metal concentration are reproduced along the 3.5-km study reach by simulating the precipitation of Fe(III) and Al solid phases and the sorption of Cu, As, and Pb onto freshly precipitated iron(III) oxides. Given this quantitative description of existing conditions, additional simulations are conducted to estimate the streamwater quality that could result from two hypothetical remediation plans. Both remediation plans involve the addition of CaCO3 to raise the pH of a small, acidic inflow from approximately 2.4 to approximately 7.0. This pH increase results in a reduced metal load that is routed downstream by the reactive transport model, thereby providing an estimate of post-remediation water quality. The first remediation plan assumes a closed system wherein inflow Fe(II) is not oxidized by the treatment system; under the second remediation plan, an open system is assumed, and Fe(II) is oxidized within the treatment system. Both plans increase instream pH and substantially reduce total and dissolved concentrations of Al, As, Cu, and Fe(II+III) at the terminus of the study reach. Dissolved Pb concentrations are reduced by approximately 18% under the first remediation plan due to sorption onto iron(III) oxides within the treatment system and stream channel. In contrast, iron(III) oxides are limiting under the second remediation plan, and removal of dissolved Pb occurs primarily within the treatment system. This limitation results in an increase in dissolved Pb concentrations over existing conditions as additional downstream sources of Pb are not attenuated by sorption.

Environmental Monitoring↗

Application of geostatistical methods to arsenic data from soil samples of the Cova dos Mouros mine (Vila Verde-Portugal).

A total of 286 soil samples were collected in the Cova dos Mouros area. All samples were dry sieved into the <200 mesh size fraction and analysed for Fe, Cu, Zn, Pb, Co, Ni, Bi and Mn by atomic absorption spectrometry (AAS) and for As, Se, Sb and Te by atomic absorption spectrometry-hydrid generation (AAS-HG). Only the results of arsenic are discussed in this paper although the survey was extended to all analysed chemical elements. The purpose of this study was to make a risk probability mapping for arsenic that would allow better knowledge about the vulnerability of the soil to arsenic contamination. To achieve this purpose, the initial variable was transformed into an indicator variable using as thresholds the risk-based standards (intervention values) for soils, as proposed by [Swartjes 1999. Risk based assessment of soil and groundwater quality in the Netherlands: Standards and remediation. J. Geochem. Explor.73 1-10]. To account for spatial structure, sample variograms were computed for the main directions of the sampling grid and a spherical model was fitted to each sample variogram (arsenic variable and indicator variables). The parameters of the spherical model fitted to the arsenic variable were used to predict arsenic concentrations at unsampled locations. A risk probability mapping was also done to assess the vulnerability of the soil towards the mining works. The parameters of the spherical model fitted to each indicator variable were used to estimate probabilities of exceeding the corresponding threshold. The use of indicator kriging as an alternative to ordinary kriging for the soil data of Cova dos Mouros produced unbiased probability maps that allowed assessment of the quality of the soil.

Arsenic↗

Genome-wide genotyping in Parkinson's disease and neurologically normal controls: first stage analysis and public release of data.

BACKGROUND: Several genes underlying rare monogenic forms of Parkinson's disease have been identified over the past decade. Despite evidence for a role for genetics in sporadic Parkinson's disease, few common genetic variants have been unequivocally linked to this disorder. We sought to identify any common genetic variability exerting a large effect in risk for Parkinson's disease in a population cohort and to produce publicly available genome-wide genotype data that can be openly mined by interested researchers and readily augmented by genotyping of additional repository subjects. METHODS: We did genome-wide, single-nucleotide-polymorphism (SNP) genotyping of publicly available samples from a cohort of Parkinson's disease patients (n=267) and neurologically normal controls (n=270). More than 408,000 unique SNPs were used from the Illumina Infinium I and HumanHap300 assays. FINDINGS: We have produced around 220 million genotypes in 537 participants. This raw genotype data has been and as such is the first publicly accessible high-density SNP data outside of the International HapMap Project. We also provide here the results of genotype and allele association tests. INTERPRETATION: We generated publicly available genotype data for Parkinson's disease patients and controls so that these data can be mined and augmented by other researchers to identify common genetic variability that results in minor and moderate risk for disease.

Aged↗

Exploiting naturally occurring DNA variation and molecular profiling data to dissect disease and drug response traits.

Identifying the key drivers of common human diseases and associated signaling pathways remains one of the primary objectives in the biomedical and life sciences. In this respect, common inbred strains of mice have played a crucial role, and recent advances in the development of genomics and bioinformatics tools have significantly enhanced their utility for this purpose. These advances have enabled a more holistic, network-oriented view of biological systems that facilitates elucidation of the underlying causes of disease and the best ways to target them. Success in reconstructing gene networks underlying disease traits (or other complex traits like drug response) and identifying the key drivers of these traits now largely rests on integrative approaches that combine data from multiple different sources. Such integrative genomics approaches that take into account genotypic, molecular profiling and clinical data in segregating mouse populations have recently been developed. Key to this integration has been the development and application of sophisticated algorithms to mine the diversity of data.

Algorithms↗