Search PubMedSearch

SEARCH · Search PubMed

Results for “data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Locations of fatal work injuries in the United States: 1980 to 1985.

Surveillance of injury at work is beset with problems of method and definition. As a result, national agencies have widely varying estimates of the number of fatal work injuries in the United States. One plausible method for identifying fatal work injuries is to use the Place of Injury variable, which is entered on all US death certificates but is not encoded by the National Center for Health Statistics. To use this method, one would assume that work injuries largely occur at "typical work sites," ie, places coded as industrial, farm, and mine and quarry. Data to test this method were derived from the National Traumatic Occupational Fatality data base maintained by the National Institute for Occupational Safety and Health. Analysis of this data base showed that work-related fatal injuries mostly occur in places where many non-work-related injuries also occur. Only about one third of fatal work injuries took place at locations coded as industrial, farm, and mine and quarry. As a method for identifying fatal work injuries, the Place of Injury variable appears to have little value.

Humans

Development of a model to aid in reconstruction of historical silica dust exposures in the taconite industry.

The ability to reconstruct employee exposure histories would be a valuable research tool for the evaluation of occupation as a factor in disease. In many cases, however, historical environmental data are available but have not been used to compute past exposures because of differences in sampling methods. This paper describes a quantitative model to convert historical environmental data (from taconite mine and mill operations) into a form consistent with current sampling methods and results and, therefore, will enable past exposure histories to be used. (Past exposure histories are to be determined in an epidemiological study.) In this study, parallel sampling results from the environmental data base were used to obtain a coefficient for the conversion of impinger-particle counts (old sampling method) to filter-respirable mass sampling results (new sampling method). Parameters in the model were estimated using multiple regression techniques. Results show that a consistent ratio exists between impinger-particle counts and filter-respirable mass concentrations for samples collected at the same locations.

Air Pollutants, Occupational

eVOC: a controlled vocabulary for unifying gene expression data.

Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.

Animals

Mortality of lead smelter workers.

To examine patterns of death in lead smelter workers, a retrospective analysis of mortality was conducted in a cohort of 1,987 males employed between 1940 and 1965 at a primary lead smelter in Idaho. Overall mortality was similar to that of the United States white male population (standardized mortality ratio (SMR) = 98). Excess mortality, however, was found from chronic renal disease (SMR = 192; confidence interval (CI) = 88-364), and the risk of death from renal disease increased with increasing duration of employment, such that after 20 years employment, the standardized mortality ratio reached 392 (CI = 107-1,004). Excess mortality was also noted for nonmalignant respiratory disease (SMR = 187, CI = 128-264). Eight of 32 deaths in this category were caused by silicosis; at least five workers who died of silicosis had been miners for a part of their lives. An additional 11 deaths resulted from tuberculosis (SMR = 139; CI = 69-249); in six of these cases, silicosis was a contributory cause of death. Cancer mortality was not increased overall (SMR = 95; CI = 78-114). An increase, however, was noted for deaths from kidney cancer (six cases; SMR = 204; CI = 75-444). Finally, excess mortality was noted for injuries (SMR = 138; CI = 104-179); 13 (23%) of the 56 deaths in this category were caused by mining injuries. The data from this study are consistent with previous reports of increased mortality from chronic renal disease in persons exposed occupationally to lead.

Actuarial Analysis

Estimates of lifetime lung cancer risks resulting from Rn progeny exposure.

Data on five mining populations exposed to Rn progeny have been used to estimate the lifetime risk of lung cancer resulting from occupational and environmental exposure under current standards. Slopes of dose-response relations for lung cancer show a tendency to decrease with increasing dose. Our best estimate of curvilinearity is given by raising dose to the power 0.92 +/- 0.07, but the improvement in fit beyond simple linearity is not significant. On the other hand, the addition of a cell-killing term significantly improves the fit of the linear model. In any event, linear extrapolation is unlikely to underestimate the excess risk at low doses by more than a factor of 1.5. However, these inferences about curvilinearity are highly subject to error from the choice of reference populations, dosimetry, and latency. Under the linear-cell-killing model, our best estimate of excess relative risk is 2.28 +/- 0.35 per 100 working level month (WLM) (a doubling dose of 44 WLM). Attributable risks in these five studies range from 3.4-17.8 per 10(6) person-yr WLM-1. Risks from Rn progeny appear to interact with age and smoking in a form intermediate between additive and multiplicative. The "relative risk" model is therefore preferable for projecting lifetime risks, but life-table projections are described for a wide variety of assumptions. Our best estimate of the effect of a 50-yr occupational exposure to 4 WLM yr-1 is 130 excess lung cancer deaths per 1000 persons (0.65 per 1000 person-WLM), with a range from 60-250 per 1000. Similar calculations for lifetime exposure to an additional 0.02 working level (WL) beyond normal background produces an estimate of 20 excess lung cancers per 1000 persons.

Adolescent

A novel isoform of tensin-1 promotes actin filament assembly for efficient erythroblast enucleation.

Mammalian red blood cells are generated via a terminal erythroid differentiation pathway culminating in cell polarization and enucleation. Actin filament (F-actin) polymerization is critical for enucleation, but the underlying molecular regulatory mechanisms remain poorly understood. We used publicly available RNA sequencing and proteomic data sets to mine for actin-regulatory factors differentially expressed during human erythroid differentiation and discovered that a focal adhesion (FA) protein, tensin-1 (TNS1), dramatically increases in expression late in differentiation. Remarkably, we found that differentiating human CD34+ cells express a novel truncated form of TNS1 (erythroid TNS1 [eTNS1]; Mr ∼125 kDa) missing the N-terminal half of the protein containing the actin-binding domain, due to an internal messenger RNA translation start site resulting in a unique exon 1E. The region upstream of eTNS1 has features of an active erythroid promoter, demonstrating increasing chromatin accessibility during terminal differentiation, paralleling increasing gene expression. Sequence comparisons across species indicate that eTNS1 is expressed in humans and nonhuman primates, but not in zebrafish, mice, or other rodents. Confocal microscopy showed that eTNS1 localized to the cytoplasm during terminal erythroid differentiation but, surprisingly, did not appear to form focal adhesions nor to colocalize with F-actin. Knockout of eTNS1 did not affect terminal differentiation or assembly of the spectrin membrane skeleton but led to reduced F-actin assembly and abnormal organization in polarized and enucleating erythroblasts, resulting in impaired enucleation efficiency. We conclude that eTNS1 is a novel regulator of F-actin during human erythroid terminal differentiation that is required for efficient enucleation.

Humans

Respirable dust exposures in U.S. surface coal mines (1982-1986).

Exposure of miners to respirable coal mine dust and to respirable quartz silica at surface coal mines in the United States during 1982-1986 were evaluated by job category using data collected by coal mine operators and Mine Safety and Health Administration (MSHA) inspectors. Average coal mine dust concentrations were usually well below the MSHA Permissible Exposure Limit (PEL) for all job categories, but at least 10% of the samples obtained from some coal preparation plant job areas and most drilling job areas had concentrations that exceeded the 2.0 mg/m3 limit. In contrast, a very high proportion of samples from surface mine driller areas exceeded the quartz PEL. Of all samples collected for highwall drill operators and helpers, 78% and 77%, respectively, were greater than the 0.1 mg/m3 quartz exposure limit (average concentrations were .32 and .36 mg/m3, respectively). Although MSHA compliance data may not be entirely adequate for assessing chronic exposure to quartz, these data and the results of other NIOSH studies nonetheless indicate excessive exposure to silica in a group of surface coal miners.

Air Pollutants, Occupational

Radiological classification of Polish underground mines and recommendations of surveillance.

The paper presents the most recent data, collected 1987-1989, on concentrations of 222Rn products in the air of all Polish underground non-uranium mines, and data on the exposure of the miners employed there. The concentrations and exposure of miners were evaluated by using 'passive' dosimeters, based on the track-etch solid state nuclear track detector, worn as small individual cassettes on helmets of representative groups in every mine for one month, four times a year (once in each season of the year). The paper contains the average annual exposure of miners in coal-, metal-ore-, and chemical raw materials oremines. The paper, also presents expected 'frequency' distributions of individual miners' exposure in particular types of mines, as well as the computer simulations of 'relative frequency' distributions of expected miner's exposure when the Annual Limit of Exposure would be adopted at the level 17, 12, 8.6, 6.9 and 3.4.10(-3) Jhm-3 (5.0, 3.5, 2.5, 2 and 1 WLM). The concept and criteria of classification of mines according to the radiation hazards are presented and discussed. According to that concept, all mines in Poland have been considered and classified into four classes of mines with a different level of radiation hazard. The appropriate radiological surveillance to the respective class of mine is proposed and discussed.

Air Pollutants, Occupational

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis

Identification and classification of ion-channels across the tree of life provide functional insights into understudied CALHM channels.

The ion channel (IC) genes encoded in the human genome play fundamental roles in cellular functions and disease and are one of the largest classes of druggable proteins. However, limited knowledge of the diverse molecular and cellular functions carried out by ICs presents a major bottleneck in developing selective chemical probes for modulating their functions in disease states. The wealth of sequence data available on ICs from diverse organisms provides a valuable source of untapped information for illuminating the unique modes of channel regulation and functional specialization. However, the extensive diversification of IC sequences and the lack of a unified resource present a challenge in effectively using existing data for IC research. Here, we perform integrative mining of available sequence, structure, and functional data on 419 human ICs across disparate sources, including extensive literature mining by leveraging advances in large language models to annotate and curate the full complement of the "channelome". We employ a well-established orthology inference approach to identify and extend the IC orthologs across diverse organisms to above 48,000. We show that the depth of conservation and taxonomic representation of IC sequences can further be translated to functional similarities by clustering them into functionally relevant groups, which can be used for downstream functional prediction on understudied members. We demonstrate this by delineating co-conserved patterns characteristic of the understudied family of the Calcium Homeostasis Modulator (CALHM) family of ICs. Through mutational analysis of co-conserved residues altered in human diseases and electrophysiological studies, we show that these evolutionarily-constrained residues play an important role in channel gating functions. Thus, by providing new tools and resources for performing large comparative analyses on ICs, this study addresses the unique needs of the IC community and provides the groundwork for accelerating the functional characterization of dark channels for therapeutic intervention.

CALHM1

Malaria among gold miners in southern Pará, Brazil: estimates of determinants and individual costs.

As malaria grows more prevalent in the Amazon frontier despite increased expenditures by disease control authorities, national and regional tropical disease control strategies are being called into question. The current crisis involving traditional control/eradication methods has broadened the search for feasible and effective malaria control strategies--a search that necessarily includes an investigation of the roles of a series of individual and community-level socioeconomic characteristics in determining malaria prevalence rates, and the proper methods of estimating these links. In addition, social scientists and policy makers alike know very little about the economic costs associated with malarial infections. In this paper, I use survey data from several Brazilian gold mining areas to (a) test the general reliability of malaria-related questionnaire response data, and suggest categorization methods to minimize the statistical influence of exaggerated responses, (b) estimate three statistical models aimed at detecting the socioeconomic determinants of individual malaria prevalence rates, and (c) calculate estimates of the average cost of a single bout of malaria. The results support the general reliability of survey response data gathered in conjunction with malaria research. Once the effects of vector exposure were controlled for, individual socioeconomic characteristics were only weakly linked to malaria prevalence rates in these very special miners' communities. Moreover, the socioeconomic and exposure links that were significant did not depend on the measure of malaria adopted. Finally, individual costs associated with malarial infections were found to be a significant portion of miners' incomes.

Brazil

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis

The derivation of estimated dust exposures for U.S. coal miners working before 1970.

A number of reports on the prevalence of coal workers' pneumoconiosis in U.S. coal miners have been published, yet very little is known about the relationship between dust exposure and pneumoconiosis levels in the U.S. This report describes the derivation of cumulative dust exposure estimates by back-extrapolation of data processed by the Mine Safety and Health Administration after 1970 by using a ratio of dust concentrations based on information collected during environmental surveys at certain U.S. mines by the Bureau of Mines between 1968 and 1969. Cumulative personal dust exposure estimates were calculated by using occupational histories obtained from the miners and job-specific estimates of dust concentration. In other reports, the resulting estimated exposures have been shown to correlate well with various measures of respiratory morbidity.

Adult