Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

ConceptDrift: leveraging spatial, temporal and semantic evolution of biomedical concepts for hypothesis generation.

MOTIVATION: Hypothesis generation is a fundamental problem in biomedical text mining that aims to generate ideas that are new, interesting, and plausible by discovering unexplored links between biomedical concepts. Despite significant advances made by existing approaches, they do not fully leverage the evolutionary properties of biomedical concepts. This is limiting because scientific knowledge continually evolves over time, with new facts being added and old ones becoming obsolete. Thus, it is crucial to capture the evolutionary properties of biomedical concepts from multiple perspectives (e.g. spatial, temporal, and semantic) to generate hypotheses that reflect the up-to-date information landscape of the biomedical domain. RESULTS: We introduce a novel framework, ConceptDrift, that models the hypothesis generation task as a sequence of temporal graphlets and simultaneously encodes spatial, temporal, and semantic change. Unlike existing approaches that treat these dimensions independently, ConceptDrift is the first to provide a holistic understanding of concept evolution by integrating them into a unified framework. Grounded in the theories of the Distributional Hypothesis and Conceptual Change, our method adapts these principles to the unique challenges of large-scale biomedical literature. We conduct extensive experiments across multiple datasets and demonstrate that ConceptDrift consistently outperforms state-of-the-art baselines in generating accurate and meaningful hypotheses. Our framework shows immediate practical benefits for web-based literature mining tools in life sciences and biomedicine, offering more robust and predictive feature representations. AVAILABILITY AND IMPLEMENTATION: https://github.com/amir-hassan25/ConceptDrift (DOI: 10.6084/m9.figshare.29975476).

Semantics↗

Lung cancer in the Schneeberg mines: a reappraisal of the data reported by Harting and Hesse in 1879.

The first description of occupational lung cancer, by Harting and Hesse in 1879, unfortunately is not readily accessible. Its account of the vicissitudes of the Schneeberg miners merits study and is therefore presented in summary and set in a historical and geological context. The authors attempted to discover the cause of the disease and made recommendations for improving the health of miners. In the course of their programme of investigations, they developed methods for measuring airborne dust and inhaled dust by personal monitoring. It was left to subsequent discovery for radon and its daughter products to be identified as the causal agents. Later generations were to discover the impact of radioactive spoils from mines situated in the mountain range in which Schneeberg was located.

Germany↗

Transport injuries in small coal mines: an exploratory analysis.

Mine Safety and Health Administration (MSHA) surveillance data were analyzed to elucidate mine characteristics or injury characteristics that distinguished mines with high rates of transport-related injuries from mines with lower transport injury rates. The results showed that most high-rate mines are small, high-rate mines have a disproportionate number of injuries involving young and less experienced workers, and injuries in high-rate mines are proportionally more severe. Further analyses of the MSHA injury data showed that smaller mines have a greater share of fatal and permanently disabling injuries, whereas larger mines have a greater share of injuries involving no lost time. Based on these results, we explored two explanations for the small mine injury risk: (1) a suggestion that differences in injury reporting between large and small mines may contribute to an apparent small mine injury risk, and (2) identification of factors contributing to a true difference in transport-related injury risk between small and large mines. Whereas it was true that most high injury rate mines were small, most small mines were actually zero-rate, having reported employment but no injuries to MSHA. An analysis employing binomial probability theory showed that a substantial proportion of small mines reported zero injuries when it was statistically probable that injuries would have occurred. This indicated that small mines may underreport injuries relative to larger mines. The possibility that reporting bias affected the associations found in this study was explored by eliminating the least severe injuries from the data set and evaluating changes in associations. This "adjustment" for reporting bias did not change previously observed relationships. Finally, MSHA injury data were analyzed in concert with mining population data collected by the Bureau of Mines. With such denominator information, the results indicated a disproportionately high risk of injury among workers in their first year at a mine and indicated that higher injury risk in small mines might be explained by the fact that workers at small mines have substantially less experience than workers at large mines. An effect of age was not found in these analyses. These results suggest the potential importance of targeted training programs for newly hired miners. Results also point to the need to explore specific factors contributing to the small mine injury risk, and to the necessity for complete and accurate reporting of injury data.

Accidents, Occupational↗

Systematic mining and quantification reveal the dominant contribution of non-HLA variations to acute graft-versus-host disease.

Human leukocyte antigen (HLA) disparity between donors and recipients is a key determinant triggering intense alloreactivity, leading to a lethal complication, namely, acute graft-versus-host disease (aGVHD), after allogeneic transplantation. Moreover, aGVHD remains a cause of mortality after HLA-matched allogeneic transplantation. Protocols for HLA-haploidentical hematopoietic cell transplantation (haploHCT) have been established successfully and widely applied, further highlighting the urgency of performing panoramic screening of non-HLA variations correlated with aGVHD. On the basis of our time-consecutive large haploHCT cohort (with a homogenous discovery set and an extended confirmatory set), we first delineated the genetic landscape of 1366 samples to quantitatively model aGVHD risk by assessing the contributions of HLA and non-HLA genes together with clinical factors. In addition to identifying multiple loss-of-function (LoF) risk variations in non-HLA coding genes, our data-driven study revealed that non-HLA genetic variations, independent of HLA disparity, contributed the most to the occurrence of aGVHD. This unexpected major effect was verified in an independent cohort that received HLA-identical sibling HCT. Subsequent functional experiments further revealed the roles of a representative non-HLA LoF gene and LoF gene pair in regulating the alloreactivity of primary human T cells. Our findings highlight the importance of non-HLA genetic risk in the new era of transplantation and propose a new direction to explore the immunogenetic mechanism of alloreactivity and to optimize donor selection strategies for allogeneic transplantation.

Humans↗

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software↗

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Acceptance of rules generated by machine learning among medical experts.

OBJECTIVES: The aim was to evaluate the potential for monotonicity constraints to bias machine learning systems to learn rules that were both accurate and meaningful. METHODS: Two data sets, taken from problems as diverse as screening for dementia and assessing the risk of mental retardation, were collected and a rule learning system, with and without monotonicity constraints, was run on each. The rules were shown to experts, who were asked how willing they would be to use such rules in practice. The accuracy of the rules was also evaluated. RESULTS: Rules learned with monotonicity constraints were at least as accurate as rules learned without such constraints. Experts were, on average, more willing to use the rules learned with the monotonicity constraints. CONCLUSIONS: The analysis of medical databases has the potential of improving patient outcomes and/or lowering the cost of health care delivery. Various techniques, from statistics, pattern recognition, machine learning, and neural networks, have been proposed to "mine" this data by uncovering patterns that may be used to guide decision making. This study suggests cognitive factors make learned models coherent and, therefore, credible to experts. One factor that influences the acceptance of learned models is consistency with existing medical knowledge.

Alzheimer Disease↗

Mining Stored-Specimen Studies for Information about Cancer Natural History.

The advent of new multicancer early detection tests and publication of early diagnostic results have generated expectations of clinical benefit from multicancer screening. The clinical benefit of a cancer screening test depends critically on disease natural history, which is typically learned from prospective screening studies. Retrospective studies of stored blood specimens are important in learning about a test's preclinical diagnostic performance but have rarely been used to infer natural history. The extent to which these studies might be harnessed to also learn natural history is discussed in the context of an article in this issue that infers the combined natural history of a range of cancers targeted by a multicancer early detection test using a case-control subsample of specimens from a large cohort study. The critical question concerns the identifiability of key transition rates in multistate models of natural history alongside state-specific sensitivities. The article suggests that these parameters are estimable within a Bayesian framework that leverages prior information about test sensitivity from diagnostic studies. We offer a heuristic discussion of identifiability in this setting and encourage formal study to determine the extent to which models with varying degrees of complexity may be learned from stored-specimen studies. See related article by Dai et al., p. 1535.

Humans↗

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus↗

An approach to the characterization of silica exposure in U.S. industry.

Quantitative evaluation of worker exposure to silica in nine Standard Industrial Classification (SIC) codes was conducted, using data derived from OSHA compliance inspections, in order to assess the silica exposure problem in the U.S. The nine SICs studied were those in which OSHA inspections were concentrated. They include: construction; chemical manufacture; stone, glass, and clay manufacturing; primary metal industries; metal fabrication; machinery; transportation; and miscellaneous manufacturing industries. High exposures to silica were documented in each industry, with the number of test samples over the permissible exposure limit ranging from 14% (aluminum foundries) to 73% (pottery). An estimation is made that 24,889 workers employed in ferrous and nonferrous foundries are at risk of silica-related pulmonary effects. The data developed in this analysis also indicate the need to investigate certain industries that had high exposures but few inspections. The limitations of the data base for estimating the scope of the silica problem, including lack of data on mining and milling, are discussed. We conclude that exposure to silica represents a continuing and significant problem in a number of U.S. industries.

Air Pollutants, Occupational↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗

Microbial genomes.

Microbial genome sequencing is driven by the need to understand and control pathogens and to exploit extremophiles and their enzymes in bioremediation and industry. It is hard for the traditional bacteriologist to grasp the scale and pace of the venture. Around two dozen microbial genomes have now been completed and, within a decade, genomes from every significant species of bacterial pathogen of humans, animals and plants will have been sequenced. Indeed, we will often have more than one sequence from a species or genus--for example, we already have sequences from two strains of Helicobacter pylori, from two strains of Mycobacterium tuberculosis and from three species of Pyrococcus. However, genome sequencing risks becoming expensive molecular stamp-collecting without the tools to mine the data and fuel hypothesis-driven laboratory-based research. Bioinformatics, twinned with the new experimental approaches forming functional genomics', provides some of the needed tools. Nonetheless, there will be an increasing need for us to explore the detailed implications of genomic findings. Microbial genome sequencing thus represents not a threat, but an exciting opportunity for molecular microbiologists.

Computational Biology↗

Large-scale analysis of the human and mouse transcriptomes.

High-throughput gene expression profiling has become an important tool for investigating transcriptional activity in a variety of biological samples. To date, the vast majority of these experiments have focused on specific biological processes and perturbations. Here, we have generated and analyzed gene expression from a set of samples spanning a broad range of biological conditions. Specifically, we profiled gene expression from 91 human and mouse samples across a diverse array of tissues, organs, and cell lines. Because these samples predominantly come from the normal physiological state in the human and mouse, this dataset represents a preliminary, but substantial, description of the normal mammalian transcriptome. We have used this dataset to illustrate methods of mining these data, and to reveal insights into molecular and physiological gene function, mechanisms of transcriptional regulation, disease etiology, and comparative genomics. Finally, to allow the scientific community to use this resource, we have built a free and publicly accessible website (http://expression.gnf.org) that integrates data visualization and curation of current gene annotations.

Animals↗

Lit-OTAR framework for extracting biological evidences from literature.

SUMMARY: The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION: The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.

Drug Discovery↗

AutoPM3: enhancing variant interpretation via LLM-driven PM3 evidence extraction from scientific literature.

MOTIVATION: Rare diseases affect over 300 million people worldwide and are often caused by genetic variants. While variant detection has become cost-effective, interpreting these variants-particularly collecting literature-based evidence like ACMG/AMP PM3-remains complex and time-consuming. RESULTS: We present AutoPM3, a method that automates PM3 evidence extraction from literatures using open-source large language models (LLMs). AutoPM3 combines a Text2SQL-based variant extractor and a retrieval-augmented generation (RAG) module, enhanced by a variant-specific retriever and fine-tuned LLM, to separately process tables and text. We curated PM3-Bench, a dataset of 1027 variant-publication evidence pairs from ClinGen. On openly accessible pairs, AutoPM3 achieved 86.1% accuracy for variant hits and 72.5% recall for in trans variants-outperforming other methods, including those using larger models. We uncovered the effectiveness of AutoPM3's key modules, especially for variant-specific retriever and Text2SQL, through the sequential ablation study. AutoPM3 located evidence in 76 s, demonstrating that open-source LLMs can offer an efficient, cost-effective solution for rare disease diagnosis. AVAILABILITY AND IMPLEMENTATION: AutoPM3 is implemented and freely available under the MIT license at https://github.com/HKU-BAL/AutoPM3.

Genetic Variation↗

Characterization of dust exposure for the study of chronic occupational lung disease: a comparison of different exposure assessment strategies.

Various exposure assessment strategies were compared in the study of the relation between dust exposure and 11-year lung function change in 1,172 miners with 36,824 concurrently measured personal dust samples available from the 1969-1981 US National Study of Coal Workers' Pneumoconiosis. A miner's average exposure was assessed by calculating average exposures based on dust samples taken from each individual and by using different job exposure matrices (JEMs) with different underlying exposure categorizations, based on occupational categories, job title, mine, and time, to obtain average exposure estimates. For each grouping procedure, intragroup and intergroup variances and the pooled standard error of the mean were calculated to assess relative efficiency. The results show that considerable variation in slopes of exposure-response relations was found using different exposure assessment strategies. Standard errors of the slopes of the exposure-response relations with exposure on an individual basis compared with JEMs. Exposure assessment on an individual basis was extremely sensitive to the number of exposure measurements per individual. The study demonstrates the advantages and disadvantages of different exposure assessment strategies and shows the need for explicit publication of exposure assessment strategies for epidemiologic studies. Careful assessment of the influence of misclassification error in the exposure assessment on exposure-response modeling is warranted.

Adult↗

Respiratory symptoms and functional status in workers exposed to silica, asbestos, and coal mine dusts.

This study aims to provide further understanding of physiologic and symptomatic changes and radiographic abnormalities due to exposure to silica, asbestos, and coal dusts. Questionnaires and pulmonary function tests were given to 220 silica, 277 asbestos, and 511 coal workers from three different industries in China. Posteroanterior chest radiographs were classified as stages 0, I, II, and III according to degree of parenchymal fibrosis. Significantly poorer pulmonary function and a higher prevalence of dyspnea and chronic cough were observed in workers with pneumoconiosis than those without, irrespective of dust type. Workers with stages II and III silicosis had worse pulmonary function and more common symptoms relative to workers with equivalent coal workers' pneumoconiosis or asbestosis. After adjusting for relevant confounders, reductions in the spirometric parameters and single breath diffusing capacity for carbon monoxide (DLCO) and the occurrence of respiratory symptoms were associated with increasing stage of silicosis, whereas lower DLCO and the occurrence of symptoms were associated with increasing stage of asbestosis and coal workers' pneumoconiosis. The study suggests that despite the differences in degree and pattern due to exposure to different fibrogenic dusts, respiratory impairments of all of the workers are associated with the presence and progression of parenchymal fibrosis and smoking.

Adult↗

Analysis of the molecular mechanism underlying di(2-ethylhexyl) phthalate-induced bladder carcinogenesis via network toxicology and molecular docking approaches: An observational study.

This study aims to investigate the toxicity of di(2-ethylhexyl) phthalate (DEHP) and the potential molecular mechanisms of DEHP-induced bladder cancer (BLCA) using network toxicology and molecular docking strategies. The toxicity of DEHP was assessed using Prox-II software, and potential targets for DEHP-induced BLCA were identified by integrating data from ChEMBL database, Search Tool for Interactions of Chemicals, SwissTargetPrediction, GeneCards, Therapeutic Target Database, Online Mendelian Inheritance in Man, and The Cancer Genome Atlas. STRING database and Cytoscape were employed to construct target networks and determine core targets. The expression levels of core targets were analyzed using R. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed on potential and core targets. Molecular docking was carried out using CB-Dock 2 to verify the interactions between DEHP and core targets. A total of 105 potential targets related to DEHP-induced BLCA were identified, from which 7 core targets were selected: cyclin-dependent kinase 1, interleukin 6, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, cyclin B2, and B-cell lymphoma 2. IL-6 and B-cell lymphoma 2 showed downregulated expression in tumor tissues, while cyclin-dependent kinase 1, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, and cyclin B2 were upregulated. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses indicated that these targets were enriched in cell signaling and cancer-related pathways. Molecular docking confirmed that DEHP interacts with these core targets. DEHP may promote the development of BLCA by interacting with key proteins and signaling pathways. This study provides a theoretical basis for understanding the molecular mechanisms of DEHP-induced BLCA and offers references for future prevention and treatment strategies.

Diethylhexyl Phthalate↗