Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Systematic mining and quantification reveal the dominant contribution of non-HLA variations to acute graft-versus-host disease.

Human leukocyte antigen (HLA) disparity between donors and recipients is a key determinant triggering intense alloreactivity, leading to a lethal complication, namely, acute graft-versus-host disease (aGVHD), after allogeneic transplantation. Moreover, aGVHD remains a cause of mortality after HLA-matched allogeneic transplantation. Protocols for HLA-haploidentical hematopoietic cell transplantation (haploHCT) have been established successfully and widely applied, further highlighting the urgency of performing panoramic screening of non-HLA variations correlated with aGVHD. On the basis of our time-consecutive large haploHCT cohort (with a homogenous discovery set and an extended confirmatory set), we first delineated the genetic landscape of 1366 samples to quantitatively model aGVHD risk by assessing the contributions of HLA and non-HLA genes together with clinical factors. In addition to identifying multiple loss-of-function (LoF) risk variations in non-HLA coding genes, our data-driven study revealed that non-HLA genetic variations, independent of HLA disparity, contributed the most to the occurrence of aGVHD. This unexpected major effect was verified in an independent cohort that received HLA-identical sibling HCT. Subsequent functional experiments further revealed the roles of a representative non-HLA LoF gene and LoF gene pair in regulating the alloreactivity of primary human T cells. Our findings highlight the importance of non-HLA genetic risk in the new era of transplantation and propose a new direction to explore the immunogenetic mechanism of alloreactivity and to optimize donor selection strategies for allogeneic transplantation.

Humans↗

CAGEcleaner: reducing genomic redundancy in gene cluster mining.

SUMMARY: Mining homologous biosynthetic gene clusters (BGCs) typically involves searching colocalised genes against large genomic databases. However, the high degree of genomic redundancy in these databases often propagates into the resulting hit sets, complicating downstream analyses and visualization. To address this challenge, we present CAGEcleaner, a Python-based pipeline with auxiliary bash scripts designed to reduce redundancy in gene cluster hit sets by dereplicating the genomes that host these hits. CAGEcleaner integrates seamlessly with widely used gene cluster mining tools, such as cblaster and CAGECAT, enabling efficient filtering and streamlining BGC discovery workflows. AVAILABILITY AND IMPLEMENTATION: Source code and documentation is hosted at GitHub (https://github.com/LucoDevro/CAGEcleaner) and Zenodo (https://doi.org/10.5281/zenodo.14726119) under an MIT license. For accessibility, CAGEcleaner is installable from Bioconda (https://anaconda.org/bioconda/cagecleaner) and PyPi (https://pypi.org/project/cagecleaner/), and is also available as a Docker image from DockerHub (https://hub.docker.com/r/lucodevro/cagecleaner).

Software↗

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Acceptance of rules generated by machine learning among medical experts.

OBJECTIVES: The aim was to evaluate the potential for monotonicity constraints to bias machine learning systems to learn rules that were both accurate and meaningful. METHODS: Two data sets, taken from problems as diverse as screening for dementia and assessing the risk of mental retardation, were collected and a rule learning system, with and without monotonicity constraints, was run on each. The rules were shown to experts, who were asked how willing they would be to use such rules in practice. The accuracy of the rules was also evaluated. RESULTS: Rules learned with monotonicity constraints were at least as accurate as rules learned without such constraints. Experts were, on average, more willing to use the rules learned with the monotonicity constraints. CONCLUSIONS: The analysis of medical databases has the potential of improving patient outcomes and/or lowering the cost of health care delivery. Various techniques, from statistics, pattern recognition, machine learning, and neural networks, have been proposed to "mine" this data by uncovering patterns that may be used to guide decision making. This study suggests cognitive factors make learned models coherent and, therefore, credible to experts. One factor that influences the acceptance of learned models is consistency with existing medical knowledge.

Alzheimer Disease↗

Mining Stored-Specimen Studies for Information about Cancer Natural History.

The advent of new multicancer early detection tests and publication of early diagnostic results have generated expectations of clinical benefit from multicancer screening. The clinical benefit of a cancer screening test depends critically on disease natural history, which is typically learned from prospective screening studies. Retrospective studies of stored blood specimens are important in learning about a test's preclinical diagnostic performance but have rarely been used to infer natural history. The extent to which these studies might be harnessed to also learn natural history is discussed in the context of an article in this issue that infers the combined natural history of a range of cancers targeted by a multicancer early detection test using a case-control subsample of specimens from a large cohort study. The critical question concerns the identifiability of key transition rates in multistate models of natural history alongside state-specific sensitivities. The article suggests that these parameters are estimable within a Bayesian framework that leverages prior information about test sensitivity from diagnostic studies. We offer a heuristic discussion of identifiability in this setting and encourage formal study to determine the extent to which models with varying degrees of complexity may be learned from stored-specimen studies. See related article by Dai et al., p. 1535.

Humans↗

Integration of genome mining and HiTES reveals secondary metabolic potential in marine-derived Aspergillus sp. WHUF0304.

AIMS: Marine-derived Aspergillus species are prolific producers of bioactive secondary metabolites, yet the majority of their biosynthetic gene clusters (BGCs) remain silent. This study aimed to integrate genome mining with high-throughput elicitor screening (HiTES) to unlock the metabolic potential of Aspergillus sp. WHUF0304 and identify elicitors that promote the accumulation of previously undetected metabolites. METHODS AND RESULTS: A high-quality genome of Aspergillus sp. WHUF0304 was assembled and annotated using multiple functional databases, revealing substantial secondary metabolic potential. antiSMASH analysis identified diverse BGCs, including NRPS/indole-related clusters potentially associated with indole diketopiperazine biosynthesis. A HiTES-inspired elicitor screening strategy was then applied to evaluate 42 small molecules for their ability to alter the metabolite profile of this strain. Among the tested elicitors, fluconazole was identified as the optimal inducer, triggering the production of several indole diketopiperazine-related differential metabolites. Subsequent activity-guided isolation led to the identification of a bioactive indole diketopiperazine dimer, cristatumin E, which exhibited antibacterial activity against Escherichia coli and Bacillus subtilis with minimum inhibitory concentrations (MICs) of 32 µg mL-1 and 256 µg mL-1, respectively. CONCLUSIONS: These findings demonstrate that integrating genomic and functional approaches effectively activates silent BGCs in marine fungi. The fluconazole-associated accumulation and subsequent isolation of cristatumin E, a bioactive indole diketopiperazine dimer, highlight the potential of elicitor-mediated activation to expand the detectable metabolite profile of Aspergillus sp. WHUF0304.

Aspergillus↗

An approach to the characterization of silica exposure in U.S. industry.

Quantitative evaluation of worker exposure to silica in nine Standard Industrial Classification (SIC) codes was conducted, using data derived from OSHA compliance inspections, in order to assess the silica exposure problem in the U.S. The nine SICs studied were those in which OSHA inspections were concentrated. They include: construction; chemical manufacture; stone, glass, and clay manufacturing; primary metal industries; metal fabrication; machinery; transportation; and miscellaneous manufacturing industries. High exposures to silica were documented in each industry, with the number of test samples over the permissible exposure limit ranging from 14% (aluminum foundries) to 73% (pottery). An estimation is made that 24,889 workers employed in ferrous and nonferrous foundries are at risk of silica-related pulmonary effects. The data developed in this analysis also indicate the need to investigate certain industries that had high exposures but few inspections. The limitations of the data base for estimating the scope of the silica problem, including lack of data on mining and milling, are discussed. We conclude that exposure to silica represents a continuing and significant problem in a number of U.S. industries.

Air Pollutants, Occupational↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗

Microbial genomes.

Microbial genome sequencing is driven by the need to understand and control pathogens and to exploit extremophiles and their enzymes in bioremediation and industry. It is hard for the traditional bacteriologist to grasp the scale and pace of the venture. Around two dozen microbial genomes have now been completed and, within a decade, genomes from every significant species of bacterial pathogen of humans, animals and plants will have been sequenced. Indeed, we will often have more than one sequence from a species or genus--for example, we already have sequences from two strains of Helicobacter pylori, from two strains of Mycobacterium tuberculosis and from three species of Pyrococcus. However, genome sequencing risks becoming expensive molecular stamp-collecting without the tools to mine the data and fuel hypothesis-driven laboratory-based research. Bioinformatics, twinned with the new experimental approaches forming functional genomics', provides some of the needed tools. Nonetheless, there will be an increasing need for us to explore the detailed implications of genomic findings. Microbial genome sequencing thus represents not a threat, but an exciting opportunity for molecular microbiologists.

Computational Biology↗

Lit-OTAR framework for extracting biological evidences from literature.

SUMMARY: The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION: The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.

Drug Discovery↗

AutoPM3: enhancing variant interpretation via LLM-driven PM3 evidence extraction from scientific literature.

MOTIVATION: Rare diseases affect over 300 million people worldwide and are often caused by genetic variants. While variant detection has become cost-effective, interpreting these variants-particularly collecting literature-based evidence like ACMG/AMP PM3-remains complex and time-consuming. RESULTS: We present AutoPM3, a method that automates PM3 evidence extraction from literatures using open-source large language models (LLMs). AutoPM3 combines a Text2SQL-based variant extractor and a retrieval-augmented generation (RAG) module, enhanced by a variant-specific retriever and fine-tuned LLM, to separately process tables and text. We curated PM3-Bench, a dataset of 1027 variant-publication evidence pairs from ClinGen. On openly accessible pairs, AutoPM3 achieved 86.1% accuracy for variant hits and 72.5% recall for in trans variants-outperforming other methods, including those using larger models. We uncovered the effectiveness of AutoPM3's key modules, especially for variant-specific retriever and Text2SQL, through the sequential ablation study. AutoPM3 located evidence in 76 s, demonstrating that open-source LLMs can offer an efficient, cost-effective solution for rare disease diagnosis. AVAILABILITY AND IMPLEMENTATION: AutoPM3 is implemented and freely available under the MIT license at https://github.com/HKU-BAL/AutoPM3.

Genetic Variation↗

Characterization of dust exposure for the study of chronic occupational lung disease: a comparison of different exposure assessment strategies.

Various exposure assessment strategies were compared in the study of the relation between dust exposure and 11-year lung function change in 1,172 miners with 36,824 concurrently measured personal dust samples available from the 1969-1981 US National Study of Coal Workers' Pneumoconiosis. A miner's average exposure was assessed by calculating average exposures based on dust samples taken from each individual and by using different job exposure matrices (JEMs) with different underlying exposure categorizations, based on occupational categories, job title, mine, and time, to obtain average exposure estimates. For each grouping procedure, intragroup and intergroup variances and the pooled standard error of the mean were calculated to assess relative efficiency. The results show that considerable variation in slopes of exposure-response relations was found using different exposure assessment strategies. Standard errors of the slopes of the exposure-response relations with exposure on an individual basis compared with JEMs. Exposure assessment on an individual basis was extremely sensitive to the number of exposure measurements per individual. The study demonstrates the advantages and disadvantages of different exposure assessment strategies and shows the need for explicit publication of exposure assessment strategies for epidemiologic studies. Careful assessment of the influence of misclassification error in the exposure assessment on exposure-response modeling is warranted.

Adult↗

Respiratory symptoms and functional status in workers exposed to silica, asbestos, and coal mine dusts.

This study aims to provide further understanding of physiologic and symptomatic changes and radiographic abnormalities due to exposure to silica, asbestos, and coal dusts. Questionnaires and pulmonary function tests were given to 220 silica, 277 asbestos, and 511 coal workers from three different industries in China. Posteroanterior chest radiographs were classified as stages 0, I, II, and III according to degree of parenchymal fibrosis. Significantly poorer pulmonary function and a higher prevalence of dyspnea and chronic cough were observed in workers with pneumoconiosis than those without, irrespective of dust type. Workers with stages II and III silicosis had worse pulmonary function and more common symptoms relative to workers with equivalent coal workers' pneumoconiosis or asbestosis. After adjusting for relevant confounders, reductions in the spirometric parameters and single breath diffusing capacity for carbon monoxide (DLCO) and the occurrence of respiratory symptoms were associated with increasing stage of silicosis, whereas lower DLCO and the occurrence of symptoms were associated with increasing stage of asbestosis and coal workers' pneumoconiosis. The study suggests that despite the differences in degree and pattern due to exposure to different fibrogenic dusts, respiratory impairments of all of the workers are associated with the presence and progression of parenchymal fibrosis and smoking.

Adult↗

Analysis of the molecular mechanism underlying di(2-ethylhexyl) phthalate-induced bladder carcinogenesis via network toxicology and molecular docking approaches: An observational study.

This study aims to investigate the toxicity of di(2-ethylhexyl) phthalate (DEHP) and the potential molecular mechanisms of DEHP-induced bladder cancer (BLCA) using network toxicology and molecular docking strategies. The toxicity of DEHP was assessed using Prox-II software, and potential targets for DEHP-induced BLCA were identified by integrating data from ChEMBL database, Search Tool for Interactions of Chemicals, SwissTargetPrediction, GeneCards, Therapeutic Target Database, Online Mendelian Inheritance in Man, and The Cancer Genome Atlas. STRING database and Cytoscape were employed to construct target networks and determine core targets. The expression levels of core targets were analyzed using R. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed on potential and core targets. Molecular docking was carried out using CB-Dock 2 to verify the interactions between DEHP and core targets. A total of 105 potential targets related to DEHP-induced BLCA were identified, from which 7 core targets were selected: cyclin-dependent kinase 1, interleukin 6, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, cyclin B2, and B-cell lymphoma 2. IL-6 and B-cell lymphoma 2 showed downregulated expression in tumor tissues, while cyclin-dependent kinase 1, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, and cyclin B2 were upregulated. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses indicated that these targets were enriched in cell signaling and cancer-related pathways. Molecular docking confirmed that DEHP interacts with these core targets. DEHP may promote the development of BLCA by interacting with key proteins and signaling pathways. This study provides a theoretical basis for understanding the molecular mechanisms of DEHP-induced BLCA and offers references for future prevention and treatment strategies.

Diethylhexyl Phthalate↗

Shared etiology of Mendelian and complex disease supports drug discovery.

BACKGROUND: Drugs targeting disease causal genes are more likely to succeed for that disease. However, complex disease causal genes are not always clear. In contrast, Mendelian disease causal genes are well-known and druggable. Here, we seek an approach to exploit the well characterized biology of Mendelian diseases for complex disease drug discovery, by exploiting evidence of pathogenic processes shared between monogenic and complex disease. One way to find shared disease etiology is clinical association: some Mendelian diseases are known to predispose patients to specific complex diseases (comorbidity). Previous studies link this comorbidity to pleiotropic effects of the Mendelian disease causal genes on the complex disease. METHODS: In previous work studying incidence of 90 Mendelian and 65 complex diseases, we found 2,908 pairs of clinically associated (comorbid) diseases. Using this clinical signal, we can match each complex disease to a set of Mendelian disease causal genes. We hypothesize that the drugs targeting these genes are potential candidate drugs for the complex disease. We evaluate our candidate drugs using information of current drug indications or investigations. RESULTS: Our analysis shows that the candidate drugs are enriched among currently investigated or indicated drugs for the relevant complex diseases (odds ratio = 1.84, p = 5.98e-22). Additionally, the candidate drugs are more likely to be in advanced stages of the drug development pipeline. We also present an approach to prioritize Mendelian diseases with particular promise for drug repurposing. Finally, we find that the combination of comorbidity and genetic similarity for a Mendelian disease and cancer pair leads to recommendation of candidate drugs that are enriched for those investigated or indicated. CONCLUSIONS: Our findings suggest a novel way to take advantage of the rich knowledge about Mendelian disease biology to improve treatment of complex diseases.

Humans↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗

Diverse RNA viruses discovered in multiple seagrass species.

Seagrasses are marine angiosperms that form highly productive and diverse ecosystems. These ecosystems, however, are declining worldwide. Plant-associated microbes affect critical functions like nutrient uptake and pathogen resistance, which has led to an interest in the seagrass microbiome. However, despite their significant role in plant ecology, viruses have only recently garnered attention in seagrass species. In this study, we produced original data and mined publicly available transcriptomes to advance our understanding of RNA viral diversity in Zostera marina, Zostera muelleri, Zostera japonica, and Cymodocea nodosa. In Z. marina, we present evidence for additional Zostera marina amalgavirus 1 and 2 genotypes, and a complete genome for an alphaendornavirus previously evidenced by an RNA-dependent RNA polymerase gene fragment. In Z. muelleri, we present evidence for a second complete alphaendornavirus and near complete furovirus. Both are novel, and, to the best of our knowledge, this marks the first report of a furovirus infection naturally occurring outside of cereal grasses. In Z. japonica, we discovered genome fragments that belong to a novel strain of cucumber mosaic virus, a prolific pathogen that depends largely on aphid vectoring for host-to-host transmission. Lastly, in C. nodosa, we discovered two contigs that belong to a novel virus in the family Betaflexiviridae. These findings expand our knowledge of viral diversity in seagrasses and provide insight into seagrass viral ecology.

RNA Viruses↗

Health care challenge in coal mines community.

The present paper depicts salient features of environment and living conditions with the comparison of various diseases prevalent among underground coal miners, surface workers, asbestos mine workers and general population of Jharia-Dhanbad coalfield as conducted by CMRS during the past few years. The investigations on coal miners' community comprise of different morbid conditions with respiratory (22%), Pneumoconiosis (11.6%), Skin (35%), Eye (29%), Intestinal parasitic infestation (44.6%), Anaemia (42%), Immunostatus (V.D.R.L. Positive-19.9%), Status of injuries and Blood pressure, Water-borne diseases, housing facilities and excreta disposal. The paper also includes the analysis of disease pattern obtained from hospital records of two coal mines which depicts 19.1%, 24.7% and 16% members of coal miners' families suffering from disorder with respiratory, gastro-intestinal and fever respectively. With speedy industrialization of the country, the mining of coal resource comes first in the chain of socio-economic development. The speedy human industrial activities are based on 80% steam, metallurgical and thermal electrical energy which hinges on coal wings. The coal has also gradually occupied all the phases of social life, our clothes, books, newspapers, cooking gas, chemical paints, dye stuff, oil phenyl, Benzene, Naphthalene, Coal tar, scents and various types of unaccountable products come out from coal derivatives and pushed to serve in the today's market for our daily exigencies. Every day one finds a new coal based industry is coming up in the area. The coal is utilized in two hundred ways in our various walks of social life.(ABSTRACT TRUNCATED AT 250 WORDS)

Air Pollution↗