Search PubMedSearch

SEARCH · Search PubMed

Results for “data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

AutoPM3: enhancing variant interpretation via LLM-driven PM3 evidence extraction from scientific literature.

MOTIVATION: Rare diseases affect over 300 million people worldwide and are often caused by genetic variants. While variant detection has become cost-effective, interpreting these variants-particularly collecting literature-based evidence like ACMG/AMP PM3-remains complex and time-consuming. RESULTS: We present AutoPM3, a method that automates PM3 evidence extraction from literatures using open-source large language models (LLMs). AutoPM3 combines a Text2SQL-based variant extractor and a retrieval-augmented generation (RAG) module, enhanced by a variant-specific retriever and fine-tuned LLM, to separately process tables and text. We curated PM3-Bench, a dataset of 1027 variant-publication evidence pairs from ClinGen. On openly accessible pairs, AutoPM3 achieved 86.1% accuracy for variant hits and 72.5% recall for in trans variants-outperforming other methods, including those using larger models. We uncovered the effectiveness of AutoPM3's key modules, especially for variant-specific retriever and Text2SQL, through the sequential ablation study. AutoPM3 located evidence in 76 s, demonstrating that open-source LLMs can offer an efficient, cost-effective solution for rare disease diagnosis. AVAILABILITY AND IMPLEMENTATION: AutoPM3 is implemented and freely available under the MIT license at https://github.com/HKU-BAL/AutoPM3.

Genetic Variation

Analysis of the molecular mechanism underlying di(2-ethylhexyl) phthalate-induced bladder carcinogenesis via network toxicology and molecular docking approaches: An observational study.

This study aims to investigate the toxicity of di(2-ethylhexyl) phthalate (DEHP) and the potential molecular mechanisms of DEHP-induced bladder cancer (BLCA) using network toxicology and molecular docking strategies. The toxicity of DEHP was assessed using Prox-II software, and potential targets for DEHP-induced BLCA were identified by integrating data from ChEMBL database, Search Tool for Interactions of Chemicals, SwissTargetPrediction, GeneCards, Therapeutic Target Database, Online Mendelian Inheritance in Man, and The Cancer Genome Atlas. STRING database and Cytoscape were employed to construct target networks and determine core targets. The expression levels of core targets were analyzed using R. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed on potential and core targets. Molecular docking was carried out using CB-Dock 2 to verify the interactions between DEHP and core targets. A total of 105 potential targets related to DEHP-induced BLCA were identified, from which 7 core targets were selected: cyclin-dependent kinase 1, interleukin 6, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, cyclin B2, and B-cell lymphoma 2. IL-6 and B-cell lymphoma 2 showed downregulated expression in tumor tissues, while cyclin-dependent kinase 1, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, and cyclin B2 were upregulated. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses indicated that these targets were enriched in cell signaling and cancer-related pathways. Molecular docking confirmed that DEHP interacts with these core targets. DEHP may promote the development of BLCA by interacting with key proteins and signaling pathways. This study provides a theoretical basis for understanding the molecular mechanisms of DEHP-induced BLCA and offers references for future prevention and treatment strategies.

Diethylhexyl Phthalate

Shared etiology of Mendelian and complex disease supports drug discovery.

BACKGROUND: Drugs targeting disease causal genes are more likely to succeed for that disease. However, complex disease causal genes are not always clear. In contrast, Mendelian disease causal genes are well-known and druggable. Here, we seek an approach to exploit the well characterized biology of Mendelian diseases for complex disease drug discovery, by exploiting evidence of pathogenic processes shared between monogenic and complex disease. One way to find shared disease etiology is clinical association: some Mendelian diseases are known to predispose patients to specific complex diseases (comorbidity). Previous studies link this comorbidity to pleiotropic effects of the Mendelian disease causal genes on the complex disease. METHODS: In previous work studying incidence of 90 Mendelian and 65 complex diseases, we found 2,908 pairs of clinically associated (comorbid) diseases. Using this clinical signal, we can match each complex disease to a set of Mendelian disease causal genes. We hypothesize that the drugs targeting these genes are potential candidate drugs for the complex disease. We evaluate our candidate drugs using information of current drug indications or investigations. RESULTS: Our analysis shows that the candidate drugs are enriched among currently investigated or indicated drugs for the relevant complex diseases (odds ratio = 1.84, p = 5.98e-22). Additionally, the candidate drugs are more likely to be in advanced stages of the drug development pipeline. We also present an approach to prioritize Mendelian diseases with particular promise for drug repurposing. Finally, we find that the combination of comorbidity and genetic similarity for a Mendelian disease and cancer pair leads to recommendation of candidate drugs that are enriched for those investigated or indicated. CONCLUSIONS: Our findings suggest a novel way to take advantage of the rich knowledge about Mendelian disease biology to improve treatment of complex diseases.

Humans

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models

Diverse RNA viruses discovered in multiple seagrass species.

Seagrasses are marine angiosperms that form highly productive and diverse ecosystems. These ecosystems, however, are declining worldwide. Plant-associated microbes affect critical functions like nutrient uptake and pathogen resistance, which has led to an interest in the seagrass microbiome. However, despite their significant role in plant ecology, viruses have only recently garnered attention in seagrass species. In this study, we produced original data and mined publicly available transcriptomes to advance our understanding of RNA viral diversity in Zostera marina, Zostera muelleri, Zostera japonica, and Cymodocea nodosa. In Z. marina, we present evidence for additional Zostera marina amalgavirus 1 and 2 genotypes, and a complete genome for an alphaendornavirus previously evidenced by an RNA-dependent RNA polymerase gene fragment. In Z. muelleri, we present evidence for a second complete alphaendornavirus and near complete furovirus. Both are novel, and, to the best of our knowledge, this marks the first report of a furovirus infection naturally occurring outside of cereal grasses. In Z. japonica, we discovered genome fragments that belong to a novel strain of cucumber mosaic virus, a prolific pathogen that depends largely on aphid vectoring for host-to-host transmission. Lastly, in C. nodosa, we discovered two contigs that belong to a novel virus in the family Betaflexiviridae. These findings expand our knowledge of viral diversity in seagrasses and provide insight into seagrass viral ecology.

RNA Viruses

Pathological findings in mine workers: II. Quality of the PATHAUT data.

To assess the feasibility of using the pathology automation system (PATHAUT) for research, the quality of the data was explored by examining who comes to autopsy, the quality of the autopsy material, interobserver variability, and repeatability of diagnoses. The data indicated that autopsy rates in the gold mining industry, especially for whites, are high and that even among blacks, gold miners are represented in proportions exceeding the relative size of the working population. Because of the perception of the autopsy service as a means of obtaining compensation, miners with occupational diseases fully compensated in life are probably underrepresented. The autopsy material submitted for full autopsy is generally better preserved than cardiorespiratory organs that are sent for examination. The gold mining industry has a high proportion of full autopsies as does the Iron and Steel Corporation of South Africa. Full autopsies are more commonly performed on older deceased miners. This was true for both blacks and whites. The allocation of material to pathologists for full autopsies and examinations of the cardiorespiratory organs were clearly not random, and this may affect comparisons among pathologists. Active tuberculosis, silicosis, and emphysema prevalences appeared fairly comparable across pathologists; however, there was wide variability in the prevalence of bronchiolitis as determined by the pathologists. Agreement between the diagnoses on PATHAUT and reclassifications by single pathologists was very good for the severity of emphysema and the histological type of bronchogenic carcinoma.

Autopsy

Underground mining, smoking, and lung cancer: a case-control study in the iron ore municipalities in northern Sweden.

A case-control study of lung cancer in males was performed in two municipalities in northern Sweden with large iron ore mines. Previous studies had revealed an increased lung cancer risk for underground workers in these mines, with all probability related to radon daughter exposure. Data concerning underground mining and smoking were obtained from questionnaires. All analyses suggested an interaction of a multiplicative type between underground mining and smoking in the causation of lung cancer in this population. The calculated population etiologic fraction was about 45% for underground mining and about 80% for smoking.

Aged

[Physiological and hygienic evaluation of controllers' work in coal mines].

The article contains data on the results of the on-the-spot studies of coal mine controllers' labour conditions and relating hygienic factors. It was established that the work at the dispatcher's control point in coal mines was considerably affected by unfavourable factors (low degree of illumination, constructional shortcomings of the control point and seat) with concomitant neuropsychic stress conditions. Revealed were specific functional shifts in CVS and CNS, neuromuscular disorders and the analyzers' malfunctioning. A set of measures was proposed for labour conditions improvement, ergometric perfection of the working place and reduction of neuroemotional tension.

Administrative Personnel

[Occupational morbidity of miners engaged in the capital building of iron ore mines].

The article contains data on the 1961-1985 occupational morbidity rates in the workers engaged in the Krivbass iron-ore mines, and narrates on the morbidity rates, structure and dynamics for different professions, age-groups and duration of work. An analysis is given of the concomitant somatic diseases, along with proposals for investigating ways and means of improving technologies in non-blasting iron-ore mines.

Adult

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus

Assessment of the respirable dust levels in the nation's underground and surface coal mining operations.

This report presents a chronological overview of the status of respirable dust exposures in underground and surface bituminous coal mines since inception of the 1969 Federal Coal Mine Health and Safety Act. Data for various intervals from 1970 through 1977, are presented for selected mining operations. Comparisons are made using data available from the mine operators' sampling program and from MSHA surveys. The data demonstrate the marked decrease that has occurred in respirable dust exposures since inception of the 1969 Act.

Coal

[The struggle against dust in the Belgian coal mines. Situation at the beginning of 1976].

The present communication gives a general view of the methods of dust control in the Belgian coal mines at the beginning of 1976. The statistical data received from the mines are presented in tabular form. The length and the output of coal faces treated by the classical methods of pre-spraying of the wall, wet cutting, water infusion and wet pneumatic picks, are given separately; in some cases two or more of these technics are used together on the same coal face. The number of stone drivages in which different methods of dust control are used, is also given.

Air Pollutants

[Dust prevention in Belgian coal mines. Situation at the beginning of 1974].

The present communication gives a general view of the methods of dust control in the Belgian coal mines at the beginning of 1974. The statistical data received from the mines are presented in tabular form. The length and the output of coal faces treated by the classical methods of pre-spraying of the wall, wet cutting, water infusion and wet pneumatic picks, are given separately; in some cases two or more of these techniques are used together on the same coal face. The number of stone drivages in which different methods of dust control are used, are also given.

Belgium

[Dust control in Belgian coal mines. Conditions at the beginning of 1979].

The present communication gives a general view of the methods of dust control in the Belgian coal mines at the beginning of 1979. The statistical data received from the mines are presented in tabular form. The length and the output of coal faces treated by the classical methods of pre-spraying of the wall, wet cutting and water infusion are given separately; in some cases two or more of these technics are used together on the same coal face. The number of stone drivages in which different methods of dust control are used, is also given.

Belgium

[Dust control in Belgian coal mines. Status at the beginning of the year 1977].

The present communication gives a general view of the methods of dust control in the Belgian coal mines at the beginning of 1977. The statistical data received from the mines are presented in tabular form. The length and the output of coal faces treated by the classical methods of pre-spraying of the wall, wet cutting, water infusion and wet pneumatic picks, are given separately; in some cases two or more of these technics are used together on the same coal face. The number of hard headings in which different methods of dust suppression are used, is also given.

Belgium

[Dust control in Belgian coal mines. Status at the beginning of 1978].

The present communication gives a general view of the methods of dust control in the Belgian coal mines at the beginning of 1978. The statistical data received from the mines are presented in tabular form. The length and the output of coal faces treated by the classical methods of pre-spraying of the wall, wet cutting, water infusion and wet pneumatic picks, are given separately; in some cases two or more of these technics are used together on the same coal face. The number of stone drivages in which different methods of dust control are used, is also given.

Coal Mining

Structuring strain data for storage and retrieval of information on fungi and yeasts in MINE, the Microbial Information Network Europe.

A distributed Microbial Information Network Europe (MINE) is being constructed by a number of major microbial culture collections in countries of the European Community, with the support of the Biotechnology Action Programme (BAP) of the Commission of the European Community. The representatives of the collections participating in MINE have agreed to adopt a general format for the computer storage and retrieval of strain data. This uniform format will facilitate the electronic combination and exchange of data from different collections in order to produce integrated catalogues and the use of identical commands to search the different databases. It is recommended to other collections who may wish to contribute data to the MINE network or between themselves. Three kinds of records can be linked to the leading 'species records': strain records, synonym records, and alternative morphonym records. A minimum data set of 30 fields (similar to the fields used for producing catalogues) is defined that facilitates the exchange of data between the national nodes and serves as a directory to strains available at other nodes. It is suggested that the full strain record comprise 99 fields, grouped in 12 blocks: internal administration--name--strain administration--status--environment and history--biological interactions--sexuality--properties (cytology, biomolecular data)--genotype and genetics--growth conditions--chemistry and enzymes--practical applications. Several fields are divided into subfields of different ranks. Delimiters are used either to separate a range of entries that have to be indexed or to divide an entry from the reference to its source or remarks that should not be indexed. The contents and structure of the fields proposed for filamentous fungi and yeasts are described and in some cases illustrated by examples. Uniformity of input is essential for indexed fields and desirable for non-indexed fields. Seven thesaurus files are envisaged to ensure consistency.

Data Collection

Scale model studies of the transport of airborne pollutants on coalfaces.

Studies have been carried out to investigate airflows in coalmine models, with special regard to the transport of airborne pollutants, and to examine how they relate to what happens at full-scale in an actual underground mine. If such models can be shown to provide data representative of actual mine ventilation engineering, then they can provide cost-effective alternatives to full-scale investigations. The work set out in the first instance to identify the properties of: (a) the bulk airflow and associated transport of airborne pollutants along a longwall coalface; and (b) the transport of material out of regions that were partially enclosed or poorly ventilated (e.g. in the cutting zone, in headings). For the former, an appropriate quantity is the dispersivity of the coalface airflow, for the latter the mean retention time. Both quantities may be rendered non-dimensional with respect to dimensions characteristic of the system and to velocities of the airflow. Their behaviour in relation to a third dimensionless quantity, the flow Reynolds' number is also important. Experiments were performed, using smoke or dust tracers, to investigate how these properties are interrelated and how they scale between small-scale and full-scale systems. They were carried out in a 1/10-scale laboratory model, in a full-scale surface model, and underground in an actual coalmine. The basis of most of the experiments was the 'tracer decay' method, in which the transport properties of the aerodynamic system under investigation were determined from observations of the changes in tracer concentration with time immediately following the removal of the tracer source. During these experiments, the feasibility of using small-scale models to investigate ventilation problems was clearly indicated and preliminary scaling relationships which may be used as an initial basis for predicting the transport and local build-up of pollutants in mines were developed. It is expected that applications of the ideas and methodology described will be relevant to other industries.

Air Movements