Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Data mining”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Lit-OTAR framework for extracting biological evidences from literature.

SUMMARY: The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION: The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.

Drug Discovery↗

AutoPM3: enhancing variant interpretation via LLM-driven PM3 evidence extraction from scientific literature.

MOTIVATION: Rare diseases affect over 300 million people worldwide and are often caused by genetic variants. While variant detection has become cost-effective, interpreting these variants-particularly collecting literature-based evidence like ACMG/AMP PM3-remains complex and time-consuming. RESULTS: We present AutoPM3, a method that automates PM3 evidence extraction from literatures using open-source large language models (LLMs). AutoPM3 combines a Text2SQL-based variant extractor and a retrieval-augmented generation (RAG) module, enhanced by a variant-specific retriever and fine-tuned LLM, to separately process tables and text. We curated PM3-Bench, a dataset of 1027 variant-publication evidence pairs from ClinGen. On openly accessible pairs, AutoPM3 achieved 86.1% accuracy for variant hits and 72.5% recall for in trans variants-outperforming other methods, including those using larger models. We uncovered the effectiveness of AutoPM3's key modules, especially for variant-specific retriever and Text2SQL, through the sequential ablation study. AutoPM3 located evidence in 76 s, demonstrating that open-source LLMs can offer an efficient, cost-effective solution for rare disease diagnosis. AVAILABILITY AND IMPLEMENTATION: AutoPM3 is implemented and freely available under the MIT license at https://github.com/HKU-BAL/AutoPM3.

Genetic Variation↗

Computational inference and experimental validation of the nitrogen assimilation regulatory network in cyanobacterium Synechococcus sp. WH 8102.

Deciphering the regulatory networks encoded in the genome of an organism represents one of the most interesting and challenging tasks in the post-genome sequencing era. As an example of this problem, we have predicted a detailed model for the nitrogen assimilation network in cyanobacterium Synechococcus sp. WH 8102 (WH8102) using a computational protocol based on comparative genomics analysis and mining experimental data from related organisms that are relatively well studied. This computational model is in excellent agreement with the microarray gene expression data collected under ammonium-rich versus nitrate-rich growth conditions, suggesting that our computational protocol is capable of predicting biological pathways/networks with high accuracy. We then refined the computational model using the microarray data, and proposed a new model for the nitrogen assimilation network in WH8102. An intriguing discovery from this study is that nitrogen assimilation affects the expression of many genes involved in photosynthesis, suggesting a tight coordination between nitrogen assimilation and photosynthesis processes. Moreover, for some of these genes, this coordination is probably mediated by NtcA through the canonical NtcA promoters in their regulatory regions.

Bacterial Proteins↗

Characterization of dust exposure for the study of chronic occupational lung disease: a comparison of different exposure assessment strategies.

Various exposure assessment strategies were compared in the study of the relation between dust exposure and 11-year lung function change in 1,172 miners with 36,824 concurrently measured personal dust samples available from the 1969-1981 US National Study of Coal Workers' Pneumoconiosis. A miner's average exposure was assessed by calculating average exposures based on dust samples taken from each individual and by using different job exposure matrices (JEMs) with different underlying exposure categorizations, based on occupational categories, job title, mine, and time, to obtain average exposure estimates. For each grouping procedure, intragroup and intergroup variances and the pooled standard error of the mean were calculated to assess relative efficiency. The results show that considerable variation in slopes of exposure-response relations was found using different exposure assessment strategies. Standard errors of the slopes of the exposure-response relations with exposure on an individual basis compared with JEMs. Exposure assessment on an individual basis was extremely sensitive to the number of exposure measurements per individual. The study demonstrates the advantages and disadvantages of different exposure assessment strategies and shows the need for explicit publication of exposure assessment strategies for epidemiologic studies. Careful assessment of the influence of misclassification error in the exposure assessment on exposure-response modeling is warranted.

Adult↗

Respiratory symptoms and functional status in workers exposed to silica, asbestos, and coal mine dusts.

This study aims to provide further understanding of physiologic and symptomatic changes and radiographic abnormalities due to exposure to silica, asbestos, and coal dusts. Questionnaires and pulmonary function tests were given to 220 silica, 277 asbestos, and 511 coal workers from three different industries in China. Posteroanterior chest radiographs were classified as stages 0, I, II, and III according to degree of parenchymal fibrosis. Significantly poorer pulmonary function and a higher prevalence of dyspnea and chronic cough were observed in workers with pneumoconiosis than those without, irrespective of dust type. Workers with stages II and III silicosis had worse pulmonary function and more common symptoms relative to workers with equivalent coal workers' pneumoconiosis or asbestosis. After adjusting for relevant confounders, reductions in the spirometric parameters and single breath diffusing capacity for carbon monoxide (DLCO) and the occurrence of respiratory symptoms were associated with increasing stage of silicosis, whereas lower DLCO and the occurrence of symptoms were associated with increasing stage of asbestosis and coal workers' pneumoconiosis. The study suggests that despite the differences in degree and pattern due to exposure to different fibrogenic dusts, respiratory impairments of all of the workers are associated with the presence and progression of parenchymal fibrosis and smoking.

Adult↗

Analysis of the molecular mechanism underlying di(2-ethylhexyl) phthalate-induced bladder carcinogenesis via network toxicology and molecular docking approaches: An observational study.

This study aims to investigate the toxicity of di(2-ethylhexyl) phthalate (DEHP) and the potential molecular mechanisms of DEHP-induced bladder cancer (BLCA) using network toxicology and molecular docking strategies. The toxicity of DEHP was assessed using Prox-II software, and potential targets for DEHP-induced BLCA were identified by integrating data from ChEMBL database, Search Tool for Interactions of Chemicals, SwissTargetPrediction, GeneCards, Therapeutic Target Database, Online Mendelian Inheritance in Man, and The Cancer Genome Atlas. STRING database and Cytoscape were employed to construct target networks and determine core targets. The expression levels of core targets were analyzed using R. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed on potential and core targets. Molecular docking was carried out using CB-Dock 2 to verify the interactions between DEHP and core targets. A total of 105 potential targets related to DEHP-induced BLCA were identified, from which 7 core targets were selected: cyclin-dependent kinase 1, interleukin 6, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, cyclin B2, and B-cell lymphoma 2. IL-6 and B-cell lymphoma 2 showed downregulated expression in tumor tissues, while cyclin-dependent kinase 1, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, and cyclin B2 were upregulated. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses indicated that these targets were enriched in cell signaling and cancer-related pathways. Molecular docking confirmed that DEHP interacts with these core targets. DEHP may promote the development of BLCA by interacting with key proteins and signaling pathways. This study provides a theoretical basis for understanding the molecular mechanisms of DEHP-induced BLCA and offers references for future prevention and treatment strategies.

Diethylhexyl Phthalate↗

Genome cluster database. A sequence family analysis platform for Arabidopsis and rice.

The genome-wide protein sequences from Arabidopsis (Arabidopsis thaliana) and rice (Oryza sativa) spp. japonica were clustered into families using sequence similarity and domain-based clustering. The two fundamentally different methods resulted in separate cluster sets with complementary properties to compensate the limitations for accurate family analysis. Functional names for the identified families were assigned with an efficient computational approach that uses the description of the most common molecular function gene ontology node within each cluster. Subsequently, multiple alignments and phylogenetic trees were calculated for the assembled families. All clustering results and their underlying sequences were organized in the Web-accessible Genome Cluster Database (http://bioinfo.ucr.edu/projects/GCD) with rich interactive and user-friendly sequence family mining tools to facilitate the analysis of any given family of interest for the plant science community. An automated clustering pipeline ensures current information for future updates in the annotations of the two genomes and clustering improvements. The analysis allowed the first systematic identification of family and singlet proteins present in both organisms as well as those restricted to one of them. In addition, the established Web resources for mining these data provide a road map for future studies of the composition and structure of protein families between the two species.

Algorithms↗

Three-dimensional quantitative structure-activity relationship analysis of human CYP51 inhibitors.

CYP51 fulfills an essential requirement for all cells, by catalyzing three sequential mono-oxidations within the cholesterol biosynthesis cascade. Inhibition of fungal CYP51 is used as a therapy for treating fungal infections, whereas inhibition of human CYP51 has been considered as a pharmacological approach to treat dyslipidemia and some forms of cancer. To predict the interaction of inhibitors with the active site of human CYP51, a three-dimensional quantitative structure-activity relationship model was constructed. This pharmacophore model of the common structural features of CYP51 inhibitors was built using the program Catalyst from multiple inhibitors (n = 26) of recombinant human CYP51-mediated lanosterol 14alpha-demethylation. The pharmacophore, which consisted of one hydrophobe, one hydrogen bond acceptor, and two ring aromatic features, demonstrated a high correlation between observed and predicted IC(50) values (r = 0.92). Validation of this pharmacophore was performed by predicting the IC(50) of a test set of commercially available (n = 19) and CP-320626-related (n = 48) CYP51 inhibitors. Using predictions below 10 microM as a cutoff indicative of active inhibitors, 16 of 19 commercially available inhibitors (84%) and 38 of 48 CP-320626-related inhibitors (79.2%) were predicted correctly. To better understand how inhibitors fit into the enzyme, potent CYP51 inhibitors were used to build a Cerius(2) receptor surface model representing the volume of the active site. This study has demonstrated the potential for ligand-based computational pharmacophore modeling of human CYP51 and enables a high-throughput screening system for drug discovery and data base mining.

Amides↗

The Connectivity Map: using gene-expression signatures to connect small molecules, genes, and disease.

To pursue a systematic approach to the discovery of functional connections among diseases, genetic perturbation, and drug action, we have created the first installment of a reference collection of gene-expression profiles from cultured human cells treated with bioactive small molecules, together with pattern-matching software to mine these data. We demonstrate that this "Connectivity Map" resource can be used to find connections among small molecules sharing a mechanism of action, chemicals and physiological processes, and diseases and drugs. These results indicate the feasibility of the approach and suggest the value of a large-scale community Connectivity Map project.

Alzheimer Disease↗

Generation, annotation, analysis and database integration of 16,500 white spruce EST clusters.

BACKGROUND: The sequencing and analysis of ESTs is for now the only practical approach for large-scale gene discovery and annotation in conifers because their very large genomes are unlikely to be sequenced in the near future. Our objective was to produce extensive collections of ESTs and cDNA clones to support manufacture of cDNA microarrays and gene discovery in white spruce (Picea glauca [Moench] Voss). RESULTS: We produced 16 cDNA libraries from different tissues and a variety of treatments, and partially sequenced 50,000 cDNA clones. High quality 3' and 5' reads were assembled into 16,578 consensus sequences, 45% of which represented full length inserts. Consensus sequences derived from 5' and 3' reads of the same cDNA clone were linked to define 14,471 transcripts. A large proportion (84%) of the spruce sequences matched a pine sequence, but only 68% of the spruce transcripts had homologs in Arabidopsis or rice. Nearly all the sequences that matched the Populus trichocarpa genome (the only sequenced tree genome) also matched rice or Arabidopsis genomes. We used several sequence similarity search approaches for assignment of putative functions, including blast searches against general and specialized databases (transcription factors, cell wall related proteins), Gene Ontology term assignation and Hidden Markov Model searches against PFAM protein families and domains. In total, 70% of the spruce transcripts displayed matches to proteins of known or unknown function in the Uniref100 database (blastx e-value < 1e-10). We identified multigenic families that appeared larger in spruce than in the Arabidopsis or rice genomes. Detailed analysis of translationally controlled tumour proteins and S-adenosylmethionine synthetase families confirmed a twofold size difference. Sequences and annotations were organized in a dedicated database, SpruceDB. Several search tools were developed to mine the data either based on their occurrence in the cDNA libraries or on functional annotations. CONCLUSION: This report illustrates specific approaches for large-scale gene discovery and annotation in an organism that is very distantly related to any of the fully sequenced genomes. The ArboreaSet sequences and cDNA clones represent a valuable resource for investigations ranging from plant comparative genomics to applied conifer genetics.

Arabidopsis↗

Shared etiology of Mendelian and complex disease supports drug discovery.

BACKGROUND: Drugs targeting disease causal genes are more likely to succeed for that disease. However, complex disease causal genes are not always clear. In contrast, Mendelian disease causal genes are well-known and druggable. Here, we seek an approach to exploit the well characterized biology of Mendelian diseases for complex disease drug discovery, by exploiting evidence of pathogenic processes shared between monogenic and complex disease. One way to find shared disease etiology is clinical association: some Mendelian diseases are known to predispose patients to specific complex diseases (comorbidity). Previous studies link this comorbidity to pleiotropic effects of the Mendelian disease causal genes on the complex disease. METHODS: In previous work studying incidence of 90 Mendelian and 65 complex diseases, we found 2,908 pairs of clinically associated (comorbid) diseases. Using this clinical signal, we can match each complex disease to a set of Mendelian disease causal genes. We hypothesize that the drugs targeting these genes are potential candidate drugs for the complex disease. We evaluate our candidate drugs using information of current drug indications or investigations. RESULTS: Our analysis shows that the candidate drugs are enriched among currently investigated or indicated drugs for the relevant complex diseases (odds ratio&#x2009;=&#x2009;1.84, p&#x2009;=&#x2009;5.98e-22). Additionally, the candidate drugs are more likely to be in advanced stages of the drug development pipeline. We also present an approach to prioritize Mendelian diseases with particular promise for drug repurposing. Finally, we find that the combination of comorbidity and genetic similarity for a Mendelian disease and cancer pair leads to recommendation of candidate drugs that are enriched for those investigated or indicated. CONCLUSIONS: Our findings suggest a novel way to take advantage of the rich knowledge about Mendelian disease biology to improve treatment of complex diseases.

Humans↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗

Diverse RNA viruses discovered in multiple seagrass species.

Seagrasses are marine angiosperms that form highly productive and diverse ecosystems. These ecosystems, however, are declining worldwide. Plant-associated microbes affect critical functions like nutrient uptake and pathogen resistance, which has led to an interest in the seagrass microbiome. However, despite their significant role in plant ecology, viruses have only recently garnered attention in seagrass species. In this study, we produced original data and mined publicly available transcriptomes to advance our understanding of RNA viral diversity in Zostera marina, Zostera muelleri, Zostera japonica, and Cymodocea nodosa. In Z. marina, we present evidence for additional Zostera marina amalgavirus 1 and 2 genotypes, and a complete genome for an alphaendornavirus previously evidenced by an RNA-dependent RNA polymerase gene fragment. In Z. muelleri, we present evidence for a second complete alphaendornavirus and near complete furovirus. Both are novel, and, to the best of our knowledge, this marks the first report of a furovirus infection naturally occurring outside of cereal grasses. In Z. japonica, we discovered genome fragments that belong to a novel strain of cucumber mosaic virus, a prolific pathogen that depends largely on aphid vectoring for host-to-host transmission. Lastly, in C. nodosa, we discovered two contigs that belong to a novel virus in the family Betaflexiviridae. These findings expand our knowledge of viral diversity in seagrasses and provide insight into seagrass viral ecology.

RNA Viruses↗

Introns regulate RNA and protein abundance in yeast.

The purpose of introns in the architecturally simple genome of Saccharomyces cerevisiae is not well understood. To assay the functional relevance of introns, a series of computational analyses and several detailed deletion studies were completed on the intronic genes of S. cerevisiae. Mining existing data from genomewide studies on yeast revealed that intron-containing genes produce more RNA and more protein and are more likely to be haplo-insufficient than nonintronic genes. These observations for all intronic genes held true for distinct subsets of genes including ribosomal, nonribosomal, duplicated, and nonduplicated. Corroborating the result of computational analyses, deletion of introns from three essential genes decreased cellular RNA levels and caused measurable growth defects. These data provide evidence that introns improve transcriptional and translational yield and are required for competitive growth of yeast.

Actins↗

Health care challenge in coal mines community.

The present paper depicts salient features of environment and living conditions with the comparison of various diseases prevalent among underground coal miners, surface workers, asbestos mine workers and general population of Jharia-Dhanbad coalfield as conducted by CMRS during the past few years. The investigations on coal miners' community comprise of different morbid conditions with respiratory (22%), Pneumoconiosis (11.6%), Skin (35%), Eye (29%), Intestinal parasitic infestation (44.6%), Anaemia (42%), Immunostatus (V.D.R.L. Positive-19.9%), Status of injuries and Blood pressure, Water-borne diseases, housing facilities and excreta disposal. The paper also includes the analysis of disease pattern obtained from hospital records of two coal mines which depicts 19.1%, 24.7% and 16% members of coal miners' families suffering from disorder with respiratory, gastro-intestinal and fever respectively. With speedy industrialization of the country, the mining of coal resource comes first in the chain of socio-economic development. The speedy human industrial activities are based on 80% steam, metallurgical and thermal electrical energy which hinges on coal wings. The coal has also gradually occupied all the phases of social life, our clothes, books, newspapers, cooking gas, chemical paints, dye stuff, oil phenyl, Benzene, Naphthalene, Coal tar, scents and various types of unaccountable products come out from coal derivatives and pushed to serve in the today's market for our daily exigencies. Every day one finds a new coal based industry is coming up in the area. The coal is utilized in two hundred ways in our various walks of social life.(ABSTRACT TRUNCATED AT 250 WORDS)

Air Pollution↗

Accounts receivable reports: underutilized mining tools.

There is gold to be found in accounts receivable reports for those willing to mine the data. The key is to know how to interpret the information buried within the numbers and use it to recover monies owed. This article identifies seven reports that should be staples in every organization committed to improving its overall collection performance. Also included are tips on understanding reports and implementing changes.

Accounts Payable and Receivable↗

Chronic pulmonary disease in South Wales coal mines: an eye-witness account of the MRC surveys (1937-1942).

In the mid-1930s reports were accumulating from the British coalfields, particularly from the anthracite area of South Wales, that coal face workers suffered a disabling lung condition that was not recognized as the (compensatable) silicosis of rock workers. The Second World War was threatening and discontent was rife. Government, through the Medical Research Council, initiated a medical and environmental investigation of chronic pulmonary disease in South Wales coalminers to make a systematic survey. The medical surveys, 1936-1942, were undertaken by a member of MRC staff, Dr Philip D'Arcy Hart assisted by Dr Edward Aslett of the Welsh National Memorial Association. One colliery (Ammanford) was intensively investigated; fifteen others less so; coal trimmers at the docks were added. The main observations were to confirm and describe radiographically the frequency of serious lung lesions apparently due to coal dust, and distinguishable from classical silicosis. Among recommendations accepted by Government, the lung condition became recognized for compensations, and the generic term pneumoconiosis of Coal Workers' was substituted for silicosis.

Causality↗