Search PubMedSearch

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Dataset Readiness Assessment With Large Language Model (DRAFT-LLM): A Multi-Axis Audit Guided by LLM.

This article details the Dataset Readiness Assessment for Training (DRAFT), a systematic method for determining whether a high-dimensional biological dataset is suitable for developing reliable, equitable (i.e., the extent to which model performance, error patterns, and potential benefits or harms are evaluated and found to be acceptably distributed across relevant demographic, biological, clinical, and contextual subgroups), and scientifically meaningful machine-learning models, and DRAFT Large Language Model (DRAFT-LLM), its optional human-in-the-loop extension for calibrating study-specific audits through structured, critically reviewed LLM guidance. Standard model validation often fails to detect when apparent performance is driven by spurious correlations, technical artifacts, or hidden stratification, leading to irreproducible and inequitable findings. DRAFT-LLM addresses this gap by shifting the focus from model tuning to structured dataset auditing, organized around Support Protocols 1 to 4 that capture the scientific intent, data structure, and governance constraints of a given study. These Support Protocols: (1) elicit and formalize investigator input into a study intake and dataset card; (2) compute standardized dataset statistics and structural summaries suitable for downstream analysis and LLM context; (3) configure the language model using form-based responses, safety guardrails, and governance rules; and (4) generate personalized instructions, prompts, and code templates for running DRAFT audits. Basic Protocols 1 to 3 are instantiated from this support layer for generalization, equity, and stability: they are reusable execution patterns whose concrete behavior is determined by the cards, statistics, and configurations defined in the Support Protocols. DRAFT-LLM and DRAFT are demonstrated in this article through an end-to-end case study on The Cancer Genome Atlas (TCGA). © 2026 Wiley Periodicals LLC. Support Protocol 1: Study intake and dataset card construction Support Protocol 2: Dataset structure and advanced summary statistics for LLM context Support Protocol 3: LLM configuration using structured form responses Support Protocol 4: Generation of personalized instructions for DRAFT audits Basic Protocol 1: Generalization audit Basic Protocol 2: Equity audit Basic Protocol 3: Stability audit.

Large Language Models

BMDx2: A Tool for Integrating Toxicogenomics-Based Dose-Dependency Analysis and AOP-Based Mechanistic Insights.

Despite the advent of mechanistic toxicology using omics data to link molecular perturbations with systemic outcomes, regulatory toxicology still lacks the application of mechanism-anchored metrics from such data. This is partially because traditional gene-centric analysis often falls short of linking molecular changes to adverse outcomes. To address this gap, BMDx2, an open-source tool that transforms multi-dose toxicogenomics datasets into quantitative, mechanistic evidence for human chemical safety assessment is developed. BMDx2 couples benchmark-dose modeling with Adverse Outcome Pathway (AOP) enrichment to derive transcriptomic-based points of departure, enabling potency ranking, chemical prioritization, and mechanistically anchored explanations of the effect of chemical exposures. BMDx2 can process a broad range of data, including DNA microarray and RNA sequencing studies. Here, case studies are used to illustrate the versatility of BMDx2 in characterizing the mechanism of action of chemicals. An initial case study on carbon nanotubes exposure applies integrative analysis of transcriptomics and genome-wide DNA methylation data, uncovering cellular reprogramming processes underlying fibrosis. A second case study on bleomycin exposure demonstrate how transcriptomic data alone can be mapped to fibrosis-related AOPs in a standardized, regulatory appropriate manner. Together, these examples show how BMDx2 supports the regulatory application of toxicogenomics and accelerates mechanism-based chemical safety evaluation.

Toxicogenetics

High-quality chromosome-level genome of three Meretrix species using Nanopore and Hi-C technologies.

Meretrix is a commercially valuable bivalve genus in Asia, but only one reference genome has hindered comprehensive genetic studies and germplasm resource evaluation. In this study, we present three reference genomes of Meretrix species: Meretrix sp. MF1, Meretrix sp. MT1, and Meretrix lamarckii JML1. Meretrix sp. MF1 was assembled at the chromosome level using Nanopore sequencing and Hi-C technologies, whereas Meretrix sp. MT1 and Meretrix lamarckii were assembled as scaffold-level assemblies. The chromosome-level genome of Meretrix sp. MF1 consists of 36 contigs, including 19 chromosomes and 17 scaffolds, with a total length of 883.3 Mb and a scaffold N50 of 46.87 Mb. Notably, the genome of Meretrix sp. MF1, a putative novel species, exhibits an Average Nucleotide Identity (ANI) of 94.33% with its closest relative, Meretrix lamarckii. These genomic resources not only provide a crucial foundation for genetic research on Meretrix but also contribute to the development of effective conservation strategies for its sustainable management.

Animals

Chromosome-level genome assembly and annotation of Spinibarbus caldwelli.

Spinibarbus caldwelli is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in southern China, the lack of genomic resources has hampered the genetic breeding and conservation. Here, we release a chromosome-level genome assembly for S. caldwelli using PacBio HiFi long-reads, Illumina short-reads, and Hi-C sequencing data. The final genome assembly is 1.77 Gb in size, with a contig N50 of 24.27 Mb. Using Hi-C scaffolding, 99.14% of the contigs were successfully anchored to 50 chromosomes, resulting in a scaffold N50 of 35.29 Mb. The final genome assembly shows a BUSCO completeness of 98.27%. The assembled genome contains 49.41% repetitive sequences and 51,505 predicted genes, 90.83% of which have been functionally annotated. This genome provides a genetic basis for S. caldwelli, facilitating the exploration of Cyprinid phylogeny, genetic improvement, and conservation efforts.

Animals

GE-IA-NAM: gene-environment interaction analysis via imaging-assisted neural additive model.

MOTIVATION: Gene-environment (G-E) interaction analysis is crucial in cancer research, offering insights into how genetic and environmental factors jointly influence cancer outcomes. Most existing G-E interaction methods are regression-based, which may lack flexibility to capture complex data patterns. Recent advances have investigated deep neural network-based G-E models. However, these methods may be more vulnerable to information deficiency due to challenges such as limited sample size and high dimensionality. Apart from genetic and environmental data, pathological images have emerged as a widely accessible and informative resource for cancer modeling, presenting its potential to enhance G-E modeling. RESULTS: We propose the pathological imaging-assisted neural additive model for G-E analysis (GE-IA-NAM). The flexible and interpretable additive network architecture is adopted to account for individualized effects associated with genetic factors, environmental factors, and their interactions. To improve G-E modeling, an assisted-learning strategy is investigated, which adopts a joint analysis to integrate information from pathological images. Simulations and the analysis of lung and skin cancer datasets from The Cancer Genome Atlas demonstrate the competitive performance of the proposed method. AVAILABILITY AND IMPLEMENTATION: Python code implementing the proposed method is available at https://github.com/Mr-maoge/NAM-IA-GE. The data that support the findings in this article are openly available in TCGA (The Cancer Genome Atlas) at https://portal.gdc.cancer.gov/.

Gene-Environment Interaction

A genome-wide association study identified 10 novel genomic loci associated with intrinsic capacity.

BACKGROUND: Intrinsic capacity (IC) is a multidimensional concept within the World Health Organization framework for healthy aging. It refers to the composite of an individual's physical and mental capacities that enable them to maintain well-being, functional ability, and engagement in valued activities throughout life. While substantial evidence supports the biological basis of IC and its subdomains, the extent to which genetic factors influence IC remains largely unexplored, with no studies currently available. METHODS: Using datasets from the UK Biobank (UKB; N = 44 631) and the Canadian Longitudinal Study on Aging (CLSA; N = 13 085), we implemented the restricted maximum likelihood method to estimate SNP-based heritability (h2snp), followed by a Genome-Wide Association Study (GWAS) to identify genetic variants associated with IC, and post-GWAS analyses to pinpoint biological implications. RESULTS: The h2snp for IC was estimated at 25.2% in UKB and 19.5% in CLSA. Our GWAS identified 38 independent SNPs for IC across 10 genomic loci and 4289 candidate SNPs, mapped to 197 genes. Post-GWAS analysis revealed the role of these genes in cellular processes such as cell proliferation, immune function, metabolism, and neurodegeneration, with high expression in muscle, heart, brain, adipose, and nerve tissues. Of the 52 traits tested, 23 showed significant genetic correlations with IC, and a higher genetic loading for IC was associated with higher IC scores. CONCLUSIONS: Overall, this study provides comprehensive evidence on the genetic architecture of IC, identifying novel genetic variants and biological pathways, advancing our current knowledge and laying the foundation for ongoing and future research on healthy aging.

Adult

Integrative multi-omics quantitative trait loci prioritize CASP7 as a candidate protective gene for cataract.

Cataracts are the leading cause of vision loss worldwide. Despite surgery being the only effective treatment, its economic burden highlights the necessity of exploring the pathogenesis of cataracts. In this study, we analyzed 4 large-scale GWAS (genome-wide association study) datasets for cataracts and performed SMR analysis along with heterogeneity in dependent instruments (HEIDI) testing to explore the effects of methylation, expression, and protein QTLs on cataracts. We further validated shared genetic variants through COLOC analysis. Additionally, we searched datasets related to cataracts from the Gene Expression Omnibus (GEO) database for differentially expressed genes (DEGs) and Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes pathway (KEGG) enrichment analyses. By integrating summary-based Mendelian randomization (SMR) results with bioinformatics findings, CASP7 showed a consistent protective-direction association with cataract risk (mQTL: OR [95% CI] = 0.959 [0.941-0.977], FDR-adjusted P = .039; eQTL: OR [95% CI] = 0.897 [0.860-0.937], FDR-adjusted P = .0046; pQTL: OR [95% CI] = 0.597 [0.483-0.738], FDR-adjusted P = .00083). GEO-based analyses provided transcriptomic support for CASP7 involvement in cataract-related lens biology. These findings prioritize CASP7 as a genetically supported candidate protective gene associated with cataract risk. Because this study is based on public summary-level and transcriptomic datasets, the results should be interpreted cautiously and require functional validation in human lens-relevant systems.

Quantitative Trait Loci

SIGEL: a context-aware genomic representation learning framework for spatial genomics analysis.

Spatial transcriptomics (ST) integrates spatial information into genomics, yet methods for generating spatially-informed gene representations are limited and computationally intensive. We present SIGEL, a cost-effective framework that derives gene manifolds from ST data by exploiting spatial genomic context. The resulting SIGEL-generated gene representations (SGRs) are context-aware, biologically meaningful, and robust across samples, making them highly effective for key downstream tasks, including imputing missing genes, detecting spatial expression patterns, identifying disease-related genes and interactions, and improving spatial clustering. Extensive experiments across diverse ST datasets validate SIGEL's effectiveness and highlight its potential in advancing spatial genomics research.

Genomics

Associations Between Walking Pace, APOE-ε4 Genotype, and Brain Health in Middle-Aged to Older Adults.

PURPOSE: This study aimed to investigate whether self-reported walking pace (a marker of physical function) and the presence of APOE-&#x3b5;4 allele interact to modify brain health outcomes. METHODS: We used data from a prospective cohort study of middle-aged to older adults from the UK Biobank who self-reported walking pace (slow or steady-to-brisk) and who were initially free of dementia ( n = 415,110). Incident all-cause dementia was obtained from hospital and death registry records, and structural brain volumes (right and left hippocampus volumes, total gray matter volume, and volume of white matter hyperintensities) were measured from a subset of participants ( n = 33,113). Cox proportional hazard models and generalized linear models were used to assess associations between exposures and outcomes. RESULTS: Slow walking pace and the presence of APOE-&#x3b5;4 allele were associated with increased dementia risk (HR = 1.79 [95% CI = 1.66-1.93], P < 0.001; HR = 3.06 [2.90-3.23], P < 0.001, respectively), and there was an interaction between these associations, indicating that the association of walking pace with dementia risk is modified by APOE-&#x3b5;4 status (reference group: HR Steady-Brisk/APOE-&#x3b5;4- = 1; HR Slow/APOE-&#x3b5;4- = 2.03 [1.84-2.25], P < 0.001; HR Steady-Brisk/APOE-&#x3b5;4+ = 3.21 [3.02-3.41], P < 0.001; HR Slow/APOE-&#x3b5;4+ = 4.99 [4.48-5.58], P < 0.001). Slow self-reported walking pace was associated with worse brain volume outcomes, and these associations were not modified by APOE-&#x3b5;4 genotype. CONCLUSIONS: These results suggest walking pace and APOE-&#x3b5;4 status independently influence brain volume outcomes, but both factors independently and jointly contribute to increased dementia risk. Individuals with both risk factors (slow walking pace and APOE-&#x3b5;4 allele) show the strongest associations with dementia risk.

Self Report

aPhyloGeo: a Python application for correlating genetic and climatic conditions.

MOTIVATION: Environmental variation and its influence on genetic diversity is a central topic in evolutionary biology and phylogeography. Accurate correlations between genetic and climatic datasets to understand the genetic adaptations of different species to specific environments. It requires integrated and reproducible workflows. RESULTS: We developed aPhyloGeo, an open-source and multiplatform application implemented in Python, for investigating correlations between genetic variation and environmental data within a phylogenetic framework. The workflow integrates multiple analytical steps, including sequence alignment, sliding window phylogenetic inference, and statistical approaches such as the Mantel test and the Procrustean randomization test. These analyses enable the identification of mutation hotspots that exhibit strong associations with environmental variables. In addition, aPhyloGeo supports multicore data processing and provides a fully reproducible pipeline for evaluating localized relationships between genomic variation and climatic distributions. AVAILABILITY AND IMPLEMENTATION: aPhyloGeo is freely available on GitHub at: https://github.com/tahiri-lab/aPhyloGeo, as both a PyPI package and as Python scripts for Linux, macOS, and Windows.

Software

Understanding and making sense of epigenetic age misalignment across different aging clocks.

The output of an epigenetic aging clock can vary depending on the training method utilized, cell type composition, the nature of the training dataset, the technology used to generate the methylomic data, acute stressors, and other factors. On an individual level, epigenetic age can fluctuate across different clocks purely due to differences in model training. Among aging clock researchers, it is well-known that the epigenetic age of a single sample can vary across different models. Based on our observations and conversations with longevity scientists and stakeholders, however, this fact is often unappreciated among non-aging clock experts. To help bring more awareness to this important topic, we highlight key literature and, as an illustrative example, use eight blood-trained clocks to show that epigenetic age is frequently misaligned in a publicly available whole blood dataset. Our simple analysis revealed that the average sample difference between the youngest and oldest predicted ages across these clocks was 17&#x2009;years. The smallest and largest individual-level differences observed were 4 and 45&#x2009;years, respectively. Clock misalignment has implications for choosing which clock to utilize, interpreting the impact of an intervention on epigenetic age, personalized tracking, and relating epigenetic age to the abstract concept of biological age.

Humans

An overview of statistical issues and methods of meta-analysis.

A meta-analysis is a statistical analysis of the data from some collection of studies in order to synthesize the results. In this paper we discuss issues that frequently arise in meta-analysis and give an overview of the methods used, with particular attention to the use of fixed- and random-effects approaches. The methods are then applied to two sample datasets.

Biopharmaceutics

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics

Improvements in protein secondary structure prediction by an enhanced neural network.

Computational neural networks have recently been used to predict the mapping between protein sequence and secondary structure. They have proven adequate for determining the first-order dependence between these two sets, but have, until now, been unable to garner higher-order information that helps determine secondary structure. By adding neural network units that detect periodicities in the input sequence, we have modestly increased the secondary structure prediction accuracy. The use of tertiary structural class causes a marked increase in accuracy. The best case prediction was 79% for the class of all-alpha proteins. A scheme for employing neural networks to validate and refine structural hypotheses is proposed. The operational difficulties of applying a learning algorithm to a dataset where sequence heterogeneity is under-represented and where local and global effects are inadequately partitioned are discussed.

Artificial Intelligence

Cutoff assignment strategies for enhancing randomized clinical trials.

The randomized clinical trial (RCT) is the preferred method for assessing the efficacy of treatments. Recent ethical and logistical criticisms suggest that new variations of the traditional RCT are needed. Some of these criticisms may be addressed with new hybrid designs that combine random assignment with assignment by one or more cutoff values on a baseline variable (e.g., severity of illness). In a simple version of such a "cutoff-based" RTC, persons scoring below a cutoff score on a baseline measure (i.e., the least severely ill) are automatically assigned to the control-treated group, those scoring above a second, higher cutoff (i.e., the most ill) are automatically assigned to the test-treated group, and those scoring in the interval between the cutoff scores (i.e., the moderately ill) are randomly assigned to either group. Depending on the baseline score, the patient is assigned to treatment either randomly or by the need-based, clinically related baseline score. Six cutoff-based design variations are studied via simulations and compared with the traditional RCT and the single-cutoff (i.e., regression-discontinuity) design. All variations yield unbiased estimates of the treatment effect but estimates differ in efficiency, with the RCT being most efficient and the single-cutoff design being least efficient. Secondary analyses of data from the Cross-National Collaborative Study of the Effects of Alprazolam (Xanax) on panic are conducted for each variation by selectivity discarding cases from the original dataset to stimulate cutoff-based assignment. The results confirm the simulations and illustrate how cutoff-based designs might look with real data.

Alprazolam

Precision Optimization of Behavioral Activation for Major and Subthreshold Depression: A Meta-Analysis of Exploring Dose-Response Relationships and Moderating Factors.

OBJECTIVE: This meta-analysis evaluated the efficacy of behavioral activation (BA) in adolescents with subthreshold depression (SD) or major depressive disorder (MDD), exploring dose-response relationships and moderating factors. METHOD: We searched PubMed, EMBASE, Web of Science, EBSCO, Scopus, and the Cochrane Library for randomized controlled trials (RCTs) through December 31, 2024. Risk of bias was assessed using RoB-2, and evidence quality with Grading of Recommendations Assessment, Development, and Evaluation (GRADE). Analyses were performed with R, using standardized mean difference (SMD) for continuous variables and meta-regression for dose-response relationships. Subgroup analyses included symptom severity, intervention setting, delivery format, and parental involvement. The primary outcome was the reduction in depressive symptoms (PROSPERO: CRD42023444273). RESULTS: A total of 14 studies were included (11 RCTs meta-analyzed, comprising 572 participants). BA demonstrated a moderate effect size compared to treatment-as-usual (SMD = -0.42) and a large effect size compared to no-treatment controls (SMD = -0.87). BA was more effective for mild depressive symptoms (SMD = -0.93) than severe symptoms (SMD = -0.43), with significant efficacy in university settings (SMD = -0.94). Intervention without parental involvement exhibited significantly larger effects than those with parental participation (SMD = -0.94 vs -0.38; p < .0001), although this finding was likely confounded by age, symptom severity, and intervention setting. BA moderately improved both behavioral activation levels and functioning (SMD = 0.49). CONCLUSION: Short-term, school-based BA is significantly beneficial for older adolescents with mild depressive symptoms. Findings provide practical guidance for optimizing BA implementation and highlight directions for future research, including the need for larger sample sizes, standardized follow-up assessments, and more representative samples. PLAIN LANGUAGE SUMMARY: This meta-analysis evaluated the effectiveness of behavioral activation (BA), a therapeutic approach focused on improving positive activities, in adolescents with subthreshold depression or major depressive disorder. Based on 14 included studies and 572 participants, BA was found to be more effective for subthreshold compared to major depression. These findings suggest a promising role for BA as an early intervention tool in adolescent depression. STUDY REGISTRATION INFORMATION: Precision Optimization of Behavioral Activation for Major and Subthreshold Depression: A Meta-Analysis of Exploring Dose-response Relationships and Moderating Factors; https://www.crd.york.ac.uk/PROSPERO/view/CRD42023444273. DIVERSITY & INCLUSION STATEMENT: We worked to ensure sex and gender balance in the recruitment of human participants. We worked to ensure race, ethnic, and/or other types of diversity in the recruitment of human participants. We worked to ensure that the study questionnaires were prepared in an inclusive way. Diverse cell lines and/or genomic datasets were not available. We actively worked to promote sex and gender balance in our author group.

Humans

Genetic analysis workshop IV: insulin dependent diabetes mellitus--summary.

The Insulin Dependent Diabetes Mellitus (IDDM) study was one of the topics of Genetic Analysis Workshop IV (GAW IV) discussed October 7 - 8 at a pre-workshop meeting and presented October 9, 1985 at the American Society of Human Genetics meeting in Salt Lake City. The aim of the study was to have different groups analyze an identical body of real data in order to compare their analytical methods and the results based on their different approaches. This summary describes the available datasets and presents the main results of the nine participating groups. The detailed descriptions of analytical methods and results are presented in the individual papers following this overview. Although differing analytical methods were used there were no discrepancies regarding the interpretation of the data. Whereas the presented gametic associations were well established before, the real gain of this workshop was the unanimous result that the recessive 2-allele-model with incomplete penetrance had to be rejected and at least a 3-allele-model with differing susceptibility alleles and differing penetrances for heterozygotes and homozygotes had to be postulated for the transmission of IDDM.

Alleles