Search PubMedSearch

SEARCH · Search PubMed

Results for “big data analysis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

17 recordsLinked to original sources

Microbial partnerships and molecular mechanisms in plant stress physiology for climate-resilient and sustainable farming.

Plant-microbial partnerships and their underlying molecular mechanisms are indispensable, natural drivers of improved nutrient acquisition and stress tolerance in the face of climate-driven environmental challenges. Modern multi-omics tools, when coupled with artificial intelligence and synthetic biology, enable the precise design of targeted bioinoculants and synthetic microbial consortia. Translating these advanced microbiome-based strategies into scalable, field-level agricultural applications provides a sustainable path toward securing global food production while maintaining soil health. Global climate change imposes multifaceted abiotic and biotic stresses on crops, disrupting physiological and molecular processes and threatening agricultural productivity. Plant-associated microbes represent an underexplored yet powerful ally in enhancing crop resilience. This review presents current knowledge of plant-microbe interactions and the molecular mechanisms governing plant stress physiology, with an emphasis on climate-resilient and sustainable farming. Hence, ever-changing environmental cues pose a significant burden on agricultural productivity, and plant-associated microbial communities modulate a cascade of physiological and molecular responses, including production of phytohormones, signaling, regulation of reactive oxygen species homeostasis, and activation of plant immune responses to help plants withstand stress and enhance productivity. Moreover, root exudates, phytohormones, and quorum sensing mediate the central communication networks, facilitating plant-microbe cross talk. Additionally, the advances in OMICs approaches aid in disentangling the molecular underpinnings of these interactions by providing mechanistic insights and potential candidate gene targets for crop improvement and stress resilience. In the post-genomic era, integrating artificial intelligence and big data analysis to optimize microbiome-based strategies for sustainable agriculture is a new frontier for disentangling plant-microbe symbiosis to improve soil health, enhance crop yields, and improve stress tolerance. Thus, by integrating the ecological, physiological, and molecular perspectives, this review highlights the transformative potential of harnessing plant-microbe symbiosis for climate-resilient and sustainable agriculture.

Stress, Physiological

Dissemination of antimicrobial resistance in Klebsiella spp. from urban aquatic environments: a multi-country genomic perspective.

INTRODUCTION: Antibiotic resistance, particularly carbapenem-resistant Klebsiella pneumoniae (CRKP), poses significant clinical and environmental threats, especially in urban aquatic ecosystems and hospital wastewaters. OBJECTIVES: This study aims to analyze the epidemiological and genomic features of CRKP isolates in urban aquatic environments and evaluate their public health and environmental impacts. METHODS AND RESULTS: Water samples were collected from 113 rivers and 3 hospitals in China, Sri Lanka, and Nepal to isolate carbapenem-resistant Klebsiella spp. isolates. Antimicrobial susceptibility testing, whole-genome sequencing, and bioinformatics analyses were performed to characterize resistance phenotypes, antibiotic resistance genes (ARGs), and evolutionary trends. Big data analysis further elucidated the genomic characteristics of CRKP in global water sources, and Galleria mellonella larvae were used to assess virulence. Statistical analysis validated the findings. A total of 192 carbapenem-resistant Klebsiella spp. isolates were identified from urban aquatic ecosystems in China (n = 60) and Nepal (n = 132), with CRKP (n = 161) being the predominant species. All CRKP isolates exhibited a multidrug-resistant phenotype, yet significant differences in resistance profiles and associated ARGs were observed between isolates from the two countries. Nine carbapenem resistance genes (CRGs) were detected, with blaNDM-1 being the most prevalent (57.8 %). Correlation analysis revealed a strong association between these CRGs and multiple Inc-type plasmids. Global genomic analysis of CRKP from water sources across eight countries identified ten distinct CRGs across 45 serotypes, with KL64 being the most predominant. Notably, carbapenem-resistant hypervirulent Klebsiella pneumoniae was detected in water samples from Nepal. CONCLUSION: Our findings highlight significant regional disparities in CRKP prevalence and ARG dissemination across urban aquatic environments, with Nepal showing the highest prevalence, particularly in untreated rivers. China exhibited lower prevalence but distinct resistance gene profiles, while no CRKP was detected in Sri Lanka, underscoring the impact of environmental management and healthcare infrastructure on ARG spread.

Humans

Big data and psychiatry: advances, constraints and future directions.

Early work in psychiatry research, often involving single sites, small samples, and limited variables, has shifted to contemporary research involving multiple sites, large samples, and many variables. Such research raises important questions, including concerns about data quality and methodological rigor, uncertainty about its key lessons, issues regarding clinical relevance, and questions about how to optimize future advances. Here we consider these questions and concerns against the context of big data work on community and register-based surveys, cohort and biobank studies, electronic health records, digital phenotyping, brain imaging, genomics and other -omics, and randomized controlled trials. The development of large datasets allowing well-powered analyses is a major milestone, but sample size alone does not guarantee more precise estimates, and ongoing attention to the quality and rigor of big data collation and analysis is needed. Big data research has fostered trans-disciplinarity and given insights into mechanisms underlying psychiatric disorders, but also emphasizes the intricacy, heterogeneity and variability of such mechanisms, and the importance of triangulating between large-scale and small-scale research. The complexity of psychiatric phenotypes and psychobiological mechanisms contributes to the difficulty in bridging from big data to clinical application; big data research reinforces the importance of holding our diagnoses of psychiatric disorders lightly and providing explanations of these conditions humbly; and future work needs to be more attentive to clinical issues. There is enormous scope for further building databases relevant to psychiatry, but advances in conceptual models and asking the right questions are equally valuable. The full impact of big data, including artificial intelligence analyses, remains to be seen, but overenthusiastic support should be tempered by a better understanding of its strengths and limitations. At its best, such work will contribute in an iterative and integrative way to advancing our knowledge of psychiatric disorders and mental health.

Big data

Evolution and applications of genome-scale metabolic models in yeast systems biology studies.

Genome-scale metabolic models (GEMs) can be used to simulate the metabolic network of an organism in a systematic and holistic way. Different yeast species, including Saccharomyces cerevisiae, have emerged as powerful cell factories for bioproduction. Recently, with the dedicated efforts from the scientific community, significant progress has been made in the development of yeast GEMs. Numerous versions of yeast GEMs and the derived multiscale models have been released, facilitating integrative omics analysis and rational strain design for different types of yeast cell factories. These advancements reflected the evolution and maturation of yeast GEMs together with a model ecosystem around them. This review will summarize the development and expansion of yeast GEMs and discuss their applications in yeast systems biology studies. It is anticipated that yeast GEMs will continue to play an increasingly important role in pioneering yeast physiological and metabolic studies in coming years.

Systems Biology

Interpreting cancer genetics through a two-step "evolutionary cascade hypothesis": bridging neutral and selective perspectives.

BACKGROUND: DNA mutations are the fundamental engines of cancer, driving its initiation and progression. The forces that fuel malignancy are also the architects of evolution, shaping life through genetic variations. Mutations, in fact, can emerge naturally from endogenous processes, such as oxidative DNA damage or errors in replication, as well as induced by external factors, including cosmic radiation and chemical carcinogens. MAIN BODY: A key question in cancer research is whether tumor evolution is primarily governed by selective bottlenecks, neutral evolution, or dynamic genetic plasticity. In this work, we examine cancer as a disease driven by evolutionary processes rooted in fundamental biological requirements, including sustained proliferation and nutrient utilization. We hypothesize that the accumulation of mutations activates an evolutionary switch, enabling tumor cells to acquire an enhanced capacity for survival, adaptation, and growth at rates far exceeding typical evolutionary timescales. We propose the "evolutionary cascade hypothesis," a unifying framework that integrates these models into a coherent sequence. At its core lies the failure of DNA repair mechanisms, representing a critical transition in cancer progression. This shift marks the transition from an initial non-Darwinian, neutral phase to a Darwinian, more deterministic phase. CONCLUSIONS: As predictive models of tumor evolution advance through genomic big data and artificial intelligence-driven analysis, the future of cancer treatment may extend beyond targeting individual mutations to disrupting the underlying evolutionary mechanisms that sustain malignancy. This paradigm shift could redefine therapeutic strategies and ultimately improve patient outcomes.

Humans

Big data analytics for CLEC5A dynamics based on single cell genomics and proteomics reveal its diverse functions in human diseases.

BACKGROUND: CLEC5A (C-type lectin domain family 5 member A) is an innate immune receptor implicated in inflammatory signaling, contributing to hyperinflammatory responses in infections and sterile inflammation. However, CLEC5A dynamics in human diseases remain to be identified. Here, we systematically characterized CLEC5A dynamics in humans across cells, tissues, and disease states, and to explore the functional significance of CLEC5A in macrophage activation based on single-cell genomics. METHODS: With multi-omics (scRNA-seq, proteomics and big data analytics), we analyzed extensive human transcriptomic datasets (>42,000 samples) to profile CLEC5A expression by cell type, tissue, and disease. Single-nucleus RNA-seq (snRNA-seq) from pediatric congenital heart disease and a virtual CLEC5A gene knockout were also performed to characterize CLEC5A dynamics in humans. RESULTS: CLEC5A is highly enriched in innate immune cells, particularly in macrophages and neutrophils. Baseline CLEC5A in most tissues is low, but it is markedly upregulated in inflammatory and infectious diseases. CLEC5A expression has sex-specific differences in certain organs. Single-cell analysis showed that CLEC5A can be considered novel marker of proinflammatory macrophages with elevated cytokine production, antigen presentation, and impaired phagocytosis. Virtual CLEC5A knockout analysis identified coordinated perturbation of immune-regulatory pathways and overlapping genes linking CLEC5A to macrophage activation networks. CONCLUSION: CLEC5A is predominantly expressed in myeloid cells and acts as a key amplifier of inflammation in human diseases. Our findings highlight CLEC5A as a potential biomarker and therapeutic target in myeloid-driven hyperinflammatory conditions, warranting further experimental and translational validation.

Humans

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework to address this challenge, which reduces runtime and memory requirements by orders of magnitude. In PSU, small subsamples of sites from a concatenated alignment are analyzed, which are expanded by upsampling before inference, and the resulting inferences are aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-alignment analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis of concatenated alignments. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical. We suggest that PSU is a general approach for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

Phylogeny

Advances in tumor subclone formation and mechanisms of growth and invasion.

Tumor subclones refer to distinct cell populations within the same tumor that possess different genetic characteristics. They play a crucial role in understanding tumor heterogeneity, evolution, and therapeutic resistance. The formation of tumor subclones is driven by several key mechanisms, including the inherent genetic instability of tumor cells, which facilitates the accumulation of novel mutations; selective pressures from the tumor microenvironment and therapeutic interventions, which promote the expansion of certain subclones; and epigenetic modifications, such as DNA methylation and histone modifications, which alter gene expression patterns. Major methodologies for studying tumor subclones include single-cell sequencing, liquid biopsy, and spatial transcriptomics, which provide insights into clonal architecture and dynamic evolution. Beyond their direct involvement in tumor growth and invasion, subclones significantly contribute to tumor heterogeneity, immune evasion, and treatment resistance. Thus, an in-depth investigation of tumor subclones not only aids in guiding personalized precision therapy, overcoming drug resistance, and identifying novel therapeutic targets, but also enhances our ability to predict recurrence and metastasis risks while elucidating the mechanisms underlying tumor heterogeneity. The integration of artificial intelligence, big data analytics, and multi-omics technologies is expected to further advance research in tumor subclones, paving the way for novel strategies in cancer diagnosis and treatment. This review aims to provide a comprehensive overview of tumor subclone formation mechanisms, evolutionary models, analytical methods, and clinical implications, offering insights into precision oncology and future translational research.

Humans

Resilience indicator traits in chickens: a systematic review.

The resilience of an animal is its ability to cope with short-term disturbances, including those caused by pathogens, through response and rapid recovery to its original state. Here, a systematic literature review was conducted to identify resilience indicators that have already been studied and implemented for chickens, along with the contexts or specific stressor under which they were examined. The literature review was based on predefined search criteria, including 'chickens' as the population of interest, a stress exposure description and the term 'resilience'. Screening of titles and abstracts was assisted by the AI-based tool ASReview, followed by full text screening and analysis. According to selection criteria, we finally identified 33 relevant publications on resilience indicator traits used in chickens. Two of these studies analyzed chicken resilience based on routinely collected big data, while the remaining majority were small or medium scale studies and trials. The majority of the studies focused on immune (n = 13) or thermal challenges (either before or after hatching; n = 14), especially heat stress. A variety of resilience indicators were studied; most studies investigated production- or performance-related parameters and/or immunity or disease-related parameters. About half of the studies included genetics- or gene expression-related indicators. Investigation of behavioral indicators of resilience - other than feed intake - was limited (n = 4). Overall, this systematic review provides a comprehensive overview of published resilience indicators in chickens and highlights gaps for future research. This review also revealed a need for the controlled use of terms like resilience, robustness, resistance, tolerance or adaptability, and the relevance of assessing the phase of recovery after a short-term disturbance in the context of resilience.

Animals

Benchmarking methods for measuring biosynthetic gene cluster similarity and determination of gene cluster families.

MOTIVATION: Natural products are often produced by a set of biosynthetic enzymes that are encoded by genes clustered together in the producer's genome, referred to as a biosynthetic gene cluster (BGC). The ability to compare and cluster BGCs is essential for several applications, including predicting which bacteria will make a known product and assessing the potential diversity of natural products produced by a set of bacteria. There are multiple methods for comparing and clustering BGCs based on their similarity, but there has been a lack of investigation into how strongly BGC similarity relates to product structural similarity and how these methods perform relative to each other. RESULTS: Using publicly available databases, we developed a benchmark dataset to assess how well different BGC similarity metrics correlate with the structural similarity of their products and how well these methods cluster BGCs. We found that all methods showed moderate correlation between BGC and structural similarity, with correlations improving for more similar BGCs and varying significantly by BGC biosynthetic class. Analysis of outliers revealed some outliers were due to mistakes or omissions in public datasets, while others represented deviation between BGC similarity and product structural similarity. All methods generally performed better on clustering metrics, with BiG-SCAPE performing the best after errors in the public datasets had been corrected. AVAILABILITY AND IMPLEMENTATION: Scripts and data required to reproduce the results are available at https://github.com/aswalker-lab/BGC-clustering-benchmark and processed similarity, clusters, and scaffolds are also available at https://huggingface.co/datasets/allie-walker/BGC-clustering-benchmark. Code is also available at Zenodo: 10.5281/zenodo.17373546.

Multigene Family

Discovery of novel diagnostic biomarkers of hepatocellular carcinoma associated with immune infiltration.

OBJECTIVE: Diagnosis of hepatocellular carcinoma (HCC) remains challenging for clinicians. Machine learning approaches and big data analyses are viable strategies for identifying HCC diagnostic markers. MATERIALS AND METHODS: In this study, we downloaded mRNA expression profiles of HCC from the GEO database and used random forest and machine learning algorithms, such as least absolute shrinkage and selection operator, to screen for reliable diagnostic genes. Disease Ontology, Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Set Enrichment Analysis enrichment analyses were performed to explore differential gene functions and disease pathways. CIBERSORT was performed to calculate the immune cell infiltration of HCC and the correlation between diagnostic genes and immune cells. Cell experiments were performed to evaluate the function of R-spondin 3 (RSPO3) in HCC cells. Immunohistochemical staining was used to evaluate the protein expression of CD138, CD206 and iNOS. RESULTS: The results indicated that extracellular matrix protein 1 (ECM1), Niemann-Pick C1-Like 1 (NPC1L1) and RSPO3 were down-regulated in HCC compared with the normal group (p&#x2009;<&#x2009;0.05), which was validated in clinical tissue samples. Moreover, ECM1, NPC1L1 and RSPO3 had high diagnostic values (AUC > 0.75) for HCC in both training and test groups. Immuno-infiltration analysis revealed that ECM1 and RSPO3 were highly positively correlated with neutrophil and macrophage M2 levels, whereas they were negatively correlated with Tregs. RSPO3-si affected cell proliferation and apoptosis in HCC. Furthermore, RSPO3 exhibited a positive correlation with tumour progression, the proportion of plasma cells and M2 macrophages in mice, while showing a negative association with M1 macrophages. CONCLUSION: The present study identified ECM1, NPC1L1 and RSPO3 as new diagnostic biomarkers for HCC based on normal and diseased samples from HCC, meanwhile the pro-oncogenic function of RSPO3 and its regulation on immune infiltration have been confirmed.

Carcinoma, Hepatocellular

Proteomics at scale: Bottlenecks and opportunities for early-career researchers in a fast developing field.

The field of proteomics has rapidly evolved over the last five years enabled by rapid advances in instrumentation and computation. At the same time, the proteomics community is also growing. This is reflected by the increasing participation in international conferences such as those organized by the European Proteomics Association and the Human Proteome Organization. These events provide early-career researchers with unique opportunities to exchange ideas, develop collaborations, and build networks that support professional development. One such network is the Young Proteomics Investigators Club, a European initiative supported by European Proteomics Association and led by early-career researchers. In this Community-Driven project, we investigate recent trends in proteomics by screening conference abstracts and evaluating the session attendance at Human Proteome Organization Congresses and European Proteomics Association conferences. Based on these analyses, we identified five areas that, from our perspective, are shaping the current trends in proteomics: clinical proteomics, proteomics of post-translational modifications, single-cell proteomics, systems biology and multi-omics, and computational proteomics. For each area, we highlight both unique challenges and identify a common theme: a shift from exploratory studies with manageable sample numbers towards large screenings and cohorts and the generation of big data, which often comes with the lack of computational support, organizational networks, and infrastructure. In this light, we describe the unique challenges and opportunities faced by early-career researchers. We point to actionable directions for enabling reproducible and transparent proteomics as well as community-driven projects and initiatives, which are often providing training and support. SIGNIFICANCE: In this perspective, the Young Proteomics Investigators Club (YPIC) discusses advances in analytical developments and computational approaches in proteomics research. Based on empirical analysis of recent European Proteomics Association conference and Human Proteome Organization congresses contributions, we identify clinical, single-cell, post-translational and systems-level proteomics as the research areas that have gained most momentum in the last three to five years. What makes this work distinctive is that it is written by and for early-career researchers, thereby uniquely identifying where momentum, challenges, and unmet needs converge for the newest generation of proteomics researchers. Rather than cataloguing advances, we examine the widening gap between what modern proteomics can generate and what individual researchers can realistically process, validate, and interpret. We describe specific structural barriers including access to high performance computing, limited formal training in scalable data analysis, the need for unified benchmarking standards and navigating clinical collaboration frameworks. We then highlight opportunities for the field, such as community-curated benchmarks, interdisciplinary mentorship models, and shared computational infrastructure. By making these challenges explicit from an early-career researchers standpoint, we aim to inform how training, funding, and community initiatives can be shaped to support the next generation of proteomics researchers.

Proteomics

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

AI-integrated digital breeding for crop improvement.

Crop breeding increasingly depends on the effective integration and interpretation of large, heterogeneous datasets spanning genomic, phenotypic, multi-omics, and environmental layers. Conventional breeding approaches are often insufficient to capture the complex relationships among these data or to support timely selection decisions. Digital breeding can help address this limitation by complementing field experimentation, mixed models, and genomic prediction with the integration of biological data and computational prediction throughout the breeding process. In particular, the rapid advancement of artificial intelligence (AI) has improved the analysis of high-dimensional datasets and broadened its application to trait prediction, selection, and breeding design. Here, we review recent developments in AI-enabled digital breeding, encompassing genomic, phenomic, and multi-omics data generation and analysis, predictive modeling, explainable and generative AI, and data-driven breeding decision support. We further discuss emerging AI applications, their current contributions to crop research and breeding, and the major considerations affecting their reliable and practical implementation. Collectively, this review provides a structured understanding of the roles of AI across the digital breeding process and offers guidance for future methodological development and practical application in crop improvement.

artificial intelligence

Differential Mutagenic Response of Rat Liver and Lung to Nicotine-Derived Nitrosamine Ketone (NNK).

Nitrosamines (NA) are chemical impurities that are present in tobacco, foods, more recently in some pharmaceuticals and are associated with genotoxicity and carcinogenicity. We evaluated the in vivo mutagenicity of nicotine-derived nitrosamine ketone (NNK) or 4-(methyl nitrosamino)-1-(3-pyridyl)-1-butanone, a model compound used as an anchor molecule to estimate carcinogenic potency of unknown nitrosamine impurities. Big Blue rats were treated with NNK at doses ranging from 0.001 to 30 mg/kg for 28 days, following which liver and lung tissue were harvested 3 days later for nuclear genomic DNA isolation. Mutations in liver and lung were assessed with the cII transgene assay and endogenous genomic loci using Duplex Sequencing (DupSeq), a highly validated error-corrected sequencing (ECS) technology. The no genotoxic effect level (NOGEL) was 1 mg/kg in liver and 0.1 mg/kg in lung while the benchmark dose (BMD) analysis for cII mutagenicity determined a BMDL50 of 1.3 mg/kg in liver and 0.12 mg/kg in lung, consistent with lung being the more sensitive target organ for carcinogenicity for NNK. ECS-derived mutagenicity was highly correlated with cII-derived mutagenicity. Interestingly, the types of mutations formed appeared to be tissue-specific with higher C > T transitions and lower T > G transversions in lung compared to liver, differences that may reflect tissue-specific DNA repair capacity and/or metabolic differences. Collectively, these data support the use of in vivo mutagenicity data&#x2500;from both TGR cII and ECS methods&#x2500;for human health and cancer risk characterization of nitrosamines and for estimating acceptable daily intakes for unknown nitrosamine drug substance related impurities.

Animals

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans