Search PubMedSearch

SEARCH · Search PubMed

Results for “model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50 K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40 kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals

A model for background selection in non-equilibrium populations.

In many taxa, levels of genetic diversity are observed to vary along their genome. The framework of background selection models this variation in terms of linkage to constrained sites, and recent applications have been able to explain a large portion of the variation in human genomes. However, these studies have also yielded conflicting results, stemming from two key limitations. First, existing models are inaccurate in a critical region of parameter space (), where the local reduction in diversity is sharpest. Second, they assume a constant population size over time. Here, we develop predictions for diversity under background selection based on the Hill-Robertson system of two-locus statistics, which allows for population size changes. We treat the joint effect of multiple selected loci independently, but we show that interference among them is well captured through local rescaling of mutation, recombination and selection in an iterative procedure that converges quickly. We further accommodate existing background selection theory to non-equilibrium demography, bridging the gap between weak and strong selection. Simulations show that our predictions are accurate across the entire range of selection coefficients. We characterize the temporal dynamics of linked selection under population size changes and demonstrate that patterns of diversity can be misinterpreted by other models. Specifically, biases due to the incorrect assumption of equilibrium carry over to downstream inferences of the distribution of fitness effects and deleterious mutation rate. Jointly modeling demography and linked selection therefore improves our understanding of the genomic landscape of diversity, which will help refine inferences of linked selection in humans and other species.

Journal Article

Statistical test to compare the linkage model and the admixture model based on central limit results.

In the Admixture Model, the probability that an individual carries a certain allele at a specific marker depends on the allele frequencies in K ancestral populations and the proportion of the individual's genome originating from these populations. The markers are assumed to be independent. The Linkage Model is a Hidden Markov Model that extends the Admixture Model by incorporating linkage between neighboring loci. We prove consistency and asymptotic normality of maximum likelihood estimators for the ancestry of individuals in the Linkage Model, complementing earlier results by (Pfaff et al., 2004; Pfaffelhuber and Rohde, 2022; Heinzel, 2025) for the Admixture Model. These results are used to prove that a statistical test that allows for model selection between the Admixture Model and the Linkage Model is an asymptotic level-α-test. Finally, we demonstrate the practical relevance of our results by applying the test to real-world data from The 1000 Genomes Project Consortium (2015).

Genetic Linkage

Interference of Competing Beneficial Mutations on Recombining Chromosomes.

Finding signatures of selective sweeps in genomes is a major goal of current population genomics, as it allows estimating the rate of beneficial mutations going to fixation and identifying the genes involved in selection. Models of recurrent selective sweeps traditionally assume that in chromosomal regions of normal recombination rates at most one beneficial allele is on the way to fixation. We review and extend here the theoretical studies on interference between closely linked beneficial mutations suggesting that this assumption may be violated. We show that interference between beneficial mutations may lead to substantially increased fixation times even in chromosomal regions of normal recombination rates. Furthermore, we discuss how interference can be detected in population genomic studies by analyzing genetic footprints of selective sweeps, and search for empirical evidence of interference in published datasets.

fixation times

On the origin of animals and placental mammals: a critique of literalist readings of the fossil record.

The fossil record is incomplete, as evidenced by the pervasive presence of ghost lineages throughout the Tree of Life. For example, across placental mammals, at least 720 Myr of basal lineages are ghost lineages, that is, lineages that have left no fossil evidence of their past history. In contrast, some studies have suggested that the fossil record is a faithful temporal archive of evolutionary history and thus the times of diversification of clades must be close to the ages of their oldest fossils. Such literalist interpretations have been contradicted by analysis of molecular datasets which, in many cases, indicate that groups including placental mammals and animals may have originated at times substantially older than their fossil records. Some of those studies have further argued that, in the case of animals and placental mammals, molecular clocks are uninformative, suffer from characteristic pathologies, and thus cannot distinguish between recent and ancient hypotheses of diversification. Here, we reexamine these two cases and show, using Bayesian model selection theory, that the explosive diversification models previously proposed for animals and placental mammals have a posterior probability of ∼0. We show the characteristic pathologies purportedly discovered do not exist, highlight errors in previous analyses, and provide advice on best practice for molecular-clock dating analysis.

Animals

Is a Win-Win possible? Achieving pareto-optimal privacy-utility balance in fine-tuned genome language model embeddings against embedding reconstruction attacks.

MOTIVATION: Genomic data is among the most sensitive categories of personal information, and the growing adoption of language models for sequence analysis raises significant privacy concerns. Prior work demonstrated that embeddings from general-purpose language models adapted for genomic sequences leak substantial single-nucleotide information under reconstruction attacks, and that fine-tuning embeddings can reduce this vulnerability at certain positions. However, three critical questions remain unaddressed: (i) whether privacy-utility tradeoffs are inherent constraints or configuration-dependent phenomena; (ii) whether genomic-specialized models such as DNABERT-base and Nucleotide Transformer exhibit different vulnerabilities than adapted general-purpose models; and (iii) how to statistically validate whether observed privacy improvements represent meaningful gains. Addressing these gaps is essential for guiding model selection in privacy-sensitive genomic applications. RESULTS: We systematically evaluated 13 transformer architectures, 9 general-purpose and 4 genomic-specialized, under position-specific embedding reconstruction attacks. We assessed the vulnerabilities of both pre-trained and fine-tuned models to the single-nucleotide inference-reconstruction attack using our new metrics, including error-based privacy gain and Pareto dominance scores, and statistically validated the results via paired t-tests. XLNet-Large achieved the best observed privacy protection among all evaluated models (+19.5% mean privacy gain) while maintaining competitive prediction performance. General-purpose models outperformed genomic-specialized models in 56% of pairwise comparisons. Tokenization strategy, rather than domain specialization, emerged as the primary determinant of the privacy-utility balance. These findings provide evidence-based guidance for selecting models in privacy-sensitive short-window genomic applications. All privacy claims in this work are specific to position-wise embedding reconstruction attacks and do not extend to other privacy risks, such as membership inference or training data extraction, which may respond differently to fine-tuning. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/AnonymousISCBConf/Win-Win-Privacy-Utility-Analysis.

Genomics

The molecular similarity landscape of preclinical cancer models to patient tumors.

Selecting appropriate preclinical models is fundamental for translational oncology, yet a large-scale, multi-omic quantitative comparison of their similarity to primary human tumors is lacking. To address this, we integrated transcriptomic, proteomic, and genomic profiles from over 10,000 primary tumors from The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC), alongside 4,000 preclinical models. Using a robust computational framework, we revealed a clear hierarchy of transcriptomic and proteomic similarity to patient tumors: with patient-dervied xenografts (PDXs) having greater transcriptomic and proteomic similarity to patient tumors (>) compared with patient-derived organoids (PDOs), which are equal in hierarchy to that of PDX-dervied organoids (PDXOs) > cell lines. We also quantified high molecular conservation (Pearson correlation coefficient = 0.96) across paired in vitro to in vivo platform (organoids to PDX) transitions. Furthermore, genomic analysis demonstrated that whole-exome sequencing (WES) outperforms RNA-seq in detecting DNA variants, and it identified a clonal complexity hierarchy (cell lines > PDXOs > PDXs > PDOs) reflecting the effect of passaging history on intratumor heterogeneity. Ultimately, this study delivers a comprehensive quantitative benchmark, establishing a population-level hierarchy of molecular similarity between preclinical models and primary tumors and providing a data-driven reference for model selection. These findings offer a data-driven framework for selecting models that balance biological representativeness with experimental practicality.

Humans

Exploration of predictive and prognostic alternative splicing signatures in lung adenocarcinoma using machine learning methods.

BACKGROUND: Alternative splicing (AS) plays critical roles in generating protein diversity and complexity. Dysregulation of AS underlies the initiation and progression of tumors. Machine learning approaches have emerged as efficient tools to identify promising biomarkers. It is meaningful to explore pivotal AS events (ASEs) to deepen understanding and improve prognostic assessments of lung adenocarcinoma (LUAD) via machine learning algorithms. METHOD: RNA sequencing data and AS data were extracted from The Cancer Genome Atlas (TCGA) database and TCGA SpliceSeq database. Using several machine learning methods, we identified 24 pairs of LUAD-related ASEs implicated in splicing switches and a random forest-based classifiers for identifying lymph node metastasis (LNM) consisting of 12 ASEs. Furthermore, we identified key prognosis-related ASEs and established a 16-ASE-based prognostic model to predict overall survival for LUAD patients using Cox regression model, random survival forest analysis, and forward selection model. Bioinformatics analyses were also applied to identify underlying mechanisms and associated upstream splicing factors (SFs). RESULTS: Each pair of ASEs was spliced from the same parent gene, and exhibited perfect inverse intrapair correlation (correlation coefficient = - 1). The 12-ASE-based classifier showed robust ability to evaluate LNM status of LUAD patients with the area under the receiver operating characteristic (ROC) curve (AUC) more than 0.7 in fivefold cross-validation. The prognostic model performed well at 1, 3, 5, and 10 years in both the training cohort and internal test cohort. Univariate and multivariate Cox regression indicated the prognostic model could be used as an independent prognostic factor for patients with LUAD. Further analysis revealed correlations between the prognostic model and American Joint Committee on Cancer stage, T stage, N stage, and living status. The splicing network constructed of survival-related SFs and ASEs depicts regulatory relationships between them. CONCLUSION: In summary, our study provides insight into LUAD researches and managements based on these AS biomarkers.

Adenocarcinoma of Lung

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals

Robotic assistance in total hip arthroplasty: a systematic review and meta-analysis of leg length, cup orientation, and early outcomes.

This review examined whether robotic assistance alters postoperative leg-length discrepancy (LLD), acetabular cup orientation, or early hip-specific outcomes relative to conventional total hip arthroplasty (THA). We searched PubMed and Web of Science through May 2026 for comparative English-language reports. Study eligibility, data extraction, and methodological appraisal were undertaken independently by two reviewers. Mean differences (MDs) and 95% confidence intervals (CIs) were calculated in Review Manager 5.4. Model selection was based on the target estimand and anticipated clinical and methodological diversity; leave-one-out and alternative-model sensitivity analyses were undertaken for heterogeneous outcomes. The protocol is registered with PROSPERO (CRD420261454043). The review included seven studies and 968 participants. Compared with conventional THA, robot-assisted THA yielded a smaller postoperative LLD (MD = -2.02, 95% CI -3.46 to -0.58; P = 0.006) and a higher Harris Hip Score (MD = 2.96, 95% CI 1.12 to 4.80; P = 0.002). Mean cup anteversion was lower in the robotic group (MD = -1.52, 95% CI -2.29 to -0.76; P < 0.0001), whereas cup inclination did not differ (MD = -0.71, 95% CI -3.26 to 1.83; P = 0.58). The robotic group also had higher Forgotten Joint Score (MD = 14.68, 95% CI 5.02 to 24.33; P = 0.003) and Oxford Hip Score values (MD = 2.61, 95% CI 0.71 to 4.51; P = 0.007). Robotic assistance was linked to a modest improvement in leg-length restoration and to higher scores on several early functional measures. The limited number of studies, predominance of nonrandomized designs, and marked heterogeneity in some analyses temper the certainty of these findings.

Humans

A multi-ancestry polygenic risk score for body mass index predicts longitudinal weight change.

BACKGROUND: Identifying individuals at risk for future weight gain is challenging, partly because associations with traditional clinical risk factors may be biased by confounding and reverse causation. Polygenic risk scores (PRS) provide a stable, lifelong measure of genetic predisposition to obesity. However, existing PRS have not been evaluated for their association with longitudinal weight change in adulthood and often lack generalizability across diverse genetic ancestry groups. METHODS: We conducted ancestry-specific genome-wide association study meta-analyses of body mass index (BMI) in populations of European, African or African American, Admixed American, East Asian, and South Asian ancestries and developed ancestry-specific PRS. A multi-ancestry polygenic risk score (MAPRS) was trained using ancestry-specific PRS in a model selection dataset (N&#x2009;=&#x2009;39,685) from the All of Us Research Program (AoU). We evaluated the MAPRS in an independent AoU model evaluation dataset (N&#x2009;=&#x2009;158,743) for BMI prediction and in a separate AoU test dataset (N&#x2009;=&#x2009;78,219) with repeated measurements over 1.5-2.5 years for weight change prediction. The outcomes included change in BMI and&#x2009;&#x2265;&#x2009;10% or&#x2009;&#x2265;&#x2009;5% total body weight (TBW) gain. We further examined the relationship between MAPRS and 12 clinical risk factors commonly comorbid with obesity in relation to weight change. RESULTS: The MAPRS captured 7.05% of the variance in measured BMI in the AoU model evaluation dataset and demonstrated improved generalizability across all non-European genetic ancestry groups. In the AoU test dataset, conditioned on baseline BMI at the second-to-last measurement, a one SD increase in MAPRS was associated with a 0.16 kg/m2 increase in future BMI (standard error&#x2009;=&#x2009;0.012 kg/m2; p-value&#x2009;=&#x2009;2.2&#x2009;&#xd7;&#x2009;10-39), 1.27-fold increased odds of experiencing&#x2009;&#x2265;&#x2009;10% TBW gain (95% CI: 1.24-1.31; p-value&#x2009;=&#x2009;1.4&#x2009;&#xd7;&#x2009;10-55), and 1.15-fold increased odds of experiencing&#x2009;&#x2265;&#x2009;5% TBW gain (95% CI: 1.13-1.18; p-value&#x2009;=&#x2009;2.8&#x2009;&#xd7;&#x2009;10-39). These associations were observed across all genetic ancestry groups and remained highly consistent after adjustment for any clinical risk factor. In contrast, most clinical risk factors demonstrated inconsistent or weaker associations with weight change outcomes. CONCLUSIONS: We developed an MAPRS for BMI that represents a robust and generalizable risk factor for longitudinal weight gain in adulthood, providing a foundation for genetically informed risk stratification and earlier, more targeted obesity prevention strategies.

Humans

Physical Appearance Anxiety and Eating Disorders Symptomatology: A Systematic Review and Meta-Analysis.

The present study aimed to assess the link between physical appearance anxiety (PAA) and eating disorder (ED) symptomatology by a meta-analysis of existing literature. Eligible studies were searched across six electronic databases up until November 20, 2025. Pooled effect sizes (r) were calculated using random-effects models. Potential variables that influence effect heterogeneity were analyzed by univariable and multivariable meta-regressions. Influence analyses and a three-parameter selection model (3PSM) were used to assess robustness of the results and publication bias. Twenty-seven effect sizes from 21 studies (N&#x2009;=&#x2009;5261) were obtained. The results indicated a strong association (i.e., r&#x2009;=&#x2009;0.559) between the two variables under consideration, which was notably stronger (i) among females compared to males; and (ii) for overall eating disorder symptoms rather than bulimic symptoms. The results of this study advocate for further investigation into the effectiveness of addressing anxiety responses related to personal body traits, particularly among females, within the context of preventing and treating eating disorders.

Humans

Decoding glioblastoma evolution and heterogeneity through mechanistic modeling: implications for clinical translation.

Glioblastoma (GBM) is one of the most aggressive and lethal primary brain tumors in adults, characterized by dynamic clonal evolution and extensive genomic, cellular, spatial, and microenvironmental heterogeneity. Multi-omics studies have revealed that GBM follows complex evolutionary trajectories involving genetic, epigenetic, transcriptional, and immune-microenvironmental remodeling as tumors grow, adapt to the brain microenvironment, and acquire therapeutic resistance. Increasing evidence suggests that GBM may originate from aberrant neural stem or progenitor cells, including those residing in the subventricular zone, and that glioblastoma stem cells (GSCs) contribute to tumor propagation, heterogeneity, and recurrence. A key conceptual challenge is to reconcile hierarchical cancer stem cell models, in which GSCs are viewed as relatively stable tumor-propagating subpopulations, with dynamic state plasticity models, in which stem-like properties can be reversibly acquired or lost during transitions among proneural-like, mesenchymal-like, invasive, and therapy-tolerant states. Recent advances in single-cell profiling, spatial transcriptomics, lineage tracing, organoid culture, 3D bioprinting, genetically engineered models, and artificial intelligence (AI)-assisted computational modeling have substantially improved the ability to study these processes. However, no currently available model fully recapitulates human GBM heterogeneity, recurrence, treatment history, and tumor-microenvironment interactions. Therefore, model selection should be guided by clearly defined mechanistic questions rather than by reliance on any single platform. This review summarizes current advances in in vitro, ex vivo, in vivo, and computational models for studying GBM evolution and heterogeneity, and discusses how integrated model pipelines may improve preclinical drug testing, treatment-response prediction, and precision neuro-oncology.

Humans

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95&#xa0;% CI 0.85-0.94; 95&#xa0;% prediction interval 0.62-0.98), with sensitivity of 0.80 (95&#xa0;% CI 0.77-0.83) and specificity of 0.87 (95&#xa0;% CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Optimal Control of Directional False Discovery Rates in Large-Scale Testing.

The high-throughput biomedical technology enables measurement of thousands of gene expression levels contemporaneously. A major task in analyzing these gene expression data is to identify both over-expressed and under-expressed genes. The popular two-group models select the non-null genes without further classifying them as overexpression or underexpression. Consequently, two-group decision rules are unable to constrain the numbers of falsely discovered over-expressed or under-expressed genes respectively. We propose a general three-group model that allows dependence between the test statistics and develop a decision rule that separately controls the two types of false discoveries. We show that the optimal decision rule in our three-group model has a special monotonic structure. By making use of this monotonic structure, we can linearize the two-directional false discovery rate constraints. We prove that our decision rule optimizes the expected number of true discoveries while controlling the proportions of falsely discovered over-expressed and under-expressed genes at desired levels simultaneously. The data-driven versions of the proposed procedures are suggested, and their consistency is established. Comparisons with state-of-the-art approaches and applications to genomic studies show that our procedures work well.

Humans

Zfp423 binds autoregulatory sites in p19 cell culture model.

Zfp423 is a 30 zinc finger transcription factor that forms regulatory complexes with EBF family members and factors targeted by canonical signaling pathways. Zfp423 mutations produce a range of developmental abnormalities in mice and humans related to the ciliopathies. Surprisingly, computational analysis of clustered Zfp423 and partner motifs in conserved genomic sequences predicts enrichment in Zfp423 and Ebf genes. In cell culture models selected for Zfp423 and EBF expression, we identify strong and reproducible occupancy of two Zfp423 intronic sites using chromatin immunoprecipitation with multiple independent antibodies. Both sites are significantly enriched in either quantitative PCR or massively parallel sequencing assays. A site in intron 5 acts as a classical enhancer in transient assays, but does not require the consensus motif for activity, suggesting a redundant or modulatory role for Zfp423 binding in this context. We speculate that Zfp423 may repress this enhancer as part of a developmental ratchet.

Animals

An individualized nomogram for predicting progression-free survival in systemic anaplastic large cell lymphoma: a multicenter, retrospective, and internally validated study.

OBJECTIVES: To develop an individualized nomogram for predicting disease progression risk in systemic anaplastic large cell lymphoma (sALCL). METHODS: Independent predictors of progression-free survival (PFS) were identified using Cox regression in a multicenter retrospective cohort of 109 sALCL patients (2010-2022). These were incorporated into a three-factor nomogram, evaluated via bootstrapped internal validation (1000 resamples), ROC analysis, C-index, decision curve analysis (DCA), and clinical impact curve (CIC). RESULTS: A total of 29 PFS events occurred during a median follow-up of 31 months. Multivariable modelling selected serum &#x3b2;2-microglobulin elevation, extranodal disease, and front-line chemotherapy choice (CHOP versus CHOPE or BV+CHP) as autonomous progression drivers. Upon internal bootstrap validation, the nomogram yielded strong prognostic accuracy, achieving AUCs of 0.81, 0.85 and 0.87 for 1-, 3- and 5-year progression-free survival, alongside a corrected C-index of 0.779 (95% CI: 0.699 - 0.861). Calibration plots showed close agreement between predicted and observed outcomes, while DCA confirmed superior net clinical benefit versus conventional IPI or Ann Arbor stratification across multiple decision thresholds. CONCLUSION: This first sALCL-specific nomogram integrates clinical and treatment variables to provide personalized PFS risk estimation. While internally validated, this exploratory, observation-based tool requires external validation and recalibration in prospective cohorts before clinical implementation.

Humans

An epigenome-wide study of selenium status and DNA methylation in the Strong Heart Study.

BACKGROUND: Selenium (Se) is an essential nutrient linked to adverse health endpoints at low and high levels. The mechanisms behind these relationships remain unclear and there is a need to further understand the epigenetic impacts of Se and their relationship to disease. We investigated the association between urinary Se levels and DNA methylation (DNAm) in the Strong Heart Study (SHS), a prospective study of cardiovascular disease (CVD) among American Indians adults. METHODS: Selenium concentrations were measured in urine (collected in 1989-1991) using inductively coupled plasma mass spectrometry among 1,357 participants free of CVD and diabetes. DNAm in whole blood was measured cross-sectionally using the Illumina MethylationEPIC BeadChip (850&#xa0;K) Array. We used epigenome-wide robust linear regressions and elastic net to identify differentially methylated cytosine-guanine dinucleotide (CpG) sites associated with urinary Se levels. RESULTS: The mean (standard deviation) urinary Se concentration was 51.8 (25.1) &#x3bc;g/g creatinine. Across 788,368 CpG sites, five differentially methylated positions (DMP) (hypermethylated: cg00163554, cg18212762, cg11270656, and hypomethylated: cg25194720, cg00886293) were significantly associated with Se in linear regressions after accounting for multiple comparisons (false discovery rate p-value: 0.10). The top hypermethylated DMP (cg00163554) was annotated to the Disco Interacting Protein 2 Homolog C (DIP2C) gene, which relates to transcription factor binding. Elastic net models selected 425 hypo- and hyper-methylated DMPs associated with urinary Se, including three sites (cg00163554 [DIP2C], cg18212762 [MAP4K2], cg11270656 [GPIHBP1]) identified in linear regressions. CONCLUSIONS: Urinary Se was associated with minimal changes in DNAm in adults from American Indian communities across the Southwest and the Great Plains in the United States, suggesting that other mechanisms may be driving health impacts. Future analyses should explore other mechanistic biomarkers in human populations, determine these relationships prospectively, and investigate the potential role of differentially methylated sites with disease endpoints.

Humans