Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “penalized regression”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

PROLONG: penalized regression for outcome guided longitudinal omics analysis with network and group constraints.

MOTIVATION: There is a growing interest in longitudinal omics data paired with some longitudinal clinical outcome. Given a large set of continuous omics variables and some continuous clinical outcome, each measured for a few subjects at only a few time points, we seek to identify those variables that co-vary over time with the outcome. To motivate this problem we study a dataset with hundreds of urinary metabolites along with Tuberculosis mycobacterial load as our clinical outcome, with the objective of identifying potential biomarkers for disease progression. For such data clinicians usually apply simple linear mixed effects models which often lack power given the low number of replicates and time points. We propose a penalized regression approach on the first differences of the data that extends the lasso + Laplacian method [Li and Li (Network-constrained regularization and variable selection for analysis of genomic data. Bioinformatics 2008;24:1175-82.)] to a longitudinal group lasso + Laplacian approach. Our method, PROLONG, leverages the first differences of the data to increase power by pairing the consecutive time points. The Laplacian penalty incorporates the dependence structure of the variables, and the group lasso penalty induces sparsity while grouping together all contemporaneous and lag terms for each omic variable in the model. RESULTS: With an automated selection of model hyper-parameters, PROLONG correctly selects target metabolites with high specificity and sensitivity across a wide range of scenarios. PROLONG selects a set of metabolites from the real data that includes interesting targets identified during EDA. AVAILABILITY AND IMPLEMENTATION: An R package implementing described methods called "prolong" is available at https://github.com/stevebroll/prolong. Code snapshot available at 10.5281/zenodo.14804245.

Humans↗

Pretreatment EBV-DNA/TLG-Based Risk Stratification Is Associated With Survival Outcomes in Nonmetastatic Nasopharyngeal Carcinoma: An Exploratory Study.

Whether combining pretreatment plasma Epstein-Barr virus DNA (EBV-DNA) with 18F-FDG PET/CT-derived total lesion glycolysis (TLG) improves prognostic stratification in nonmetastatic nasopharyngeal carcinoma (NPC) is unclear, particularly in nonendemic populations. We retrospectively analyzed 86 eligible nonmetastatic NPC patients treated with definitive radiotherapy (2010-2024) at a single nonendemic-region institution. EBV-DNA (prespecified cutoff 3500 copies/mL) and TLG (cutoff 200, ROC-derived within this cohort) were dichotomized. Both were available in 59/86 patients (68.6%), who differed from the rest in nodal and overall stage and in RT technique. Baseline PET/CT was in-house in 57 of 86 patients, and a robustness analysis in that subgroup is reported. Given limited events (13 PFS, 9 OS), Cox analyses are exploratory and were supplemented with penalized regression and bootstrap validation. At a median follow-up of 75.5 months, 5-year PFS and OS for the whole cohort (n = 86) were 81.1% and 85.9%. The EBV-DNAhigh/TLGhigh subgroup remained associated with inferior PFS after adjustment in an exploratory model (adjusted HR = 3.97, 95% CI: 1.32-11.93) and, in a single-variable model, with inferior OS (HR = 4.13, 95% CI: 1.10-15.52). Discrimination was comparable to the individual-biomarker model for PFS and lower for OS. Five-year PFS fell monotonically across the four risk groups in the complete-case cohort (n = 59; 89.7%-58.3%). OS differed across groups (log-rank p = 0.044) but was not strictly monotonic, with wide, overlapping confidence intervals. This two-biomarker model is hypothesis-generating and needs prospective, multicenter validation before any consideration of risk-adapted treatment.

Epstein–Barr virus DNA↗

Identification of a novel human gut microbes and microbial metabolites related genes signature for prognostic implication in head and neck squamous carcinomas.

BACKGROUND: The gut microbiota acts as a critical driver influencing the pathogenesis, therapeutic response, and clinical outcomes across various cancer types. This study aimed to investigate the prognostic value of human gut microbes and microbial metabolites related genes (HGMMMRGs) in head and neck squamous cell carcinoma (HNSCC). METHODS: We constructed a prognostic risk model comprising 19 core HGMMMRGs using LASSO penalized regression and a multivariate Cox proportional hazards model. The predictive performance of the model was evaluated through Kaplan-Meier analysis, receiver operating characteristic (ROC) curves, nomograms, and concordance index. In addition, functional enrichment analysis was performed on the differentially expressed risk genes. Furthermore, the relationship between the immune microenvironment of HNSCC and the risk diagnostic model was examined. Western blot analysis was used to assess the expression levels of IL10 in both HNSCC tissues and adjacent normal tissues. Finally, the correlation between IL10 and the gut microbiota was explored. RESULTS: This study developed a risk score model integrating 19 HGMMMRG genes, which can serve as a tool to guide prognosis and immune microenvironment assessment in HNSCC patients. Survival analysis showed that patients in the high-risk group had significantly worse outcomes (P&#x2009;<&#x2009;0.05). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis revealed significant enrichment of differentially expressed genes (DRLs) and immune-related pathways. Western blot analysis further confirmed that IL10 was highly expressed in HNSCC, and the abundance of Faecalibacterium prausnitzii and Enterococcus durans colonies was correlated with IL10 expression. CONCLUSION: We developed a prognostic model for HGMMMRGs that can be effectively used to predict OS in patients with HNSCC. Second, Faecalibacterium prausnitzii and Enterococcus durans can influence the prognosis of patients with HNSCC by mediating the expression IL10 and thereby affecting the prognosis of HNSCC patients. Thus, human gut microbes and microbial metabolite-related genes may be another promising strategy for the treatment of patients with HNSCC.

HNSCC↗

Multi-Ancestry Genome-Wide Association with Fine-Mapping Identifies Novel Loci for Pigment Dispersion Syndrome and Pigmentary Glaucoma.

PURPOSE: Pigment dispersion syndrome and pigmentary glaucoma are important causes of ocular hypertension and glaucomatous optic neuropathy, yet their genetic determinants remain incompletely defined, particularly across diverse ancestries. This study aimed to use a large multi-ancestry cohort from the All of Us Research Program to investigate the genetic basis of pigment dispersion syndrome and pigmentary glaucoma. DESIGN: Case-control study. PARTICIPANTS: In total, 572 cases and 37&#x2009;808 controls with array genotyping and 537 cases and 35&#x2009;493 controls with whole-genome sequencing. METHODS: Using electronic health record phenotyping in the All of Us Research Program, we performed multi-ancestry genome-wide association analyses using both array-based data and whole-genome sequencing-based data, comparing patients with pigment dispersion syndrome or pigmentary glaucoma to those without either condition. We also performed Firth penalized regression and Fisher analyses, and we performed principal component analyses to assess effect sizes across genetic ancestries. We applied statistical fine-mapping, examined for cross-trait overlap, and assessed expression quantitative trait locus associations for lead variants. MAIN OUTCOME MEASURES: P values and odds ratios of lead loci from genome-wide association analyses; size of credible sets determined from fine-mapping; allele frequency of lead variants in cases, controls, and the general population; expression quantitative trait loci effect size and P values linking lead variants to gene expression. RESULTS: We identified 4 loci reaching genome-wide significance across analyses, including signals near EPHA7 (which mediates cell-cell signaling), within TYR (involved in melanin synthesis and replicated from prior studies), within LINC01138, and near OTX2. Statistical fine-mapping refined 3 of these loci to single-variant 95% credible sets and narrowed the TYR locus to small credible sets, prioritizing possible causal variants. Effect estimates were broadly consistent across genetic ancestry clusters. Lead variants showed regulatory evidence in expression quantitative trait locus, including reduced EPHA7 expression. CONCLUSIONS: These findings implicate both melanogenesis and cell-cell adhesion and signaling pathways in pigment dispersion syndrome and pigmentary glaucoma. FINANCIAL DISCLOSURE(S): Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Genome-wide association study↗

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance↗

Predicting Weight Loss After Vertical Sleeve Gastrectomy Using a Whole-genome Sequencing-derived Polygenic Risk Score in the All of Us Cohort.

OBJECTIVE: To create a genome-wide polygenic risk score (PRS) to improve prediction of a 12-month percentage weight loss (WL) after vertical sleeve gastrectomy (VSG). BACKGROUND: Variability in post-VSG WL is not well explained by clinical factors. The All of Us program provides access to a 414,830 short-read whole-genome sequencing resource, enabling unbiased discovery of genetic predictors after VSG. METHODS: VSG counts, demographic, anthropomorphic and vital sign information were obtained from the linked electronic health record. The discovery cohort (DC) included participants from version 7 carried into version 8 while the validation cohort (VC) included those newly added to v8. We defined good responders and nonresponders as having WL&#xb1;1SD from the mean. Following quality filtering, we applied a 2-stage penalized-regression, followed by elastic-net logistic regression, to identify 1583 stable variants and derive &#x3b2;-weights. We then tested this PRS on the DC into a prediction model. RESULTS: We identified 395 participants in the DC and 336 participants in the VC, respectively. Of these, VSG, 44 were classified as good responders (&#x2265;37% WL) and 55 as nonresponders (&#x2264;19% WL). In the VC, 55 were classified as good responders and 48 as nonresponders. Adding the PRS to models to clinical predictors increased the area under the curve following logistic regression by 0.03; P <4.3 &#xd7; 10 -14 , random forest by 0.03; P <9.1 &#xd7; 10 -7 , decision tree by 0.05; P = 1.2 &#xd7; 10 -3 , and gradient boosting by 0.08; P <8.3 &#xd7; 10 -10 . CONCLUSIONS: Use of short-read whole-genome sequencing from All of Us (AoU) can be effectively used to generate PRS to enhance predictive WL accuracy. This work has implications for outcomes of both bariatric surgery and other surgical procedures.

Humans↗

A SuperLearner-based pipeline for the development of DNA methylation-derived predictors of phenotypic traits.

BACKGROUND: DNA methylation (DNAm) provides a window to characterize the impacts of environmental exposures and the biological aging process. Epigenetic clocks are often trained on DNAm using penalized regression of CpG sites, but recent evidence suggests potential benefits of training epigenetic predictors on principal components. METHODOLOGY/FINDINGS: We developed a pipeline to simultaneously train three epigenetic predictors; a traditional CpG Clock, a PCA Clock, and a SuperLearner PCA Clock (SL PCA). We gathered publicly available DNAm datasets to generate i) a novel childhood epigenetic clock, ii) a reconstructed Hannum adult blood clock, and iii) as a proof of concept, a predictor of polybrominated biphenyl exposure using the three developmental methodologies. We used correlation coefficients and median absolute error to assess fit between predicted and observed measures, as well as agreement between duplicates. The SL PCA clocks improved fit with observed phenotypes relative to the PCA clocks or CpG clocks across several datasets. We found evidence for higher agreement between duplicate samples run on alternate DNAm arrays when using SL PCA clocks relative to traditional methods. Analyses examining associations between relevant exposures and epigenetic age acceleration (EAA) produced more precise effect estimates when using predictions derived from SL PCA clocks. CONCLUSIONS: We introduce a novel method for the development of DNAm-based predictors that combines the improved reliability conferred by training on principal components with advanced ensemble-based machine learning. Coupling SuperLearner with PCA in the predictor development process may be especially relevant for studies with longitudinal designs utilizing multiple array types, as well as for the development of predictors of more complex phenotypic traits.

DNA Methylation↗

Penalized partial likelihood regression for right-censored data with bootstrap selection of the penalty parameter.

The Cox proportional hazards model is often used for estimating the association between covariates and a potentially censored failure time, and the corresponding partial likelihood estimators are used for the estimation and prediction of relative risk of failure. However, partial likelihood estimators are unstable and have large variance when collinearity exists among the explanatory variables or when the number of failures is not much greater than the number of covariates of interest. A penalized (log) partial likelihood is proposed to give more accurate relative risk estimators. We show that asymptotically there always exists a penalty parameter for the penalized partial likelihood that reduces mean squared estimation error for log relative risk, and we propose a resampling method to choose the penalty parameter. Simulations and an example show that the bootstrap-selected penalized partial likelihood estimators can, in some instances, have smaller bias than the partial likelihood estimators and have smaller mean squared estimation and prediction errors of log relative risk. These methods are illustrated with a data set in multiple myeloma from the Eastern Cooperative Oncology Group.

Antineoplastic Agents↗

Penalized likelihood in Cox regression.

In a Cox regression model, instability of the estimated regression coefficients can be reduced by maximizing a penalized partial log-likelihood, where a penalty function of the regression coefficients is substracted from the partial log-likelihood. In this paper, we choose the optimal weight of the penalty function by maximizing the predictive value of the model, as measured by the crossvalidated partial log-likelihood. Our methods are illustrated by a study of ovarian cancer survival and by a study of centre-effects in kidney graft survival.

Clinical Trials as Topic↗

Identifying Single-Cell Expression Quantitative Trait Loci Using a Bootstrap Penalized Hurdle Model.

BACKGROUND: Expression quantitative trait loci (eQTL) analysis links genetic variants to gene expression levels, helping to uncover how genetic variation contributes to gene regulation. While traditional eQTL analyses rely on bulk RNA-seq data, recent advances in single-cell RNA sequencing (scRNA-seq) have made it possible to detect cell-type-specific eQTLs. However, the inherent sparsity and heterogeneity of scRNA-seq data present major challenges for standard modeling approaches. METHODS: In this paper, we propose a novel statistical framework, Bootstrap Penalized Hurdle regression model (BPHurdle), designed specifically for scRNA-seq data. BPHurdle employs a hurdle modeling framework, where a logistic component accounts for the excess zeros in single-cell expression data, and a Poisson component jointly evaluates the effects of multiple SNPs on positive gene expression levels. RESULTS: Through simulation studies, we show that BPHurdle achieves high accuracy and robustness in identifying regulatory variants. We further demonstrate its utility on a real dataset through a case study focusing on a subset of differentially expressed genes, where it successfully identifies reliable cell-type-specific eQTLs. CONCLUSIONS: Overall, BPHurdle offers an advanced and flexible approach for single-cell eQTL mapping, providing deeper insight into the genetic regulation of gene expression at cellular resolution.

Quantitative Trait Loci↗

Modelling the effects of biological intervention in a dynamical gene network.

Cellular response to environmental and internal signals can be modeled by dynamical gene regulatory networks (GRN). In the literature, three main classes of gene network models can be distinguished: (1) non-quantitative (or data-based) models which do not describe the probability distribution of gene expressions; (2) quantitative models which fully describe the probability distribution of all genes co-expression; and (3) mechanistic models which allow for a causal interpretation of gene interactions. We propose two rigorous frameworks to model gene alteration in a dynamical GRN, depending on whether the network model is quantitative or mechanistic. We explain how these models can be used for design of experiment, or, if additional alteration data are available, for validation purposes or to improve the parameter estimation of the original model. We apply these methods to the Gaussian graphical model, which is quantitative but non-mechanistic, and to mechanistic models of Bayesian networks and penalized linear regression.

Gene Regulatory Networks↗

Safety profiles of CAR-T cell therapy in systematic autoimmune diseases: a systematic review and analysis.

BACKGROUND: Chimeric antigen receptors (CARs)-T cell therapy is emerging as a potent approach for autoimmune diseases. However, its application in autoimmune conditions remains limited, and safety outcomes observed in malignancies can't reliably serve as a reference. Therefore, it's necessary to summarize the safety profiles in autoimmune diseases to provide evidence for future expanding trials. METHODS: A systematic review was conducted to analyze the CAR-T therapy safety in rheumatic diseases via database searches up to December 2025. Studies reporting safety data were included, while abstracts, reviews, and cases with malignancies were excluded. Factors associated with cytokine release syndrome (CRS) were analyzed using Firth's penalized logistic regression. RESULTS: This study included 38 studies, involving a total of 115 patients with autoimmune disease. Severe adverse events were rare. CRS and immune effector cell-associated neurotoxicity syndrome (ICANS) occurred in 70.4% and 4.3% of patients, respectively. Most CRS were low-grade. Multivariate analysis identified BCMA-targeted therapy and allogeneic CAR-T products may as independent factors associated with a reduced risk of CRS. Transient hematologic toxicity and hypogammaglobulinemia were frequently reported, with infections occurring in nearly half of the patients. However, prolonged cytopenia and severe infection were infrequent. CONCLUSION: Based on the current available evidence, CAR-T therapy appears to have a generally manageable safety profile in autoimmune diseases, supporting its potential as a promising treatment option for patients with relapsed or refractory autoimmune diseases. However, these findings remain preliminary, and further expanded studies are warranted in the future to provide higher-level evidence.

Humans↗

Proteomic signatures for sudden cardiac death and related intermediate phenotypes.

BACKGROUND: Novel markers for sudden cardiac death (SCD) are needed. OBJECTIVE: This study aimed to explore whether a protein risk score derived from a large-scale proteomics dataset improves risk prediction of SCD in the general population. METHODS: A total of 52,705 individuals with 1459 unique plasma protein measurements were included from the UK Biobank Pharma Proteomics Project. A protein risk score was developed using lasso-penalized Cox regression on 40,722 participants enrolled at the English centers and validated on 11,983 participants enrolled at the remaining centers. RESULTS: The protein risk score formula developed from the derivation set comprised 64 unique plasma proteins including latent-transforming growth factor beta-binding protein 2, protein tyrosine phosphatase receptor sigma, and spondin-1. In the test set, a per standard deviation increase in protein risk score was associated with a hazard ratio of 2.60 (95% confidence interval [CI] 2.12-3.18) for SCD. Adding a protein risk score to SCD clinical risk factors resulted in a concordance index increase of 0.063 (95% CI 0.037-0.105) for SCD. For ventricular arrhythmia-mediated SCDs, an increase in concordance index when a protein risk score was added to SCD clinical risk factors was 0.070 (95% CI 0.010-0.188). A protein risk score added to SCD clinical risk factors resulted in a risk reclassification of 16.9% (95% CI 9.0-24.7) at a 10-year risk threshold of 5%. A protein risk score was significantly associated with intermediate phenotypes of SCD including corrected QT prolongation, an increase in left ventricular mean myocardial thickness, and a decrease in left ventricular global longitudinal strain. CONCLUSION: A protein risk score derived from a single plasma sample significantly improved risk prediction of SCD and related intermediate phenotypes.

Humans↗

Prevalence and Prognostic Significance of Exercise-Accentuated J-Point Elevation in Brugada Syndrome.

BACKGROUND: The clinical significance of accentuation of precordial J-point elevation during exercise in Brugada syndrome (BrS) remains unclear. OBJECTIVES: This study sought to determine the prevalence and prognostic significance of accentuation of J-point elevation during exercise in a large single-center BrS cohort. METHODS: In this retrospective study, 141 consecutive patients referred for BrS evaluation (95 with type 1 BrS pattern-BrS1 cohort, 46 without type 1 pattern but with loss-of-function sodium voltage-gated channel alpha subunit 5 variants-SCN5A cohort) who underwent exercise stress testing (EST) from January 1, 2000, through October 31, 2025 were included. Two blinded cardiologists reviewed all tracings. An exercise-accentuated BrS phenotype was defined as J-point elevation increase &#x2265;1 mm in V1/V2 during exercise. Cardiac events included arrhythmic syncope, cardiac arrest, and appropriate implantable cardioverter-defibrillator shocks. Firth penalized logistic regression was used for unadjusted and adjusted analyses. RESULTS: Overall, 41 patients (29%) demonstrated an exercise-accentuated BrS phenotype, emerging near peak exercise (median 90% age-predicted maximum heart rate). The phenotype was highly reproducible on serial testing (88% of follow-up ESTs). Exercise-accentuated phenotype was not associated with overall cardiac events (unadjusted OR: 1.83 [0.84-3.99]; P = 0.13; adjusted OR: 1.22 [0.47-3.21]; P = 0.69). However, it was strongly associated with exertion-triggered cardiac events (unadjusted OR: 10.50 [2.94-37.50]; P < 0.001; adjusted OR: 10.04 [2.76-36.53]; P < 0.001), independent of sex, exercise workload, and baseline type 1 pattern. Consistent results were noted in SCN5A variant-positive patients. CONCLUSIONS: Exercise-induced accentuation of J-point elevation reproducibly identifies a subset of BrS patients at risk for exertional cardiac events. These findings support the inclusion of EST in the evaluation of patients with a clinical diagnosis or genetic susceptibility to BrS and may inform exercise-related risk counseling.

Brugada syndrome↗

Early epidemiologic and genomic insights from Sierra Leone's first mpox cases, 2025.

INTRODUCTION: early characterization of outbreak cases supports rapid decisions. We conducted real-time analysis during the initial outbreak phase in mid-March 2025 of Sierra Leone's first 44 laboratory-confirmed Mpox cases (10th January to 5th March 2025) to generate epidemiologic and genomic intelligence for response. METHODS: we summarized surveillance data and sequenced 18 early cases with Oxford Nanopore. Firth-penalized logistic regression was employed to explore risk factors for severe disease. RESULTS: median age was 27 years (interquartile range 22 to 35); 68.2% were male. Most cases reported no recent international travel (95.5%). All cases presented with rash; fever occurred in 72.7%. Six cases (13.6%) met severe criteria; no deaths occurred (0 of 44; 95% confidence interval 0 to 8.0). Household secondary attack rate was 7.3% during the study window and 8.2% after completion of follow-up. Sequencing identified two co-circulating sub-lineages consistent with A.2.2 and B.1.6 and a strong APOBEC3 pattern (142 of 212 guanine-to-adenine in thymine-cytosine versus 70 of 212 cytosine-to-thymine in guanine-adenine; exact binomial p approximately 8.6x10-7; X2= 24.5, df = 1, p&#x2248; 7.6x10-7). Root-to-tip analysis showed weak temporal signal (R2=0.31; date randomization p= 0.18), so we did not interpret clock estimates. CONCLUSION: real-time analysis during the initial outbreak phase showed community transmission, quantified household spread, documented two sub-lineages with an APOBEC3 signature, and generated severity hypotheses that immediately informed surveillance and vaccine prioritization. Findings require validation in larger cohorts.

Adult↗

Identification of candidate variants in plasma associated with early versus late disease progression under anti-PD-1 therapy in metastatic NSCLC.

BACKGROUND: Immune checkpoint inhibitors (ICIs), including anti-programmed cell death protein 1 (anti-PD-1) antibodies, have significantly improved outcomes in patients with metastatic non-small cell lung cancer (mNSCLC). However, substantial heterogeneity exists in clinical benefit, with some patients exhibiting early progression (EP) and others late progression (LP). To date, no biomarkers of EP versus LP disease have been implemented in clinical practice. Circulating tumor DNA (ctDNA) analysis represents a minimally invasive strategy for identifying such biomarkers. In this proof-of-concept study, we evaluated the performance of the TruSight Oncology 500 ctDNA (TSO500 ctDNA) panel and explored its feasibility to identify candidate variants associated with early and late disease progression under anti-PD-1 therapy. METHODS: Baseline ctDNA from eight mNSCLC patients treated with pembrolizumab was extracted and sequenced using the TSO500 ctDNA assay, a 523-gene targeted next-generation sequencing panel. Patients were classified according to their response as LP or EP. Variant calling was performed using the DRAGEN Bio-IT platform, and variants were annotated and clinically interpreted using the Clinical Genomics Workspace (CGW; PierianDx) according to Association for Molecular Pathology (AMP)/American Society of Clinical Oncology (ASCO)/College of American Pathologists (CAP) guidelines. Survival outcomes were assessed using Kaplan-Meier and log-rank tests. Performance of ctDNA variants was evaluated using receiver operating characteristic (ROC) curve analysis, and multi-gene models were assessed using leave-one-out cross-validation with penalized logistic regression. RESULTS: All patients harbored detectable variants, including SNVs (100%), MNVs (87.5%), deletions (75%), and insertions (62.5%). Tier I variants were identified in 37.5% of patients, while all cases showed tier II and multiple tier III alterations. TP53 variants were associated with poorer outcomes under anti-PD-1 therapy. Individual gene alterations in TP53, ERBB3, SMC1A or LATS1 showed moderate discriminatory performance between LP and EP patients; however, combination of mutated genes improved apparent discrimination. Notably, specific two-gene combinations (SMC1A + LATS1 or ERBB3 + LATS1) showed the highest discriminatory performance between LP and EP patients in this exploratory cohort. CONCLUSIONS: This study demonstrates the feasibility and analytical performance of the TSO500 ctDNA panel and provides hypothesis-generating evidence that plasma gene variants may be useful to evaluate early versus late disease progression in patients with mNSCLC receiving immunotherapy.

TruSight Oncology 500↗

Nonparametric regression sinogram smoothing using a roughness-penalized Poisson likelihood objective function.

We develop and investigate an approach to tomographic image reconstruction in which nonparametric regression using a roughness-penalized Poisson likelihood objective function is used to smooth each projection independently prior to reconstruction by unapodized filtered backprojection (FBP). As an added generalization, the roughness penalty is expressed in terms of a monotonic transform, known as the link function, of the projections. The approach is compared to shift-invariant projection filtering through the use of a Hanning window as well as to a related nonparametric regression approach that makes use of an objective function based on weighted least squares (WLS) rather than the Poisson likelihood. The approach is found to lead to improvements in resolution-noise tradeoffs over the Hanning filter as well as over the WLS approach. We also investigate the resolution and noise effects of three different link functions: the identity, square root, and logarithm links. The choice of link function is found to influence the resolution uniformity and isotropy properties of the reconstructed images. In particular, in the case of an idealized imaging system with intrinsically uniform and isotropic resolution, the choice of a square root link function yields the desirable outcome of essentially uniform and isotropic resolution in reconstructed images, with noise performance still superior to that of the Hanning filter as well as that of the WLS approach.

Algorithms↗

Composition-on-composition regression analysis for multi-omics integration of metagenomic data.

MOTIVATION: Compositional data are frequently encountered in many disciplines, such as in next-generation sequencing experiments widely used in biomedical studies. Regression analysis with compositional data as either responses or predictors has been well studied. However, when both responses and predictors are compositional, the inventory of analysis tools is surprisingly limited, especially in the high-dimensional setting. Among the few existing methods, most of them rely on a log-ratio transformation to move compositional data from the simplex to real numbers. Yet, a serious weakness of these methods is their failure to handle the substantial fraction of zeroes observed in data collected from next-generation sequencing experiments. RESULTS: To investigate associations between two high-dimensional multi-omics compositions, we propose a composition-on-composition (COC) regression analysis method which does not require log-ratio transformations and hence can handle zeroes in the data. To account for high dimensionality, we estimate regression coefficients using a penalized estimation equation approach. Finally, inference procedures for COC regression are also proposed. Superior performance of COC is demonstrated through both comprehensive numerical simulations and case studies. AVAILABILITY AND IMPLEMENTATION: Source R codes to implement COC method is available at https://github.com/nrios4/COC.

Regression Analysis↗