Search PubMedSearch

SEARCH · Search PubMed

Results for “genomic features”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Robust prioritization of genomic features with stability selection.

MOTIVATION: The heterogeneity of complex diseases including cancer leads to heavy-tailed distributions in the disease traits. In such settings, non-robust variable selection methods are inherently susceptible to data contamination and can yield unstable or misleading results. This vulnerability becomes more severe for recently proposed approaches that introduce pseudo-features as negative controls, as these methods further amplify the curse of dimensionality by expanding the genotype matrix in the presence of outliers and high-dimensional genomic features. RESULTS: We develop a robust variable selection framework with stability selection to prioritize genomic features in the presence of contamination. In contrast to existing approaches that rely on pseudo-features for error control, the proposed method achieves double robustness. First, it adopts least absolute deviation (LAD) LASSO to ensure robustness against outliers and heavy-tailed errors in disease traits. Second, it avoids augmenting the genotype matrix with pseudo-features, thereby mitigating the curse of dimensionality that is particularly problematic in high-dimensional genomic data. The proposed method has been extensively evaluated in simulation studies to demonstrate its effectiveness over multiple competing methods for variable selection. In addition, we have applied the proposed method and competing approaches to two real-data case studies: the The Cancer Genome Atlas (TCGA) Skin Cutaneous Melanoma (SKCM) dataset and an eQTL dataset. The results demonstrate that the proposed method achieves superior performance by identifying genomic features with higher reproducibility. AVAILABILITY AND IMPLEMENTATION: The source code for implementing the proposed methods is publicly available at https://github.com/cenwu/RSS with an archival DOI https://doi.org/10.6084/m9.figshare.32306883.

Genomics

Identification of genomic features that uniquely impact estrogen receptor alpha binding and its effects on gene expression in endometrial cancer.

Estrogen receptor 1 (ESR1, also known as estrogen receptor alpha or ER) is an established oncogenic transcription factor in breast and endometrial cancer; however, more is known about the mechanisms controlling ER behavior in breast cancer, and therapies targeting ER have been much more successful in breast cancer. To address this disparity, we characterize the genomic features that control ER in endometrial cancer and determine to what extent these factors differ from those in breast cancer. We focus on the locations of estrogen response elements (EREs), ER's preferred DNA-binding motif, throughout the human genome. To identify factors that predict ER genomic binding and effects on target gene expression, we apply machine learning to genomic data for each ERE in Ishikawa cells (ER-positive endometrial cancer) and T-47D cells (ER-positive breast cancer). Many of these factors, such as chromatin accessibility and histone modifications, are predictive of ER activity in both cell lines. However, the transcription factors that predict ER activity are cell type specific, including FOXA1 and GATA3 in T-47D cells and ETV4 and SOX17 in Ishikawa cells. In addition, the features that predict ER binding and effects on gene expression differ, with transcription at EREs in the absence of estrogen being predictive of ER regulatory activity. A CRISPR knockout screen in Ishikawa cells, as well as follow-up experiments, confirms the discovery that SOX17 controls ER activity in endometrial cancer cells. These results identify important genomic features of ER binding and regulatory activity and how these features differ between endometrial cancer and breast cancer cells.

Humans

Integrating clinical and genomic features to predict response to neoadjuvant therapy in microsatellite-stable rectal cancer.

BACKGROUND: Neoadjuvant therapy (NAT) has shifted rectal cancer management toward organ preservation. However, achieving a complete response (CR) for "watch-and-wait" strategies is hindered by high response heterogeneity. Although immunotherapy-combined NAT has expanded the candidate pools, the predictive significance of molecular alterations remains unclear. OBJECTIVES: This study aimed to evaluate clinical and genomic profiles of rectal cancer patients undergoing NAT to identify response predictors and to develop a nomogram for estimating CR probability. DESIGN: Retrospective, single-center cohort study. METHODS: This study included 437 patients with rectal adenocarcinoma at Fudan University Shanghai Cancer Center between December 2019 and March 2023. Patients underwent paired tumor and germline genomic sequencing (887-gene panel) before NAT. Logistic and Cox regression analyses were performed to identify clinical and genetic risk factors associated with tumor response and long-term survival. RESULTS: Of the 437 patients, 96.6% had microsatellite-stable (MSS) tumors. In the MSS locally advanced rectal cancer cohort (N = 307), the CR rate was 35.5%. Multivariate analysis identified immunotherapy-combined NAT (iTNT) (OR 4.41, 95% CI: 2.42-8.27), SYNE1 mutation (OR 2.12, 95% CI: 1.06-4.26), negative mesorectal fascia (MRF) status (OR 0.34, 95% CI: 0.17-0.66), and lower tumor location (OR 0.48, 95% CI: 0.27-0.84) as independent predictors of CR. KRAS mutation was the sole independent predictor of reduced disease-free survival (DFS; HR 1.93, 95% CI: (1.11-3.36), p = 0.020). KRAS G12D subtype was associated with the worst 2-year distant metastasis-free survival (71.3%) and exhibited a distinct predilection for lung metastasis. The clinical-genomic nomogram yielded strong discrimination (AUC = 0.705) and calibration, with favorable DCA net benefit. CONCLUSION: Clinical and genomic features jointly determine outcomes in MSS rectal cancer. SYNE1 mutation serves as a novel biomarker for CR, while KRAS mutations, especially the G12D subtype, identify patients at high risk for systemic relapse. The clinical-genomic nomogram facilitates individualized selection for organ-preservation strategies.

biomarker

High MGMT expression identifies aggressive colorectal cancer with distinct genomic features and immune evasion properties.

INTRODUCTION: The epigenetic silencing of O6-methylguanine DNA methyltransferase (MGMT) is associated with reduced DNA repair capacity, carcinogenesis and increased sensitivity to alkylating chemotherapy. However, the biological role and clinical significance of MGMT overexpression in cancer remains poorly understood. METHODS: Using multiplexed quantitative immunofluorescence we measured the localized levels of MGMT protein, γH2AX and CD8+ T cells in multiple retrospective colorectal cancer (CRC) cohorts. Genomic and transcriptomic features of selected cases were also studied with whole exome DNA sequencing and genome-wide methylation analysis. MGMT-methylated human CRC cells SW620 were transfected with an MGMT-containing plasmid and co-cultured with allogeneic peripheral blood mononuclear cells. RESULTS: A subset of CRCs showed MGMT protein upregulation associated with lower γH2AX, reduced CD8+ tumor infiltrating lymphocytes (TILs), mismatch repair proficient (pMMR) status and shorter survival. CD8+ TILs were more distant from MGMT-expressing cells than MGMT-negative cells and the MGMT promoter methylation status did not highly correlate with MGMT protein levels in CRC. In genomic/transcriptomic analysis, high MGMT expression was associated with a lower nonsynonymous somatic mutational burden, higher transition-to-transversion mutation ratio, increased deleterious TP53 variants and distinct transcriptomic profiles. The exogenous expression of MGMT in SW620 CRC cells reduced the number of spontaneous nonsynonymous mutations, reproduced mutational features of MGMT-high CRC and limited the in vitro T-cell-mediated killing of malignant cells induced by proinflammatory cytokines in tumor/immune cell co-cultures. CONCLUSIONS: MGMT overexpression identifies a previously undescribed subset of CRCs with distinct biological and clinical properties including reduced mutagenesis, adaptive immune evasion, predominantly pMMR phenotype and aggressive clinical course. Direct, quantitative assessment of MGMT protein expression using spatially resolved analysis is more reliable than inference of MGMT expression by promoter methylation status in CRC.

Humans

Clinical and genomic features of mitis group streptococcal bacteremia in patients with febrile neutropenia.

BACKGROUND: Viridans group streptococci (VGS) can cause the life-threatening viridans streptococcal shock syndrome (VSSS) in patients with febrile neutropenia (FN). The Mitis group, a major subgroup of VGS, is frequently implicated in these severe infections, but its specific clinical and genomic characteristics remain incompletely characterized, particularly in patients with FN. This study aimed to systematically describe these features in this population. METHODS: In this single-center retrospective study, we compared the clinical data and whole-genome sequencing (WGS) results of Mitis group streptococcal isolates from patients with and without FN. Virulence-associated and antimicrobial resistance genes were initially screened using a reference-based approach, followed by assembly-based reanalysis and manual sequence validation. RESULTS: Compared with the non-FN cohort (n = 34), the FN cohort (n = 61) was significantly younger, had a higher prevalence of hematologic malignancy, and more frequently presented with primary bacteremia. VSSS occurred exclusively in the FN group (11.5%) and was associated with high mortality (14-day mortality, 42.9%), which did not correlate with in vitro antimicrobial susceptibility. Genomic analyses revealed marked diversity among isolates. Initial screening suggested variable detection of several virulence-associated loci, including pavA, slrA, and rfb-related loci; however, subsequent assembly-based analyses indicated that many apparent absences were attributable to extreme allelic divergence rather than true gene loss. No single virulence determinant clearly segregated with clinical severity. CONCLUSIONS: Mitis group bacteremia in patients with FN appears to be characterized by distinct clinical features and marked genomic diversity. Our findings suggest that the development of severe disease, including VSSS, may not be explained by microbial factors alone and potentially reflects complex host-pathogen interactions. CLINICAL TRIAL: Not applicable.

Humans

Genomic Features Do Not Account for Differences in Multiple Myeloma Risk by Ancestry.

UNLABELLED: Studies have reported conflicting findings regarding the contribution of germline variants or somatic genomic drivers to racial disparities in multiple myeloma. To comprehensively investigate somatic drivers in relation to inherited genetics in multiple myeloma, we combined newly sequenced whole-genome sequencing data with publicly available datasets (total n = 1,286). Overall, we did not identify germline or somatic genomic differences that explain the different risk of developing multiple myeloma between patients with genetic similarity to African (AFR) or European (EUR) reference populations. A difference in the detectability and timing of APOBEC-associated and germinal center mutational activity was observed. Integrating epidemiologic data and mutational signature-based temporal estimates, we challenge the assumption that individuals in the AFR group develop multiple myeloma at a younger age. Finally, we demonstrate that, with equal access to efficacious therapies, patients in the AFR and EUR groups have equivalent clinical outcomes. SIGNIFICANCE: Multiple myeloma is reported to occur at higher rates in individuals who self-identify as non-Hispanic Black. In this large dataset, genomic drivers occur at the same rate among ancestry groups, except for APOBEC mutagenesis. With equivalent therapy, clinical outcomes did not differ for patients grouped by genetic ancestry similarity.

Humans

Clinical outcomes and genomic features of uncommon EGFR exon 19 deletion subtypes in osimertinib-treated non-small cell lung cancer.

BACKGROUND: Epidermal growth factor receptor (EGFR) exon 19 deletion subtypes may be associated with differential survival outcomes following EGFR-tyrosine kinase inhibitor treatment. However, evidence remains scarce, particularly regarding osimertinib, and the underlying biological mechanisms are poorly understood. We aimed to compare survival outcomes among EGFR exon 19 deletion subtypes in patients with non-small cell lung cancer (NSCLC) treated with osimertinib. METHODS: In this multicenter retrospective study, patients with NSCLC were stratified according to exon 19 deletion subtypes. Whole-exome sequencing data from the American Association for Cancer Research Genomics Evidence Neoplasia Information Exchange registry and Memorial Sloan Kettering Clinicogenomic Harmonized Oncologic Real-World Dataset were analyzed to investigate co-occurring genomic alterations. RESULTS: Overall, 111 patients with advanced EGFR exon 19 deletion-positive NSCLC were analyzed and 86.5% received osimertinib as first-line therapy. Patients with non-E746_A750del (n&#xa0;=&#xa0;25) had shorter progression-free survival (PFS) than those with E746_A750del (n&#xa0;=&#xa0;86) (median: 14.3 vs. 20.6&#xa0;months; p&#xa0;<&#xa0;0.05). Among non-E746_A750del subtypes, L747_A750delinsP (n&#xa0;=&#xa0;4) had a particularly poor prognosis, with significantly worse survival than those with E746_A750del (median PFS: 3.5 vs. 20.6&#xa0;months; p&#xa0;<&#xa0;0.001, and median overall survival: 11.8 vs. 48.5&#xa0;months; p&#xa0;<&#xa0;0.001). In public database analyses, non-E746_A750del had a higher rate of RBM10 co-mutations, whereas L747_A750delinsP was characterized by frequent CDKN2A/B homozygous deletions and MYC amplifications. CONCLUSIONS: Non-E746_A750del was associated with poorer outcomes, with L747_A750delinsP potentially being a high-risk subtype. Differences in co-occurring genomic alterations may contribute to the prognostic heterogeneity among exon 19 deletion subtypes.

Humans

Genomic features, metabolism, and biotechnological applications of Candida tropicalis and other non-albicans Candida species.

The production of bio-based products by yeasts from agroindustrial byproducts is a key strategy for advancing circular bioeconomy. While Saccharomyces species remain the predominant industrial yeasts, their limited ability to assimilate lactose, pentoses, and glycerol, as well as their sensitivity to lignocellulose-derived inhibitors, restricts their efficient application in bioprocesses based on using industrial byproducts as fermentation media. In contrast, several non-albicans Candida species exhibit broad substrate utilization capacities and enhanced tolerance to industrial stresses, making them attractive candidates for the bioconversion of agroindustrial residues. This review critically examines recent advances in the genomic, metabolic, and physiological characterization of promising non-albicans Candida species, including Candida tropicalis, Candida parapsilosis, Candida viswanathii, Candida sojae, and Candida maltosa. Emphasis is given to genome-scale metabolic models, carbon assimilation pathways, stress-response mechanisms, and metabolic engineering approaches aiming at the production of value-added compounds. By identifying current achievements, knowledge gaps, and biotechnological bottlenecks, this review highlights the potential of these yeasts as emerging platforms for sustainable bioprocesses within a circular bioeconomy framework.

Biotechnology

Whole-Exome Sequencing Identifies Candidate Genomic Features Associated with Response to Platinum-Based Chemotherapy and Ixabepilone-Based Treatment in Ovarian Cancer.

Carboplatin/paclitaxel (CP) chemotherapy is the cornerstone of therapy for advanced stage ovarian cancer (OC). However, despite initial sensitivity, this regimen cannot avoid the emergence of resistance. Ixabepilone &#xb1; bevacizumab (IB) is a combination recently added to NCCN guidelines for the treatment of platinum-resistant OC. It would be desirable to identify biomarkers able to differentiate patients who are resistant to CP and IB, and biomarkers that identify which patients may benefit from IB treatment. We analyzed whole-exome-sequencing (WES) data from 49 OC patients exposed to CP, including 28 platinum-sensitive vs. 21 platinum-resistant, and 31 additional platinum-resistant patients, including 16 responders (i.e., CR/PR) vs. 15 non-responders (SD/PD) to ixabepilone &#xb1; bevacizumab. Comprehensive genetic analyses were performed to identify alterations correlated with resistance to CP and IB. WES analysis of CP responders vs. non-responders revealed differences in HRD-signatures (p < 0.05), OS (p < 0.005) and gain/loss-of-function in multiple genes associated with tumor growth/progression including but not limited to ACVR2A, INHBA, MAP3K7, ATG5, SGK1, FYN, RSPO3, NOD1 and LRRK2. WES analysis of platinum-resistant IB-treated patients revealed additional nominally significant genes and deranged pathways including gains in the DROSHA and SDHA genes in responders vs. non-responders (p < 0.05). Patients harboring HRD-signatures showed significantly higher sensitivity to CP and prolonged survival compared to HRD-negative patients. Alterations in genes associated with tumor growth/progression correlated with resistance to CP regimen and may represent novel "druggable" candidate biomarkers for the targeted treatment of CP/IB-resistant patients. Further validation in independent cohorts and preclinical experiments in CP/IB-resistant models are warranted to establish the clinical utility of these findings.

Humans

Clinical and molecular landscape of metastatic extramammary Paget's disease.

BACKGROUND: Extramammary Paget's disease (EMPD) is a rare malignancy without established systemic therapy. EMPD shares molecular features with breast cancer, such as human epidermal growth factor receptor 2 (HER2) and hormone receptor (HR) expression, but their clinical relevance remains unclear. MATERIALS AND METHODS: Tumors from 20 metastatic invasive EMPD cases were analyzed for molecular and biological features. Genomic features, transcriptomic profiles, and HER2 and HR expression status were investigated using immunohistochemistry, fluorescence in situ hybridization, and targeted-genome next-generation sequencing and nCounter BC360 panels. Metastatic breast cancer samples were used as a comparison to clarify metastatic EMPD's clinical relevance. RESULTS: Estrogen receptor expression was observed in 45% of EMPD tumors, while only 10% expressed progesterone receptor. HER2 was overexpressed in 30% of cases, and HER2-directed therapies were durably effective. Among 8 patients with NGS data, 63% (5/8) harbored oncogenic ERBB2 alterations independent of HER2 expression. BC360 profiling revealed biological differences between EMPD and breast cancer, particularly poor biological compatibility for HR-positive tumors. Immune profiling showed that a subset of EMPD tumors exhibited CD8+ T-cell signatures and PD-1/PD-L1 gene expression comparable to triple-negative breast cancer. The median overall survival was 22.1&#x2009;months (95% CI, 12.0-42.2), with 16 patients (80%) treated with systemic therapy, including anti-HER2 therapy, hormonal therapy, or cytotoxic therapies based on their molecular features. CONCLUSIONS: This study highlights the unique molecular and biological features of metastatic EMPD, emphasizing the need for tailored treatment approaches. This information should be used to guide future clinical strategies for metastatic EMPD.

Humans

Using the DNA language model, GROVER, to parse effects of sequence, chromatin and regulatory features on genome stability.

MOTIVATION: Genome stability is shaped by DNA sequence and chromatin context, but their relative contributions to double-strand break (DSB) sensitivity remain unclear. RESULTS: We show that the DNA language model, GROVER, can infer DSB location based on sequence. DSB hotspots tend to contain GC-rich sequences that belong to promoters, genes and short interspersed nuclear elements (SINEs). Additionally, we identified several specific short sequences (tokens) that are associated with modulating DSB sensitivity. Another model using chromatin and genome regulatory features outperforms the sequence-only model, highlighting complementary and cell-type specific information. Integrating sequence and genome biological features yields the best performance, demonstrating their synergy. Analyzing this model revealed that, dependent on the sample, genome stability information encoded in H3K36me3 and DNase-seq can be learned from the sequence, but not H3K27ac or H3K9me3. Embedding chromatin data directly into the GROVER architecture enabled cell-type specific modeling with performance matching the full chromatin feature model. Our results suggest that while chromatin and regulatory context provides important information, such as cell-type specificity, much of the information shaping DSB patterns is already encoded in the DNA sequence itself. Our integrative modeling approach not only reveals DSB patterns but also provides a generalizable strategy for tracing predictions in genomic data. AVAILABILITY: Data, models, and a tutorial are available on Zenodo.

Chromatin

Deep-Learning Model for Tumor-Type Prediction Using Targeted Clinical Genomic Sequencing Data.

UNLABELLED: Tumor type guides clinical treatment decisions in cancer, but histology-based diagnosis remains challenging. Genomic alterations are highly diagnostic of tumor type, and tumor-type classifiers trained on genomic features have been explored, but the most accurate methods are not clinically feasible, relying on features derived from whole-genome sequencing (WGS), or predicting across limited cancer types. We use genomic features from a data set of 39,787 solid tumors sequenced using a clinically targeted cancer gene panel to develop Genome-Derived-Diagnosis Ensemble (GDD-ENS): a hyperparameter ensemble for classifying tumor type using deep neural networks. GDD-ENS achieves 93% accuracy for high-confidence predictions across 38 cancer types, rivaling the performance of WGS-based methods. GDD-ENS can also guide diagnoses of rare type and cancers of unknown primary and incorporate patient-specific clinical information for improved predictions. Overall, integrating GDD-ENS into prospective clinical sequencing workflows could provide clinically relevant tumor-type predictions to guide treatment decisions in real time. SIGNIFICANCE: We describe a highly accurate tumor-type prediction model, designed specifically for clinical implementation. Our model relies only on widely used cancer gene panel sequencing data, predicts across 38 distinct cancer types, and supports integration of patient-specific nongenomic information for enhanced decision support in challenging diagnostic situations. See related commentary by Garg, p. 906. This article is featured in Selected Articles from This Issue, p. 897.

Humans

Isolation, identification, and genomic characterization of Staphylococcus aureus phage vB_SauL_202595 and its bacteriostatic application in dairy products.

Staphylococcus aureus is an important pathogen associated with bovine mastitis and dairy product contamination, posing economic and public health risks through the food chain. In this study, a temperate phage, vB_SauL_202595, was isolated from a dairy farm environmental sample using S. aureus SHZ-0127 as the host, and its biological characteristics, genomic features, and antibacterial activity in dairy matrices were evaluated. vB_SauL_202595 lysed 18 of 66 tested S. aureus strains, with a lysis susceptibility rate of 27.3%, including 5 highly susceptible strains, indicating a relatively limited host range. The optimal multiplicity of infection was 0.01, the latent period was approximately 30 min, and the burst size was approximately 316 PFU/cell. The phage remained stable at 4&#xb0;C-37&#xb0;C and pH 6-10. Genome analysis showed that vB_SauL_202595 belongs to the class Caudoviricetes, has a genome of 44,503 bp with 33.59% GC content, and encodes 63 predicted proteins. No typical antibiotic resistance genes or major virulence factors were detected; however, integrase and repressor genes were identified, supporting its temperate nature. vB_SauL_202595 inhibited S. aureus SHZ-0127 growth, reduced mature biofilm biomass, and decreased viable bacterial counts in milk and yogurt, with reductions of 1.23 and 1.42 log10 CFU/mL under representative conditions, respectively. From a One Health perspective, these findings provide foundational evidence for reducing S. aureus contamination and related antimicrobial resistance risks along the dairy chain. Overall, vB_SauL_202595 represents a candidate phage resource for dairy-associated S. aureus biocontrol research, but its limited host range and lysogeny-related genes require further safety assessment before food-related applications.IMPORTANCEStaphylococcus aureus is a major pathogen associated with bovine mastitis and a common contaminant in dairy products, causing economic losses and public health risks through the food chain. Although phage-based biocontrol has emerged as a promising strategy for controlling S. aureus contamination in dairy products, systematic evidence regarding phage activity in actual dairy matrices remains limited. In this study, we isolated and characterized a dairy farm environment-derived temperate phage, vB_SauL_202595, and evaluated its biological characteristics, genomic features, host range, stability, biofilm removal ability, and antibacterial performance in milk and yogurt. These findings provide foundational experimental evidence for phage-based dairy biocontrol against S. aureus. However, due to its limited host range and lysogeny-related genomic features, vB_SauL_202595 should be considered a candidate phage resource for further study. Broader validation, including phage-cocktail testing, long-term storage assays, product quality assessment, and regulatory safety evaluation, is needed before practical application.

Staphylococcus aureus

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning

Multimodal deep learning for immunotherapy response prediction and biomarker discovery in non-small cell lung cancer.

OBJECTIVE: Immunotherapy has emerged as a promising treatment for advanced non-small cell lung cancer (NSCLC), but accurately predicting which patients will benefit from it remains a major clinical challenge. To address this, we aim to develop a novel multimodal method, DeepAFM, that integrates histopathology, genomic features, and clinical information to predict patient responses to anti-PD-(L)1 immunotherapy. MATERIALS AND METHODS: A total of 93 patients with advanced NSCLC were included in this study. Histopathological whole-slide images were processed using a self-supervised VQVAE2 for representation learning. PCA and K-means clustering were then applied for dimensionality reduction and feature grouping. Key regions of interest were visualized through permutation importance evaluation and color-coding techniques. The extracted histopathological features, along with genomic alterations and clinical variables, were integrated into the DeepAFM multimodal prediction model. RESULTS: The DeepAFM achieved a high predictive performance with an area under the curve (AUC) of 0.77 (95% confidence interval: 0.69-1.00). Attention-based heatmaps revealed that the model could identify critical pathological patterns, genomic mutations, and clinical indicators associated with patient responses to immunotherapy. DISCUSSION: The integration of multimodal data enabled the model to capture complex interactions among pathology, genomics, and clinical characteristics, enhancing the interpretability and predictive power of immunotherapy response prediction. The visualization techniques facilitated the identification of biologically meaningful features and potential biomarkers. CONCLUSION: This study demonstrates the effectiveness of the DeepAFM in predicting responses to immunotherapy in advanced NSCLC. The approach not only improves prediction accuracy but also provides valuable insights for personalized treatment strategies and biomarker discovery.

Humans

Complete telomere-to-telomere genome assembly of Guazuma ulmifolia uncovers evolutionary mechanisms, drought adaptation, and flavonoid biosynthesis.

The first T2T reference genome of Guazuma ulmifolia is reported, which serves as a core genomic resource for stress adaptation research and stress-tolerant breeding in cacao wild relatives. Climate change, particularly increased incidence of drought, poses a major threat to food security. Understanding the genomic basis of environmental adaptation in crop wild relatives can provide valuable resources for improving stress resilience. Guazuma ulmifolia, a wild relative of Theobroma cacao with important ecological and medicinal value, lacks high-quality reference genomic resources. Here, we report the first telomere-to-telomere (T2T) chromosome-level genome assembly of G. ulmifolia, with a genome size of 311.31&#xa0;Mb, contig N50 of 35.19&#xa0;Mb, and 98.70% BUSCO completeness. Repetitive sequences constitute 27.43% of the G. ulmifolia genome, with LTR retrotransposons as the predominant class. Comparative genomic analyses revealed that genome-size variation among Malvaceae species is associated with differences in polyploidization history and TE dynamics. Ancestral karyotype reconstruction identified five lineage-specific chromosome fusion events distinguishing G. ulmifolia from T. cacao. Comparative analyses further identified tandem duplication-associated expansion of stress-related LEA and GST gene families, suggesting potential genomic features associated with stress responses. Flavonoid biosynthesis genes were largely conserved in copy number but showed tissue-specific expression patterns, providing candidate genes for investigating secondary metabolism. Together, this study establishes a high-quality T2T genome resource for exploring genome evolution, chromosome organization, and stress-related genomic features in Malvaceae.

Genome, Plant

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24&#x2009;months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical

Bridging-driven condensation by eukaryotic SMC complexes is a conserved feature of genome organization.

The Structural Maintenance of Chromosome (SMC) protein family plays a central role in higher-order genome organization through ATP-dependent DNA loop extrusion by cohesin and condensin and other processes. Whether these activities fully account for the complexity of chromosome architecture remains unknown. Here, we uncover a conserved ATP-independent mechanism of chromatin condensation by SMC complexes, occurring via biomolecular condensation. Using single-molecule fluorescence imaging, we show that a variety of SMCs form dynamic DNA-bound condensates that exhibit key features of biomolecular condensates, including droplet coalescence, fluorescence recovery after photobleaching, and rapid exchange with free SMC complexes. Atomic force microscopy analysis of human cohesin-DNA assemblies reveals DNA-length-dependent clustering, providing evidence for bridging-driven condensation. Analyses of&#xa0;in vivo super-resolution imaging and high-throughput chromosome conformation capture (Hi-C) data indicate that these condensates form chromatin-associated clusters with multi-loop structures. Together, our results establish that SMC complexes employ ATP-independent phase condensation as well as ATP-dependent activities to shape genome architecture. This work reveals a broadly conserved principle of chromosomal organization across eukaryotes.

Chromosomal Proteins, Non-Histone