Search PubMedSearch

SEARCH · Search PubMed

Results for “phenotypic clustering”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Transcriptome-wide analysis reveals potential roles of CFD and ANGPTL4 in fibroblasts regulating B cell lineage for extracellular matrix-driven clustering and novel avenues for immunotherapy in breast cancer.

BACKGROUND: The remodeling of the extracellular matrix (ECM) plays a pivotal role in tumor progression and drug resistance. However, the compositional patterns of ECM in breast cancer and their underlying biological functions remain elusive. METHODS: Transcriptome and genome data of breast cancer patients from TCGA database was downloaded. Patients were classified into different clusters by using non-negative matrix factorization (NMF) based on signatures of ECM components and regulators. Weighted Gene Co-expression Network Analysis (WGCNA) was used to identify core genes related to ECM clusters. Additional 10 independent public cohorts including Metabric, SCAN_B, GSE12276, GSE16446, GSE19615, GSE20685, GSE21653, GSE58644, GSE58812, and GSE88770 were collected to construct Training or Testing cohort, following machine learning calculating ECM correlated index (ECI) for survival analysis. Pathway enrichment and correlation analysis were used to explore the relationship among ECM clusters, ECI and TME. Single-cell transcriptome data from GSE161529 was processed for uncovering the differences among ECM clusters. RESULTS: Using NMF, we identified three ECM clusters in the TCGA database: C1 (Neuron), C2 (ECM), and C3 (Immune). Subsequently, WGCNA was employed to pinpoint cluster-specific genes and develop a prognostic model. This model demonstrated robust predictive power for breast cancer patient survival in both the Training cohort (n = 5,392, AUC = 0.861) and the Testing cohort (n = 1,344, AUC = 0.711). Upon analyzing the tumor microenvironment (TME), we discovered that fibroblasts and B cell lineage were the core cell types associated with the ECM cluster phenotypes. Single-cell RNA sequencing data further revealed that angiopoietin like 4 (ANGPTL4)+ fibroblasts were specifically linked to the C2 phenotype, while complement factor D (CFD)+ fibroblasts characterized the other ECM clusters. CellChat analysis indicated that ANGPTL4+ and CFD+ fibroblasts regulate B cell lineage via distinct signaling pathways. Additionally, analysis using the Kaplan-Meier Plotter website showed that CFD was favorable for immunotherapy response, whereas ANGPTL4 negatively impacted the outcomes of cancer patients receiving immunotherapy. CONCLUSION: We identified distinct ECM clusters in breast cancer patients, irrespective of molecular subtypes. Additionally, we constructed an effective prognostic model based on these ECM clusters and recognized ANGPTL4+ and CFD+ fibroblasts as potential biomarkers for immunotherapy in breast cancer.

Humans

Quantitative natural history modeling of HPDL-related disease based on cross-sectional data reveals genotype-phenotype correlations.

PURPOSE: Biallelic HPDL variants have been identified as the cause of a progressive childhood-onset movement disorder, with a broad clinical spectrum from severe neurodevelopmental disorder to juvenile-onset pure hereditary spastic paraplegia type 83. This study aims at delineating the geno- and phenotypic spectra of patients with HPDL-related disease, quantitatively modeling the natural history, and uncovering genotype-phenotype associations. METHODS: A cross-sectional analysis of 90 published and 1 novel case was performed, using a Human-Phenotype-Ontology-based approach. Unsupervised phenotypic clustering was used alongside in silico analyses to identify distinct patient subgroups. RESULTS: The study models the natural history of the HPDL-related disease in a global cohort, clarifying the molecular and phenotypic spectrum and identifying 3 distinct subgroups characterized by differences in onset, clinical trajectories, and survival. It establishes genotype-phenotype associations, showing that the presence of moderately pathogenic missense variants in 1 allele leads to a milder, spastic paraplegic phenotype with later disease onset, whereas biallelic, highly pathogenic missense or truncating variants are associated with a more severe phenotype and reduced life span. CONCLUSION: Quantitative and unbiased natural history modeling in HPDL-related disease reveals significant genotype-phenotype associations, providing a foundation for variant interpretation, anticipatory guidance, and choice of outcome measures in future prospective and functional studies.

Humans

Genome-sequencing-based benchmarking of antimicrobial resistance, treatment outcomes and healthcare transmission events for Clostridioides difficile infection in Australian hospitals.

BACKGROUND: Clostridioides difficile infection (CDI) remains a priority for infection prevention and control in health care, particularly with the emergence of hypervirulent strains and antimicrobial resistance (AMR). AIM: To characterize the genomic epidemiology and AMR profiles of culture-confirmed CDI cases within tertiary hospitals in Australia. METHODS: A total of 155 C. difficile isolates from 142 patients with CDI diagnosed in four hospitals between 2023 and 2025 were studied. Data collected included patient demographics, severity of infection, antibiotic treatment and clinical outcomes at 8 weeks. Phenotypic susceptibility to vancomycin, fidaxomicin, metronidazole, moxifloxacin, meropenem, tetracycline and rifaximin were determined by agar dilution. Isolates underwent whole-genome sequencing (WGS) for genotyping and resistome assessment. FINDINGS: WGS differentiated 39 distinct sequence types among CDI isolates across different healthcare services. In total, 100 isolates were singletons and 55 (35% clustering rate) isolates were considered to be genomically related (difference of two or fewer single-nucleotide polymorphisms). Of these, 12 patients (8.5%) with close hospital contact formed six epidemiologically linked clusters. Phenotypic susceptibility results were obtained for 134 (86.4%) CDI isolates. There was no phenotypic resistance to vancomycin [minimum inhibitory concentration required to inhibit the growth of 90% of isolates (MIC90) 1 mg/L], metronidazole (MIC90 0.5 mg/L) or fidaxomicin (MIC90 0.5 mg/L). There was no association in the study cohort between the presence of resistance genes or reduced phenotypic susceptibility and CDI recurrence. CONCLUSION: Genomic analysis of C. difficile isolates did not identify any outbreaks or an association between the sequence type or presence of a resistance gene and clinical outcomes. High-resolution characterization and identification of antibiotic resistance, CDI clinical relapse and recent transmission offered by genome sequencing can provide important benchmarks for hospital infection control.

Antibiotic resistance

Multi-omic underpinnings of heterogeneous aging across multiple organ systems.

Aging is the main determinant of chronic diseases and mortality, yet organ-specific aging trajectories vary, and the molecular basis underlying this heterogeneity remains unclear. To elucidate this, we integrated genomic, epigenomic, transcriptomic, proteomic, and metabolomic data, employing post-genome-wide association study methodologies to systematically investigate the molecular mechanisms of nine organ-specific aging clocks and four blood-based epigenetic clocks. We uncovered genetic correlations and specific phenotypic clusters among these aging-related traits, identified prioritized genetic drug targets for heterogeneous aging, and elucidated downstream proteomic and metabolomic effects mediated by heterogeneous aging. We constructed a cross-layer molecular interaction network of heterogeneous aging across multiple organ systems and characterized detectable biomarkers of this heterogeneity. Integrating these findings, we developed an R/Shiny-based framework that provides a comprehensive multi-omic molecular landscape of heterogeneous aging, thereby advancing the understanding of aging heterogeneity and informing precision medicine strategies to delay organ-specific aging and prevent or treat its associated chronic diseases.

Aging

Coagulation activation is associated with genomic-instability-related features in TP53-mutated AML and MDS: routine laboratory patterns beyond classical disseminated intravascular coagulation.

BACKGROUND: Disseminated intravascular coagulation (DIC) is a serious complication of acute myeloid leukemia (AML) associated with poor prognosis. In TP53-mutated AML and myelodysplastic syndrome (MDS), however, the classical ISTH criteria rarely identify overt DIC, although bleeding and thrombotic complications are well documented in acute leukaemia. We hypothesized that these patients exhibit a lower-grade, subclinical coagulation activation that is associated with the underlying genomic-instability-related features of TP53-mutant disease. METHODS: We retrospectively analyzed 107 consecutive patients with TP53-mutated AML (n = 52) or high-risk MDS (MDS, n = 55), median age 65 years, diagnosed and initially evaluated at our centre between 2018 and 2025. Seven routine coagulation markers and 46 co-mutated genes were evaluated for associations with overall survival (OS) using univariate and multivariable Cox regression, continuous dose-response modeling, and unsupervised k-means clustering. Internal validity was assessed by 1000 bootstrap resamples. RESULTS: Overt DIC according to ISTH criteria was rare (15%). Subclinical activation was common: 50% of patients had a D-dimer &#x2265;1&#xa0;&#x3bc;g/mL, 41% a fibrinogen &#x2265;4&#xa0;g/L, and 29% an INR &#x2265;1.2. In univariate analysis, D-dimer, fibrinogen, INR, prothrombin time, and activated partial thromboplastin time were each associated with OS (HR 1.33-1.38 per SD; all p < 0.05). Complex karyotype correlated with higher D-dimer (median 1.39 vs. 0.60&#xa0;&#x3bc;g/mL, p = 0.022) and fibrinogen (3.91 vs. 2.53&#xa0;g/L, p = 0.007), while TP53 variant allele frequency (VAF) showed modest positive correlations with D-dimer (&#x3c1; = 0.21), INR (&#x3c1; = 0.27), and PT (&#x3c1; = 0.27; all p < 0.05). Clustering identified three coagulation phenotypes: Silent (51%), Thrombo-inflammatory (31%), and Consumption-like (18%), showing a graded but statistically non-significant gradient in molecular features and a stepwise decline in median OS (14, 10 and 8 months; log-rank p = 0.041). After adjustment for complex karyotype, TP53 VAF, and favorable co-mutation count, the Consumption-like phenotype was associated with a non-significant increased risk (HR 1.83, 95% CI 0.92-3.65, p = 0.084), whereas favorable co-mutation pathways remained independently protective (HR 0.56, 95% CI 0.35-0.90, p = 0.016). CONCLUSION: In TP53-mutated AML/MDS, coagulation activation intensity is associated with the degree of genomic instability. The three phenotypes may add biological resolution beyond classical DIC and cytogenetic risk groups, but represent laboratory patterns rather than validated bleeding or thrombosis prediction tools. However, after accounting for genomic features, phenotypes were not independent predictors of outcome, with complex karyotype, TP53 VAF, and favorable co-mutation count driving prognosis. Because treatment intensity and other clinical confounders were not available, these survival associations are hypothesis-generating. Coagulation profiling remains inexpensive, widely accessible, and offers a practical window into disease biology that warrants prospective validation.

TP53

Unsupervised characterization of 100,272 EHR patients identifies high-risk groups and comorbidities linked to premature aging.

Electronic health records (EHRs) contain extensive multidimensional patient data, presenting challenges for the discovery of novel and meaningful clinical patterns. Unsupervised clustering of high-dimensional clinical data holds great potential for identifying novel clinical patterns. Here, we performed unsupervised clustering and characterized 100,272 patients in the Electronic Medical Records and GEnomics (eMERGE) Network. We identified 70 clusters defined by distinct comorbidity patterns. Meanwhile, age and sex are also strongly associated with patient stratification, influencing phenotype prevalence and onset time. Notably, phenotype onset time accurately predicted chronological age and was significantly associated with overall mortality risk. Besides age and sex, we assessed the contribution of genetic variation to phenotype development and observed evidence of cross-phenotype associations influencing cluster membership and comorbidity patterns. However, the role of genetics recedes during aging. We also identified several high-risk clusters with elevated Charlson Comorbidity Index (CCI) scores and validated these findings in an independent cohort. Further analysis of these clusters revealed phenotypes linked to premature aging and highlighted a survival selection among older participants in observational studies. Overall, this study enables phenome-wide unsupervised patient stratification for multimorbidity discovery in largely unannotated clinical data, offering valuable insights into patient stratification, comorbidity analysis, aging, and health outcomes.

Journal Article

Multi-locus allelic architecture underlying natural variation in leaf rolling in japonica rice.

Leaf rolling is a key component of rice canopy architecture that affects light interception, microclimate formation, and planting density. The contribution of naturally occurring allelic variation to quantitative variation in leaf rolling within cultivated rice remains poorly understood, while extreme leaf rolling caused by loss-of-function mutations often results in detrimental pleiotropic effects. Herein, we examined how multi-locus allelic variation contributes to natural variation in leaf rolling within japonica rice. Leaf rolling was quantified based on the leaf rolling index (LRI) using a panel of 201 japonica accessions. The phenotype was transformed using the Yeo-Johnson method to reduce strong right skewness and improve the distributional properties of the data, thereby facilitating subsequent regression modeling. Haplotype analyses were performed for previously reported leaf rolling-associated genes and genome-wide association study (GWAS) lead loci, leading to the identification of five loci exhibiting substantial haplotype-dependent phenotypic variation. Phenotypically defined allelic groups represented these loci were subsequently evaluated using multiple linear regression (MLR), with the first two principal components derived from genome-wide SNP data included as covariates to account for population structure. The final MLR model identified four loci (qALR1, OsYABBY1, OsSLL2, and OsSRL10) as the independent contributors to leaf rolling variation, collectively explaining 21% of the variance in the transformed phenotype after accounting for population structure. Model diagnostics and ten-fold cross-validation supported the statistical validity of the framework and indicated stable model performance across validation folds. Analysis of multi-locus allelic combinations showed 13 distinct configurations that clustered into three phenotypically differentiated groups. This reflected the cumulative dosage of high-leaf rolling alleles. Thus, the natural variation in leaf rolling in japonica rice is governed by the additive effects of multiple moderate-impact loci. The multi-locus allelic framework established here provides a statistically sound and biologically interpretable basis for dissecting polygenic canopy traits and practical guidance for developing genetic materials aimed at optimizing rice plant architecture.

cross-validation

Construction of a molecular diagnostic system for neurogenic rosacea by combining transcriptome sequencing and machine learning.

Patients with neurogenic rosacea (NR) frequently demonstrate pronounced neurological manifestations, often unresponsive to conventional therapeutic approaches. A molecular-level understanding and diagnosis of this patient cohort could significantly guide clinical interventions. In this study, we amalgamated our sequencing data (n&#x2009;=&#x2009;46) with a publicly accessible database (n&#x2009;=&#x2009;38) to perform an unsupervised cluster analysis of the integrated dataset. The eighty-four rosacea patients were partitioned into two distinct clusters. Neurovascular biomarkers were found to be elevated in cluster 1 compared to cluster 2. Pathways in cluster 1 were predominantly involved in neurotransmitter synthesis, transmission, and functionality, whereas cluster 2 pathways were centered on inflammation-related processes. Differential gene expression analysis and WGCNA were employed to delineate the characteristic gene sets of the two clusters. Subsequently, a diagnostic model was constructed from the identified gene sets using linear regression methodologies. The model's C index, comprising genes PNPLA3, CUX2, PLIN2, and HMGCR, achieved a remarkable value of 0.9683, with an area under the curve (AUC) for the training cohort's nomogram of 0.9376. Clinical characteristics from our dataset (n&#x2009;=&#x2009;46) were assessed by three seasoned dermatologists, forming the NR validation cohort (NR, n&#x2009;=&#x2009;18; non-neurogenic rosacea, n&#x2009;=&#x2009;28). Upon application of our model to NR diagnosis, the model's AUC value reached 0.9023. Finally, potential therapeutic candidates for both patient groups were predicted via the Connectivity Map. In summation, this study unveiled two clusters with unique molecular phenotypes within rosacea, leading to the development of a precise diagnostic model instrumental in NR diagnosis.

Humans

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n&#x202f;=&#x202f;38, 74%). Hierarchical clustering (n&#x202f;=&#x202f;20) and K-means clustering (n&#x202f;=&#x202f;14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Characterization of the genotypic and phenotypic spectrum of TCF7L2-related neurodevelopmental disorder (TRND).

PURPOSE: TCF7L2 (OMIM 602228; HGNC:11641) is a transcription factor and a critical effector of the Wnt/ &#x3b2;-Catenin pathway. In 2021, 11 pediatric patients with monoallelic predicted loss-of-function (pLOF) TCF7L2 variants and syndromic features were observed. Characterization of patients with pLOF TCF7L2 variants and neurodevelopmental features-herein referred to as TCF7L2-related neurodevelopmental disorder-is urgently needed. METHODS: We leveraged multiple methods (eg, GeneMatcher, DECIPHER, literature review, and public/private repositories) to identify an international cohort of 76 patients with pLOF TCF7L2 variants and neurodevelopmental features and phenotypically characterized them. We also retrospectively searched for an independent cohort of adults with pLOF TCF7L2 variants (n = 11) from more than 60,000 PennMedicine BioBank patients. RESULTS: Among 76 patients with pLOF TCF7L2 variants, speech delay (95.3%), craniofacial dysmorphisms (73.3%), ophthalmologic conditions (65.5%), autism (62.1%), and orthopedic abnormalities (52.6%) were the most commonly observed. Phenotypic differences did not cluster by variant type or genomic locus. Among PennMedicine BioBank patients, an association of nominal significance with type 2 diabetes with renal manifestations (odds ratio = 5.8; P = .03) was detected, warranting further investigation. CONCLUSION: This study represents the most comprehensive characterization of TCF7L2-related neurodevelopmental disorder to date, a novel neurodevelopmental disorder, defining its genotypic and phenotypic spectra. We opened a Simons Searchlight natural history study that is now available for patient enrollment to enhance the understanding of this condition.

Neurodevelopmental syndrome

Integrative multi-omics analysis proposes a metabolic classification of gliomas: distinct metabolic states, immune infiltration, and prognosis.

BACKGROUND: The tumor microenvironment (TME) of glioma harbors diverse cell types; however, cell metabolic heterogeneity remains to be explored. This study aims to characterize the metabolic features of different cell types in the TME by integrating multiple datasets, including genomics, bulk and single-cell transcriptomics, and metabolomics. METHODS: Unsupervised machine learning was used to construct an energy metabolic classifier based on the metabolic pathways identified from bulk RNA-seq of gliomas in the TCGA dataset. The classifier was externally validated using multiple datasets, including genomics, bulk RNA-seq, snRNA-seq, and the metabolomics data. Furthermore, metabolic heterogeneity associated with the classifier was further characterized at single-cell resolution. RESULTS: The energy metabolism-based classifier stratified patients into two prognostic clusters: patients in cluster 1 were characterized by high pathway activity of glycolysis, the pentose phosphate pathway (PPP), and fatty acid oxidation (FAO), whereas patients in cluster 2 exhibited higher activity in glutaminolysis. This metabolic classifier revealed both intratumoral and intertumoral metabolic heterogeneity, and the complexity was further validated by the metabolomics profiling and snRNA-seq data from the CPTAC dataset. Notably, OSMR, highly expressed in cluster 1, showed significant co-expression with key glycolytic enzyme genes. The OSM/OSMR/JAK1/STAT3 axis potently drives malignant progression of glioma cells, specially enhancing their invasive and migratory capabilities. Single-cell resolution analyses demonstrated that tumor metabolic heterogeneity is primarily driven by malignant cells rather than non-malignant components, while tumor microenvironment (TME) factors were also found to modulate malignant cell metabolism. Significantly, glycolytic activity in glioma cells increased during the phenotypic transition from PN (proneural) to MES (mesenchymal), with cluster 1 metabolic phenotypes predominating in the tumor core. Compared to cluster 2, cluster 1 patients exhibited higher mRNA expression of immunosuppressive checkpoint genes, which correlated with pronounced immunosuppression in the TME. Furthermore, various immune cells demonstrated distinct metabolic preferences at single-cell resolution. CONCLUSIONS: This study developed an energy metabolic-based classifier for gliomas with prognostic and therapeutic potential. Metabolic reprogramming was linked with the PN-to-MES transition of glioma cells and immunosuppression in the tumor microenvironment. Multi-omics data, especially snRNA-seq, offered insights into metabolism heterogeneity at single-cell resolution, enabling personalized treatment strategies.

Humans

Pilot study identifying distinct circulating proteomic profiles associated with longitudinal CT-defined fibrotic and inflammatory sarcoidosis.

INTRODUCTION: Pulmonary sarcoidosis exhibits heterogeneous clinical trajectories ranging from self-limited disease resolution to chronic progressive fibrosis, yet reliable biomarkers capable of distinguishing these disease patterns remain lacking. Whether longitudinal CT-defined sarcoidosis phenotypes are associated with distinct circulating molecular signatures remains unknown. METHODS: We performed high-throughput plasma proteomics (SomaScan 11K) in participants with pulmonary sarcoidosis classified into longitudinal chest CT-defined progressive fibrosis, progressive nodular inflammatory disease, or resolving disease trajectories, along with healthy controls. CT phenotypes were assigned based on predefined longitudinal changes in reticulation, traction bronchiectasis, nodular involvement, and mediastinal lymphadenopathy across serial CT scans. One plasma sample per participant was selected from the study visit corresponding to the CT time point at which criteria for the assigned longitudinal phenotype were met. Principal component analysis, hierarchical clustering, pathway enrichment, and correlation-based analyses linking protein expression to quantitative CT features were used to evaluate whether distinct longitudinal CT phenotypes were associated with divergent proteomic signatures. RESULTS: Principal component analysis and hierarchical clustering suggested partial segregation by CT-defined phenotype. Longitudinal CT phenotypes were associated with distinct pathway-level proteomic signatures, with progressive fibrosis enriched for epithelial-mesenchymal transition signaling, and progressive nodular inflammatory disease enriched for mTORC1, MYC, oxidative phosphorylation, adipogenesis, and fatty acid metabolism pathways. Correlation analyses showed coordinated protein-expression patterns associated with fibrotic CT features and mediastinal lymph node enlargement. DISCUSSION: These findings suggest that longitudinal CT-defined fibrotic and inflammatory sarcoidosis phenotypes are associated with distinct pathway-level proteomic signatures. This pilot study provides preliminary proof-of-concept evidence that integrating longitudinal CT imaging phenotypes with plasma proteomics may serve as a framework for future mechanistic studies and biomarker discovery in pulmonary sarcoidosis.

Humans

FuNTB: a functional network clustering tool for the analysis of genome-wide genetic variants in Mycobacterium tuberculosis.

MOTIVATION: Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mtb), still claims around 1.25 million lives each year. The growing threat of drug resistance-often driven by single&#x2011;nucleotide polymorphisms (SNPs) in Mtb genomes underscores the need for high&#x2011;quality genomic data and powerful bioinformatics tools. We present FuNTB, a python&#x2011;based pipeline that detects non&#x2011;synonymous SNPs in Mtb and builds functional network clusters to reveal genotype-phenotype relationships. RESULTS: FuNTB profiles non&#x2011;synonymous SNPs at the gene level across user&#x2011;defined phenotypes, pinpointing both shared and unique mutations. It ingests annotated Variant Call Format (VCF) files or MTBseq outputs and merges them with clinical metadata to produce network&#x2011;XML files compatible with Cytoscape and Gephi. When applied to the CRyPTIC Mtb collection, FuNTB rapidly recovered established resistance genes and surfaced novel candidates, validating its utility for mapping genotype-phenotype associations. AVAILABILITY AND IMPLEMENTATION: FuNTB is implemented in Python 3.8+ and is freely available under the MIT license at https://doi.org/10.5281/zenodo.15399917.

Mycobacterium tuberculosis

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome

Heterogeneous effects of genetic variants and traits associated with fasting insulin on cardiometabolic outcomes.

Elevated fasting insulin levels (FI), indicative of altered insulin secretion and sensitivity, may precede type 2 diabetes (T2D) and cardiovascular disease onset. In this study, we group FI-associated genetic variants based on their genetic and phenotypic similarities and identify seven clusters with distinct mechanisms contributing to elevated FI levels. Clusters fall into two types: "non-diabetogenic hyperinsulinemia," where clusters are not associated with increased T2D risk, and "diabetogenic hyperinsulinemia," where T2D associations are driven by body fat distribution, liver function, circulating lipids, or inflammation. In over 1.1 million multi-ancestry individuals, we demonstrated that diabetogenic hyperinsulinemia cluster-specific polygenic scores exhibit varying risks for cardiovascular conditions, including coronary artery disease, myocardial infarction (MI), and stroke. Notably, the visceral adiposity cluster shows sex-specific effects for MI risk in males without T2D. This study underscores processes that decouple elevated FI levels from T2D and cardiovascular risk, offering new avenues for investigating process-specific pathways of disease.

Humans

APAV: An advanced pangenome analysis and visualization toolkit.

Traditional pangenome analysis focuses on gene presence/absence variations (gene PAVs). However, the current methods for gene PAV analysis are insensitive to detect small but valuable mutations within gene regions, and they overlook variations in intergenic regions. Additionally, the visual inspection of PAVs is an important but time-consuming step for pangenome analysis and result interpretation. To address these issues, we present APAV, an advanced toolkit designed for comprehensive PAV analysis and visualization. It integrates gene element-level PAV analysis and provides PAV analysis for arbitrary given regions in a genome. The resulted PAV profile can be visualized and investigated interactively with reports in HTML format, enabling researchers to conveniently verify sequencing read depth, target region coverage, and intervals of absence for each PAV. Furthermore, APAV offers various subsequent analysis and visualization functions based on the PAV profile table, including basic statistics, sample clustering, genome size estimation, and phenotype association analysis. We demonstrated the capability of APAV with pangenome analysis of tumor genomes and rice genomes. Performing PAV analysis at the element level not only provides more accurate information about the variations but also uncovers a larger number of variations for the phenotype-genotype association studies. In the rice genome analysis, we identified over twenty thousand distributed genes and more than fifty thousand distributed genetic elements. In the tumor genome analysis, element-level analysis revealed approximately three times as many phenotype-related genes as gene-level analysis. This indicates that altering the PAV unit from genes to smaller segments or elements can lead to more biological insights.

Software

T cell subsets of urine-derived lymphocytes (UDLs) serve as an indicator of TILs and reflect immunological sex differences in bladder cancer.

BACKGROUND: Bladder cancer is unique among visceral malignancies in that urine, which can be easily obtained, has prolonged contact with bladder tumors. Urinary biomarkers offer the potential to provide insight into the host and tumor immune microenvironment to guide therapeutic strategies. We evaluated the immune cellular composition of urine (urine-derived lymphocytes (UDLs)) versus tumor (tumor-infiltrating lymphocytes (TILs)). METHODS: We employed high-dimensional flow cytometry analyses on immune cells from tumors (TILs), urine (UDLs), and peripheral blood (peripheral blood mononuclear cells) among patients with bladder cancer. We performed multiplexed immunofluorescence (mIF) of matched tumors to provide spatial context to our findings, comparing deep/invasive and superficial/urine-facing regions of matched tumors. RESULTS: Our findings suggest that the CD4+ and CD8+ T cell subsets of UDLs characterized by flow cytometry had similar phenotypic profiles to those found in TILs (cell clusters quantified by multidimensional scaling and differentiation states). Results of mIF imaging with a panel of phenotypic and functional T cell markers suggested that UDLs reflected TILs in both superficial and deep tumor sections. We also found sex-dependent patterns in TILs and UDLs, indicating the male bladder cancer tumor microenvironment is enriched in exhausted CD4+ and CD8+ T cells, while the female bladder cancer microenvironment is enriched for activated T cells. CONCLUSIONS: Assessment of UDLs opens avenues of non-invasive biomarker development in clinical settings where bladder cancer TILs are hypothesized to predict clinical response. UDLs may also reflect sex-based differences in antitumor immunity.

Humans

Osteoarthritis phenotypes: advancing precision medicine through clinical, structural, and molecular stratification.

PURPOSE: Osteoarthritis (OA) is now understood as a heterogeneous syndrome driven by diverse biological, biomechanical, metabolic, genetic, and molecular mechanisms. This variability explains differences in disease progression and treatment response, challenging the traditional "one-size-fits-all" approach. This review highlights OA phenotyping as a key step toward precision medicine, focusing on clinical, structural, and molecular classifications that inform individualized care. METHODS: A narrative review was conducted using a non-systematic search of major databases and Osteoarthritis Research Society International sources (2010-2026). Evidence was thematically synthesized across clinical, imaging, and molecular domains to characterize OA phenotypes and their potential relevance to precision medicine. RESULTS: Multiple OA phenotypes were identified: inflammatory, metabolic, biomechanical, cartilage-subchondral, pain-sensitization, and aging/senescence. These exhibit distinct clinical features, risk factors, and therapeutic responses. Imaging-based phenotypes (e.g., inflammatory, meniscus-cartilage, subchondral bone, atrophic, hypertrophic) and molecular endotypes (low turnover, structural damage, systemic inflammation) further refine stratification. Pain-structure discordance is notable in sensitization phenotypes and may predict poorer surgical outcomes. Joint-specific variations and emerging genomic and epigenetic insights underscore disease complexity. Advances in imaging, biomarkers, and machine learning may enable earlier detection and patient clustering, though clinical application remains limited. CONCLUSION: Phenotype- and endotype-based classification represents a critical advancement toward precision OA management. Tailored interventions based on stratification hold promise for improving outcomes; however, clinical translation remains limited by overlapping phenotypes, lack of validated biomarkers, and inconsistent results from phenotype-driven trials. Wider clinical adoption requires standardized definitions, validation across joints, and integration of multimodal diagnostic tools into routine practice.

Humans