Search PubMedSearch

SEARCH · Search PubMed

Results for “UMAP”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

7 recordsLinked to original sources

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins

Yomix: an interactive tool for the exploration of low-dimensional embeddings in omics data.

SUMMARY: In the analysis of diverse omics data, a common and important preliminary step involves computing low-dimensional embeddings using techniques such as PCA, UMAP, t-SNE, or variational autoencoders. These embeddings provide a global overview of sample distributions and their relationships, often serving as the basis for formulating biological hypotheses. To facilitate rapid and intuitive exploration of such low-dimensional embeddings, we developed Yomix, an interactive omics-agnostic visualization and data exploration tool. Yomix enables users to flexibly define subsets of interest using a lasso selection tool, instantly compute their feature signatures, and compare their distributions. Yomix is a fast and efficient tool for interactive exploration of diverse omics datasets. AVAILABILITY AND IMPLEMENTATION: Yomix and its documentation are publicly available at https://github.com/perrin-isir/yomix.

Software

Systematic identification of oscillatory gene expression in single cell types.

Many biological cycles are driven by oscillatory gene expression coordinated across cell types. For example, larval development in Caenorhabditis elegans involves coordinated cyclic changes in cell division, behavior, and growth, the latter requiring production of a structured extracellular matrix called the cuticle. Here, we combine single-cell RNA sequencing and novel computational approaches to identify oscillatory gene expression in individual cell types. We find that many cell types exhibit looping structures in PCA and UMAP space that correspond to transcriptional oscillations at each larval stage. Oscillatory gene expression is found in all cuticle-producing cell types, including glia, but not detected in neurons or muscle. We develop rigorous statistical approaches for de novo identification of oscillatory genes and cell types, yielding >5,000 genes. While many oscillatory genes relate to cuticle production, each cell type expresses largely distinct genes, suggesting that cuticle production is a patchwork of cell-type-specific programs. Finally, we derive a potential set of regulatory transcription factors that can explain coordinated oscillatory gene expression and find that shared upstream factors likely control gene timing across cell types. Together, our results suggest that shared regulators control cell-type-specific oscillatory gene expression, including in previously overlooked cell types such as glia.

Journal Article

The antimicrobial gut resistome of the Wayampi reveals a shared background of antibiotic and metal resistance genes with industrialized populations, underscoring the "robust-yet-fragile" architecture of human gut microbiomes.

BACKGROUND: Metagenomics enables detailed profiling of genes encoding antimicrobial resistance. However, most studies focus exclusively on antibiotic resistance genes (ARGs), excluding those associated with non-antibiotic antimicrobials (metals, biocides), and often rely on methods with low-sensitivity and low-specificity. Furthermore, they rarely examine populations exposed to minimal anthropogenic pollution. We analyzed fecal resistomes of 95 Wayampi individuals, an Indigenous community in remote French Guiana, using a targeted metagenomic capture platform covering 8667 genes, including ARGs, metal resistance genes (MRGs) and biocide resistance genes (BRGs) (PMID: 29335005). Resistome profiles were compared with those of Europeans to assess population-level differences. RESULTS: ARG richness was similar between groups (259 in Wayampi vs. 264 in Europeans, 159 shared), but MRGs&#x2009;+&#x2009;BRGs gene richness was significantly higher in Wayampi (11,930 vs. 7419). Most genes appeared in a minority of individuals (mean 5% for ARGs, 2% for MRGs&#x2009;+&#x2009;BRGs), but several ARGs for tetracyclines [tet(32), tet(40), tet(O), tet(Q), tet(W), tet(X), tetAB(P)], aminoglycosides (ant6'-I, aph3-III), macrolides (ermB, ermF, mefA), and sulfonamides (sul2) were present in all individuals. Tetracycline resistance genes predominated overall, while beta-lactam resistance genes were more common in Wayampi, and genes conferring resistance to aminoglycosides, amphenicols, and folate inhibitors were more frequent in Europeans. Among MRGs, copper and arsenic resistance genes prevailed in both groups, followed by those for zinc, iron, cobalt, and nickel. Up to 76% of Wayampiis carried acquired MRGs for copper (pcoABCDRS and tcrB), silver (silACFPRS), arsenic (ars), and mercury (mer) detoxification. Shannon diversity indices were similar for ARGs, MRGs, and BRGs, but composition and evenness differed significantly. UMAP and ADONIS analyses distinguished cohorts based on ARG profiles (p&#x2009;<&#x2009;0.001), but not on MRGs or BRGs. Correlation analysis revealed conserved gene-sharing networks and introgression of acquired ARGs and MRGs within both gut microbiomes. CONCLUSIONS: The diverse and balanced Wayampi resistome reflects a less perturbed microbiome compared to industrialized populations, and reveals a background of "core" and "shell" acquired ARGs and MRGs, consistent with the "robust-yet-fragile" architecture of scale-free networks. The patchy yet resilient gene distribution suggests varying levels of conserved gene sharing highways among populations, likely shaped by long-term microbial-human evolution, and supports a broader view on acquired antimicrobial resistance. Video Abstract.

Humans

A High-Resolution Stereo-Seq Spatial Transcriptomic Resource for Adult Holstein Cattle Liver.

The bovine liver is a highly compartmentalized organ that plays essential roles in continuous gluconeogenesis and nitrogen recycling; however, its spatial molecular architecture has remained largely uncharacterized due to the limitations of traditional bulk and single-cell approaches. To address this gap, Spatial Enhanced Resolution Omics-sequencing (Stereo-seq) was utilized to generate a subcellular-resolution (500 nm) transcriptomic map of an adult Holstein cattle liver, and a refined reference-guided workflow was implemented to overcome standard annotation limitations in livestock. Raw sequencing data were processed using the Stereo-seq Analysis Workflow and analyzed with Stereopy, Seurat, SingleR, and reference-guided workflows. Spatial aggregation was evaluated at Bin20, Bin50, Bin100, Bin150, and Bin200. Increasing bin size increased molecular identifier counts and detected-gene complexity while progressively reducing spatial granularity. Bin50, corresponding to 50 &#xd7; 50 DNA nanoballs and an approximate nominal footprint of 25 &#xd7; 25 &#xb5;m, was therefore selected as a practical intermediate aggregation level for the primary analyses. Quality-control assessment, Leiden clustering, UMAP visualization, reference-based cell-type annotation, cluster-marker analysis, and spatial mapping of canonical hepatic genes demonstrated preservation of biologically interpretable liver transcriptional organization. Raw sequencing data processed spatial matrices, annotated objects, and analysis code are publicly available to support reanalysis and computational benchmarking. In summary, we present a Stereo-seq spatial transcriptomic resource generated from liver tissue of an adult Holstein cow. This initial resource provides a valuable foundation for future studies of bovine liver biology, comparative genomics, and the spatial basis of livestock health and production traits.

Animals

Charting the phenotypic landscape of mitochondrial diseases through a systematic evaluation of pathogenic mitochondrial DNA and nuclear gene variants.

PURPOSE: Primary mitochondrial diseases (PMD) arise from variants in the mitochondrial or nuclear genomes. Phenotype-based recognition of specific PMD genotypes remains difficult, prolonging the diagnostic odyssey. We expanded the MitoPhen database to characterize phenotypic variation across PMD more systematically. METHODS: Individual-level data on mitochondrial DNA disorders, nuclear-encoded mitochondrial diseases, and single large-scale mitochondrial DNA deletions were manually curated with Human Phenotype Ontology (HPO) terms to produce MitoPhen v2. Principal-component analysis summarized system-level abnormalities; HPO-level enrichment and mean phenotype-similarity scores were then used to distinguish common PMD genotypes. RESULTS: MitoPhen v2 adds 3940 individuals to the original release, now encompassing 1597 publications, 10,626 individuals, and 117 genotypes. Among 7586 affected cases, 72,861 HPO terms were recorded. Principal-component analysis revealed 6 phenotype dimensions capturing most system-level variance. At the HPO level, we observed genotype-specific enrichments and identified 111 gene-phenotype links absent from the current HPO database. Using MT-TL1, single large-scale mitochondrial DNA deletions, and POLG as exemplars, phenotype-similarity scores reliably separated individuals with these genotypes from those without. CONCLUSION: MitoPhen v2 enabled systematic, genotype-aware analysis of heterogeneous PMD phenotypes and highlighted the diagnostic value of structured, individual-level data. Phenotype-similarity metrics from such data sets can refine variant interpretation in large rare-disease cohorts and provide a transferable framework for other phenotypically complex genetic disorders.

Humans

Meningioma transcriptomic landscape demonstrates novel subtypes with regional associated biology and patient outcome.

Meningiomas, although mostly benign, can be recurrent and fatal. World Health Organization (WHO) grading of the tumor does not always identify high-risk meningioma, and better characterizations of their aggressive biology are needed. To approach this problem, we combined 13 bulk RNA sequencing (RNA-seq) datasets to create a dimension-reduced reference landscape of 1,298 meningiomas. The clinical and genomic metadata effectively correlated with landscape regions, which led to the identification of meningioma subtypes with specific biological signatures. The time to recurrence also correlated with the map location. Further, we developed an algorithm that maps new patients onto this landscape, where the nearest neighbors predict outcome. This study highlights the utility of combining bulk transcriptomic datasets to visualize the complexity of tumor populations. Further, we provide an interactive tool for understanding the disease and predicting patient outcomes. This resource is accessible via the online tool Oncoscape, where the scientific community can explore the meningioma landscape.

Meningioma