Search PubMedSearch

SEARCH · Search PubMed

Results for “Mixture of experts”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

7 recordsLinked to original sources

Unifying multimodal single-cell data with a mixture-of-experts β-variational autoencoder framework.

Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. Here, we present Unified Variational Inference (UniVI), a scalable mixture-of-experts β-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or preannotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein (CITE-seq) and RNA-chromatin (10x Genomics Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin-a nonhematopoietic tissue with continuous differentiation hierarchies-UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to trimodal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA-protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, trimodal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

Journal Article

Breast Cancer Recurrence Status Assessment in 5 Years Using Multimodal Integrated Learning: A Feasibility Study.

Despite advances in breast cancer detection and treatment, recurrence after curative therapy continues to impact long-term survival and quality of life. Therefore, early identification of high-risk patients is crucial to guide personalized treatment and follow-up strategies. Although genomic assays provide valuable prognostic insights, their high cost and limited accessibility hinder widespread adoption in clinical practice. Recent machine learning or deep learning approaches leveraging clinical, imaging, or multimodal data have shown promise but do not reflect real-world clinical scenarios. This study proposes a deep learning-based multimodal framework for predicting 5-year breast cancer recurrence using routinely collected clinical data. The framework consists of three main components. First, we adopted automated tumor segmentation with MedSAM to extract the tumor region from ultrasound images. The radiomics features are extracted from those tumor regions. Second, report features are extracted using a Med-Contrastive Pre-trained Transformers (MedCPT)-based approach incorporating predefined, clinically informed queries. Third, a multimodal integration model jointly processes image, radiomics, clinical features, and report features through modality-specific branches. The image branch employs the Ultrasound Foundation Model (USFM) as the backbone, while structured tabular data is processed using the FT-Transformer architecture. The features of all branches are fused using a mixture-of-experts (MoE)-based classifier, and the entire model is trained using a progressive fusion training strategy. Experimental results confirm the feasibility of using ultrasound images with tumor mask integration for recurrence prediction and demonstrate the additive value of integrating multiple data modalities through the proposed multimodal integration model. The final model for recurrence prediction achieved an AUC of 0.7540, accuracy of 74.61%, sensitivity of 70.41%, and specificity of 76.44%. This feasibility study's findings underscore the potential of the proposed multimodal deep learning framework to provide accessible, accurate, and generalizable recurrence risk prediction using routinely available clinical data, potentially supporting more informed treatment decisions and personalized post-treatment monitoring in real-world clinical practice.

Breast cancer recurrence

CodonMoE: DNA language models for codon-dependent mRNA prediction.

MOTIVATION: Genomic language models (gLMs) face a fundamental efficiency challenge: one must either maintain separate specialized models for each biological modality (DNA and RNA) or develop large multimodal architectures. Both approaches impose significant computational burdens-modality-specific models require redundant infrastructure despite inherent biological connections, while multi-modal architectures demand increased parameter counts and extensive cross-modality pretraining. RESULTS: To address this limitation, we introduce CodonMoE (Adaptive Mixture of Codon Reformative Experts), a lightweight adapter that transforms DNA language models into effective RNA analyzers without RNA-specific pretraining. Our theoretical analysis establishes CodonMoE as a universal approximator at the codon level, capable of mapping arbitrary functions from codon sequences to codon-dependent RNA properties given sufficient expert capacity. Across four RNA prediction tasks spanning stability, expression, and regulation, DNA models augmented with CodonMoE significantly outperform their unmodified counterparts, with the HyenaDNA+CodonMoE series achieving state-of-the-art results using 80% fewer parameters than specialized RNA models. By maintaining sub-quadratic complexity while achieving superior performance, our approach provides a principled path toward unifying genomic language modeling, leveraging more abundant DNA data and reducing computational overhead while preserving modality-specific performance advantages. AVAILABILITY AND IMPLEMENTATION: Source code for the method and to reproduce the results is available at https://github.com/Kingsford-Group/CodonMoE.

Codon

Expert opinion elicitation for assisting deep learning based Lyme disease classifier with patient data.

BACKGROUND: Diagnosing erythema migrans (EM) skin lesion, the most common early symptom of Lyme disease, using deep learning techniques can be effective to prevent long-term complications. Existing works on deep learning based EM recognition only utilizes lesion image due to the lack of a dataset of Lyme disease related images with associated patient data. Doctors rely on patient information about the background of the skin lesion to confirm their diagnosis. To assist deep learning model with a probability score calculated from patient data, this study elicited opinions from fifteen expert doctors. To the best of our knowledge, this is the first expert elicitation work to calculate Lyme disease probability from patient data. METHODS: For the elicitation process, a questionnaire with questions and possible answers related to EM was prepared. Doctors provided relative weights to different answers to the questions. We converted doctors' evaluations to probability scores using Gaussian mixture based density estimation. We exploited formal concept analysis and decision tree for elicited model validation and explanation. We also proposed an algorithm for combining independent probability estimates from multiple modalities, such as merging the EM probability score from a deep learning image classifier with the elicited score from patient data. RESULTS: We successfully elicited opinions from fifteen expert doctors to create a model for obtaining EM probability scores from patient data. CONCLUSIONS: The elicited probability score and the proposed algorithm can be utilized to make image based deep learning Lyme disease pre-scanners robust. The proposed elicitation and validation process is easy for doctors to follow and can help address related medical diagnosis problems where it is challenging to collect patient data.

Humans

Emerging protein sequencing technologies: proteomics without mass spectrometry?

INTRODUCTION: Liquid chromatography-tandem mass spectrometry (LC-MS/MS) has been a leading method for proteomics for 30 years. Advantages provided by LC-MS/MS are offset by significant disadvantages, including cost. Recently, several non-mass spectrometric methods have emerged, but little information is available about their capacity to analyze the complex mixtures routine for mass spectrometry. AREAS COVERED: We review recent non-mass-spectrometric methods for sequencing proteins and peptides, including those using nanopores, sequencing by degradation, reverse translation, and short-epitope mapping, with comments on bioinformatics challenges, fundamental limitations, and areas where new technologies will be more or less competitive with LC-MS/MS. In addition to conventional literature searches, instrument vendor websites, patents, webinars, and preprints were also consulted to give a more up-to-date picture. EXPERT OPINION: Many new technologies are promising. However, demonstrations that they outperform mass spectrometry in terms of peptides and proteins identified have not yet been published, and astute observers note important disadvantages, especially relating to the dynamic range of single-molecule measurements of complex mixtures. Still, even if the performance of emerging methods proves inferior to LC-MS/MS, their low cost could create a different kind of revolution: a dramatic increase in the number of biology laboratories engaging in new forms of proteomics research.

Proteomics

Bayesian classification of OXPHOS deficient skeletal myofibres.

Mitochondria are organelles in most human cells which release the energy required for cells to function. Oxidative phosphorylation (OXPHOS) is a key biochemical process within mitochondria required for energy production and requires a range of proteins and protein complexes. Mitochondria contain multiple copies of their own genome (mtDNA), which codes for some of the proteins and ribonucleic acids required for mitochondrial function and assembly. Pathology arises from genetic defects in mtDNA and can reduce cellular abundance of OXPHOS proteins, affecting mitochondrial function. Due to the continuous turn-over of mtDNA, pathology is random and neighbouring cells can possess different OXPHOS protein abundance. Estimating the proportion of cells where OXPHOS protein abundance is too low to maintain normal function is critical to understanding disease severity and predicting disease progression. Currently, one method to classify single cells as being OXPHOS deficient is prevalent in the literature. The method compares a patient's OXPHOS protein abundance to that of a small number of healthy control subjects. If the patient's cell displays an abundance which differs from the abundance of the controls then it is deemed deficient. However, due to the natural variation between subjects and the low number of control subjects typically available, this method is inflexible and often results in a large proportion of patient cells being misclassified. These misclassifications have significant consequences for the clinical interpretation of these data. We propose a single-cell classification method using a Bayesian hierarchical mixture model, which allows for inter-subject OXPHOS protein abundance variation. The model accurately classifies an example dataset of OXPHOS protein abundances in skeletal muscle fibres (myofibres). When comparing the proposed and existing model classifications to manual classifications performed by experts, the proposed model results in estimates of the proportion of deficient myofibres that are consistent with expert manual classifications.

Oxidative Phosphorylation

'Truthsets' for clinical validation of large-scale functional assays: Practice recommendations from Cancer Variant Interpretation Group UK (CanVIG-UK).

BACKGROUND: Large-scale functional assays, including multiplex assays of variant effect, have substantial potential to resolve variants of uncertain significance (VUS), particularly for rare missense variants where clinical and population evidence are limited. The ClinGen assay-level clinical validation framework described by Brnich et al provided baseline guidance for the use of functional data for variant classification. However, clear consensus regarding construction of variant 'truthsets' by which to clinically validate functional data remains lacking. METHODS: CanVIG-UK developed consensus recommendations for truthset construction through an iterative national consultation process involving the CanVIG Steering Advisory Group (CStAG), wider CanVIG-UK membership, and engagement with international functional genomics experts. Consultation was based on previous analyses of 2,120 truthset constructions examining the impact of truthset composition on evidence point allocation within the ClinGen assay-level clinical validation framework. RESULTS: Across several consultations, CanVIG-UK established nine guiding principles and seven best-practice recommendations for assay-level clinical validation, using the assumed context of an assay for a cancer susceptibility gene where loss-of-function is the mechanism of pathogenicity. The principal recommendation stipulates, where assays are intended for use in interpretation of largely missense variants, the truthset used to validate should comprise only missense variants. Rather than mixtures of different variant types which may serve to over-estimate assay performance. Additional recommendations support option for relaxation of truthset stringency to improve power, augmentation of benign missense truthsets with systematically derived 'proxy-clinical' benign variants, independent clinical validation separate from assayist-defined validation, and careful evaluation of missense score distributions against that of protein-truncating and synonymous variants. Guidance is also provided for scenarios with limited pathogenic truthset availability and for assays reporting multiple deleterious zones or readouts. CONCLUSIONS: The CanVIG-UK principles and recommendations for truthset construction upon the ClinGen assay-level clinical validation framework, while aiming to form a baseline for future discussion regarding other functional and disease contexts and helping to address the gap between publication of new data and routine clinical implementation.

Journal Article