Computer simulation model provides design framework.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Biomolecular condensates exhibit spontaneous electrochemical microenvironments characterized by asymmetric ion distributions and pH gradients that emerge from protein-sequence-dependent charge regulation. Despite their biological importance, mechanistic understanding of these microenvironments has been constrained by the absence of computationally tractable frameworks capable of treating proton exchange, counterion partitioning, and buffer equilibria on consistent thermodynamic footing. Here, we introduce the buffered Charge-Regulation Monte Carlo (b-CR-MC) framework, which couples grand-canonical exchange of ions and buffer species with explicit charge regulation of titratable residues. By extending the CR-MC ion-merging strategy to multicomponent reservoirs and employing the restricted primitive model, b-CR-MC achieves computational efficiency while maintaining thermodynamic rigor, achievingquantitative agreement with the more expensive generalized grand-reaction Monte Carlo approach. Applied to full-length FUS (net positive) and PGL-3 (net negative) under physiological conditions, the framework reveals sequence-dependent pH gradients: the dense phase of FUS exhibits an alkaline shift, while that of PGL-3 exhibits an acidic shift, in both cases driving the condensate interior toward the protein's isoelectric point. Slab-geometry simulations further resolve the Donnan potential and continuous ion profiles across the condensate interface, confirming the direction of these electrochemical shifts. Additionally, we identify spatially resolved buffer depletion within dense phases, establishing that dynamic charge regulation is a primary determinant rather than a secondary correction to condensate electrochemistry. By establishing a sequence-resolved, thermodynamically consistent computational platform, b-CR-MC enables quantitative prediction of how mutations and post-translational modifications reprogram condensate microenvironments across biological and pathophysiological contexts.
Tissue function emerges from coordinated interactions among diverse cell populations, whereas disruption of these interactions can lead to dysfunction. Recent advances in single-cell and spatial genomics have not only cataloged cellular diversity but also revealed how tissues are organized as dynamic multicellular ecosystems. Moving beyond descriptive cell atlases toward functional, system-level representations represents a major frontier in tissue biology. In this review, we outline conceptual and methodological frameworks for dissecting multicellular coordination, highlight recurrent multicellular ecosystems across physiological and pathological contexts, and explore translational opportunities such as patient stratification, therapeutic reprogramming, and regenerative strategies. Viewing tissues through an ecosystem lens provides a unifying framework that links cellular diversity to emergent tissue function and informs strategies for disease intervention.
Genetic contributions to complex traits are often mediated through coordinated gene-gene interaction networks, yet most existing association frameworks focus on marginal single-gene effects and overlook higher-order dependency structures. Direct modeling of interactions remains challenging due to combinatorial complexity and statistical instability. We introduce Interaction-Bridged Association Study (IBAS), a general framework that incorporates pathway-level interaction patterns into genotype-phenotype association analysis without explicitly enumerating interactions. IBAS leverages transcriptomic reference data to construct low-dimensional representations of pathway activity, which guide SNP-weighting and gene-level association testing within a kernel-based framework. In perturbation-based simulations, IBAS demonstrates improved stability and reproducibility compared to conventional TWAS and gene-based methods, while maintaining well-calibrated Type I error under phenotype permutation. Application to the WTCCC datasets identifies both known and novel genes across multiple complex diseases, including candidates with modest marginal effects missed by standard approaches. These findings are supported by replication in an independent cohort, and analyses across multiple reference tissues revealing both shared and tissue-specific signals. Overall, IBAS provides a statistically robust and computationally tractable framework for incorporating interaction effects into association mapping, extending beyond the single-gene paradigm and enabling more comprehensive characterization of complex trait. IBAS is available on GitHub at: https://github.com/QingrunZhangLab/IBAS.
Accurate diagnosis and progression prediction of Alzheimer's disease (AD) remain challenging due to the heterogeneous nature of the disease, which involves structural brain degeneration, electrophysiological dysfunction, and molecular dysregulation. Most existing deep learning approaches rely on a single modality or limited multimodal combinations, thereby failing to capture the complex cross-domain interactions underlying AD progression. Furthermore, the scarcity of large-scale datasets containing synchronized neuroimaging, electrophysiological, and genomic measurements restricts the development of comprehensive multimodal diagnostic systems. To address these challenges, this study proposes NeuroOmics-Net, a multimodal deep learning framework for Alzheimer's disease analysis that integrates structural magnetic resonance imaging (sMRI), electroencephalography (EEG), and gene expression data. The proposed framework combines a Hierarchical Multi-View Encoder (HME) for modality-specific feature extraction, a Cross-Omics Attention Fusion (CAF) module for adaptive integration of complementary biomarkers, and a Disease Progression Graph Learning (DPGL) module for modeling progression-related relationships across biological domains. To facilitate cross-modal integration from independent cohorts, Regularized Canonical Correlation Analysis (RCCA) is employed to align heterogeneous feature representations within a shared latent space. Experiments were conducted using publicly available datasets from ADNI, PhysioNet, and GEO repositories comprising 1120 diagnosis-aligned samples. The proposed framework achieved 94.3% classification accuracy and an AUC of 0.975 for distinguishing normal controls (NC), mild cognitive impairment (MCI), and Alzheimer's disease subjects, while attaining 93.7% accuracy for predicting conversion from stable mild cognitive impairment (sMCI) to progressive mild cognitive impairment (pMCI). However, a fairness sensitivity analysis using stratified demographic reweighting revealed accuracy ranging from 90.8% (low-education, high-comorbidity proxy subgroup) to 96.1% (low-risk, high-reserve proxy subgroup), a demographic parity gap of 5.3 percentage points, indicating that overall accuracy reflects a performance ceiling in a relatively homogeneous research cohort rather than a realistic estimate for demographically diverse clinical populations. Comparative evaluations demonstrated consistent improvements over state-of-the-art unimodal and multimodal deep learning models. Interpretability analysis further identified clinically relevant biomarkers, including hippocampal and entorhinal atrophy, theta-alpha EEG alterations, and APOE-associated molecular pathways. Because sMRI, EEG, and gene expression data were sourced from separate, unpaired cohorts with no subjects possessing all three synchronized measurements, all reported cross-modal associations reflect population-level statistical correspondence across diagnosis-matched groups rather than within-subject physiological coupling; no claim of intra-individual causal cross-modal interaction is made. These findings demonstrate that NeuroOmics-Net provides an effective computer-aided framework for multimodal biomedical data processing and Alzheimer's disease analysis. By integrating neuroimaging, electrophysiological, and genomic information, the proposed approach enables accurate diagnosis, progression prediction, and biologically interpretable decision support for clinical and translational applications.
Microbial colony growth is shaped by the physics of biomass propagation and nutrient diffusion and by the metabolic reactions that organisms activate as a function of the surrounding environment. While microbial colonies have been explored using minimal models of growth and motility, full integration of biomass propagation and metabolism is still lacking. Here, building upon our framework for computation of microbial ecosystems in time and space (COMETS), we combine dynamic flux balance modeling of metabolism with collective biomass propagation and demographic fluctuations to provide nuanced simulations of E. coli colonies. Simulations produced realistic colony morphology, consistent with our experiments. They characterize the transition between smooth and furcated colonies and the decay of genetic diversity. Furthermore, we demonstrate that under certain conditions, biomass can accumulate along "metabolic rings" that are reminiscent of coffee-stain rings but have a completely different origin. Our approach is a key step toward predictive microbial ecosystems modeling. A record of this paper's transparent peer review process is included in the supplemental information.
The rapid global expansion of antimicrobial resistance (AMR) threatens to undermine decades of progress in infectious disease management and highlights the limitations of conventional antibiotic-centered therapeutic strategies. Although emerging technologies-including antimicrobial peptides, bacteriophage therapy, CRISPR-based antimicrobials, microbiome therapeutics, anti-virulence approaches, nanotechnology-enabled drug delivery, and artificial intelligence (AI)-have individually demonstrated considerable promise, they are predominantly being developed as independent interventions rather than as coordinated components of an integrated therapeutic strategy. This Perspective proposes the Intelligent Anti-Infective Ecosystem (IAIE) as a conceptual systems-level framework that computationally integrates multimodal diagnostics, pathogen genomics, microbiome profiling, AI-assisted decision support, programmable precision therapeutics, ecological monitoring, and longitudinal clinical feedback within a continuously learning dynamically optimized workflow. Unlike existing paradigms that primarily optimize individual technologies or therapeutic decisions, IAIE emphasizes closed-loop coordination among complementary antimicrobial approaches to support precision-guided infection management while preserving microbiome integrity and mitigating resistance selection pressure. We further outline the core components, operational principles, translational challenges, and technology readiness of the major therapeutic platforms that could contribute to such an ecosystem, while distinguishing clinically established interventions from emerging experimental strategies. Importantly, IAIE should be interpreted as a prospective conceptual architecture rather than an existing clinical platform. Its proposed clinical value remains to be established through sequential computational, preclinical, and prospective clinical investigations using standardized microbiological, ecological, and patient-centered outcome measures. By framing antimicrobial innovation within an responsive systems perspective, IAIE provides a roadmap for future multidisciplinary research aimed at integrating artificial intelligence and systems microbiology to enable sustainable management of antimicrobial resistance.
Explore the source record for details and available documents.
MOTIVATION: Wikipedia is a vital open educational resource in computational biology; however, a significant knowledge gap exists between English and non-English Wikipedias. Reducing this knowledge gap via intensive editing events, or "editathons," would be beneficial in reducing language barriers that disadvantage learners whose native language is not English. Results: We present a framework to guide educators in organizing editathons for learners to improve and create relevant Wikipedia articles. As a case study, we present the results of an editathon held at the 2024 ISCB Latin America conference, in which ten new articles were created for the Spanish-language edition of Wikipedia. We also present a web tool, "compbio-on-wiki," which identifies relevant English Wikipedia articles missing in other languages. We demonstrate the value of editathons to expand the accessibility and visibility of computational biology content in multiple languages. AVAILABILITY AND IMPLEMENTATION: Source code for the compbio-on-wiki Toolforge site is available at: https://github.com/lubianat/compbio-on-wiki.
Protein-based antibacterials such as bacteriophage endolysins offer a targeted therapeutic strategy against Gram-positive pathogens. However, prioritizing the most effective candidates from the large sequence diversity available remains a significant challenge. Here we present a standardized computational-experimental benchmarking framework that evaluates seven phage-derived endolysin variants (E1, E2, E3, E7, E10, E12, and E15) identified from Bacillus genomes. We combined molecular docking and residue-level interaction mapping against muramyl dipeptide (MDP), a minimal conserved peptidoglycan motif, with 1000-ns molecular dynamics simulations, MM/PBSA binding free-energy estimation, and matched functional inhibition assays against Staphylococcus aureus and Micrococcus luteus. Computational analyses revealed generally favorable MDP recognition across variants, albeit with notable differences in contact patterns and complex stability profiles. Experimental screening identified E2 as the most potent antibacterial agent against both species, while E7 and E1 performed strongly in selected computational metrics. Integrated analysis showed only modest correlations between computational descriptors of fragment recognition/stability and observed antibacterial performance. This study establishes a practical comparative benchmarking platform for endolysin candidate prioritization, nominates E2 and E7 as promising candidates for further development, and highlights E1 as a potential structural scaffold for rational engineering, while explicitly demonstrating both the utility and the current limitations of using minimal peptidoglycan fragments as proxies for full cell-wall recognition in lysin benchmarking.
We describe the preverbal system of counting and arithmetic reasoning revealed by experiments on numerical representations in animals. In this system, numerosities are represented by magnitudes, which are rapidly but inaccurately generated by the Meck and Church (1983) preverbal counting mechanism. We suggest the following. (1) The preverbal counting mechanism is the source of the implicit principles that guide the acquisition of verbal counting. (2) The preverbal system of arithmetic computation provides the framework for the assimilation of the verbal system. (3) Learning to count involves, in part, learning a mapping from the preverbal numerical magnitudes to the verbal and written number symbols and the inverse mappings from these symbols to the preverbal magnitudes. (4) Subitizing is the use of the preverbal counting process and the mapping from the resulting magnitudes to number words in order to generate rapidly the number words for small numerosities. (5) The retrieval of the number facts, which plays a central role in verbal computation, is mediated via the inverse mappings from verbal and written numbers to the preverbal magnitudes and the use of these magnitudes to find the appropriate cells in tabular arrangements of the answers. (6) This model of the fact retrieval process accounts for the salient features of the reaction time differences and error patterns revealed by experiments on mental arithmetic. (7) The application of verbal and written computational algorithms goes on in parallel with, and is to some extent guided by, preverbal computations, both in the child and in the adult.
Accurate coordinates for a represented protein state do not, by themselves, establish activity or any other condition-specific function. This article defines the Dynamic Protein Structure Paradox (DPSP) as the apparent conflict between structural accuracy and functional underdetermination and develops it as an integrative evidentiary assessment framework rather than a new theory or paradigm. The underlying problem has been longstanding, since structural genomics, function annotation, allostery, and disorder research each established that fold does not determine function and that function does not determine fold. DPSP consolidates those results into one endpoint-conditioned rule. Once a measurable endpoint is defined, it assesses four coupled dimensions: relevant-state completeness, context completeness, ensemble or kinetic dependence, and chemical dependence. A rubric rates each dimension as adequate, uncertain, or missing, and a materiality test determines which gaps influence the stated decision. The outcome is one of three mutually exclusive modes of utilization: geometry-led, conditional, or function-measured. The deliverable is a concise evidence statement delineating what the structure supports, which decisive variable remains unmeasured, and what corroboration is necessary. DPSP complements, rather than replaces, existing structural, ensemble, and computational approaches. The framework remains unvalidated, its thresholds are provisional, and the studies necessary to confirm or refute it are specified.
Small heat-shock proteins (sHSPs, the HSP20 family) are ATP-independent molecular chaperones that hold partially unfolded substrates and protect the proteome during heat and other abiotic stresses; every member is defined by a conserved α-crystallin domain (ACD). Finger millet (Eleusine coracana) is a climate-resilient, calcium-rich allotetraploid cereal of the semi-arid tropics whose HSP20 repertoire had not been catalogued. The present study is an entirely computational (in silico) analysis of the chromosome-scale reference genome of finger millet (NCBI GenBank assembly GCA_032690845.1, cultivar KNE 796-S). Mining the predicted proteome with the ACD profile (Pfam PF00011) and confirming every candidate by NCBI CD-search recovered 76 non-redundant ACD-bearing HSP20 genes (EcHSP20-1-EcHSP20-76). Based on phylogeny and TargetP-predicted localization, the members were classified into ten subfamilies: seven cytosolic/nuclear classes (C-I to C-VII, 60 members) together with chloroplastic (11), mitochondrial (3) and endoplasmic-reticulum (2) groups. The proteins ranged from 110 to 355 amino acids (12.1-39.2 kDa) with theoretical pI of 4.85-9.69. The 76 loci were distributed over 14 of the 18 chromosomes and were conspicuously absent from chromosomes 8 A, 8B, 9 A and 9B, with pronounced clustering on chromosomes 1, 2, 3 and 6. Duplication analysis detected 149 paralogous pairs (49 homoeologous, 80 segmental/dispersed and 18 tandem); 147 of 148 pairs for which substitution rates could be calculated returned Ka/Ks < 1 (mean 0.20), indicating strong purifying selection consistent with retention after whole-genome/allopolyploid duplication. Promoter analysis (PlantCARE) revealed enrichment of abscisic-acid-responsive (ABRE), MYB/MYC drought-related, STRE, DRE, low-temperature (LTR) and methyl-jasmonate/salicylic-acid elements, whereas canonical heat-shock elements (HSE) were not recovered. Expression profiling against a public drought transcriptome (SRP081350) showed that about half of the genes (39 of 76) are transcribed in leaf tissue, the expressed fraction being dominated by the cytosolic class C-I. This first finger-millet HSP20 catalogue provides a verified, reproducible framework and nominates computationally predicted candidate genes for future functional work on thermotolerance in cereals.
In commercial pig production, many important traits are recorded as binary phenotypes. For such traits, threshold models offer an appropriate framework but are computationally intensive. Thus, linear models are widely used to obtain genomic estimated breeding values (GEBV); however, these are on the observed scale (phenotypic). This creates the need for a robust method to approximate GEBV from linear models to the liability scale. A recently proposed approximation showed good concordance for low-prevalence traits (<5%) but has not yet been tested for a wider range of prevalence values and for models with more than one random effect. We aimed to evaluate the performance of this approximation for pig binary traits with prevalences ranging from <5% to >86%, in both animal and maternal animal models. Data were available for five fitness traits (FT1-FT5), with up to 233k animals with phenotypes, of which 204k animals were genotyped with a 25k SNP array. Variance component estimates were obtained using threshold models. Classical animal models were used for FT1-FT3, and maternal animal models for FT4 and FT5. Variance components on the observed scale were then obtained by multiplying estimates from a threshold model by the square of the height of the standard normal density evaluated at the threshold. GEBV were predicted using single-step genomic best linear unbiased prediction under both linear and threshold models. The approximation tested involved scaling the GEBV using the height of the ordinate of the standard normal distribution evaluated at the threshold as a scaling factor. The agreement between GEBV from the scaled linear model and the threshold model on the probability scale was evaluated using Pearson and Spearman correlations, mean squared error (MSE), regression parameters, overlapping coefficient (OVL), distribution overlap, and classification accuracy (CACC). Correlations between linear and threshold GEBV ranged from 0.94 (low-prevalence traits) to 0.99 (high-prevalence traits) for the direct GEBV and were 0.99 for the maternal GEBV. MSE were close to zero. The OVL exceeded 0.83 for all traits. CACC ranged from 95.10% to 98.33% for the direct GEBV and from 92.54% to 97.42% for the maternal GEBV. Regardless of model and trait prevalence, this approximation yielded GEBV that are highly consistent with threshold model GEBV, providing a reliable, practical approach for large-scale pig genetic evaluations for binary traits using linear models.
MOTIVATION: Identifying promising therapeutic targets from thousands of genes in transcriptomic studies remains a major bottleneck in biomedical research. While large language models (LLMs) show potential for gene prioritization, they suffer from hallucination and lack systematic validation against expert knowledge. RESULTS: The framework identified 609 sepsis-relevant genes with >94% filtering efficiency, demonstrating strong enrichment for inflammatory pathways including TNF-α signaling, complement activation, and interferon responses. Literature validation yielded 30 ultra-high confidence therapeutic candidates, including both established sepsis genes (IL10, TREM1, S100A9, NLRP3) and novel targets warranting investigation. Benchmark validation against expert-curated databases achieved 71.2% recall, with systematic correlation between computational confidence and evidence quality. The final candidate set balanced discovery (11 novel genes) with validation (19 known genes), maintaining biological coherence throughout the filtering process. This framework demonstrates that rigorous methodology can transform unreliable LLM outputs into systematically validated biological insights. By combining computational efficiency with literature grounding, the approach provides a practical tool for prioritizing experimental validation efforts. The modular design enables adaptation to other diseases through knowledge base substitution, offering a systematic approach to literature-guided biomarker discovery. AVAILABILITY AND IMPLEMENTATION: We developed a two-stage computational framework that combines LLM-based screening with literature validation for systematic gene prioritization. Starting with 10 824 genes from the BloodGen3 repertoire, we applied multi-criteria evaluation for sepsis relevance, followed by retrieval-augmented generation using 6346 curated sepsis publications. A novel faithfulness evaluation system verified that LLM predictions aligned with retrieved literature evidence. Source code and implementation details are available at https://github.com/taushifkhan/llm-geneprioritization-framework, vector database at https://doi.org/10.5281/zenodo.15802241, and Interactive demonstration at https://llm-geneprioritization.streamlit.app/.
This comprehensive narrative review examines recent advances in multi-omics research for Systemic Lupus Erythematosus (SLE), emphasizing integrated approaches over single-omics studies. The review critically evaluates technological advancements, methodological innovations, and clinical applications while identifying current limitations and future research directions. We conducted a comprehensive narrative review following SANRA guidelines, searching PubMed, Web of Science, Scopus, and Embase, covering publications from January 2018 to June 2025. The review focuses on studies integrating two or more omics layers in SLE research, with emphasis on computational methods, biomarker validation, and clinical applications. Multi-omics integration has revealed critical insights into SLE pathogenesis, including immune cell heterogeneity, gene-environment interactions, and metabolic dysregulation. However, significant challenges remain in data integration methodologies, small sample sizes, and biomarker reproducibility. Current computational approaches include early integration (concatenation), intermediate integration (joint dimensionality reduction), and late integration (ensemble methods). While multi-omics approaches offer unprecedented insights into SLE complexity, standardized integration protocols and robust validation frameworks are urgently needed. Small sample sizes and heterogeneity issues limit reproducibility, particularly affecting biomarker discovery and clinical translation. Multi-omics integration represents a paradigm shift toward precision medicine in SLE, but realizing this potential requires addressing current methodological limitations, standardizing validation processes, and developing robust computational frameworks for reliable clinical applications.
BACKGROUND: African swine fever virus (ASFV) and porcine epidemic diarrhea virus (PEDV) differ in viral biology and cellular tropism, yet both pathogens suppress macrophage-mediated immune responses in pigs. OBJECTIVE: To identify a conserved macrophage suppression module shared by ASFV and PEDV and evaluate quantum computing as an independent framework for biological network validation. METHODS: Integrated analysis of publicly available GEO datasets (GSE231435 for ASFV and GSE306895) identified 471 shared downregulated genes. A network- and multi-omics-informed 20-gene core was selected and encoded as a 20-qubit modularity-based Quadratic Unconstrained Binary Optimization (QUBO) problem. Community detection was benchmarked using the Quantum Approximate Optimization Algorithm (QAOA) on both the IBM Quantum Aer simulator and the 156-qubit IBM Fez (Heron r2) quantum processor and compared with brute-force enumeration and simulated annealing. RESULTS: A conserved macrophage suppression module shared by ASFV and PEDV was identified. For the STRING protein-protein interaction network, QAOA at circuit depth p = 3 reproduced the brute-force optimum with an approximation ratio of 1.000. In contrast, performance progressively declined in the denser co-expression network with increasing circuit depth, consistent with noise accumulation under current Noisy Intermediate-Scale Quantum (NISQ) conditions. Multi-run consensus analysis identified stable hub genes, including MMP9 and SLA-DOA, as well as genes exhibiting variable community assignments. CONCLUSION: These findings reveal a conserved macrophage suppression module shared between ASFV and PEDV and demonstrate that quantum computing can serve as an independent validation framework for biologically meaningful host-response networks. Network topology emerged as a key determinant of QAOA performance on real NISQ hardware.
BACKGROUND: Despite major advances in serologic testing, extended phenotyping, and blood group genomics, clinically similar transfusion exposures may result in markedly different immune and clinical outcomes. Existing compatibility strategies do not fully explain this biological variability. OBJECTIVES: To examine transfusion compatibility as an emergent donor-recipient biological state and propose a systems-level conceptual framework that integrates established biological determinants into a testable model for future precision transfusion medicine. METHODS: This narrative review critically synthesizes current evidence from blood group genomics, recipient immunobiology, inflammation, disease-specific biology, transfusion medicine, and computational prediction. The proposed framework distinguishes Compatibility Intelligence Theory (CIT) as a biological interpretation from Precision Transfusion Intelligence (PTI) as its potential clinician-supervised translational application. RESULTS: The review argues that transfusion compatibility is shaped by interactions among donor genetics, recipient immune biology, inflammatory physiology, disease context, transfusion history, and longitudinal adaptation rather than by antigen matching alone. CIT provides an organizational framework for integrating these determinants, whereas PTI describes a possible clinician-supervised translation. To address current feasibility, the revised framework separates variables into routinely measurable, contextually available but incompletely standardized, and research-stage domains, and proposes a staged strategy for deriving rather than assuming their quantitative weights. Any clinical implementation would require comparative validation against current serologic, phenotypic, and genotype-based practice. CONCLUSIONS: Compatibility Intelligence Theory offers a testable systems-level framework for understanding transfusion compatibility without replacing established transfusion practices. The framework is not presented as a ready-to-use score: currently measurable variables can be organized for structured risk review, whereas inflammatory, immunogenetic, and multi-omic inputs require prospective standardization and validation. If future studies demonstrate incremental predictive and patient-centered benefit, CIT-informed PTI could support an adaptive, evidence-based extension of current precision transfusion practice.