Search PubMedSearch

SEARCH · Search PubMed

Results for “computational frameworks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

A scalable HPC framework for bioinformatics in resource-limited settings: design principles, implementation, and sustainability from the UVRI experience.

MOTIVATION: Building and sustaining High-Performance Computing (HPC) infrastructure for bioinformatics research in resource-limited settings presents significant technical, financial and operational challenges. Institutions in low-and middle-income regions often face constraints such as limited technical expertise, unstable infrastructure and restricted funding which can hinder the deployment of large-scale computational platforms necessary for modern genomics and bioinformatics analyses. RESULTS: We present a scalable and modular HPC framework developed at the Uganda Virus Research Institute (UVRI) to support large-scale genomics and other omics data analyses in resource-limited settings. The framework integrates open-source HPC management tools, infrastructure automation, and reproducible configuration management to enable reliable deployment and maintenance. Optimized storage and networking configurations combined with a phased capacity-building strategy support high-throughput genomic workflows while strengthening local technical expertise. From our implementation experience, we derive ten practical design and operational rules that provide a transferable methodology for establishing and sustaining in-house HPC infrastructure. These rules emphasize strategic investment in human capacity, structured planning, leveraging collaborations, adoption of open-source technologies and service management practices to improve operational resilience and long-term sustainability. AVAILABILITY: The design principles, automation strategies and implementation guidelines described in this work are applicable to institutions seeking to establish sustainable HPC resources for bioinformatics research in resource-constrained environments.

Computational Biology

Information technology and computer-based decision support in diabetic management.

This paper describes the application of computer-based techniques within an intelligent, knowledge-based framework to the management of diabetes. The objectives are to structure data collection and storage so that the relevant patient-specific data are collected and made accessible as needed, and to provide clinical decision support on either a day-by-day or longer timescale as appropriate; these objectives relating to both hospital clinic and general practice. For longer-term management, a prototype rule set (greater than 500 rules) has been developed (coded in Sigma PROLOG), validated and tested on patient data. The data collection programs (written in SCULPTOR) to feed the ruleset have been tested in the hospital clinic and compared with the resident data collection system for usability, and impact on the running of the clinic. Links between the data collection programs and the ruleset program have been written and tested. The computer system will also incorporate a module, combining knowledge-based advisory system and glucose/insulin model as patient simulator, that can be tested as a potential decision aid for adjusting insulin dosage on a daily basis.

Data Collection

Collective posterior inference from highly variable empirical replicates.

High-throughput experimental platforms now routinely generate data from dozens or hundreds of independent observations. Simulation-based inference (SBI) offers a powerful framework for estimating model parameters from such complex datasets, but standard methods struggle to scale to the noisy multiple-replicates regime without incurring prohibitive computational costs or careful hyperparameter tuning. Here, we introduce a new method for fast and robust collective posterior inference from multiple independent replicates using a robust product-of-experts aggregation scheme that automatically mitigates the influence of outliers. Evaluating it on synthetic and empirical evolutionary datasets, we find it achieves state-of-the-art estimation accuracy and computational efficiency, including inference from noisy observations. Our method is compatible with any SBI framework, providing a scalable, plug-and-play solution for inference from noisy multiple-replicate datasets.

Computational Biology

FIERCE: reconstructing dynamic trajectories from the differentiation potency of single cells.

MOTIVATION: Since the introduction of single-cell RNA sequencing (scRNA-seq), numerous computational approaches have been developed to reconstruct dynamic cellular processes from static transcriptional profiles. These methods order cells along continuous trajectories by assessing their similarity in the gene-expression space. However, they rely on several assumptions, such as prior knowledge of the structure and directionality of the expected genealogy. These assumptions can limit their application to complex cellular systems with poorly understood developmental paths. RESULTS: To address this challenge, we introduce FIERCE (Framework for InfERence of the veloCity of Entropy), a novel computational pipeline designed to predict the changes in the differentiation potency of single cells during dynamic processes. Through a fully unsupervised approach, FIERCE enables the inference of cell lineages directly on the differentiation landscape of the biological system, thus eliminating the need for prior specification of developmental parameters. We demonstrate the efficacy of FIERCE by reconstructing three well-known mouse differentiation systems and by quantifying its accuracy on simulated data. AVAILABILITY AND IMPLEMENTATION: The FIERCE R package is available on GitHub at https://github.com/bicciatolab/FIERCE.

Cell Differentiation

Evaluation of a generalized model of human postural dynamics and control in the sagittal plane.

A two-step identification method is used to evaluate a generalized model of human postural control in the sagittal plane. Postural dynamics are represented as a planar open-chain linkage system supported by a triangular foot. The control mechanism is modeled as a state feedback element in which the torque acting at a given link is an arbitrary function of the state variables--angles and angular velocities. To validate the approach, six normal subjects underwent two series of experiments. The first series were used to determine an appropriate model of the system dynamics. The second series were used to estimate the parameters of the feedback model. A computer simulation of the complete system shows that the model predictions closely match the observed responses. These results suggest that the proposed model provides a useful framework for analysis of postural control mechanisms.

Computer Simulation

CancerOmicsStudio (CoS): a web server for integrative and interpretable analysis of multi-omics cancer data.

MOTIVATION: Large-scale omics resources, including The Cancer Genome Atlas, Genomics of Drug Sensitivity in Cancer, and the Cancer Dependency Map, have become essential for cancer research. However, these datasets are distributed across different platforms, formats and analysis frameworks, which limits their practical use by researchers without extensive computational expertise. RESULTS: We developed CancerOmicsStudio (CoS), a web server for integrative and interpretable analysis of multi-omics cancer data across 33 cancer types. CoS provides five major modules: CosAI, Traditional Analysis, Drug Sensitivity, CRISPR Dependency and Single-Cell Tumor Microenvironment. The Traditional Analysis module supports expression comparison, diagnostic evaluation, survival analysis, enrichment analysis and gene correlation. The Drug Sensitivity and CRISPR Dependency modules enable systematic evaluation of gene-drug response associations and gene essentiality in cancer cell lines. The Single-Cell Tumor Microenvironment module supports tumor microenvironment analysis at single-cell resolution. In total, approximately 1.23 million results have been precomputed to enable rapid retrieval. CosAI further allows users to submit natural-language queries and obtain results through a Real-time Analysis as Retrieval framework, with responses summarized by a lightweight language model. AVAILABILITY AND IMPLEMENTATION: CancerOmicsStudio is freely available at Zenodo (doi: 10.5281/zenodo.18744990) and https://cos.wanglab.bio.

Humans

Testing between the TRACE model and the fuzzy logical model of speech perception.

The TRACE model of speech perception (McClelland & Elman, 1986) is contrasted with a fuzzy logical model of perception (FLMP) (Oden & Massaro, 1978). The central question is how the models account for the influence of multiple sources of information on perceptual judgment. Although the two models can make somewhat similar predictions, the assumptions underlying the models are fundamentally different. The TRACE model is built around the concept of interactive activation, whereas the FLMP is structured in terms of the integration of independent sources of information. The models are tested against test results of an experiment involving the independent manipulation of bottom-up and top-down sources of information. Using a signal detection framework, sensitivity and bias measures of performance can be computed. The TRACE model predicts that top-down influences from the word level influence sensitivity at the phoneme level, whereas the FLMP does not. The empirical results of a study involving the influence of phonological context and segmental information on the perceptual recognition of a speech segment are best described without any assumed changes in sensitivity. To date, not only is a mechanism of interactive activation not necessary to describe speech perception, it is shown to be wrong when instantiated in the TRACE model.

Adult

Surgical pathology of cancer of the oral cavity and oropharynx.

A study was designed to determine the influence of certain surgical pathologic findings on tumor spread and survival in patients with cancer of the oral cavity and oropharynx. All patients with the histopathological diagnosis of carcinoma of the oral cavity or oropharynx from 1955 to 1983 were included in the study. Using the Head and Neck Tumor Registry of the department of otolaryngology of the Washington University School of Medicine, information was obtained regarding preoperative evaluation, staging, classification, diagnosis, treatment, surgical pathology parameters, and outcome results. The patient populations consisted of 545 patients with oral cavity cancer and 224 patients with oropharynx cancer, all of whom were eligible for 3-year follow-up. Information from a retrospective analysis of the pretreatment examination records regarding site and size of the primary tumor and neck dissection, and specific treatment, and from surgical pathology reports regarding site, size, tumor spread and resection margins, was correlated with treatment outcome. The database file was analyzed using dbase III and its companion program Framework, and SAS PC (Statistical Analysis Systems for personal computers).

Cause of Death

Optimizing staff scheduling by Monte-Carlo simulation.

DOCS is a computer program which generates the staff schedule. An accounting framework is combined with an optimization technique that searches for a schedule in which all accounts are simultaneously in balance. The search is accomplished using a Monte-Carlo process which shuffles staff within the schedule. The shuffling is biased according to each staffer's account balance: the staffer who owes the most is most likely to be scheduled.

Algorithms

Energy filters, motion uncertainty, and motion sensitive cells in the visual cortex: a mathematical analysis.

Energy filters are tuned to space-time frequency orientations. In order to compute velocity it is necessary to use a collection of filters, each tuned to a different space-time frequency. Here we analyze, in a probabilistic framework, the properties of the motion uncertainty. Its lower bound, which can be explicitly computed through the Cramér-Rao inequality, will have different values depending on the filter parameters. We show for the Gabor filter that, in order to minimize the motion uncertainty, the spatial and temporal filter sizes cannot be arbitrarily chosen; they are only allowed to vary over a limited range of values such that the temporal filter bandwidth is larger than the spatial bandwidth. This property is shared by motion sensitive cells in the primary visual cortex of the cat, which are known to be direction selective and are tuned to space-time frequency orientations. We conjecture that these cells have larger temporal bandwidth relative to their spatial bandwidth because they compute velocity with maximum efficiency, that is, with a minimum motion uncertainty.

Animals

The evaluation of artificial intelligence systems in medicine.

This paper discusses the underlying issues in the evaluation of computer systems which apply artificial intelligence in medicine. Three different levels of evaluation are described: the subjective evaluation of the research contribution of a developmental prototype, the validation of a system's knowledge and performance, and the evaluation of the clinical efficacy of an operational system. The paper outlines a number of evaluation issues at each level, and discusses how previous artificial intelligence in medicine evaluations fit into this framework.

Artificial Intelligence

Proteomics at scale: Bottlenecks and opportunities for early-career researchers in a fast developing field.

The field of proteomics has rapidly evolved over the last five years enabled by rapid advances in instrumentation and computation. At the same time, the proteomics community is also growing. This is reflected by the increasing participation in international conferences such as those organized by the European Proteomics Association and the Human Proteome Organization. These events provide early-career researchers with unique opportunities to exchange ideas, develop collaborations, and build networks that support professional development. One such network is the Young Proteomics Investigators Club, a European initiative supported by European Proteomics Association and led by early-career researchers. In this Community-Driven project, we investigate recent trends in proteomics by screening conference abstracts and evaluating the session attendance at Human Proteome Organization Congresses and European Proteomics Association conferences. Based on these analyses, we identified five areas that, from our perspective, are shaping the current trends in proteomics: clinical proteomics, proteomics of post-translational modifications, single-cell proteomics, systems biology and multi-omics, and computational proteomics. For each area, we highlight both unique challenges and identify a common theme: a shift from exploratory studies with manageable sample numbers towards large screenings and cohorts and the generation of big data, which often comes with the lack of computational support, organizational networks, and infrastructure. In this light, we describe the unique challenges and opportunities faced by early-career researchers. We point to actionable directions for enabling reproducible and transparent proteomics as well as community-driven projects and initiatives, which are often providing training and support. SIGNIFICANCE: In this perspective, the Young Proteomics Investigators Club (YPIC) discusses advances in analytical developments and computational approaches in proteomics research. Based on empirical analysis of recent European Proteomics Association conference and Human Proteome Organization congresses contributions, we identify clinical, single-cell, post-translational and systems-level proteomics as the research areas that have gained most momentum in the last three to five years. What makes this work distinctive is that it is written by and for early-career researchers, thereby uniquely identifying where momentum, challenges, and unmet needs converge for the newest generation of proteomics researchers. Rather than cataloguing advances, we examine the widening gap between what modern proteomics can generate and what individual researchers can realistically process, validate, and interpret. We describe specific structural barriers including access to high performance computing, limited formal training in scalable data analysis, the need for unified benchmarking standards and navigating clinical collaboration frameworks. We then highlight opportunities for the field, such as community-curated benchmarks, interdisciplinary mentorship models, and shared computational infrastructure. By making these challenges explicit from an early-career researchers standpoint, we aim to inform how training, funding, and community initiatives can be shaped to support the next generation of proteomics researchers.

Proteomics

Inferring Gene Regulatory Networks in Stem Cells: Methods and Applications.

Gene regulatory networks (GRNs) represent the complex interplay of transcription factors, regulatory elements, and target genes that orchestrate cellular identity and function, playing a crucial role in the differentiation and maintenance of stem cells. This chapter provides an overview of experimental and computational methodologies for inferring GRNs, with particular emphasis on single-cell approaches. We first review key experimental techniques for detecting transcription factor binding sites, chromatin accessibility, and DNA motifs, alongside essential databases that support GRN reconstruction. We then introduce computational inference methods that can be categorized into four principal frameworks: correlation-based approaches, regression and machine learning models, probabilistic and deep learning methods, and integrative or message-passing frameworks. To illustrate practical application, we present a case study applying the pySCENIC workflow to a peripheral blood mononuclear cell single-cell RNA sequencing dataset from mouse, demonstrating how regulon-based analysis can reveal cell-type-specific regulatory programs. This chapter aims to serve as a practical guide for researchers seeking to understand and implement GRN inference methodologies in stem cell biology and related fields.

Gene Regulatory Networks

A model for rouleaux pattern formation of red blood cells.

Human red blood cells (RBCs) in a solution form rouleaux patterns under various conditions. The degree of rouleaux formation depends on, for example, the concentration and molecular weight of added large molecules. We present a two-dimensional discrete cellular space model in which an RBC is represented by a rectangle and differential adhesion is assumed among the longer (a-site), the shorter (b-site) sides of the rectangle and the solvent. The total sum of the adhesion energy is assumed to guide the step-by-step change of the model cell configuration and also define absolutely stable patterns. We compare the set of absolutely stable patterns and cell aggregate patterns for both actual and computer-simulated cases to obtain the basic validity of our framework. Then we proceed to assess the effects of added high polymers to the adhesion parameters. We first note that under suitable conditions, decrease in a-site-solvent affinity is necessary to have complex patterns rather than increase of a-a affinity. The hypothesis that addition of high polymers reduce the a-site-solvent affinity is concomitant with a newly proposed osmotic stress theory. The parameter fitting results for the experimental phase change curves can also be interpreted as supporting more the new theory than existing traditional explanations.

Cell Adhesion

Phylogenomic subsampling and upsampling for efficient evolutionary analyses of big data.

Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting analyses are then aggregated to obtain evolutionary estimates. PSU harnesses the fact that the computational cost of maximum likelihood analysis is strongly influenced by the number of distinct site patterns in the concatenated alignment, whereas statistical power depends primarily on the amount of evolutionary information represented by the total number of sites and substitutions. By reducing the former while restoring the latter through upsampling, PSU can approximate many full-data analyses at substantially lower computational cost. Analysis of simulated and empirical datasets shows that PSU can accurately estimate bootstrap support values, select the optimal substitution model, test evolutionary hypotheses, and infer branch lengths, divergence times, and associated uncertainty measures, while reducing runtime and memory requirements by orders of magnitude. PSU also provides distributions of inferred clade support across independent subsamples, enabling detection of conflicting phylogenetic signals that may remain hidden in conventional bootstrap analysis. Automated tuning of subsample size, the number of subsamples, and the number of upsampling replicates make PSU practical across diverse datasets. We suggest that PSU is a general strategy for scalable phylogenomic inference using a broad range of statistical methods. By enabling analyses of genome-scale alignments on commodity hardware, PSU broadens research access and reduces environmental and infrastructural costs of big-data phylogenomics.

confidence limits

transFusion: a novel comprehensive platform for integration analysis of single-cell and spatial transcriptomics.

MOTIVATION: Understanding spatial organization, intercellular interactions, and regulatory networks within the spatial context of tissues is crucial for uncovering complex biological processes and disease mechanisms. Spatial transcriptomics technologies have revolutionized this field by enabling the spatially resolved profiling of gene expression. 10× Visium has emerged as the predominant spatial technology, but its low resolution and the complexity of integrating multimodal datasets present significant analytical challenges, particularly for researchers with limited computational and statistical expertise. Current spatial transcriptomics analysis platforms generally fall short of effectively integrating multimodal data and maximizing the utility of spatial information-such as uncovering complex cellular spatial dependencies, multimodal gradient patterns, and spatial coexpression of ligand-receptor pairs and regulatory networks related to disease or biological states-thereby limiting their ability to provide comprehensive end-to-end analytical workflows when analyzing 10× Visium data. RESULTS: To address these limitations, we developed transFusion, a novel, advanced web-based platform specializing in the most comprehensive and effective integration analysis of scRNA-seq and 10× Visium spatial transcriptomics data. transFusion offers 12 key functions, from basic visualization to advanced analyses, including intercellular dependency analysis, ligand-receptor coexpression identification and visualization, and spatial multimodal gradient variation patterns. Two case studies were used to demonstrate transFusion's capabilities in exploring tissue architecture, intercellular communication, dependency networks, and multimodal gradient variation patterns with minimal computational skills and statistical expertise. transFusion provides a flexible and powerful framework for multimodal data integration analysis. AVAILABILITY AND IMPLEMENTATION: transFusion is freely available at https://github.com/WQLin8/transFusion.

Spatial Transcriptomics

Computers in the cognitive rehabilitation of brain-injured persons.

Currently there is a rapid expansion of work among rehabilitation professionals in applying computer-assisted instructional technology to remediate cognitive deficits resulting from brain injury. The present article presents a framework for relating the various theoretical, empirical, and clinical challenges raised by the coalescence of two such recently emerging disciplines as computer-assisted instruction and cognitive rehabilitation. These challenges are presented from the perspectives of diverse disciplines, including cognitive science, rehabilitation professions, and computer science. A set of guiding principles are derived for evaluating the potential efficacy of currently existing programs and for directing future developmental work in software design, evaluation research, and service delivery.

Animals