Search PubMedSearch

SEARCH · Search PubMed

Results for “computational frameworks”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Biological Parts in Yeast Synthetic Biology: From Regulatory Elements to Predictive Design Platforms.

Yeasts, particularly Saccharomyces cerevisiae, are important eukaryotic chassis for synthetic biology because of their tractable genetics, versatile toolkits, and broad utility in metabolic engineering and functional genomics. Progress in this field has been driven by biological parts that enable programmable control of gene expression and cellular behavior. Early efforts focused mainly on promoters, terminators, and other regulatory elements for tuning individual genes. However, as engineering expanded to multigene pathways, genetic circuits, and dynamic regulatory systems, the limits of part-centric design became clear. Part performance is often shaped by genomic context, chromatin state, host physiology, and interactions with other components, which restricts modularity and predictability. In response, yeast synthetic biology is shifting toward integrated design frameworks combining multilayer regulation, standardized assembly, automated experimentation, and computational modeling. This review provides an integrated perspective on the evolution of biological parts across DNA-, RNA-, and protein-level regulation, connecting these advances with assembly frameworks, biofoundries, and machine learning to trace the trajectory from part-centric engineering toward predictive, system-level design in yeast synthetic biology.

Biofoundry

Characterizing the regulatory logic of transcriptional control at the DNA sequence level by ensembles of thermodynamic models.

MOTIVATION: Understanding how the genome encodes the regulatory logic of transcription is a main challenge of the post-genomic era, and can be overcome with the aid of customized computational tools. RESULTS: We report an automated framework for analyzing an ensemble of fits to data of a thermodynamics-based sequence-level model for transcriptional regulation. The fits are clustered accordingly with their intrinsic regulatory logic. A multiscale analysis enables visualization of quantitative features resulting from the deconvolution of the regulatory profile provided by multiple transcription factors interacting with the locus of a gene. Quantitative experimental data on reporters driven by the whole locus of the even-skipped gene in the blastoderm of Drosophila embryos was used for validating our approach. A few clusters of highly active DNA binding sites within the enhancers collectively modulate even-skipped gene transcription. Analysis of variable enhancers' length shows the importance of bound protein-protein interactions for transcriptional regulation. The interplay between activation and quenching enables function conservation of enhancers despite length variations. AVAILABILITY AND IMPLEMENTATION: The transcription factor level data used for performing the reported study is accessible in the input files in Zenodo and GitHub as well the full code. Additional data from formerly FlyEx database will be available under request.

Thermodynamics

CaXML: Chemistry-informed machine learning explains mutual changes between protein conformations and calcium ions in calcium-binding proteins using structural and topological features.

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of CaXML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

Machine Learning

An algorithm for approximating conditional probabilities.

When diagnostic programs are constructed within a probabilistic framework, it is often the case that computation of joint probabilities of exhaustive combinations of events is easy, but computation of the kind of conditional probabilities the user wishes to know, is hard. This paper describes a simple algorithm for computing the required values, and then suggests several heuristic optimizations that may enable suitable approximations to be obtained in a feasible time when the task is otherwise intractable. An account is given of a specific application of the method in the construction of a medical diagnostic program, which is described in more detail elsewhere.

Algorithms

Leveraging Interradiomic Feature Relationships for Enhanced Prediction of Distant Metastasis and Characterization of Heterogeneity in Head and Neck Cancer.

PURPOSE: Distant metastasis remains a major cause of treatment failure in head and neck (HN) cancer, highlighting the need for more accurate early risk stratification. This study developed and validated a deep radiomics framework to characterize tumor heterogeneity from pretreatment computed tomography (CT) images and improve prediction of distant metastasis-free survival (DMFS). METHODS AND MATERIALS: This multicenter study included 3421 patients with HN cancer from 4 cohorts across 12 institutions. Radiomics features were extracted from primary tumors and transformed into OmicsMaps, a structured representation that spatially organizes interfeature relationships to facilitate learning of complex prognostic patterns. A convolutional neural network was trained to derive prognostic signatures, which were integrated with key clinical variables to construct an OmicsMap-clinical fusion model for patient risk stratification. Model performance was assessed using the concordance index (C-index) and time-dependent area under the receiver operating characteristic curve (AUC) in the CT Images from Large Head and Neck Cohort (RADCURE), HEAD-NECK-RADIOMICS-HN1 (HN1), and Head-Neck-Positron Emission Tomography-Computed Tomography (HN-PET-CT) cohorts. Radiogenomic analyses using RNA-seq data were conducted in the Cancer Genome Atlas Head-Neck Squamous Cell Carcinoma (TCGA-HNSC) cohort to investigate biological characteristics associated with the imaging-defined risk groups. RESULTS: The OmicsMap achieved C-index values of 0.742, 0.768, and 0.671 in the RADCURE, HN1, and HN-PET-CT cohorts, outperforming the conventional radiomics approach by 5.40%-6.37%. Incorporating clinical variables further improved generalizability, yielding a C-index of 0.864 (HN1) and 0.730 (HN-PET-CT), with time-dependent AUC of 0.727-0.895. The fusion model consistently stratified patients into distinct high- and low-risk groups for both DMFS and overall survival across cohorts (P <.01). Radiogenomic analyses revealed enrichment of immune-related pathways in the low-risk group, whereas the high-risk group exhibited a more aggressive phenotype enriched for proliferation, hypoxia, and epithelial-mesenchymal transition pathways, along with a fibrosis-prone tumor microenvironment characterized by extracellular matrix remodeling. CONCLUSIONS: Modeling interradiomic feature relationships using the OmicsMap representation substantially improves CT-based prediction of DMFS and characterization of tumor heterogeneity in HN cancer, supporting precision risk stratification in clinical oncology.

Journal Article

WinPCA: a package for windowed principal component analysis.

SUMMARY: With chromosomal reference genomes and population-scale whole genome-sequencing becoming increasingly accessible, contemporary studies often include characterizations of the genomic landscape as it varies along chromosomes, commonly termed genome scans. While traditional summary statistics like FST and dXY between pre-assigned populations remain integral to characterizing the genomic divergence profile, PCA differs by providing single-sample resolution, thereby supporting the identification of polymorphic inversions, introgression and other types of divergent sequence that may not be fully aligned with global population structure. Here, we introduce WinPCA, a user-friendly package to compute, polarize and visualize genetic principal components in windows along the genome. To accommodate low-coverage whole genome-sequencing datasets, WinPCA can optionally make use of PCAngsd methods to compute principal components in a genotype likelihood framework. WinPCA accepts variant data in either VCF or BEAGLE format and can generate rich plots for interactive data exploration and downstream presentation. AVAILABILITY AND IMPLEMENTATION: WinPCA is implemented in Python and freely available at https://github.com/MoritzBlumer/winpca and https://doi.org/10.5281/zenodo.15614979.

Software

Risk assessment extrapolations and physiological modeling.

The process of assessing the risk associated with human exposure to environmental chemicals inevitably relies on a number of assumptions, estimates and rationalizations. One of the more challenging aspects of risk assessment involves the need to extrapolate beyond the range of conditions used in experimental animal studies to predict anticipated human risks. The most obvious extrapolation required is that from the tested animal species to humans; but others are also generally required, including extrapolating from high dose to low dose, from one route of exposure to another and from one exposure timeframe to another. Several avenues are available for attempting these extrapolations, ranging from the assumption of strict correspondence of dose to the use of statistical correlations. One promising alternative for conducting more scientifically sound extrapolations is that of using physiologically based pharmacokinetic models that contain sufficient biological detail to allow pharmacokinetic behavior to be predicted for widely different exposure scenarios. In recent years, successful physiological models have been developed for a variety of volatile and nonvolatile chemicals, and their ability to perform the extrapolations needed in risk assessment has been demonstrated. Techniques for determining the necessary biochemical parameters are readily available, and the computational requirements are now within the scope of even a personal computer. In addition to providing a sound framework for extrapolation, the predictive power of a physiologically based pharmacokinetic model makes it a useful tool for more reliable dose selection before beginning large-scale studies, as well as for the retrospective analysis of experimental results.

Animals

Physical basis of charge pairing in mitochondria.

The postulate of charge pairing in the mitochondrial inner membrane is justified by applying a formula due to Fuoss to calculate the probability density for the distance between a positive and a negative charge. For dielectric constants 10 or less pairing is absolute, for 20 there is some tendency towards pairing, and at 78 it is nonexistent. Pairing, partner exchange or charge substitution, inhibition, and antiport uncoupling can be rationalized within this framework.

Computers

Decision analysis: a framework for critical care decision assistance.

The ultimate goal of medical computer systems is to help clinicians make good decisions. Such systems must be based on sound principles. Decision analysis is a 25-year-old discipline that provides the needed rigorous foundation for decision assistance. Decision analysis comprises the philosophy, procedures, and tools that can correct the flaws in existing critical care decision-making practice. Intelligent decision systems--computer-based systems that automate decision analysis--make it practical to apply decision analysis to critical care. Orchestra is a pilot intelligent decision system (now under development) that coordinates the efforts of the critical care specialist, the bedside physician, and the bedside nurse in building decision models that can provide recommendations and insight for ventilator management decisions. Decision analysis delivered by intelligent decision systems has great potential for improving critical care decision-making.

Algorithms

AI-genomics synergy for drug repurposing in breast cancer: an interpretability-driven framework.

Breast cancer's genomic heterogeneity complicates drug discovery, making repurposing an attractive but challenging strategy. Advances in artificial intelligence now enable integration of multi-omics data to reveal drug-gene-disease relationships and generate subtype-specific repurposing hypotheses. In this Review, we examine AI-driven computational approaches from signature-based to multi-modal frameworks and propose an integrated interpretability-driven framework linking mechanistic validation with clinical translation toward more transparent and actionable precision oncology.

Journal Article

Comparative and Subtractive Genomics Analysis of Multidrug-Resistant Klebsiella pneumoniae Strains for Novel Target Identification and Drug Repurposing Strategies.

The rapid rise of multidrug-resistant (MDR) Klebsiella pneumoniae has created a major global health challenge due to the limited availability of conserved therapeutic targets effective across diverse resistant strains. In this study, an integrative computational target-discovery and drug-repurposing framework was applied to six clinically relevant K. pneumoniae strains. Comparative genomic analysis identified 3012 conserved genes, which were subsequently filtered to nine essential, non-host homologous proteins. Among these, three conserved cytoplasmic proteins (accD, cpxR, and mraZ) were prioritized for functional analysis, with acetyl-CoA carboxylase subunit beta (accD) emerging as the most promising therapeutic target based on sequence conservation, predicted essentiality, subcellular localization, and pathway association. Structural assessment supported the reliability of the predicted accD model, whereas consensus binding-site analysis identified key residues suitable for ligand interaction. Virtual screening of FDA-approved drugs followed by molecular docking identified several compounds with favorable binding profiles toward accD. Subsequent molecular dynamics simulations, including root mean square deviation (RMSD), root mean square fluctuation (RMSF), radius of gyration (Rg), hydrogen-bond occupancy, principal component analysis (PCA), and PCA-based free energy landscape (FEL) analyses, consistently identified tenapanor, micafungin, deferoxamine, and cobicistat as the most stable protein-ligand complexes, with tenapanor exhibiting the most favorable overall structural and thermodynamic stability profile. These findings identify accD as a promising therapeutic target in MDR K. pneumoniae and suggest several FDA-approved compounds as potential candidates for drug repurposing. Although experimental validation is needed to confirm their biological activity and therapeutic potential, this study demonstrates the potential of integrating comparative genomics with molecular dynamics analyses to support antimicrobial target identification and drug repurposing against MDR bacterial pathogens.

Klebsiella pneumoniae

Cluster analysis and related techniques in medical research.

In this paper we review methods of cluster analysis in the context of classifying patients on the basis of clinical and/or laboratory type observations. Both hierarchical and non-hierarchical methods of clustering are considered, although the emphasis is on the latter type, with particular attention devoted to the mixture likelihood-based approach. For the purposes of dividing a given data set into g clusters, this approach fits a mixture model of g components, using the method of maximum likelihood. It thus provides a sound statistical basis for clustering. The important but difficult question of how many clusters are there in the data can be addressed within the framework of standard statistical theory, although theoretical and computational difficulties still remain. Two case studies, involving the cluster analysis of some haemophilia and diabetes data respectively, are reported to demonstrate the mixture likelihood-based approach to clustering.

Algorithms

Mosaic architecture of the somatic sensory-recipient sector of the cat's striatum.

The striatum is known to have a compartmental organization in which histochemically defined zones called striosomes form branched 3-dimensional labyrinths embedded within the surrounding matrix. We explored how fiber projections from cortical somatic sensory areas representing cutaneous and deep-receptor inputs are organized in relation to this striatal architecture. Areas SI and 3a were mapped electrophysiologically, and distinguishable anterograde tracers (wheat germ agglutinin-HRP and 35S-methionine) were injected into physiologically identified loci. Primary somatic sensory corticostriatal projections were confined to a small, well-defined sector in the dorsolateral corner of the ipsilateral striatum. The somatic sensory afferents were arranged according to a coherent global body map in which rostral body parts were represented more laterally than caudal body parts. Single cortical loci innervated branched and clustered striatal zones that were reminiscent of the striosomes in their range of sizes and shapes yet lay strictly within the extrastriosomal matrix. In contrast to the global orderliness of the striatal body map, there were clear examples of locally complex patterns in which functionally distinct inputs interdigitated with each other. These patterns were often, but not always, produced when corticostriatal afferents carrying different submodality types were labeled. These findings demonstrate the existence of striosome-like striatal compartments within the seemingly uniform extrastriosomal matrix. The principle of mosaic organization thus holds throughout the tissue of the somatic sensory striatum. The striatal architecture delineated here could provide the anatomical substrate for computations requiring cross-modality comparisons within the framework of an overall somatotopy. If a similar multicompartmental architecture also characterizes other striatal regions, as seems likely, it may set general constraints on the nature of associative processing within the striatum as a whole.

Afferent Pathways

Generalized linear models with random effects; salamander mating revisited.

In recent years much effort has been devoted to extending regression methodology to non-Gaussian data, where responses are not independent. These methods for dependent responses are suitable for data from longitudinal studies or nested designs. However, use of these methods for crossed designs seems to have serious limitations due to the intensive computations involved because of the intractable nature of the joint distribution. In this paper, we cast the problem in a Bayesian framework and use a Monte Carlo method, the Gibbs sampler, to avoid current computational limitations. The flexibility of this approach is illustrated by analyzing the interesting salamander mating data reported by McCullagh and Nelder (1989, Generalized Linear Models, 2nd edition, London: Chapman and Hall).

Analysis of Variance

[Deviation and rotation of the larynx in computer tomography].

Many authors described the clinical importance of asymmetry of the laryngeal framework. However, its pathogenesis is generally unknown. In this study, CT images of 315 Japanese subjects were investigated to define the laryngeal position relative to the midline of the cervical vertebra. The CT slice of each subject within 5 mm cephalad of the cricoarytenoid joint was traced. Then, the deviation and rotation angles were measured using our method. Seventy one percent of the subjects' larynges deviated and/or rotated to the right side, while 17% to the left side. Six percent showed neither deviation nor rotation. As to the rest of 6%, deviation and rotation were in opposite directions. Besides, the length of the thyroid alae were measured in 282 subjects. Left ala was longer in 55%, and right was in 23%, and almost equal in 22%. The conclusions are as follows, 1. The majority of the subjects' CT images showed deviation and/or rotation of the laryngeal framework to the right side. 2. So called idiopathic laryngeal deviation is a case which observed in those cases with remarkable deviation and/or rotation of the laryngeal framework. 3. Aging seemed to be an important factor in acceleration of the laryngeal deviation and rotation. 4. The type of diseases and the side of mass lesions had no statistical significance in deviation and rotation of the larynx.

Adult

A scalable HPC framework for bioinformatics in resource-limited settings: design principles, implementation, and sustainability from the UVRI experience.

MOTIVATION: Building and sustaining High-Performance Computing (HPC) infrastructure for bioinformatics research in resource-limited settings presents significant technical, financial and operational challenges. Institutions in low-and middle-income regions often face constraints such as limited technical expertise, unstable infrastructure and restricted funding which can hinder the deployment of large-scale computational platforms necessary for modern genomics and bioinformatics analyses. RESULTS: We present a scalable and modular HPC framework developed at the Uganda Virus Research Institute (UVRI) to support large-scale genomics and other omics data analyses in resource-limited settings. The framework integrates open-source HPC management tools, infrastructure automation, and reproducible configuration management to enable reliable deployment and maintenance. Optimized storage and networking configurations combined with a phased capacity-building strategy support high-throughput genomic workflows while strengthening local technical expertise. From our implementation experience, we derive ten practical design and operational rules that provide a transferable methodology for establishing and sustaining in-house HPC infrastructure. These rules emphasize strategic investment in human capacity, structured planning, leveraging collaborations, adoption of open-source technologies and service management practices to improve operational resilience and long-term sustainability. AVAILABILITY: The design principles, automation strategies and implementation guidelines described in this work are applicable to institutions seeking to establish sustainable HPC resources for bioinformatics research in resource-constrained environments.

Computational Biology

Information technology and computer-based decision support in diabetic management.

This paper describes the application of computer-based techniques within an intelligent, knowledge-based framework to the management of diabetes. The objectives are to structure data collection and storage so that the relevant patient-specific data are collected and made accessible as needed, and to provide clinical decision support on either a day-by-day or longer timescale as appropriate; these objectives relating to both hospital clinic and general practice. For longer-term management, a prototype rule set (greater than 500 rules) has been developed (coded in Sigma PROLOG), validated and tested on patient data. The data collection programs (written in SCULPTOR) to feed the ruleset have been tested in the hospital clinic and compared with the resident data collection system for usability, and impact on the running of the clinic. Links between the data collection programs and the ruleset program have been written and tested. The computer system will also incorporate a module, combining knowledge-based advisory system and glucose/insulin model as patient simulator, that can be tested as a potential decision aid for adjusting insulin dosage on a daily basis.

Data Collection