Search PubMed⌕ Search

Biomedical subjects

Niko Beerenwinkel

Publications and source records attributed to Niko Beerenwinkel.

18 recordsLinked to original sources

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans↗

Wastewater-based sequencing of respiratory syncytial virus to investigate lineage dynamics and antigenic site mutations: a retrospective genomic epidemiology study.

BACKGROUND: Respiratory syncytial virus (RSV) infections pose a substantial health burden, particularly for clinically vulnerable populations such as infants and older adults. Although novel immunoprophylactic interventions show promise in providing protection, many countries may not have robust surveillance systems to monitor circulating RSV lineages and detect mutations that might reduce the effectiveness of these new interventions. We aimed to assess the diversity and temporal dynamics of circulating RSV lineages in urban populations through amplicon-based sequencing and analysis of wastewater extracts. METHODS: In this prospective observational wastewater-based genomic surveillance study, 32 raw influent 24-h composite samples were collected during the 2022-23 and 2023-24 RSV seasons from both Zurich and Geneva, Switzerland. We applied an RSV subtype-specific amplicon-based sequencing approach to obtain RSV-A and RSV-B sequences from all 64 samples. Mutations relative to reference genomes were identified at positions with read depth above 30. Relative abundances of RSV lineages were estimated from frequencies of lineage-signature mutations, present in greater than 90% of publicly available sequences of that lineage. FINDINGS: Relative abundances of RSV-B (2022-23) and RSV-A (2023-24) lineages were estimated over the two RSV seasons. During the 2022-23 season, the RSV-B B.D.E.1 lineage prevailed in both cities. In the 2023-24 season, multiple RSV-A lineages cocirculated, including A.D.1, A.D.3, A.D.5, and their sub-lineages. Identification and frequency estimation of mutations showed low-frequency, non-synonymous mutations in antigenic sites on the fusion gene of both RSV-A and RSV-B, some of which have not been reported in clinical sequences. The primary outcome was identification and relative abundance of RSV lineages in wastewater samples. INTERPRETATION: These findings show the potential of wastewater-based genomic surveillance to identify and track circulating RSV lineages and clinically relevant mutations. As novel RSV immunoprophylaxis measures are introduced in upcoming RSV seasons, wastewater-derived genomic RSV data provide a valuable baseline for understanding RSV diversity and future viral evolution under increased immunological pressure. FUNDING: This study was funded by the Swiss National Science Foundation and in part by the National Institute Of Allergy And Infectious Diseases of the National Institutes of Health. Funding for sample collection and processing was provided by the Swiss Federal Office of Public Health.

Humans↗

Bayesian inference of fitness landscapes via tree-structured branching processes.

MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.

Bayes Theorem↗

Single-cell copy number calling and event history reconstruction.

MOTIVATION: Copy number alterations are driving forces of tumour development and the emergence of intra-tumour heterogeneity. A comprehensive picture of these genomic aberrations is therefore essential for the development of personalised and precise cancer diagnostics and therapies. Single-cell sequencing offers the highest resolution for copy number profiling down to the level of individual cells. Recent high-throughput protocols allow for the processing of hundreds of cells through shallow whole-genome DNA sequencing. The resulting low read-depth data poses substantial statistical and computational challenges to the identification of copy number alterations. RESULTS: We developed SCICoNE, a statistical model and MCMC algorithm tailored to single-cell copy number profiling from shallow whole-genome DNA sequencing data. SCICoNE reconstructs the history of copy number events in the tumour and uses these evolutionary relationships to identify the copy number profiles of the individual cells. We show the accuracy of this approach in evaluations on simulated data and demonstrate its practicability in applications to two breast cancer samples from different sequencing protocols. AVAILABILITY AND IMPLEMENTATION: SCICoNE is available at https://github.com/cbg-ethz/SCICoNE.

Single-Cell Analysis↗

Evolution on distributive lattices.

We consider the directed evolution of a population after an intervention that has significantly altered the underlying fitness landscape. We model the space of genotypes as a distributive lattice; the fitness landscape is a real-valued function on that lattice. The risk of escape from intervention, i.e., the probability that the population develops an escape mutant before extinction, is encoded in the risk polynomial. Tools from algebraic combinatorics are applied to compute the risk polynomial in terms of the fitness landscape. In an application to the development of drug resistance in HIV, we study the risk of viral escape from treatment with the protease inhibitors ritonavir and indinavir.

Bayes Theorem↗

A mutagenetic tree hidden Markov model for longitudinal clonal HIV sequence data.

RNA viruses provide prominent examples of measurably evolving populations. In human immunodeficiency virus (HIV) infection, the development of drug resistance is of particular interest because precise predictions of the outcome of this evolutionary process are a prerequisite for the rational design of antiretroviral treatment protocols. We present a mutagenetic tree hidden Markov model for the analysis of longitudinal clonal sequence data. Using HIV mutation data from clinical trials, we estimate the order and rate of occurrence of seven amino acid changes that are associated with resistance to the reverse transcriptase inhibitor efavirenz.

Alkynes↗

Evolution of HIV resistance during treatment interruption in experienced patients and after restarting a new therapy.

BACKGROUND: To analyse the evolution of resistance patterns in patients undergoing treatment interruption (TI) and re-initiating highly active anti-retroviral therapy (HAART). METHODS: HIV-RT and -PR gene-sequences were analysed in 14 patients (>5 failing prior drugs) before and during TI and under a new HAART. Genotypes were interpreted using two bioinformatics systems. Additionally, virus load (VL) and CD4(+)-T-cell counts were measured. RESULTS: Six patients (42%) achieved sustained undetectable VL up to one year after TI (responders), while 8 (57%) maintained VL of more than 2,000 copies/mL (non-responders). Different patterns of resistance-mutations evolution were detected. During TI loss of all mutations was observed in three patients, a reduction of mutations was detected in seven patients, and no alteration was seen in four patients. In the responders, 87.5% of protease inhibitor (PI)-resistance mutations waned during TI and remained undetectable under the new treatment. In contrast, in the non-responder group most PI-resistance mutations continued noticeable under the new therapy. Loss of primary PI-resistance mutations and the presence of one fully active PI in the new regimen significantly correlated with success of subsequent treatment (p=0.028). In two patients new reverse transcriptase associated mutations were detected during TI, G190A (NNRTI mutation) and K70R (NRTI mutation). Appearance of K70R could be explained by a reverse direction of a previously described pathway of thymidin analogues mutation resistance development, while G190A could be due to prolonged subinhibitory drug levels after cessation of NNRTIs. CONCLUSION: In the evolution of HAART-resistance, different patterns were observed in responders and non-responders during but not before TI. Absence of PI-resistance associated mutations during and after TI and administration of a predicted fully active PI for the new therapy correlated with success. Newly detected mutations during TI may indicate reversibility of previously described mutational pathways.

Anti-Retroviral Agents↗

Computational methods for the design of effective therapies against drug resistant HIV strains.

The development of drug resistance is a major obstacle to successful treatment of HIV infection. The extraordinary replication dynamics of HIV facilitates its escape from selective pressure exerted by the human immune system and by combination drug therapy. We have developed several computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genomic data.

Database Management Systems↗

ROCR: visualizing classifier performance in R.

UNLABELLED: ROCR is a package for evaluating and visualizing the performance of scoring classifiers in the statistical language R. It features over 25 performance measures that can be freely combined to create two-dimensional performance curves. Standard methods for investigating trade-offs between specific performance measures are available within a uniform framework, including receiver operating characteristic (ROC) graphs, precision/recall plots, lift charts and cost curves. ROCR integrates tightly with R's powerful graphics capabilities, thus allowing for highly adjustable plots. Being equipped with only three commands and reasonable default values for optional parameters, ROCR combines flexibility with ease of usage. AVAILABILITY: http://rocr.bioinf.mpi-sb.mpg.de. ROCR can be used under the terms of the GNU General Public License. Running within R, it is platform-independent. CONTACT: tobias.sing@mpi-sb.mpg.de.

Computer Graphics↗

Estimating HIV evolutionary pathways and the genetic barrier to drug resistance.

BACKGROUND: The evolution of drug-resistant viruses challenges the management of human immunodeficiency virus (HIV) infections. Understanding this evolutionary process is important for the design of effective therapeutic strategies. METHODS: We used mutagenetic trees, a family of probabilistic graphical models, to describe the accumulation of resistance-associated mutations in the viral genome. On the basis of these models, we defined the genetic barrier, a quantity that summarizes the difficulty for the virus to escape from the selective pressure of the drug by developing escape mutations. RESULTS: From HIV reverse-transcriptase sequences that had been obtained from treated patients, we derived evolutionary models for zidovudine, zidovudine plus lamivudine, and zidovudine plus didanosine. The genetic barriers to resistance to zidovudine, stavudine, lamivudine, and didanosine, for the above 3 regimens, were computed and analyzed. We found both the mode and the rate of development of resistance to be heterogeneous. The genetic barrier to zidovudine resistance was increased if lamivudine was added to zidovudine but was decreased for didanosine. The barrier to lamivudine resistance was maintained with zidovudine plus didanosine, whereas the barrier to didanosine resistance was reduced most with zidovudine plus lamivudine. CONCLUSION: Mutagenetic trees provide a quantitative picture of the evolution of drug resistance. The genetic barrier is a useful tool for design of effective treatment strategies.

Anti-HIV Agents↗

Estimating cancer survival and clinical outcome based on genetic tumor progression scores.

MOTIVATION: In cancer research, prediction of time to death or relapse is important for a meaningful tumor classification and selecting appropriate therapies. Survival prognosis is typically based on clinical and histological parameters. There is increasing interest in identifying genetic markers that better capture the status of a tumor in order to improve on existing predictions. The accumulation of genetic alterations during tumor progression can be used for the assessment of the genetic status of the tumor. For modeling dependences between the genetic events, evolutionary tree models have been applied. RESULTS: Mixture models of oncogenetic trees provide a probabilistic framework for the estimation of typical pathogenetic routes. From these models we derive a genetic progression score (GPS) that estimates the genetic status of a tumor. GPS is calculated for glioblastoma patients from loss of heterozygosity measurements and for prostate cancer patients from comparative genomic hybridization measurements. Cox proportional hazard models are then fitted to observed survival times of glioblastoma patients and to times until PSA relapse following radical prostatectomy of prostate cancer patients. It turns out that the genetically defined GPS is predictive even after adjustment for classical clinical markers and thus can be considered a medically relevant prognostic factor. AVAILABILITY: Mtreemix, a software package for estimating tree mixture models, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de. The raw cancer datasets and R code for the analysis with Cox models are available upon request from the corresponding author.

Biomarkers, Tumor↗

Mtreemix: a software package for learning and using mixture models of mutagenetic trees.

SUMMARY: Mixture models of mutagenetic trees constitute a class of probabilistic models for describing evolutionary processes that are characterized by the accumulation of permanent genetic changes. They have been applied to model the accumulation of chromosomal gains and losses in tumor development and the development of drug resistance-associated mutations in the HIV genome.Mtreemix is a software package for estimating mutagenetic trees mixture models from observed cross-sectional data and for using these models for predictions. We provide programs for model fitting, model selection, simulation, likelihood computation and waiting time estimation. AVAILABILITY: Mtreemix, including source code, documentation, sample data files and precompiled Solaris and Linux binaries, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de/

Algorithms↗

Patients with high-grade gliomas harboring deletions of chromosomes 9p and 10q benefit from temozolomide treatment.

Surgical cure of glioblastomas is virtually impossible and their clinical course is mainly determined by the biologic behavior of the tumor cells and their response to radiation and chemotherapy. We investigated whether response to temozolomide (TMZ) chemotherapy differs in subsets of malignant glioblastomas defined by genetic lesions. Eighty patients with newly diagnosed glioblastoma were analyzed with comparative genomic hybridization and loss of heterozygosity. All patients underwent radical resection. Fifty patients received TMZ after radiotherapy (TMZ group) and 30 patients received radiotherapy alone (RT group). The most common aberrations detected were gains of parts of chromosome 7 and losses of 10q, 9p, or 13q. The spectrum of genetic aberrations did not differ between the TMZ and RT groups. Patients treated with TMZ showed significantly better survival than patients treated with radiotherapy alone (19.5 vs 9.3 months). Genomic deletions on chromosomes 9 and 10 are typical for glioblastoma and associated with poor prognosis. However, patients with these aberrations benefited significantly from TMZ in univariate analysis. In multivariate analysis, this effect was pronounced for 9p deletion and for elderly patients with 10q deletions, respectively. This study demonstrates that molecular genetic and cytogenetic analyses potentially predict responses to chemotherapy in patients with newly diagnosed glioblastomas.

Adult↗

Geno2pheno: Estimating phenotypic drug resistance from HIV-1 genotypes.

Therapeutic success of anti-HIV therapies is limited by the development of drug resistant viruses. These genetic variants display complex mutational patterns in their pol gene, which codes for protease and reverse transcriptase, the molecular targets of current antiretroviral therapy. Genotypic resistance testing depends on the ability to interpret such sequence data, whereas phenotypic resistance testing directly measures relative in vitro susceptibility to a drug. From a set of 650 matched genotype-phenotype pairs we construct regression models for the prediction of phenotypic drug resistance from genotypes. Since the range of resistance factors varies considerably between different drugs, two scoring functions are derived from different sets of predicted phenotypes. Firstly, we compare predicted values to those of samples derived from 178 treatment-naive patients and report the relative deviance. Secondly, estimation of the probability density of 2000 predicted phenotypes gives rise to an intrinsic definition of a susceptible and a resistant subpopulation. Thus, for a predicted phenotype, we calculate the probability of membership in the resistant subpopulation. Both scores provide standardized measures of resistance that can be calculated from the genotype and are comparable between drugs. The geno2pheno system makes these genotype interpretations available via the Internet (http://www.genafor.org/).

Anti-HIV Agents↗

Methods for optimizing antiviral combination therapies.

MOTIVATION: Despite some progress with antiretroviral combination therapies, therapeutic success in the management of HIV-infected patients is limited. The evolution of drug-resistant genetic variants in response to therapy plays a key role in treatment failure and finding a new potent drug combination after therapy failure is considered challenging. RESULTS: To estimate the activity of a drug combination against a particular viral strain, we develop a scoring function whose independent variables describe a set of antiviral agents and viral DNA sequences coding for the molecular targets of the respective drugs. The construction of this activity score involves (1) predicting phenotypic drug resistance from genotypes for each drug individually, (2) probabilistic modeling of predicted resistance values and integration into a score for drug combinations, and (3) searching through the mutational neighborhood of the considered strain in order to estimate activity on nearby mutants. For a clinical data set, we determine the optimal search depth and show that the scoring scheme is predictive of therapeutic outcome. Properties of the activity score and applications are discussed.

Algorithms↗

Tenofovir resistance and resensitization.

Human immunodeficiency viruses in 321 samples from tenofovir-naïve patients were retrospectively evaluated for resistance to this nucleotide analogue. All virus strains with insertions between amino acids 67 and 70 of the reverse transcriptase (n = 6) were highly resistant. Virus strains with the Q151M mutation were divided into susceptible (n = 12) and highly resistant (n = 8) viruses. This difference was due to the absence or presence of the K65R mutation, which was confirmed by site-directed mutagenesis. Viral clones with various combinations of the mutations M41L, K70R, L210W, and T215F or T215Y were analyzed for cross-resistance induced by thymidine analogue mutations (TAMs). The levels of increased resistance induced by single, double, and triple mutations at the indicated positions could be ranked as follows: for mutants with single mutations, mutations at positions 41 > 215 > 70; for mutants with double mutations, mutations at positions 41 and 215 > 70 and 215 = 210 and 215 > 41 and 70; for mutants with triple mutations, mutations at positions 41, 210, and 215 > 41, 70, and 215. Viral clones with M184V or M184I exhibited slightly increased susceptibilities to tenofovir (0.7-fold). Almost all clones with TAM-induced resistance were resensitized when M184V was present (P < 0.001). Among the viruses in the clinical samples, the rate of tenofovir resistance significantly increased with the number of TAMs both in the samples with 184M and in those with 184V (P = 0.005 and P = 0.003, respectively). A resensitizing effect of M184V was confirmed for all samples exhibiting at least one TAM (P = 0.03). However, accumulation of at least two TAMs resulted in more than 2.0-fold reduced susceptibility to tenofovir, irrespective of the presence of M184V. Decision tree building, a classical machine learning technique, was used to generate models for the interpretation of mutations with respect to tenofovir resistance. The application of previously proposed cutoffs for a reduced response to therapy and treatment failure demonstrated the central roles of positions 215 and 65 for 1.5- and 4.0-fold reduced susceptibilities, respectively. Thus, clinically relevant resistance may be conferred by the accumulation of TAMs, and the resensitizing effect of M184V should be considered only minor.

Adenine↗

Diversity and complexity of HIV-1 drug resistance: a bioinformatics approach to predicting phenotype from genotype.

Drug resistance testing has been shown to be beneficial for clinical management of HIV type 1 infected patients. Whereas phenotypic assays directly measure drug resistance, the commonly used genotypic assays provide only indirect evidence of drug resistance, the major challenge being the interpretation of the sequence information. We analyzed the significance of sequence variations in the protease and reverse transcriptase genes for drug resistance and derived models that predict phenotypic resistance from genotypes. For 14 antiretroviral drugs, both genotypic and phenotypic resistance data from 471 clinical isolates were analyzed with a machine learning approach. Information profiles were obtained that quantify the statistical significance of each sequence position for drug resistance. For the different drugs, patterns of varying complexity were observed, including between one and nine sequence positions with substantial information content. Based on these information profiles, decision tree classifiers were generated to identify genotypic patterns characteristic of resistance or susceptibility to the different drugs. We obtained concise and easily interpretable models to predict drug resistance from sequence information. The prediction quality of the models was assessed in leave-one-out experiments in terms of the prediction error. We found prediction errors of 9.6-15.5% for all drugs except for zalcitabine, didanosine, and stavudine, with prediction errors between 25.4% and 32.0%. A prediction service is freely available at http://cartan.gmd.de/geno2pheno.html.

Computational Biology↗

Learning multiple evolutionary pathways from cross-sectional data.

We introduce a mixture model of trees to describe evolutionary processes that are characterized by the ordered accumulation of permanent genetic changes. The basic building block of the model is a directed weighted tree that generates a probability distribution on the set of all patterns of genetic events. We present an EM-like algorithm for learning a mixture model of K trees and show how to determine K with a maximum likelihood approach. As a case study, we consider the accumulation of mutations in the HIV-1 reverse transcriptase that are associated with drug resistance. The fitted model is statistically validated as a density estimator, and the stability of the model topology is analyzed. We obtain a generative probabilistic model for the development of drug resistance in HIV that agrees with biological knowledge. Further applications and extensions of the model are discussed.

Algorithms↗