Search PubMedSearch

SEARCH · Search PubMed

Results for “Epistasis”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Evaluation of epistasis detection methods for quantitative phenotypes.

MOTIVATION: Epistasis, or genetic interaction, plays a crucial role in shaping complex traits and has been increasingly recognized for its widespread influence in genetic architectures. While epistasis detection has been extensively evaluated in case-control studies, its performance with quantitative phenotypes remains comparatively understudied. RESULTS: We identified and evaluated six epistasis detection methods applicable to quantitative trait analysis: EpiSNP, Matrix Epistasis, MIDESP, PLINK Epistasis, QMDR, and REMMA. Using the EpiGEN simulator, we generated synthetic datasets modeling four classes of pairwise SNP interactions-dominant, multiplicative, recessive, and XOR. We also assessed BOOST and MDR algorithms using discretized (case-control) versions of the same datasets. Performance varied notably by interaction type: REMMA achieved the highest overall detection rate (55%), particularly excelling with dominant interactions (100%). MDR excelled with multiplicative (57%) and XOR (69%) interactions. Meanwhile, EpiSNP attained the best performance for recessive interactions (67%). All methods except BOOST produced F1 scores below 0.05 for most interaction types. We further evaluated the methods using a real-world dataset. When applied to the Adolescent Brain Cognitive Development dataset to analyse the externalizing behavior phenotype, both PLINK Epistasis and PLINK BOOST identified SNPs within the DRD2 and DRD4 genes, consistent with previously reported genetic associations. Given the variability in tool performance across interaction types, no single method provides optimal detection across all scenarios. Leveraging multiple detection algorithms may therefore yield more comprehensive insights into epistatic effects in quantitative trait analyses. AVAILABILITY AND IMPLEMENTATION: All relevant code and simulated datasets can be found at github.com/staslist/Epistasis_Review repository.

Epistasis, Genetic

Interstrain Recombinants of Human Cytomegalovirus Reveal Complex Genetic Correlates and Epistasis Influencing Glycoprotein Display, Virion Infectivity and Spread Characteristics.

Most of the nucleotide diversity in the human cytomegalovirus (HCMV) genome is due to approximately 17 genes with 2-14 alleles each. These allelic genes are interspersed among longer stretches of highly conserved sequences with signatures of extensive recombination that would shuffle the allelic genes into a vast number of allelic haplotypes. Bacterial artificial chromosome clones derived from 3 independent clinical isolates (TB40/e (TB), TR and Merlin (ME)) display dramatic differences in the abundance of entry-mediating glycoproteins gH/gL/gO and gH/gL/UL128-131, virion infectivity and efficiency of cell-free and cell-to-cell modes of spread. Of these, TB and ME are the most phenotypically different and share only 2 of the 17 allelic genes. A set of recombinant HCMV was generated by coinfecting cells with TB and ME and restriction fragment length polymorphism (RFLP) analyses demonstrated complex crossover patterns. Most recombinants were either "TB-like" with much more gH/gL/gO than gH/gL/UL128-131, or "ME-like" with much more gH/gL/UL128-131. This correlated with a TB or ME UL128 sequence, consistent with a G/T polymorphism affecting UL128 pre-mRNA splicing. One recombinant had a gH/gL/gO:gH/gL/UL128-131 ratio of 0.8, suggesting genetic determinants beyond UL128. Virion infectivity correlated with TB versus ME-like glycoprotein display, but intragroup variability indicated additional factors and variability in spread efficiency and the contribution of cell-free and cell-to-cell spread modes indicated an influence of characteristics beyond virion infectivity. Results suggest that the relationships among these three phenotypes are not strictly causal and that all three phenotypes are genetically complex and influenced by epistasis among polymorphic loci across the genome.

Journal Article

Epistasis of ERAP1 With 4 Major Histocompatibility Complex Class I Alleles in Frontal Fibrosing Alopecia: A Genome-Wide Association Study Meta-Analysis.

IMPORTANCE: Frontal fibrosing alopecia (FFA) is an inflammatory and scarring form of hair loss of increasing prevalence that most commonly affects women. An improved understanding of the genetic basis of FFA will support the identification of pathogenic mechanisms and therapeutic targets. OBJECTIVE: To identify novel genomic loci at which common genetic variation affects FFA susceptibility and assess nonadditive effects on genetic risk between susceptibility loci. DESIGN, SETTING, AND PARTICIPANTS: Four genome-wide association studies were combined using an SE-weighted meta-analysis. Within the major histocompatibility complex (MHC) locus, stepwise conditional analysis was undertaken to determine independently associated classical MHC class I alleles. Statistical tests for epistatic interaction were performed between risk alleles at the MHC and endoplasmic reticulum aminopeptidase 1 (ERAP1) loci. MAIN OUTCOMES AND MEASURES: Genome-wide significant locus associated with FFA and nonadditive effects on genetic risk between susceptibility loci. RESULTS: Of 6668 included patients, there were 1585 European female individuals with FFA and 5083 controls. Genome-wide significant associations were identified at 4 genomic loci, including a novel susceptibility locus at 5q15, and the association signal could be fine-mapped to a single nucleotide substitution (rs10045403) in the 5' untranslated region of ERAP1 (rs10045403; odds ratio, 1.30; 95% CI, 1.19-1.43; P = 3.6 × 10-8). Within the MHC, FFA risk was statistically independently associated with HLA-A*11:01, HLA-A*33:01, HLA-B*07:02, and HLA-B*35:01. FFA risk was affected by genetic variation at the ERAP1 locus only in individuals who carried at least 1 of the MHC class I risk alleles. CONCLUSIONS AND RELEVANCE: In this genome-wide meta-analysis, a supra-additive effect of genetic variation was found that affected peptide trimming and antigen presentation on FFA susceptibility. Patients with FFA may benefit from emerging therapeutic approaches that modulate ERAP-mediated processes.

Female

Epistasis and the changing fitness landscapes of SARS-CoV-2.

Since its emergence in late 2019, millions of SARS-CoV-2 genomes have been generated as part of global efforts to monitor the evolution and spread of the virus. This unprecedented volume of data provides a unique opportunity to study viral evolution at unparalleled resolution. In particular, individual genomic sites can be observed to have mutated independently thousands of times. These mutation counts have been used to estimate site-specific mutation rates and fitness effects for most mutations across the viral genome. Here, we use these data to investigate how the landscape of mutational fitness costs has changed over the course of the pandemic. SARS-CoV-2 evolution over the past 6 years has been characterized by the emergence of distinct variants separated by long branches corresponding to evolutionary saltations involving up to 50 mutations. We compare inferred fitness landscapes of the Spike protein across these variants and find that shifts in the estimated effects of non-synonymous mutations are linked to genetic differences between them. Sites with altered fitness costs are enriched near positions where the genetic backgrounds differ. To explain the observed changes, we introduce a model with pairwise epistatic interactions between mutations and residues that differ between variants. This model is able to explain about half of the variance in the shifts of fitness effects and suggests that each mismatch between variants substantially alters mutation effects at typically 1 to 3 additional positions.

SARS-CoV-2

Pleiotropic mutational effects on function and stability constrain the antigenic evolution of influenza hemagglutinin.

The evolution of human influenza virus hemagglutinin (HA) involves simultaneous selection to acquire antigenic mutations that escape population immunity while preserving protein function and stability. Epistasis shapes this evolution, as an antigenic mutation that is deleterious in one genetic background may become tolerated in another. However, the extent to which epistasis can alleviate pleiotropic conflicts between immune escape and protein function/stability is unclear. Here, we measure how all amino acid mutations in the HA of a recent human H3N2 influenza strain affect its cell entry function, acid stability, and neutralization by human serum antibodies. We find that epistasis has entrenched certain mutations so that reverting to the ancestral amino acid identity in earlier strains is no longer tolerated. Epistasis has also enabled the emergence of antigenic mutations that were detrimental to HA's cell entry function in earlier strains. However, epistasis appears insufficient to overcome the pleiotropic costs of antigenic mutations that impair HA's stability, explaining why some mutations that strongly escape human antibodies never fix in nature. Our results refine our understanding of the mutational constraints that shape recent H3N2 influenza evolution: epistasis can enable antigenic change, but pleiotropic effects can restrict its trajectory.

Journal Article

Testing for Genetic Interactions in Complex Disease With Distance Correlation.

Understanding epistasis (genetic interaction) may shed some light on the genomic basis of common diseases, including disorders of maximum interest due to their high socioeconomic burden, like schizophrenia. Distance correlation is an association measure that characterizes general statistical independence between random variables, not only the linear one. Here, we propose distance correlation as a novel tool for the detection of epistasis from case-control data of single-nucleotide polymorphisms. On the methodological side, we highlight the derivation of the explicit asymptotic null distribution of the test statistic. We show that this is the only way to obtain enough computational speed for the method to be used in practice, in a scenario where the resampling techniques found in the literature are impractical. Our simulations show satisfactory calibration of significance, as well as comparable or better power than existing methodology. We conclude with the application of our technique to a schizophrenia genetics dataset, obtaining biologically sound insights.

Epistasis, Genetic

Dual functional genomics reveals a broad and convergent landscape of asciminib resistance in BCR::ABL1.

BACKGROUND: Drug resistance is a constantly evolving challenge. The allosteric inhibitor asciminib is a novel therapy for chronic myelogenous leukemia (CML) that targets the myristoyl pocket of the BCR::ABL1 kinase. While it can overcome resistance to active-site inhibitors like imatinib, new resistance mutations to asciminib are emerging. The complete landscape of these mutations, particularly those outside the kinase domain or those arising from epistatic interactions between mutations, are not well understood. METHODS: This study employed a dual functional genomics approach in CML cell line models. A high-throughput adenosine base editing (ABE) screen was used to identify broad hotspots of asciminib resistance across the entire BCR::ABL1 protein. Deep mutational scanning (DMS) was then used to create a high-resolution map of all possible amino acid changes within these hotspots. An "edit-on-edit" screen was performed to investigate epistasis by introducing a library of mutations into a cell line that was pre-edited to incorporate the common imatinib-resistance mutation, Y253H. Finally, a novel Förster resonance energy transfer (FRET) biosensor was developed to measure the conformational state of BCR::ABL1 in live cells and link it to drug sensitivity. RESULTS: The screens identified 279 asciminib resistance mutations and revealed resistance hotspots distributed across the SH3, SH2, and kinase domains, in contrast to imatinib resistance, which is largely confined to the kinase domain. The study uncovered a potent epistatic interaction between a mutation in the SH3 domain (V73A) and a mutation in the kinase domain P-loop (Y253H), which synergistically conferred high-level resistance. The FRET biosensor demonstrated that asciminib resistance mutations tend to destabilize the "closed" inactive conformation of the ABL1 kinase. CONCLUSIONS: The landscape of asciminib resistance is broader and more complex than previously appreciated, involving mutations across multiple domains that disrupt ABL1 autoinhibition. Epistasis between mutations acquired during sequential therapies can create unexpected and potent resistance. However, these diverse genetic resistance mechanisms converge on a single biophysical measurement of the openness of the active ABL1 conformation. This provides a unified framework for understanding asciminib resistance and underscores the need for routine clinical resistance monitoring to include the SH3 and SH2 domains in first line and later line therapy.

Fusion Proteins, bcr-abl

The consequences of using inadequate testers in the simplified triple test-cross.

The genetical consequences of common alleles in the L1 and L2 testers of a simplified version of the triple test-cross which is applicable to populations of inbred lines are examined. The test for epistasis under these circumstances becomes ambiguous and can spuriously detect non-allelic interactions when they may not exist although it still provides a test for epistasis and the adequacy of the testers simultaneously. The tests of significance and the estimates of additive variation are biased to an extent related to the dominance and dominance x additive effects of the common loci while the significance and estimates of dominance variation are deflated because they reflect the dominance effects at the non-common loci only. The covariance of sums and differences is also underestimated for the same reasons. These expectations are illustrated by analysing the 190 simplified triple test-crosses that could be extracted from a 20 x 20 diallel set of crosses between pure-breeding lines of Nicotiana rustica.

Animals

Deconstructing empirical fitness seascapes across scales of granularity.

The fitness landscape metaphor remains resonant in evolutionary theory and has facilitated the birth of newer concepts, like the fitness seascape, that consider the role of environmental context in shaping the dynamics of evolution. Since its emergence, the seascape has appeared in numerous studies examining how different and fluctuating environments shape evolutionary outcomes. Despite growing interest, we lack comprehensive examinations of how environmental context shapes features of fitness seascapes. In this study, we address this gap by deconstructing empirical fitness seascapes across scales of granularity: loci, locus interactions (epistasis), alleles, trajectories, and entire seascapes. For each, we examine how environmental context influences qualitative and quantitative aspects of seascapes, and find that they change appreciably, with patterns specific to individual systems of study. We also quantify how much each scale varies across environments, and find that certain scales tend to be more sensitive to context than others. In summary, we reflect on the implications of the seascape metaphor for the incorporation of environmental effects into theoretical population genetics, for understanding how the environment shapes evolution in disease systems, and for contemporary bioengineering efforts.

Genetic Fitness

Bayesian inference of fitness landscapes via tree-structured branching processes.

MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.

Bayes Theorem

Compensatory Evolution Following Deleterious Episodes of GC-biased Gene Conversion in Rodents.

GC-biased gene conversion (gBGC) is a widespread evolutionary force associated with meiotic recombination that favors the accumulation of deleterious AT to GC substitutions in proteins, moving them away from their fitness optimum. In many mammals, recombination hotspots have a rapid turnover, leading to episodic gBGC, with the accumulation of deleterious mutations stopping when the recombination hotspot dies. Selection is therefore expected to act to repair the damage caused by gBGC episodes through compensatory evolution. However, this process has never been studied or quantified so far. Here, we analyzed the nucleotide substitution pattern in coding sequences of a highly diversified group of Murinae rodents. Using phylogenetic analyses of about 70,000 coding exons, we identified numerous exon-specific, lineage-specific gBGC episodes, characterized by a clustering of synonymous AT to GC substitutions and by an increasing rate of nonsynonymous AT to GC substitutions, many of which are potentially deleterious. Analyzing the molecular evolution of the affected exons in downstream lineages, we found evidence for pervasive compensatory evolution after deleterious gBGC episodes. Compensation appears to occur rapidly after the end of the episode and to be driven by the standing genetic variation rather than new mutations. Our results demonstrate the impact of gBGC on the evolution of amino-acid sequences and underline the key role of epistasis in protein adaptation. This study contributes to a growing body of literature emphasizing that adaptive mutations, which arise in response to environmental changes, are just 1 subset of beneficial mutations, alongside mutations resulting from oscillations around the fitness optimum.

Gene Conversion

Variations in carbapenem resistance associated with the VIM-1 metallo-β-lactamase across the Enterobacterales.

The VIM-1 metallo-β-lactamase enzyme, encoded within class 1 integrons, is found in Gram-negative clinical isolates worldwide and has been linked to outbreaks of bacterial pathogens in nosocomial settings. Six vim-1+ clinical isolates, from the genera Escherichia, Klebsiella and Enterobacter, were obtained from Kingston, Ontario, Canada. Whole-genome sequencing revealed that vim-1 was plasmid-borne in all strains and situated as the first gene in In916 or In110 integrons. Analysis of related plasmids suggested that these vim-1-containing plasmids are globally disseminated and have spread via horizontal gene transfer and autochthonous vertical spread within Ontario. Interestingly, the MICs of ertapenem and meropenem, two clinically relevant carbapenem antibiotics, against these six isolates varied more than tenfold, suggesting that the effects of VIM-1 are dependent on the genomic content of the host microbe. Introducing vim-1 into three common Enterobacterales laboratory strains was not sufficient to confer resistance to ertapenem and meropenem. Instead, adaptive laboratory evolution of the vim-1 + laboratory strains revealed that vim-1-mediated carbapenem resistance in these strains was dependent on epistatic interactions with ompC mutations, likely due to decreased outer membrane permeability to these antibiotics. Together, these results provide additional support for the role of gene epistasis in modulating the antimicrobial resistance phenotypes of acquired resistance genes, as well as previous results suggesting that the presence of a β-lactamase gene is insufficient to confer strong resistance to carbapenems without being paired with reduced outer membrane permeability.

beta-Lactamases

A gene-based model of fitness and its implications for genetic variation: Linkage disequilibrium.

A widely used model of the effects of mutations on fitness (the "sites" model) assumes that heterozygous recessive or partially recessive deleterious mutations at different sites in a gene complement each other, similarly to mutations in different genes. However, the general lack of complementation between major effect allelic mutations suggests an alternative possibility, which we term the "gene" model. This assumes that a pair of heterozygous deleterious mutations in trans behave effectively as homozygotes, so that the fitnesses of trans heterozygotes are lower than those of cis heterozygotes. We examine the properties of the two different models, using both analytical and simulation methods. We show that the gene model predicts positive linkage disequilibrium (LD) between deleterious variants within the coding sequence, under conditions when the sites model predicts zero or slightly negative LD. We also show that focussing on rare variants when examining patterns of LD, especially with Lewontin's´ measure, is likely to produce misleading results with respect to inferences concerning the causes of the sign of LD. Synergistic epistasis between pairs of mutations was also modeled; it is less likely to produce negative LD under the gene model than the sites model. The theoretical results are discussed in relation to patterns of LD in natural populations of several species.

complementation

Inference and visualization of complex genotype-phenotype maps with gpmap-tools.

Understanding how biological sequences give rise to observable traits, that is, how genotype maps to phenotype, is a central goal in biology. Yet our knowledge of genotype-phenotype maps in natural systems is limited due to the high dimensionality of sequence space and the context-dependent effects of mutations. The emergence of Multiplex assays of variant effect (MAVEs), along with large collections of natural sequences, offer new opportunities to empirically characterize these maps at an unprecedented scale. However, tools for statistical and exploratory analysis of these high-dimensional data are still needed. To address this gap, we developed gpmap-tools (https://github.com/cmarti/gpmap-tools), a python library that integrates a series of models for inference, phenotypic imputation, and error estimation from MAVE data or collections of natural sequences in the presence of genetic interactions of every possible order. gpmap-tools also provides methods for summarizing patterns of epistasis and visualization of genotype-phenotype maps containing up to millions of genotypes. To demonstrate its utility, we used gpmap-tools to infer genotype-phenotype maps containing 262,144 variants of the Shine-Dalgarno sequence from both genomic 5'UTR sequences and experimental MAVE data. Visualization of the inferred landscapes consistently revealed high-fitness ridges that link core motifs at different distances from the start codon. In summary, gpmap-tools provides a flexible, interpretable framework for studying complex genotype-phenotype maps, opening new avenues for understanding the architecture of genetic interactions and their evolutionary consequences.

Gaussian process

Renal albumin excretion: twin studies identify influences of heredity, environment, and adrenergic pathway polymorphism.

Albumin excretion marks early glomerular injury in hypertension. This study investigated heritability of albumin excretion in twin pairs and its genetic determination by adrenergic pathway polymorphism. Genetic associations used single nucleotide polymorphisms at adrenergic pathway loci spanning catecholamine biosynthesis, storage, catabolism, receptor action, and postreceptor signal transduction. We studied 134 single nucleotide polymorphisms at 46 loci for a total of >51,000 genotypes. Albumin excretion heritability was 45.2+/-7.4% (P=2x10(-7)), and the phenotype aggregated significantly with adrenergic, renal, metabolic, and hemodynamic traits. In the adrenergic system, excretions of both norepinephrine and epinephrine correlated with albumin. In the kidney, albumin excretion correlated with glomerular and tubular traits (Na(+) and K(+) excretion; fractional excretion of Na(+) and Li(+)). Albumin excretion shared genetic determination (genetic covariance) with epinephrine excretion, and environmental determination with glomerular filtration rate and electrolyte intake/excretion. Albumin excretion associated with polymorphisms at multiple points in the adrenergic pathway: catecholamine biosynthesis (tyrosine hydroxylase), catabolism (monoamine oxidase A), storage/release (chromogranin A), receptor target (dopamine D1 receptor), and postreceptor signal transduction (sorting nexin 13 and rho kinase). Epistasis (gene-by-gene interaction) occurred between alleles at rho kinase, tyrosine hydroxylase, chromogranin A, and sorting nexin 13. Dopamine D1 receptor polymorphism showed pleiotropic effects on both albumin and dopamine excretion. These studies establish new roles for heredity and environment in albumin excretion. Urinary excretions of albumin and catecholamines are highly heritable, and their parallel suggests adrenergic mediation of early glomerular permeability alterations. Albumin excretion is influenced by multiple adrenergic pathway genes and is, thus, polygenic. Such functional links between adrenergic activity and glomerular injury suggest novel approaches to its prediction, prevention, diagnosis, and treatment.

Adolescent

Beyond antibiotic resistance: The whiB7 transcription factor coordinates an adaptive response to alanine starvation in mycobacteria.

Pathogenic mycobacteria are a significant cause of morbidity and mortality worldwide. The conserved whiB7 stress response reduces the effectiveness of antibiotic therapy by activating several intrinsic antibiotic resistance mechanisms. Despite our comprehensive biochemical understanding of WhiB7, the complex set of signals that induce whiB7 expression remain less clear. We employed a reporter-based, genome-wide CRISPRi epistasis screen to identify a diverse set of 150 mycobacterial genes whose inhibition results in constitutive whiB7 expression. We show that whiB7 expression is determined by the amino acid composition of the 5' regulatory uORF, thereby allowing whiB7 to sense amino acid starvation. Although deprivation of many amino acids can induce whiB7, whiB7 specifically coordinates an adaptive response to alanine starvation by engaging in a feedback loop with the alanine biosynthetic enzyme, aspC. These findings describe a metabolic function for whiB7 and help explain its evolutionary conservation across mycobacterial species occupying diverse ecological niches.

Transcription Factors

A modified triple test-cross analysis to test and allow for inadequate testers.

When the population under investigation consists of highly inbred lines the full triple test-cross of Kearsey and Jinks (1968) supplemented by the selfed progenies of the population allows unambiguous and independent tests for epistasis and the adequacy of the pure-breeding testers, L1 and L2. This can also be achieved by supplementing the simplified triple test-cross of Jinks, Perkins and Breese (1969) with the selfed progenies of the L1i and L2i families. If the L1 and L2 testers prove to be inadequate due to the presence of common loci, modifications of the analyses are proposed which correct the resulting biases in the genetical components of variation.

Animals