Identifying changes in viral fitness using population genetic structure.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
A widely used model of the effects of mutations on fitness (the "sites" model) assumes that heterozygous recessive or partially recessive deleterious mutations at different sites in a gene complement each other, similarly to mutations in different genes. However, the general lack of complementation between major effect allelic mutations suggests an alternative possibility, which we term the "gene" model. This assumes that a pair of heterozygous deleterious mutations in trans behave effectively as homozygotes, so that the fitnesses of trans heterozygotes are lower than those of cis heterozygotes. We examine the properties of the two different models, using both analytical and simulation methods. We show that the gene model predicts positive linkage disequilibrium (LD) between deleterious variants within the coding sequence, under conditions when the sites model predicts zero or slightly negative LD. We also show that focussing on rare variants when examining patterns of LD, especially with Lewontin's´ measure, is likely to produce misleading results with respect to inferences concerning the causes of the sign of LD. Synergistic epistasis between pairs of mutations was also modeled; it is less likely to produce negative LD under the gene model than the sites model. The theoretical results are discussed in relation to patterns of LD in natural populations of several species.
Understanding the genetic basis of rapid adaptation is key to predicting species' evolutionary responses to environmental change. However, it is still debatable whether many small-effect mutations or a few large-effect mutations underlie rapid adaptation, and how this knowledge can predict population survival or extinction. To address this question, we performed a series of ecologically grounded forward-in-time genetic simulations to study rapid adaptation and extinction with increasing magnitudes of environmental change. These simulations were seeded with genomic variation of the plant Arabidopsis thaliana to have a realistic genomic structure, with one (monogenic) to 1,000 (polygenic) variants with varying heritabilities contributing to an environmental adaptive trait. Our results revealed two distinct scenarios of rapid adaptation and population rescue. Under small-to-moderate environmental shifts, high polygenic traits increased evolutionary rescue probability. Under extreme environmental shifts, high polygenic traits lead predictably to extinction, yet monogenic traits sometimes produce one-off winning adaptive genotypes. We interpret our rapid evolutionary rescue findings in terms of the fundamental theorem of natural selection, where trait polygenicity shapes the distribution of genetic variance in fitness across replicates and, in turn, the probability of population survival, with polygenic architectures producing more stable and predictable fitness variance and monogenic architectures generating highly skewed and variable outcomes. These results highlight the insights genomics gives us into the (un)predictability of species' evolutionary responses to global change, with management implications for assisted adaptation and conservation.
The fitness landscape metaphor remains resonant in evolutionary theory and has facilitated the birth of newer concepts, like the fitness seascape, that consider the role of environmental context in shaping the dynamics of evolution. Since its emergence, the seascape has appeared in numerous studies examining how different and fluctuating environments shape evolutionary outcomes. Despite growing interest, we lack comprehensive examinations of how environmental context shapes features of fitness seascapes. In this study, we address this gap by deconstructing empirical fitness seascapes across scales of granularity: loci, locus interactions (epistasis), alleles, trajectories, and entire seascapes. For each, we examine how environmental context influences qualitative and quantitative aspects of seascapes, and find that they change appreciably, with patterns specific to individual systems of study. We also quantify how much each scale varies across environments, and find that certain scales tend to be more sensitive to context than others. In summary, we reflect on the implications of the seascape metaphor for the incorporation of environmental effects into theoretical population genetics, for understanding how the environment shapes evolution in disease systems, and for contemporary bioengineering efforts.
PURPOSE: We investigated potentially causal associations between genetically predicted aerobic fitness and multiple health phenotypes using a two-stage phenome-wide Mendelian randomization (MR) study. METHODS: Genetically determined aerobic fitness, as operationalized by Cai et al., served as the exposure instrument. We screened 712 health-related phenotypes as outcomes using publicly available European-ancestry genome-wide association studies (GWAS) summary statistics from OpenGWAS (Discovery GWAS n > 5000), prioritizing non-UK Biobank/non-FinnGen datasets for Discovery when available and selecting an independent GWAS for validation. Associations were estimated using the MR-Robust Adjusted Profile Score method, controlled for multiple testing (5% false discovery rate) and unaffected by violations of MR assumptions (directional concordance between discovery and validation; no evidence of horizontal pleiotropy across inverse-variance weighted, MR-Egger, weighted-median, and weighted-mode methods; negative control analysis on hair color). RESULTS: We identified 108 discovery associations, of which 34 remained valid and statistically significant after validation. Higher genetically determined aerobic fitness was associated with lower lacunar stroke risk, lower arterial stiffness, higher heart rate variability, lower diastolic blood pressure, more favorable anthropometric measures, lower use of antidiabetic drugs, lower asthma risk, lower C-reactive protein, higher bone mineral density, favorable liver function biomarkers, favorable platelet-related traits, multiple blood count-derived hematological cell indices and counts, as well as higher years of schooling. Adverse associations were confined to atrial fibrillation, valvular heart disease, and systolic blood pressure. CONCLUSIONS: Genetically determined aerobic fitness is linked to a broad pattern of favorable cardiometabolic, inflammatory, musculoskeletal, respiratory, hepatic, and hematological phenotypes, alongside a narrow set of potential cardiovascular hazards.
Bacteria utilize diverse defence systems to protect against harmful foreign DNA such as bacteriophages1,2, but how these systems coordinate with each other remains poorly understood. Here we uncover CRISIS (CRISPR-supervised immune system), a widespread regulatory paradigm whereby type I CRISPR-Cas loci embed and transcriptionally modulate diverse innate defences. Small non-canonical CRISPR RNA (crRNA)-like RNAs guide the I-C CRISPR-associated complex for antiviral defence (Cascade) effector complex to inhibit promoters of diverse immune cassettes-including composite multi-system clusters-enabling their basal expression for antiviral activity while mitigating fitness costs associated with hyperactivation, such as host growth impairment or exclusion of beneficial plasmids. When CRISPR-Cas is compromised by mutation or anti-CRISPR proteins, there is a burst in transcription of these embedded defence systems, leading to higher-level innate immunity at the expense of host fitness. Together, adaptive CRISPR-Cas systems orchestrate diverse innate immune systems into a layered defence network, comprising a prokaryotic 'immunity guard' strategy.
Mutations in gene regulatory regions have been shown to play a role in rapid adaptation, but the factors determining their contribution are largely unknown. Here, using the metabolic enzyme cytosine deaminase of budding yeast, we examine whether adaptation to 5-fluorocytosine, which requires reduced cytosine deamination and can readily arise from amino acid substitutions, may be reached by single promoter mutations. We generated all single-nucleotide substitutions and indels in the FCY1 promoter and assayed the resulting mutants in presence of 5-fluorocytosine. This revealed that no promoter mutation is sufficient for adaptation to occur. We next investigated how this inaccessibility of adaptation arises by combining large-scale expression measurements with the experimental characterization of the corresponding expression-fitness function. These experiments showed that the shape of this function precludes single promoter mutations from being adaptive. Although 24% of mutations significantly affect expression, the fitness curve is flat around wild-type level. As such, adaptation can only emerge from a severe reduction of expression, which cannot occur from a single mutation in the promoter. Our results show that the contribution of regulatory mutations to rapid adaptation depends not only on the distribution of mutational effect sizes on expression level but also on the shape of the function linking fitness to expression levels.
Often, more pollen grains land on recipient flowers than there are ovules to fertilize. Consequently, the haploid male gametophyte engages in post-pollination competition, one way that pollen genotype can influence inheritance. The maize (Zea mays subsp. mays L.) inflorescence (ear), with its elongated stigma and style structures (silks), has a conspicuous spatial heterogeneity, with longer silks at the base of the ear than at the apex. To evaluate the hypothesis that alleles with reduced pollen fitness influence the spatial distribution of progeny genotypes along the ear, we developed an updated phenotyping platform that maps fluorescently marked mutant (Ds-GFP) kernel phenotypes on the ear via an implementation of the Faster R-CNN machine vision model (EarVision.v2) and a statistical pipeline that evaluates the relationship between kernel position and transmission ratio (EarScape). Our dataset (1384 ears) represents 58 Ds-GFP insertion alleles. None of the 48 alleles with Mendelian inheritance showed any significant spatial trend. In contrast, 50% of alleles with a pollen-specific transmission defect (5/10) exhibited significant spatial effects. An insertional mutant of the gene encoding a putative actin-binding protein, base-to-apex gradient1* (bag1*), is associated with decreased mutant transmission at the ear base relative to the apex. Surprisingly, a mutant allele of another pollen-expressed gene (Zm00001eb236740) generates the opposite trend, decreased mutant transmission toward the ear apex; and two mutant alleles of the sperm cell attachment factor gamete expressed2 (gex2) can produce ears with transmission highest at both base and apex. We conclude that pollen fitness mutants cause unexpectedly diverse spatial patterns of progeny genotypes.
MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.
Inbreeding depression is a widespread phenomenon that reflects the burden of deleterious effects hidden in heterozygosis in non-inbred populations but exposed in homozygosis in inbred individuals, known as inbreeding load (B). This load can be due to partially or fully recessive deleterious mutations (dominance model) or to heterozygote advantage (overdominance model, where both homozygotes are deleterious relative to the heterozygote). There are many studies addressing the changes in inbreeding load in finite populations assuming the dominance model. However, the contribution of overdominance to inbreeding depression has been focused on infinite-size populations. We carried out computer simulations to investigate the joint impact of dominant and pure overdominant mutations on inbreeding load, both for self-fertilizing populations and for panmictic populations suffering from a drastic bottleneck. We found that the overdominant inbreeding load can be substantially reduced by drift even for symmetrical overdominance, at least when considering mutations of small effect. For panmictic bottlenecked populations, the reduction in inbreeding load under dominance and overdominance loci cannot be easily distinguished. However, while purging depletes inbreeding load from dominant loci, slowing inbreeding depression and leading to partial fitness recovery, for overdominant loci fitness declines monotonically.
MOTIVATION: Intratumor heterogeneity arises from ongoing somatic evolution and complicates cancer diagnosis, prognosis, and treatment. Reconstructing evolutionary dynamics typically requires spatiotemporal samples, which are often unavailable in clinical settings. Computational approaches that can infer tumor evolutionary history from single-timepoint bulk sequencing data remain limited. RESULTS: We present estimating evolutionary events through single-timepoint sequencing (TEATIME), a novel computational framework that models tumors as mixtures of two competing cell populations: an ancestral clone with baseline fitness and a derived subclone with elevated fitness. Using cross-sectional bulk sequencing data, TEATIME estimates mutation rates, timing of subclone emergence, relative fitness, and number of generations of growth. To quantify intratumor fitness asymmetries, we introduce a novel metric-fitness diversity-which captures the imbalance between competing cell populations and serves as a measure of functional intratumor heterogeneity. Applying TEATIME to 33 tumor types from The Cancer Genome Atlas, we revealed divergent as well as convergent evolutionary patterns. Notably, we found that immune-hot microenvironments constraint subclonal expansion and limit fitness diversity. Moreover, we detected temporal dependencies in mutation acquisition, where early driver mutations in ancestral clones epistatically shape the fitness landscape, predisposing specific subclones to selective advantages. These findings underscore the importance of intratumor competition and tumor-microenvironment interactions in shaping evolutionary trajectories, driving intratumor heterogeneity. Lastly, we demonstrate that TEATIME-derived evolutionary parameters and fitness diversity offer novel prognostic insights across multiple cancer types. AVAILABILITY AND IMPLEMENTATION: R implementation of TEATIME is available on GitHub (https://github.com/liliulab/TEATIME) and Zenodo (https://zenodo.org/records/17422174).
Drosophila melanogaster Down Syndrome cell adhesion molecule 1 (Dscam1) gene encodes 38,016 diverse cell surface receptor proteins via alternative splicing, which have both nervous and immune functions. However, it remains elusive why organisms have evolved such an astonishing diversity of isoforms. Here, we show that fitness and immunity properties have driven the modern evolution of Dscam1 isoform diversity. We assess multiple aspects of fly fitness in deletion mutants harboring exon 4, 6, or 9 clusters, respectively, reducing ectodomain isoform diversity stepwise from 18,612 to 396. All fitness-related traits generally improved as the potential number of isoforms increased; however, the magnitude of the changes varied remarkably in a variable cluster-specific manner. Correlation analysis revealed that fitness-related traits were much more sensitive to reductions in Dscam1 diversity compared to canonical neuronal self/non-self discrimination. We conclude that the role of Dscam1 isoforms in canonical neuronal self-avoidance and self/non-self discrimination is mediated by a small fraction of all isoforms (<1/10), whereas a separate role essential for other developmental contexts and resistances, likely in fitness and immunity, requires almost full isoform diversity. Thus, fitness and immunity properties, rather than canonical neuronal functions, are the dominant drivers during the modern diversification of the Dscam1 isoform. Our findings suggest that Dscam1 diversity is closely linked to adaptation and species diversification in arthropods.
A small percentage of species in the fungal kingdom can cause devastating infections in humans, with Candida albicans reigning as a leading cause of systemic disease. One of the key virulence phenotypes for pathogenic fungi is the ability to survive at host body temperature; however, a comprehensive understanding of the mechanisms that orchestrate thermal adaptation in fungi remains incomplete. In this study, we expand the largest functional genomics resource in C. albicans, reaching 71.3% coverage of the entire genome, and perform screens under six different temperatures to identify genes important for temperature-dependent fitness. We describe the function of genes involved in translation (GAR1), splicing (C1_11680C or YSF3), and cell cycle progression (C6_00110C or RHT1) in enabling fungal survival at both low and high temperatures. Through experimental evolution, we also show that C. albicans can rapidly overcome deleterious mutations and adapt to extreme temperature environments. Overall, our study highlights the transformative potential of genome-wide functional genomics to uncover critical vulnerabilities in pathogenic fungi.
Nitrogen-fixing microbes are a primary contributor of this important nutrient to the global nitrogen cycle. Biological nitrogen fixation (BNF) through the enzyme nitrogenase requires extensive energy that in whole cells is generally studied during the oxidation of carbohydrates such as sugars. The nitrogen-fixing bacterium Azotobacter vinelandii is a model diazotroph for the study of aerobic BNF. Much is known about metabolism in A. vinelandii when cultured on a simple medium where energy is provided primarily in the form of sucrose or glucose. Outside of the laboratory, this soil bacterium grows on metabolites primarily derived from plant root exudates or from the degradation of dead plant matter. In this work, we expand on previous studies looking at genes that are essential to BNF in A. vinelandii when grown on sucrose medium using transposon sequencing (Tn-seq). We applied Tn-seq to determine the genes essential to growth when the medium was shifted to acetate, succinate or glycerol as the primary carbon and energy source to fuel both growth and BNF. A global overview of the genes of central metabolism and those directing substrates toward central metabolism, along with a selection of unexpected genes that were essential for specific growth substrates, is provided.
The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.
Why is there so much non-neutral genetic variation segregating in natural populations? We dissect function and evolution of a near-cryptic quantitative trait locus (QTL) for defense metabolites in Arabidopsis using the CRISPR/Cas9 system and nucleotide polymorphism patterns. The QTL is explained by genetic variation in a family of 4 tightly linked indole-glucosinolate O-methyltransferase genes. Some of this variation appears to be maintained by balancing selection, some appears to be generated by non-reciprocal transfer of sequence, also known as ectopic gene conversion (EGC), between functionally diverged gene copies. Here, we elucidate how EGC, as an inevitable consequence of gene duplication, could be a general mechanism for generating genetic variation for fitness traits.
IMPORTANCE: Treatment to lower high levels of low-density lipoprotein cholesterol (LDL-C) reduces incident coronary artery disease (CAD) risk but modestly increases the risk for incident type 2 diabetes (T2D). The extent to which genetic factors across the cholesterol spectrum are associated with incident T2D is not well understood. OBJECTIVE: To investigate the association of genetic predisposition to increased LDL-C levels with incident T2D risk. DESIGN, SETTING, AND PARTICIPANTS: In this large prospective, population-based cohort study, UK Biobank participants who underwent whole-exome sequencing and genome-wide genotyping were included. Participants were separated into 7 groups with familial hypercholesterolemia (FH), predicted loss of function (pLOF) in APOB or PCSK9 variants, and LDL-C polygenic risk score (PRS) quintiles. Data were collected between 2006 and 2010, with a median follow-up of 13.7 (IQR, 12.9-14.5) years. Data were analyzed from March 1 to November 1, 2024. EXPOSURES: LDL-C level, LDL-C PRS, FH, or pLOF variant status. MAIN OUTCOMES AND MEASURES: Cox proportional hazards regression models adjusted for age, sex, genotyping array, lipid-lowering medication use, and the first 10 genetic principal components were fitted to assess the association between LDL-C genetic factors and incident T2D and CAD risks. RESULTS: Among the 361 082 participants, mean (SD) age was 56.8 (8.0) years, 194 751 (53.9%) were female, and mean (SD) baseline LDL-C level was 138.0 (33.6) mg/dL. During the follow-up period, 22 619 (6.3%) participants developed incident T2D and 17 966 (5.0%) developed incident CAD. The hazard ratio for incident T2D was lowest in the FH group (0.65; 95% CI, 0.54-0.77), while the highest risk was in the pLOF group (1.48; 95% CI, 1.18-1.86). The association between LDL-C PRS and incident T2D was 0.72 (95% CI, 0.66-0.79) for very high LDL-C PRS, 0.87 (95% CI, 0.84-0.90) for high LDL-C PRS, 1.13 (95% CI, 1.09-1.17) for low LDL-C PRS, and 1.26 (95% CI, 1.15-1.38) for very low LDL-C PRS. CAD risk increased directly with the LDL-C PRS. CONCLUSIONS AND RELEVANCE: In this cohort study, LDL-C and T2D risks were inversely associated across genetic mechanisms for LDL-C variation. Further elucidation of the mechanisms associating low LDL-C risk with increased risk of T2D is warranted.
How genetic variance for fitness is maintained is incompletely understood. Mutation-selection balance and single-locus overdominance cannot account for the large variance observed. Recent work suggests that antagonistic balancing selection, favoring different alleles in different contexts and involving beneficial dominance reversals, might contribute to maintaining fitness variance. However, while this mechanism is plausible, evidence for dominance reversals remains scarce. Here, we study how In(3R)Payne, a balanced inversion polymorphism in Drosophila melanogaster, affects gene expression and chromatin accessibility by using RNA-seq and ATAC-seq (assay for transposase-accessible chromatin with sequencing). We find that, in embryos, the inverted (INV) arrangement tends to have dominant effects, while the standard (STD) arrangement behaves like a recessive Mendelian allele. Yet, in wing discs, this pattern is reversed: STD has mostly dominant effects, whereas INV behaves recessively. Since this shift in the dominance of the INV "allele" between developmental contexts affects the expression of suites of genes in a concerted manner, it might be mediated by a dominance modifier, for example, a transcription factor. In favor of this idea, 25% of the differentially expressed genes between INV and STD encode transcription factors. Interestingly, while only four differentially expressed genes are shared between embryos and wing discs, one of them is HP1c, a chromatin-binding protein and major transcriptional regulator, and thus a promising candidate for mediating the context-dependent change in dominance. Although the relationship between these patterns and fitness is presently unknown, our observations are consistent with a potential role of reversals (or, more generally, shifts) of dominance in maintaining inversion polymorphism.