Search PubMedSearch

SEARCH · Search PubMed

Results for “fitness landscapes”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Bayesian inference of fitness landscapes via tree-structured branching processes.

MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.

Bayes Theorem

Epistasis and the changing fitness landscapes of SARS-CoV-2.

Since its emergence in late 2019, millions of SARS-CoV-2 genomes have been generated as part of global efforts to monitor the evolution and spread of the virus. This unprecedented volume of data provides a unique opportunity to study viral evolution at unparalleled resolution. In particular, individual genomic sites can be observed to have mutated independently thousands of times. These mutation counts have been used to estimate site-specific mutation rates and fitness effects for most mutations across the viral genome. Here, we use these data to investigate how the landscape of mutational fitness costs has changed over the course of the pandemic. SARS-CoV-2 evolution over the past 6 years has been characterized by the emergence of distinct variants separated by long branches corresponding to evolutionary saltations involving up to 50 mutations. We compare inferred fitness landscapes of the Spike protein across these variants and find that shifts in the estimated effects of non-synonymous mutations are linked to genetic differences between them. Sites with altered fitness costs are enriched near positions where the genetic backgrounds differ. To explain the observed changes, we introduce a model with pairwise epistatic interactions between mutations and residues that differ between variants. This model is able to explain about half of the variance in the shifts of fitness effects and suggests that each mismatch between variants substantially alters mutation effects at typically 1 to 3 additional positions.

SARS-CoV-2

Deconstructing empirical fitness seascapes across scales of granularity.

The fitness landscape metaphor remains resonant in evolutionary theory and has facilitated the birth of newer concepts, like the fitness seascape, that consider the role of environmental context in shaping the dynamics of evolution. Since its emergence, the seascape has appeared in numerous studies examining how different and fluctuating environments shape evolutionary outcomes. Despite growing interest, we lack comprehensive examinations of how environmental context shapes features of fitness seascapes. In this study, we address this gap by deconstructing empirical fitness seascapes across scales of granularity: loci, locus interactions (epistasis), alleles, trajectories, and entire seascapes. For each, we examine how environmental context influences qualitative and quantitative aspects of seascapes, and find that they change appreciably, with patterns specific to individual systems of study. We also quantify how much each scale varies across environments, and find that certain scales tend to be more sensitive to context than others. In summary, we reflect on the implications of the seascape metaphor for the incorporation of environmental effects into theoretical population genetics, for understanding how the environment shapes evolution in disease systems, and for contemporary bioengineering efforts.

Genetic Fitness

Protein overabundance is driven by growth robustness.

Protein expression levels optimize cell fitness: Too low an expression level of essential proteins will slow growth by compromising essential processes; whereas overexpression slows growth by increasing the metabolic load. This trade-off naïvely predicts that cells maximize their fitness by sufficiency, expressing just enough of each essential protein for function. We test this prediction in the naturally-competent bacterium Acinetobacter baylyi by characterizing the proliferation dynamics of essential-gene knockouts at a single-cell scale (by imaging) as well as at a genome-wide scale. In these experiments, cells proliferate for multiple generations as target protein levels are diluted from their endogenous levels. This approach facilitates a proteome-scale analysis of the fitness landscape with respect to protein abundance. We find that most essential proteins are subject to a threshold-like fitness landscape: growth is independent of protein abundance above a critical threshold and arrests below that threshold. We have recently analyzed the implications of this landscape for growth robustness. Confirming signature predictions of this model, we find that (i) roughly 70% of essential proteins are overabundant, (ii) overabundance increases as the expression level decreases and (iii) the lowest abundance proteins are in vast excess (>10×) of what is required for growth in the typical cell. These results reveal that robustness plays a fundamental role in determining the expression levels of essential genes and that overabundance is a key mechanism for ensuring robust growth.

Journal Article

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals

Competing subclones and fitness diversity shape tumor evolution across cancer types.

MOTIVATION: Intratumor heterogeneity arises from ongoing somatic evolution and complicates cancer diagnosis, prognosis, and treatment. Reconstructing evolutionary dynamics typically requires spatiotemporal samples, which are often unavailable in clinical settings. Computational approaches that can infer tumor evolutionary history from single-timepoint bulk sequencing data remain limited. RESULTS: We present estimating evolutionary events through single-timepoint sequencing (TEATIME), a novel computational framework that models tumors as mixtures of two competing cell populations: an ancestral clone with baseline fitness and a derived subclone with elevated fitness. Using cross-sectional bulk sequencing data, TEATIME estimates mutation rates, timing of subclone emergence, relative fitness, and number of generations of growth. To quantify intratumor fitness asymmetries, we introduce a novel metric-fitness diversity-which captures the imbalance between competing cell populations and serves as a measure of functional intratumor heterogeneity. Applying TEATIME to 33 tumor types from The Cancer Genome Atlas, we revealed divergent as well as convergent evolutionary patterns. Notably, we found that immune-hot microenvironments constraint subclonal expansion and limit fitness diversity. Moreover, we detected temporal dependencies in mutation acquisition, where early driver mutations in ancestral clones epistatically shape the fitness landscape, predisposing specific subclones to selective advantages. These findings underscore the importance of intratumor competition and tumor-microenvironment interactions in shaping evolutionary trajectories, driving intratumor heterogeneity. Lastly, we demonstrate that TEATIME-derived evolutionary parameters and fitness diversity offer novel prognostic insights across multiple cancer types. AVAILABILITY AND IMPLEMENTATION: R implementation of TEATIME is available on GitHub (https://github.com/liliulab/TEATIME) and Zenodo (https://zenodo.org/records/17422174).

Neoplasms

Surface architecture of the bacterial envelope determines phage adsorption route in pathogenic Escherichia coli O157:H7.

UNLABELLED: The outermost surface layers of Gram-negative bacteria determine phage access to terminal receptors, yet their genetic basis has been mapped almost exclusively in laboratory strains that lack them. Here we apply genome-wide RB-TnSeq fitness profiling to four Escherichia coli O157:H7 strains from distinct phylogenetic clades sharing the O157 O-antigen, using 38 phages with terminal receptors previously mapped in E. coli K-12 strain. RB-TnSeq fitness landscapes across all four pathogenic backgrounds were mostly similar, and dominated by surface-associated loci, including the gfc-etk group 4 capsule operon, O-antigen biosynthesis genes, LPS core assembly genes and outer membrane proteins. Disruption of gfc-etk abolished infection in 11 genetically diverse myoviruses, establishing the O-antigen capsule as a widespread required primary recognition substrate. O-antigen loci generated two classes of fitness score patterns. For 10 phages, disruption increased infectivity, indicating it is a barrier to receptor access; for 3 others, disruption abolished infectivity, demonstrating it can also be a primary recognition substrate. Outer membrane protein receptor identity was conserved across laboratory and pathogenic backgrounds, with the same proteins recognized in both K-12 and O157:H7, while glycan layer state determines whether these receptors are reached. These results demonstrate that outer surface glycan layers can act as primary and optional recognition substrates for phage infection, or as physical barriers preventing terminal receptor access. Extending the ability to probe phage-targeted receptors beyond outer membrane proteins provides a framework for incorporating glycan layer state into predictive models of phage-host interactions. IMPORTANCE: Bacteriophage-based interventions for controlling Escherichia coli O157:H7, a major foodborne pathogen responsible for tens of thousands of illnesses annually in the United States, require a mechanistic understanding of the factors governing strain-level susceptibility. Predictive frameworks developed in laboratory model strains lacking O-antigen and capsular polysaccharides can map the terminal protein receptors that phages bind, but are currently limited in their ability to determine whether those receptors are accessible in pathogenic isolates carrying full outer surface complexity. This study provides the first genome-scale, functional genetic map of phage susceptibility determinants in O157:H7 and demonstrates that the state of the outer surface layers, specifically the O-antigen and the gfc-etk capsule, determines whether phages can reach conserved terminal receptors. This finding explains differences in phage susceptibility between strains sharing nearly identical gene content, and identifies the molecular layers that must be characterized to predict phage host interaction in pathogenic E. coli backgrounds.

Journal Article

Pneumococcal population structure influences the effects of air pollution on invasive disease risk in South Africa.

Streptococcus pneumoniae is highly diverse, comprising over 100 serotypes and hundreds of genomic lineages amid widespread vaccination. While it can cause invasive pneumococcal disease (IPD) which exhibits pronounced seasonal spikes, the interplay between pneumococcal diversity and environmental drivers remains unexplored. Here we analysed 59,017 IPD cases over 19 years from South Africa, incorporating 4,350 genome-sequenced isolates, using Bayesian spatiotemporal models to link environmental exposure and pneumococcal diversity. Cumulatively, across an 8-week period, moderate relative humidity (33-49%) and cold minimum temperatures (4-10 °C) increased IPD risk by 5% and 4%, respectively. Conversely, warm maximum temperatures (27-38 °C) were associated with up to a 10% increased risk within a week of exposure. There was a positive association between air pollution (PM2.5) and IPD, although it varied by age, disease presentation, and most notably serotype and lineage. Specifically, the lag time between PM2.5 exposure and disease onset varied by serotype, with only serotypes 4, 8 and 23F conferring an immediate IPD risk. High prevalence of GPSC21 lineage (serotype 19F) also modified the pollution response, shifting the lag structure to produce immediate risk of disease following high PM2.5 exposure. Our results demonstrate that pneumococcal population structure shapes air quality risk which in turn can shape the fitness landscape of microbial populations. Integration of these data may inform public health policy.

Journal Article

Domestication as gene-culture coevolution.

Human preferences can shape the genetic evolution of other species via conservation practices, public health actions, and domestication. While the dynamics of domestication have been explored in depth through empirical and theoretical analyses, few studies have analyzed models for the coevolution of human cultural preferences with the genetics of a domesticate population. Humans shape the fitness landscape of domesticate populations both intentionally and unconsciously, by selecting for desirable traits and modifying environments; in turn, changes in domesticate phenotypes can affect the cultural preferences in the domesticator population. We present a model for the dynamics of domestication which includes interactions between genetic evolution, cultural transmission, and selective pressures. The model includes forms of selection due to culturally transmitted domesticator preferences that can affect the dynamics of domesticate genetic variants, which then affect the dynamics of domesticators. Equilibria with simultaneous genetic and cultural polymorphisms may exist, and may occur under apparent heterozygote disadvantage in the domesticate. Stable quasiperiodic cycles in both domesticates and domesticators are also possible.

Humans

The genetic code at the balance point of error and demand.

The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.

Genetic Code

EprX associates with concurrent shifts in antimicrobial resistance and virulence in clinical bloodstream E. coli: a putative adaptive node for bacterial fitness.

Bloodstream infections (BSIs) caused by E. coli represent a growing global threat, driven by escalating antimicrobial resistance (AMR) and sustained virulence. However, the regulatory mechanisms linking these two phenotypes remain poorly understood. Here, we identify EprX, a previously uncharacterized YjbI-type pentapeptide repeat protein (PRP), a locus that our data suggest may influence metabolic and transcriptional profiles in clinical BSI E. coli isolates. Genomic screening of 85 clinical BSI strains reveals that eprX is present in 21.2% of isolates, often within distinct genomic contexts suggestive of mobile acquisition. Using λ-Red recombineering, we constructed eprX knockout mutants. Loss of eprX is associated with altered antimicrobial resistance profiles, increasing susceptibility to gentamicin, ciprofloxacin, and levofloxacin. This phenotype is consistent with upregulation of outer membrane porin genes (ompC, ompF) and downregulation of multidrug efflux pump genes (macB, mdtC, emrB) and two-component regulatory system genes. eprX deficiency also appears to correlate with attenuated virulence in our assays, as evidenced by improved survival of Galleria mellonella larvae (65-95% at 72 h post-infection vs. 40-60% for wild-type strains) and reduced adhesion to and invasion of human HeLa cells. Transcriptomic profiling reveals that eprX carriage is associated with broad, coordinated shifts in the expression of genes involved in LPS transport (lptG/lptF), type ;II secretion system components (gspD/gspE/gspF), autotransporter adhesins (ag43), and flagellar assembly, suggesting potential disruptions in outer-membrane integrity, biofilm formation, and virulence programs. Our data suggests that eprX is a genetic locus whose presence correlates with concurrent shifts in resistance maintenance and virulence traits, representing a putative adaptive node within the E. coli fitness landscape.

Animals

Inference and visualization of complex genotype-phenotype maps with gpmap-tools.

Understanding how biological sequences give rise to observable traits, that is, how genotype maps to phenotype, is a central goal in biology. Yet our knowledge of genotype-phenotype maps in natural systems is limited due to the high dimensionality of sequence space and the context-dependent effects of mutations. The emergence of Multiplex assays of variant effect (MAVEs), along with large collections of natural sequences, offer new opportunities to empirically characterize these maps at an unprecedented scale. However, tools for statistical and exploratory analysis of these high-dimensional data are still needed. To address this gap, we developed gpmap-tools (https://github.com/cmarti/gpmap-tools), a python library that integrates a series of models for inference, phenotypic imputation, and error estimation from MAVE data or collections of natural sequences in the presence of genetic interactions of every possible order. gpmap-tools also provides methods for summarizing patterns of epistasis and visualization of genotype-phenotype maps containing up to millions of genotypes. To demonstrate its utility, we used gpmap-tools to infer genotype-phenotype maps containing 262,144 variants of the Shine-Dalgarno sequence from both genomic 5'UTR sequences and experimental MAVE data. Visualization of the inferred landscapes consistently revealed high-fitness ridges that link core motifs at different distances from the start codon. In summary, gpmap-tools provides a flexible, interpretable framework for studying complex genotype-phenotype maps, opening new avenues for understanding the architecture of genetic interactions and their evolutionary consequences.

Gaussian process

A new MRR1 gain-of-function mutation involved in cross-resistance to antifungal agents in the fungal priority pathogen Candida parapsilosis.

OBJECTIVES: Candida parapsilosis is a leading cause of invasive candidiasis globally, with rising reports of fluconazole resistance threatening its clinical management. Among the mechanisms involved, gain-of-function mutations in the MRR1 gene have emerged as key drivers of antifungal resistance. We aimed to investigate a novel amino acid substitution (G982E) in the Mrr1 zinc cluster transcription factor, identified in a fluconazole-resistant C. parapsilosis isolate from a patient exposed to fluconazole. METHODS: Using CRISPR-Cas9 genome editing, we introduced the G982E variant into two fluconazole-susceptible C. parapsilosis genetic backgrounds. The antifungal susceptibility of the engineered mutants was assessed in vitro against a broad panel of systemic antifungal agents. A Galleria mellonella infection model was also used to evaluate the impact of the G982E variant on antifungal treatment efficacy and virulence in vivo. RESULTS: Acquisition of the G982E substitution dramatically altered the antifungal susceptibility profile, particularly for fluconazole for which the MIC increased to >256 µg/mL. However, the magnitude of the MIC increase varied by azole, with the greatest increase seen for fluconazole (>9-10-fold), followed by voriconazole (5-fold), isavuconazole (3-fold), but also flucytosine (1.5-fold). In contrast, susceptibility to posaconazole remained largely unchanged. In vivo, this new variant conferred fluconazole treatment failure but was associated with a significant reduction in virulence. CONCLUSIONS: The G982E is a novel Mrr1 gain-of-function mutation driving high-level fluconazole resistance in C. parapsilosis. These findings reinforce the central role of Mrr1 in antifungal resistance, underscore the functional diversity of its mutational landscape, with potential implications for fungal fitness and transcriptional regulation.

Candida parapsilosis

A model for background selection in non-equilibrium populations.

In many taxa, levels of genetic diversity are observed to vary along their genome. The framework of background selection models this variation in terms of linkage to constrained sites, and recent applications have been able to explain a large portion of the variation in human genomes. However, these studies have also yielded conflicting results, stemming from two key limitations. First, existing models are inaccurate in a critical region of parameter space (), where the local reduction in diversity is sharpest. Second, they assume a constant population size over time. Here, we develop predictions for diversity under background selection based on the Hill-Robertson system of two-locus statistics, which allows for population size changes. We treat the joint effect of multiple selected loci independently, but we show that interference among them is well captured through local rescaling of mutation, recombination and selection in an iterative procedure that converges quickly. We further accommodate existing background selection theory to non-equilibrium demography, bridging the gap between weak and strong selection. Simulations show that our predictions are accurate across the entire range of selection coefficients. We characterize the temporal dynamics of linked selection under population size changes and demonstrate that patterns of diversity can be misinterpreted by other models. Specifically, biases due to the incorrect assumption of equilibrium carry over to downstream inferences of the distribution of fitness effects and deleterious mutation rate. Jointly modeling demography and linked selection therefore improves our understanding of the genomic landscape of diversity, which will help refine inferences of linked selection in humans and other species.

Journal Article

EscaPRRS-ORF5: a structure-aware evolutionary framework for prioritizing immune escape-prone variants in porcine reproductive and respiratory syndrome virus.

MOTIVATION: Porcine Reproductive and Respiratory Syndrome Virus (PRRSV) is a rapidly evolving RNA virus causing significant economic losses, posing a formidable challenge to vaccine efficacy due to its high mutational variability and immune escape. As the viral mutants evolve, their ability to sustain in population is driven by a range of host biology factors such as receptor binding, fusion, and uncoating. Existing tools that predict viral fitness and escape propensities rely heavily on extensive, up-to-date sequence data and lack integration of biochemical host interactions, limiting mechanistic understanding of the mutational landscape. We introduce Esca, a sequence-only toolchain framework that identifies immune escape-prone residues by exhaustively scanning each residue position for all amino acid substitutions using a Bayesian Variational Autoencoder (VAE) trained on protein language model embeddings. We demonstrate Esca on the GP5(ORF5) glycoprotein of PRRSV (EscaPRRS-ORF5) by training on ESM-2 embeddings of 32 146 GP5 sequences (2015-2022) spanning 140 sub-lineages. RESULTS: Despite being trained only on GP5 sequence data, EscaPRRS-ORF5 recovered 85.7% of the surface-exposed receptor binding interfaces as escape-prone regions. We use a mutation-sensitive fitness scoring scheme that goes beyond Hamming distances, to predict antibody escape tendencies, supporting surveillance of (re) emerging PRRSV variants. We do not claim that ORF5 alone captures PRRSV evolution or serves as a surveillance endpoint; rather, Esca offers a scalable path toward whole-genome, structure-aware surveillance. AVAILABILITY AND IMPLEMENTATION: EscaPRRS-ORF5 is freely available at https://doi.org/10.6084/m9.figshare.32661033 with an interactive Colab notebook at https://colab.research.google.com/drive/1TEgzAhPwvNAZ01VXeJbIFibfri2jnDA5? usp=sharing.

Porcine respiratory and reproductive syndrome viru

In vivo CRISPRi screens reveal Escherichia coli functional adaptations in the mouse gut.

Escherichia coli exhibits remarkable genetic diversity that enables it to adapt to the intestinal environment. Here we establish an in vivo CRISPR interference platform that leverages bacterial gene fitness as a high-resolution functional reporter of E. coli adaptations within mice harbouring a defined minimal microbial community (OligoMM12). The screen revealed that diet profoundly shapes the metabolic landscape of E. coli and the essential gene profile identified cross-feeding interactions. Comparison between a laboratory strain (MG1655), a uropathogenic strain (CFT073) and an adherent-invasive E. coli (AIEC LF82) identified distinct genetic requirements for intestinal colonization, highlighting divergent motility, stress response and respiration strategies. In a host inflammatory environment, we found that AIEC LF82 preferably colonized the small intestine with a mobile genetic element, Gally prophage, playing an important role in modulating fitness. These findings provide a high-resolution genetic atlas of E. coli's functional adaptation and demonstrate the utility of functional genomics to probe the gut environment itself.

Journal Article

NumSimEX: A method using EXX hydrogen exchange mass spectrometry to map the energetics of protein folding landscapes.

Hydrogen exchange mass spectrometry (HXMS) is a powerful tool to understand protein folding pathways and energetics. However, HXMS experiments to date have used exchange conditions termed EX1 or EX2 which limit the information that can be gained compared to the more general EXX exchange regime. If EXX behavior could be understood and analyzed, a single HXMS timecourse on an intact protein could fully map its folding landscape without requiring denaturation. To address this challenge, we developed a numerical simulation method called NumSimEX that models EXX exchange for arbitrarily complex folding pathways. NumSimEx fits protein folding dynamics to experimental HXMS data by iteratively comparing the simulated and experimental timecourses, allowing for determination of both kinetic and thermodynamic protein folding parameters. After analytically verifying NumSimEX's accuracy, we demonstrated its power on HXMS data from beta-2 microglobulin (β2M), a protein involved in dialysis-related amyloidosis. In particular, using NumSimEX, we identified three-state kinetics that near-perfectly matched experimental observation. This proof-of-principle application of NumSimEX sets the stage for harnessing HXMS to expand our understanding of proteins currently excluded from traditional protein folding methods. NumSimEX is freely available at https://github.com/JaswalLab/NumSimEX_Public.

Protein Folding

Evolution of maize recombination landscape during domestication.

Despite the plethora of knowledge about the benefits of meiotic recombination and numerous theoretical studies examining how recombination rates evolve, there is a general lack of empirical support and consensus across species. To fill this knowledge gap, we characterized the evolution of recombination landscape in maize during its domestication from teosinte and related the observed changes to established theoretical frameworks. Through examining recombination in experimental populations of maize and teosinte and the population genomics approach of identifying historical recombination events using ancestral recombination graph inference to generate saturated maize and teosinte recombination maps, we found that during domestication, maize experienced a 12% increase in its genome-wide recombination rate. Furthermore, maize evolved higher recombination rates on the long arms of chromosomes in regions closer to centromeres, where recombination is generally very low. The repatterning of crossover events came from changes in global crossover positioning rather than alterations in cis-acting chromatin factors. Consequently, we found evidence of selection acting on trans-acting recombination modifiers affecting crossover interference and controlling the interference-dependent class I crossover pathway. We show that CO repatterning was likely beneficial for maize fitness, as significant recombination rate increases were predominantly in gene-rich regions, which harbor domestication-related variation. This work suggests genomic and mechanistic processes leading to the evolution of meiotic recombination landscape in response to directional selection pressure and provides evidence for the evolutionary advantage of recombination.

Zea mays