Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “selective sweeps”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Prospects for identifying functional variation across the genome.

The genetic factors contributing to complex trait variation may reside in regulatory, rather than protein-coding portions of the genome. Within noncoding regions, SNPs in regulatory elements are more likely to contribute to phenotypic variation than those in nonregulatory regions. Thus, it is important to be able to identify and annotate noncoding regulatory elements. DNA conservation among diverged species successfully identifies noncoding regulatory regions. However, because rapidly evolving regulatory regions will not generally be conserved across species, these will not detected by using purely conservation-based methods. Here we describe additional approaches that can be used to identify putative regulatory elements via signatures of nonneutral evolution. An examination of the pattern of polymorphism both within and between populations of Drosophila melanogaster, as well as divergence with its sibling species Drosophila simulans, across 24.2 kb of noncoding DNA identifies several nonneutrally evolving regions not identified by conservation. Because different methods tag different regions, it appears that the methods are complementary. Patterns of variation at different elements are consistent with the action of selective sweeps, balancing selection, or population differentiation. Together with regions conserved between D. melanogaster and Drosophila pseudoobscura, we tag 5.3 kb of noncoding DNA as potentially regulatory. Ninety-seven of the 408 common noncoding SNPs surveyed are within putatively regulatory regions. If these methods collectively identify the majority of functional noncoding polymorphisms, genotyping only these SNPs in an association mapping framework would reduce genotyping effort for noncoding regions 4-fold.

Animals↗

Sweeps in Space: Leveraging Geographic Data to Identify Beneficial Alleles in Anopheles gambiae.

As organisms adapt to environmental changes, natural selection modifies the frequency of nonneutral alleles. For beneficial mutations, the outcome of this process may be a selective sweep, in which an allele rapidly increases in frequency and perhaps reaches fixation within a population. Selective sweeps have well-studied effects on patterns of local genetic variation in panmictic populations, but much less is known about the dynamics of sweeps in continuous space. In particular, because limited movement across a landscape leads to unique patterns of population structure, spatial dynamics may influence the trajectory of selected mutations. Here, we use forward-in-time, individual-based simulations in continuous space to study the impact of space on beneficial mutations as they sweep through a population. In particular, we show that selection changes the joint distribution of allele frequency and geographic range occupied by a focal allele and demonstrate that this signal can be used to identify selective sweeps. We then leverage this signal to identify in-progress selective sweeps within the malaria vector Anopheles gambiae, a species under strong selection pressure from vector control measures. By considering space, we identify multiple previously undescribed variants with potential phenotypic consequences, including mutations impacting known IR-associated genes and altering protein structure and properties. Our results demonstrate a novel signal for detecting selection in spatial population genetic data that may have implications for genomic surveillance and understanding geographic patterns of genetic variation.

Animals↗

Sweeps in space: leveraging geographic data to identify beneficial alleles in Anopheles gambiae.

As organisms adapt to environmental changes, natural selection modifies the frequency of non-neutral alleles. For beneficial mutations, the outcome of this process may be a selective sweep, in which an allele rapidly increases in frequency and perhaps reaches fixation within a population. Selective sweeps have well-studied effects on patterns of local genetic variation in panmictic populations, but much less is known about the dynamics of sweeps in continuous space. In particular, because limited movement across a landscape leads to unique patterns of population structure, spatial dynamics may influence the trajectory of selected mutations. Here, we use forward-in-time, individual-based simulations in continuous space to study the impact of space on beneficial mutations as they sweep through a population. In particular, we show that selection changes the joint distribution of allele frequency and geographic range occupied by a focal allele and demonstrate that this signal can be used to identify selective sweeps. We then leverage this signal to identify in-progress selective sweeps within the malaria vector Anopheles gambiae , a species under strong selection pressure from vector control measures. By considering space, we identify multiple previously undescribed variants with potential phenotypic consequences, including mutations impacting known IR-associated genes and altering protein structure and properties. Our results demonstrate a novel signal for detecting selection in spatial population genetic data that may have implications for genomic surveillance and understanding geographic patterns of genetic variation.

Journal Article↗

Interference of Competing Beneficial Mutations on Recombining Chromosomes.

Finding signatures of selective sweeps in genomes is a major goal of current population genomics, as it allows estimating the rate of beneficial mutations going to fixation and identifying the genes involved in selection. Models of recurrent selective sweeps traditionally assume that in chromosomal regions of normal recombination rates at most one beneficial allele is on the way to fixation. We review and extend here the theoretical studies on interference between closely linked beneficial mutations suggesting that this assumption may be violated. We show that interference between beneficial mutations may lead to substantially increased fixation times even in chromosomal regions of normal recombination rates. Furthermore, we discuss how interference can be detected in population genomic studies by analyzing genetic footprints of selective sweeps, and search for empirical evidence of interference in published datasets.

fixation times↗

Soft sweeps III: the signature of positive selection from recurrent mutation.

Polymorphism data can be used to identify loci at which a beneficial allele has recently gone to fixation, given that an accurate description of the signature of selection is available. In the classical model that is used, a favored allele derives from a single mutational origin. This ignores the fact that beneficial alleles can enter a population recurrently by mutation during the selective phase. In this study, we present a combination of analytical and simulation results to demonstrate the effect of adaptation from recurrent mutation on summary statistics for polymorphism data from a linked neutral locus. We also analyze the power of standard neutrality tests based on the frequency spectrum or on linkage disequilibrium (LD) under this scenario. For recurrent beneficial mutation at biologically realistic rates, we find substantial deviations from the classical pattern of a selective sweep from a single new mutation. Deviations from neutrality in the level of polymorphism and in the frequency spectrum are much less pronounced than in the classical sweep pattern. In contrast, for levels of LD, the signature is even stronger if recurrent beneficial mutation plays a role. We suggest a variant of existing LD tests that increases their power to detect this signature.

Alleles↗

Controlling the false-positive rate in multilocus genome scans for selection.

Rapid typing of genetic variation at many regions of the genome is an efficient way to survey variability in natural populations in an effort to identify segments of the genome that have experienced recent natural selection. Following such a genome scan, individual regions may be chosen for further sequencing and a more detailed analysis of patterns of variability, often to perform a parametric test for selection and to estimate the strength of a recent selective sweep. We show here that not accounting for the ascertainment of loci in such analyses leads to false inference of natural selection when the true model is selective neutrality, because the procedure of choosing unusual loci (in comparison to the rest of the genome-scan data) selects regions of the genome with genealogies similar to those expected under models of recent directional selection. We describe a simple and efficient correction for this ascertainment bias, which restores the false-positive rate to near-nominal levels. For the parameters considered here, we find that obtaining a test with the expected distribution of P-values depends on accurately accounting both for ascertainment of regions and for demography. Finally, we use simulations to explore the utility of relying on outlier loci to detect recent selective sweeps. We find that measures of diversity and of population differentiation are more effective than summaries of the site-frequency spectrum and that sequencing larger regions (2.5 kbp) in genome-scan studies leads to more power to detect recent selective sweeps.

Demography↗

Accessible, realistic genome simulation with selection using stdpopsim.

Selection is a fundamental evolutionary force that shapes patterns of genetic variation across species. However, simulations incorporating realistic selection along heterogeneous genomes in complex demographic histories are challenging, limiting our ability to benchmark statistical methods aimed at detecting selection and to explore theoretical predictions. stdpopsim is a community-maintained simulation library that already provides an extensive catalog of species-specific population genetic models. Here we present a major extension to the stdpopsim framework that enables simulation of various modes of selection, including background selection, selective sweeps, and arbitrary distributions of fitness effects (DFE) acting on annotated subsets of the genome (for instance, exons). This extension maintains stdpopsim's core principles of reproducibility and accessibility while adding support for species-specific genomic annotations and published DFE estimates. We demonstrate the utility of this framework by comparing methods for demographic inference, DFE estimation, and selective sweep detection across several species and scenarios. Our results demonstrate the robustness of demographic inference methods to selection on linked sites, reveal the sensitivity of DFE-inference methods to model assumptions, and show how genomic features, like recombination rate and functional sequence density, influence power to detect selective sweeps. This extension to stdpopsim provides a powerful new resource for the population genetics community to explore the interplay between selection and other evolutionary forces in a reproducible, user-friendly framework.

Journal Article↗

Non-African origin of a local beneficial mutation in D. melanogaster.

It is well understood that the out-of-Africa habitat expansion of D. melanogaster was associated with the fixation of many beneficial mutations. Nevertheless, it is not clear yet whether these beneficial mutations segregated already in Africa or originated outside of Africa. In this article, we describe an ongoing selective sweep specific to one European population. One microsatellite allele has increased in a population from The Netherlands to a frequency of 18%, whereas it is virtually absent in 12 other European populations. The selective sweep resulted in a genomic region of more than 600 kb that is identical by descent. This is probably the first evidence of a beneficial mutation that has arisen outside of Africa and has resulted in a selective sweep localized in a population from The Netherlands.

Adaptation, Biological↗

Palliating the impact of fixation of a major gene on the genetic variation of artificially selected polygenes.

Selective sweeps of variation caused by fixation of major genes may have a dramatic impact on the genetic gain from background polygenic variation, particularly in the genome regions closely linked to the major gene. The response to selection can be restrained because of the reduced selection intensity and the reduced effective population size caused by the increase in frequency of the major gene. In the context of a selected population where fixation of a known major gene is desired, the question arises as to which is the optimal path of increase in frequency of the gene so that the selective sweep of variation resulting from its fixation is minimized. Using basic theoretical arguments we propose a frequency path that maximizes simultaneously the effective population size applicable to the selected background and the selection intensity on the polygenic variation by minimizing the average squared selection intensity on the major gene over generations up to a given fixation time. We also propose the use of mating between carriers and non-carriers of the major gene, in order to promote the effective recombination between the major gene and its linked polygenic background. Using a locus-based computer simulation assuming different degrees of linkage, we show that the path proposed is more effective than a similar path recently published, and that the combination of the selection and mating methods provides an efficient way to palliate the negative effects of a selective sweep.

Computer Simulation↗

DNA diversity in sex-linked and autosomal genes of the plant species Silene latifolia and Silene dioica.

The relatively recent origin of sex chromosomes in the plant genus Silene provides an opportunity to study the early stages of sex chromosome evolution and, potentially, to test between the different population genetic processes likely to operate in nonrecombining chromosomes such as Y chromosomes. We previously reported much lower nucleotide polymorphism in a Y-linked gene (SlY1) of the plant Silene latifolia than in the homologous X-linked gene (SlX1). Here, we report a more extensive study of nucleotide diversity in these sex-linked genes, including a larger S. latifolia sample and a sample from the closely related species Silene dioica, and we also study the diversity of an autosomal gene, CCLS37.1. We demonstrate that nucleotide diversity in the Y-linked genes of both S. latifolia and S. dioica is very low compared with that of the X-linked gene. However, the autosomal gene also has low DNA polymorphism, which may be due to a selective sweep. We use a single individual of the related hermaphrodite species Silene conica, as an outgroup to show that the low SlY1 diversity is not due to a lower mutation rate than that for the X-linked gene. We also investigate several other possibilities for the low SlY1 diversity, including differential gene flow between the two species for Y-linked, X-linked, and autosomal genes. The frequency spectrum of nucleotide polymorphism on the Y chromosome deviates significantly from that expected under a selective-sweep model. However, we detect population subdivision in both S. latifolia and S. dioica, so it is not simple to test for selective sweeps. We also discuss the possibility that Y-linked diversity is reduced due to highly variable male reproductive success, and we conclude that this explanation is unlikely.

DNA, Plant↗

Polaris: Polarization of ancestral and derived polymorphic alleles for inferences of extended haplotype homozygosity in human populations.

SUMMARY: Statistical methods that measure the extent of haplotype homozygosity on chromosomes have been highly informative for identifying episodes of recent selection. For example, the integrated haplotype score (iHS) and the extended haplotype homozygosity (EHH) statistics detect long-range haplotype structure around derived and ancestral alleles indicative of classic and soft selective sweeps, respectively. However, to our knowledge, there are currently no publicly available methods that classify ancestral and derived alleles in genomic datasets for the purpose of quantifying the extent of haplotype homozygosity. Here, we introduce the Polaris package, which polarizes chromosomal variants into ancestral and derived alleles and creates corresponding genetic maps for analysis by selscan and HaploSweep, two versatile haplotype-based programs that perform scans for selection. With the input files generated by Polaris, selscan and/or HaploSweep can produce the appropriate sign (either positive or negative) for outlier iHS statistics, enabling users to distinguish between selection on derived or ancestral alleles. In addition, Polaris can convert the numerical output of these analyses into graphical representations of selective sweeps, increasing the functionality of our software. RESULTS: To demonstrate the utility of our approach, we applied the Polaris package to Chromosome 2 in the European Finnish, Middle Eastern Bedouin, and East African Maasai populations. More specifically, we examined the regulatory sequence in intron 13 of the MCM6 gene associated with lactase persistence (i.e. the ability to digest the lactose sugar present in fresh milk), a region of intense interest to human evolutionary geneticists. Our analyses showed that derived alleles (at known enhancers for lactase expression) sit on an extended haplotype background in the Finnish, Bedouin, and Maasai consistent with a classic selective sweep model as determined by iHS and EHH statistics. Importantly, we were able to immediately identify this target allele under selection based on the information generated by our software. We also explored outlier statistics across Chromosome 2 in two distinct datasets from these populations: (i) one containing polarized alleles generated with Polaris and (ii) the other containing unpolarized alleles in the original phased vcf file. Here, we found an excess of outlier statistics on Chromosome 2 in the unpolarized datasets, raising the possibility that a subset of these "hits" of selection may be unreliable. Overall, Polaris is a versatile package that enables users to efficiently explore, interpret, and report signals of recent selection in genomic datasets. AVAILABILITY AND IMPLEMENTATION: The Polaris package is free and open source on GitHub (https://github.com/alisi1989/Polaris) and DropBox (https://www.dropbox.com/scl/fo/mlxizft5267vem9u62qkn/AAnM0qX923zPzQBlPX8iteM?rlkey=uezrp4t2waffpj0nmo1evr320&e=1&st=jaodccws&dl=0).

Haplotypes↗

Multiple signatures of positive selection downstream of notch on the X chromosome in Drosophila melanogaster.

To identify genomic regions affected by the rapid fixation of beneficial mutations (selective sweeps), we performed a scan of microsatellite variability across the Notch locus region of Drosophila melanogaster. Nine microsatellites spanning 60 kb of the X chromosome were surveyed for variation in one African and three non-African populations of this species. The microsatellites identified an approximately 14-kb window for which we observed relatively low levels of variability and/or a skew in the frequency spectrum toward rare alleles, patterns predicted at regions linked to a selective sweep. DNA sequence polymorphism data were subsequently collected within this 14-kb region for three of the D. melanogaster populations. The sequence data strongly support the initial microsatellite findings; in the non-African populations there is evidence of a recent selective sweep downstream of the Notch locus near or within the open reading frames CG18508 and Fcp3C. In addition, we observe a significant McDonald-Kreitman test result suggesting too many amino acid fixations species wide, presumably due to positive selection, at the unannotated open reading frame CG18508. Thus, we observe within this small genomic region evidence for both recent (skew toward rare alleles in non-African populations) and recurring (amino acid evolution at CG18508) episodes of positive selection.

Animals↗

High-Density Genome-Wide Association Mapping Identifies Candidate Loci Associated with Maize Stalk Cell Wall Composition.

Maize (Zea mays L.) stalk cell wall composition is a key determinant of forage digestibility, lodging resistance, and biomass utilization efficiency. Although previous genome-wide association studies (GWAS) have identified loci associated with lignin (LIG), cellulose (CEL), and hemicellulose (HC), advances in genomic resources provide an opportunity to revisit existing phenotypic datasets at substantially higher resolution. Here, we re-analyzed a maize association panel consisting of 341 diverse inbred lines using an expanded genotype dataset containing 10.77 million SNPs, two derived compositional indices (CEL/HC and [LIG/(CEL + HC)], and six complementary GWAS models. Across all traits and models, we identified 855 unique significant SNPs associated with 579 candidate genes. Among the traits examined, LIG/(CEL + HC) yielded the greatest number of associations, suggesting that indices representing the relative balance among cell wall components may better capture the genetic architecture of cell wall composition than individual component measurements alone. Integration of multiple GWAS models with functional enrichment, haplotype, and selective sweep analyses prioritized three biologically relevant candidate genes encoding a MYB58 transcription factor, the glycosyltransferase Xt9, and a putative xyloglucan 6-xylosyltransferase. Haplotype analysis revealed significant effects of Xt9 and the xyloglucan 6-xylosyltransferase on cell wall composition, while selective sweep analysis identified Xt9 as a target of repeated selection during maize domestication, ecological adaptation, and modern breeding. Although these candidate genes provide promising targets for future investigation, the associations identified here are based on a single association panel and require functional and independent population validation. Collectively, our results demonstrate how high-density genotyping combined with complementary GWAS models can refine candidate associations and generate testable hypotheses from existing phenotypic datasets.

cell wall composition↗

The structure of a local population of phytopathogenic Pseudomonas brassicacearum from agricultural soil indicates development under purifying selection pressure.

Among the isolates of a bacterial community from a soil sample taken from an agricultural plot in northern Germany, a population consisting of 119 strains was obtained that was identified by 16S rDNA sequencing and genomic fingerprinting as belonging to the recently described species Pseudomonas brassicacearum. Analysis of the population structure by allozyme electrophoresis (11 loci) and random amplified polymorphic DNA-polymerase chain reaction (RAPD-PCR; four primers) showed higher resolution with the latter method. Both methods indicated the presence of three lineages, one of which dominated strongly. Stochastic tests derived from the neutral theory of evolution (including Slatkin's exact test, Watterson's homozygosity test and the Tajima test) indicated that the population had developed under strong purifying selection pressure. The presence of strains clearly divergent from the majority of the population can be explained by in situ evolution or by influx of strains as a result of migration or both. Phytopathogenicity of a P. brassicacearum strain determined with tomato plants reached the level obtained with the type strain of the known pathogen Pseudomonas corrugata. The results show that a selective sweep was identified in a local population. Previously, a local selective sweep had not been seen in several populations of different bacterial species from a variety of environmental habitats.

Agriculture↗

Malaria's Eve: evidence of a recent population bottleneck throughout the world populations of Plasmodium falciparum.

We have analyzed DNA sequences from world-wide geographic strains of Plasmodium falciparum and found a complete absence of synonymous DNA polymorphism at 10 gene loci. We hypothesize that all extant world populations of the parasite have recently derived (within several thousand years) from a single ancestral strain. The upper limit of the 95% confidence interval for the time when this most recent common ancestor lived is between 24,500 and 57,500 years ago (depending on different estimates of the nucleotide substitution rate); the actual time is likely to be much more recent. The recent origin of the P. falciparum populations could have resulted from either a demographic sweep (P. falciparum has only recently spread throughout the world from a small geographically confined population) or a selective sweep (one strain favored by natural selection has recently replaced all others). The selective sweep hypothesis requires that populations of P. falciparum be effectively clonal, despite the obligate sexual stage of the parasite life cycle. A demographic sweep that started several thousand years ago is consistent with worldwide climatic changes ensuing the last glaciation, increased anthropophilia of the mosquito vectors, and the spread of agriculture. P. falciparum may have rapidly spread from its African tropical origins to the tropical and subtropical regions of the world only within the last 6,000 years. The recent origin of the world-wide P. falciparum populations may account for its virulence, as the most malignant of human malarial parasites.

Africa↗

Molecular evolution of daphnia immunity genes: polymorphism in a gram-negative binding protein gene and an alpha-2-macroglobulin gene.

Studies of DNA polymorphism have shown that some immune system genes of mammals and plants are exceptionally diverse, indicating that coevolution between these taxa and their parasites mediates positive selective sweeps and/or balancing selection. The genes of the arthropod immune system remain comparatively unstudied. We isolated two putative immune system genes from the cladoceran crustacean Daphnia and examined DNA sequence diversity. For one gene, encoding a putative gram-negative binding protein, we found evidence of only purifying selection, indicating that this gene is under strong functional constraint and that selection acts to eliminate amino acid variation. For another gene, encoding a putative alpha-2-macroglobulin, we found evidence of positive selection, indicating the possible involvement of this gene in a host-parasite arms race. We discuss the assumed function of these genes and offer speculation regarding which components of the arthropod immune system might experience diversifying adaptive evolution.

Amino Acid Sequence↗

Understanding the overdispersed molecular clock.

Rates of molecular evolution at some protein-encoding loci are more irregular than expected under a simple neutral model of molecular evolution. This pattern of excessive irregularity in protein substitutions is often called the "overdispersed molecular clock" and is characterized by an index of dispersion, R(T) > 1. Assuming infinite sites, no recombination model of the gene R(T) is given for a general stationary model of molecular evolution. R(T) is shown to be affected by only three things: fluctuations that occur on a very slow time scale, advantageous or deleterious mutations, and interactions between mutations. In the absence of interactions, advantageous mutations are shown to lower R(T); deleterious mutations are shown to raise it. Previously described models for the overdispersed molecular clock are analyzed in terms of this work as are a few very simple new models. A model of deleterious mutations is shown to be sufficient to explain the observed values of R(T). Our current best estimates of R(T) suggest that either most mutations are deleterious or some key population parameter changes on a very slow time scale. No other interpretations seem plausible. Finally, a comment is made on how R(T) might be used to distinguish selective sweeps from background selection.

Alleles↗

The influence of linkage and inbreeding on patterns of nucleotide sequence diversity at duplicate alcohol dehydrogenase loci in wild barley (Hordeum vulgare ssp. spontaneum).

Patterns of nucleotide sequence diversity are analyzed for three duplicate alcohol dehydrogenase loci (adh1-adh3) within a species-wide sample of 25 accessions of wild barley (Hordeum vulgare ssp. spontaneum). The adh1 and adh2 loci are tightly linked (recombination fraction <0.01) while the adh3 locus is inherited independently. Wild barley is predominantly self-fertilizing (approximately 98%), and as a consequence, effective recombination is restricted by the extreme reduction in heterozygosity. Large reductions in effective recombination, in turn, widen the conditions for linkage to influence nucleotide sequence diversity through the action of selective sweeps or background selection. These considerations would appear to predict (1) homogeneity in patterns of nucleotide sequence diversity, especially between closely linked loci, and (2) extensive linkage disequilibrium relative to random-mating species. In contrast to these expectations, the wild barley data reveal heterogeneity in patterns of nucleotide sequence diversity and levels of linkage disequilibrium that are indistinguishable from those observed at adh1 in maize, an outbreeding grass species.

Alcohol Dehydrogenase↗