Search PubMedSearch

SEARCH · Search PubMed

Results for “population genetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

plinkQC: an integrated tool for ancestry inference, sample selection, and quality control in population genetics.

MOTIVATION: Population genetic analyses rely on high quality datasets that pass rigorous controls for sample and marker quality. Many analyses also require additional processing including identification of ancestry and sample relatedness. A software package that addresses all these common, yet crucial tasks is missing. RESULTS: We have developed plinkQC, an R/CRAN package that combines these functionalities into a single software package with detailed vignettes for example applications. plinkQC determines the ancestry of study samples via a pre-trained random forest classifier that reaches 98% performance accuracy with just 5% of marker overlap between reference and user data. To obtain the maximal set of unrelated study samples, we developed a graph-based pruning method, taking both relationship estimates and sample quality into account. We demonstrate optimal sample selection on the 1000 Genomes project, where we retain an additional 71 samples compared to publicly available exclusion lists. Finally, plinkQC bundles these results together with per-individual and per-marker quality control checks into three simple functions and returns both the quality controlled dataset and quality control report about each step of the analysis. AVAILABILITY AND IMPLEMENTATION: plinkQC is available as an R/CRAN package. The documentation and code are available on github: https://meyer-lab-cshl.github.io/plinkQC/ and https://github.com/meyer-lab-cshl/plinkQC_manuscript.

Software

Differences in structural color and population genetic structure of Western and Central Palearctic Polyommatus icarus populations.

The blue structural coloration of male Polyommatus icarus butterflies functions as a sexual signaling trait and exhibits remarkable spectral stability within populations despite being generated by highly complex photonic nanoarchitectures. The correlation of the blue sexual signaling color and population genetic variation of the butterflies was investigated across the Western and Central Palearctic regions. Dorsal wing reflectance spectra was measured for 95 male specimens and compared with the population genetic structure revealed in 99 specimens by 18 recently developed microsatellites. Reflectance measurements indicated a clear separation between the European and Central Asian populations, consistent with our previous findings, while the intermediate populations near the Ural Mountains exhibited distinct European spectral characteristics. In contrast, genetic variation showed limited structuring and correlated primarily with geographic distance, as indicated by a significant isolation-by-distance pattern. Thus, although both reflectance and genetic variations are geographically structured, spectral properties are only weakly correlated with genetic differentiation. Populations near the Ural Mountains exhibited genetic ancestry linked to Central Palearctic groups, while displaying distinct Western Palearctic coloration, suggesting that the focal species' sexual signaling is strongly influenced by local factors. These findings suggest that sexual signaling coloration may evolve at least partially independently of the neutral genetic background, offering additional insight into evolutionary divergence across broad geographic scales.

Animals

Genetics of Latin American Diversity Project: Insights into population genetics and association studies in admixed groups in the Americas.

Latin Americans are underrepresented in genetic studies, increasing disparities in personalized genomic medicine. Despite available genetic data from thousands of Latin Americans, accessing and navigating the bureaucratic hurdles for consent or access remains challenging. To address this, we introduce the Genetics of Latin American Diversity (GLAD) Project, compiling genome-wide information from 53,738 Latin Americans across 39 studies representing 46 geographical regions. Through GLAD, we identified heterogeneous ancestry composition and recent gene flow across the Americas. Additionally, we developed GLAD-match, a simulated annealing-based algorithm, to match the genetic background of external samples to our database, sharing summary statistics (i.e., allele and haplotype frequencies) without transferring individual-level genotypes. Finally, we demonstrate the potential of GLAD as a critical resource for evaluating statistical genetic software in the presence of admixture. By providing this resource, we promote genomic research in Latin Americans and contribute to the promises of personalized medicine to more people.

Humans

A guide to understanding tumour evolution through the lens of population genetics.

Every cancer carries the history of its own evolution, hidden in its genome. Modern DNA sequencing can catalogue millions of mutations and profile tumours across space and time, but sequencing alone struggles to answer the questions that matter most: when did key adaptations emerge, how strongly were they selected, why do some tumours relapse whereas others do not, and how will the cancer evolve next? The reason is fundamental: sequencing is a snapshot, whereas evolution is a dynamic process. Bridging this gap requires moving beyond descriptive cancer genomics towards quantitative evolutionary inference. In this Review, we argue that population genetics provides the mathematical framework needed to extract evolutionary dynamics from cancer genomes. We show how models of mutation, selection and drift transform allele frequencies from descriptive measurements into quantitative estimates of clonal fitness and evolutionary timings. We discuss how these principles extend to epigenetic inheritance, plasticity and ecological interactions within the tumour ecosystem, and examine the assumptions and limitations for their application to modern sequencing data. By reframing cancer genomes as quantitative records of evolutionary processes rather than catalogues of mutations, researchers have used population genetics to provide a foundation for understanding - and ultimately predicting - the trajectories of cancer evolution.

Journal Article

Sugar kelp (Saccharina latissima) population genetics map onto geographic distance and oceanographic features across coastal Maine.

Sugar kelp (Saccharina latissima; order Laminariales) plays a vital role in kelp forest ecosystems, as well as an expanding kelp aquaculture industry, in the Gulf of Maine, United States. However, ocean warming is eroding the resilience of Maine's kelp forests and may be compromising their local genetic diversity, with impacts on population structure and gene flow. Here, we used genome-wide single nucleotide polymorphism (SNP) data to assess the genetic diversity, structure, and connectivity of S. latissima populations at 11 outer coastal sites spanning the historical range of kelp forests in Maine. Our analyses identified moderate genetic diversity and limited inbreeding within sites (average heterozygosity: 0.27). Further, they revealed that three clusters comprising four genetically distinct populations exist across the study region. Population structure was strongly associated with geographic distance and oceanographic features, as supported by principal coordinate analysis, FST calculations, Bayesian clustering, and spore dispersal modeling. Lastly, our outlier analysis identified genes potentially under selection. Thus, our findings highlight distinct, genetically unique kelp populations along Maine's coast and emphasize the need for regional management strategies that support both ecosystem resilience and sustainable aquaculture under climate change.

Gulf of Maine

Genome sequencing and population genetics provide insights into local adaptation of Opisthopappus species on cliff environments of Taihang Mountains.

Local adaptation represents a pivotal theme in evolutionary biology. The Opisthopappus genus, comprising Opisthopappus longilobus and O. taihangensis, thrives on the cliffs of the Taihang Mountains. During their evolutionary history, two species are hypothesized to have locally adapted to their cliff habitats. In the present study, we employed a combined approach of whole-genome sequencing of O. taihangensis and population genomic analysis from both species to gain deeper insights into their patterns of local adaptation. Our results revealed that the expansive genome of O. taihangensis (3010.18 Mb), a consequence of a whole-genome duplication (WGD) event, coupled with a high proportion of repetitive sequences (82.70%), was postulated as one of its adaptive strategies. A clear differentiation between O. taihangensis and O. longilobus was observed, with the two species diverging approximately 17.57 million years ago (Mya), with O. longilobus serving as the ancestor. Since their divergence, limited gene flow was observed between the two species. Post-divergence, the effective population sizes of both species expanded, yet underwent a dramatic reduction at approximately 0.07 Mya. Furthermore, a total of 798 adaptive genes were identified, of which 207 overlapped with expanded genes, and eight genes were found to be under positive selection. These genes primarily regulated the growth and development of both species via pathways such as oxidation-reduction and ubiquitin-proteasome, enabling them to withstand climate changes. These findings provide profound insights into the local adaptation of Opisthopappus species to the cliff environments and offer valuable clues for further exploring the local adaptation among various cliff-dwelling organisms.

Adaptation, Physiological

Influence of recombination and niche separation on the population genetic structure of the pathogen Streptococcus pyogenes.

The throat and skin of the human host are the principal reservoirs for the bacterial pathogen Streptococcus pyogenes. The emm locus encodes structurally heterogeneous surface fibrils that play numerous roles in virulence, depending on the strain. Isolates harboring the emm pattern A-C marker exhibit a strong tendency to cause throat infection, whereas emm pattern D strains are usually recovered from impetigo lesions; as a group, emm pattern E organisms fail to display obvious tissue tropisms. The peak incidence for streptococcal pharyngitis and impetigo varies with season and locale, leading to wide spatial and temporal distances between throat and skin strains. To assess any impact of niche separation on genetic variation, the extent of recombinational exchange between emm pattern A-C, D, and E subpopulations was evaluated. Analysis of nucleotide sequence data for internal portions of seven housekeeping loci from 212 isolates provides evidence of extensive recombination between strains belonging to different emm pattern subpopulations. Furthermore, no fixed nucleotide differences were found between emm pattern A-C and D strains. Thus, despite some niche separation created by distinct epidemiological trends and innate tissue tropisms there is little evidence for neutral gene divergence between throat and skin strains. Maintenance of a relationship between emm pattern and tissue tropism in the face of underlying recombination suggests that tissue tropism is associated with emm or a closely linked gene.

Alleles

Genetic analysis in African ancestry populations reveals genetic contributors to lung cancer susceptibility.

Striking disparities in lung cancer exist, with Black/African American individuals disproportionately affected by lung cancer, yet the genetic architecture in African ancestry individuals is poorly understood. We aimed to address this by performing a comprehensive genetic association study of lung cancer, incorporating local ancestry, across 6,490 African ancestry individuals (2,390 individuals with lung cancer and 4,100 control subjects). We identified a single genome-wide significant (p < 5 &#xd7; 10-8) locus, 15q25.1 (lead SNP rs17486278, OR [95% CI] = 1.34 [1.23-1.45], p = 4.52 &#xd7; 10-12), that has consistently shown a strong association with lung cancer across populations. Additionally, we identified nine suggestive (p < 1 &#xd7; 10-6) loci. Four of these loci (3p12.1, 8q22.2, 14q11.2, and 18q22.3) have no prior reported associations with lung cancer. We performed a multi-ancestry lung cancer meta-analysis using prior large-scale summary statistics from European and Asian ancestry populations, incorporating our African ancestry results. The meta-analysis identified 17 genome-wide significant loci, including an association with locus 4q35.2 (p = 1.22 &#xd7; 10-8), a genomic region that has been previously linked to forced expiratory volume. Genome-wide SNP-based heritability for lung cancer was 16% among African ancestry individuals. Follow-up in silico functional analyses identified genetically regulated gene expression (GReX) of nine genes (AC012184.3, ADK, CCDC12, CHRNA3, EML4, PSMA4, SNRNP200, TMEM50A, and ZYG11A) associated with lung cancer risk and biological pathways relevant to cancer and lung function. Cumulatively, these findings further elucidate the genetic architecture of lung cancer in African ancestry individuals, confirming prior loci and revealing new loci.

Female

Genetic investigation of population structure in Atlantic chub mackerel, Scomber colias Gmelin, 1789 along the West African coast.

Sustainable management of transboundary fish stocks hinges on accurate delineation of population structure. Genetic analysis offers a powerful tool to identify potential subpopulations within a seemingly homogenous stock, facilitating the development of effective, coordinated management strategies across international borders. Along the West African coast, the Atlantic chub mackerel (Scomber colias) is a commercially important and ecologically significant species, yet little is known about its genetic population structure and connectivity. Currently, the stock is managed as a single unit in West African waters despite new research suggesting morphological and adaptive differences. Here, eight microsatellite loci were genotyped on 1,169 individuals distributed across 33 sampling sites from Morocco (27.39&#xb0;N) to Namibia (22.21&#xb0;S). Bayesian clustering analysis depicts one homogeneous population across the studied area with null overall differentiation (F ST = 0.0001ns), which suggests panmixia and aligns with the migratory potential of this species. This finding has significant implications for the effective conservation and management of S. colias within a wide scope of its distribution across West African waters from the South of Morocco to the North-Centre of Namibia and underscores the need for increased regional cooperation in fisheries management and conservation.

Animals

Comparing ARG Inference Methods Under Transmission of Reproductive Success: Tree Imbalance Matters.

Inferring coalescent trees from genomic data has become a major subject in population genetics, particularly with the recent advances in tree sequence reconstruction methods. However, it remains unclear how well these methods perform for imbalanced genealogies. Such imbalances can arise from processes such as cultural transmission of reproductive success (CTRS) or positive selection. Using simulated genomic data, we benchmarked three major software packages, SINGER, Relate, and tsinfer, by comparing the imbalance of reconstructed trees by these methods with that of the true simulated trees, for three indices that quantify this imbalance. The three methods performed well under scenarios yielding balanced trees. However, their accuracy declined as imbalance increased. Performances also varied with mutation rate, recombination rate, and sample size. This study opens possibilities for applying these methods to infer CTRS or positive selection in large-scale genomic datasets, using simulation-based inference such as approximate Bayesian computation.

Models, Genetic

Recent Adaptation in a Threatened Salmonid Revealed by Museum Genomics.

Steelhead/rainbow trout (Oncorhynchus mykiss) is an imperilled salmonid with two main life history strategies: migrate to the ocean or remain in freshwater. Domesticated hatchery forms of this species have been stocked into almost all California waterways, possibly resulting in introgression into natural populations and altered population structure. We compared whole-genome sequence data from contemporary populations against a set of museum population samples of steelhead from the same locations that were collected prior to most hatchery stocking. We observed minimal introgression and few steelhead-hatchery trout hybrids despite a century of extensive stocking. Our historical data show signals of introgression with a sister species and indications of an early hatchery facility. Finally, we found that migration-associated haplotypes have become less frequent over time, a likely adaptation to decreased opportunities for migration. Since contemporary migration-associated haplotype frequencies have been used to guide species management, we consider this to be a rare example of shifting baseline syndrome that has been validated with historical data. We suggest cautious optimism that a century of hatchery stocking has had minimal impact on California steelhead population genetic structure, but we note that continued shifts in life history may lead to further declines in the ocean-going form of the species.

Animals

Whole-Genome Sequencing Reveals Population Structure, Genetic Diversity, and Selection Signatures in Kazakh Dromedary and Bactrian Camels.

Understanding the genomic basis of environmental adaptation is essential for the conservation and genetic improvement of domestic camels. In this study, we investigated the population structure, genetic diversity, and genomic variation potentially associated with environmental adaptation of Kazakh dromedary and Bactrian camels using whole-genome sequencing. Whole-genome sequencing data were generated for Kazakh camels (15 dromedaries and 16 Bactrian camels) and integrated with 131 publicly available genomes representing camel populations from the Arabian Peninsula, Iran, Xinjiang, Inner Mongolia, and Mongolian wild camels. Population structure, genetic diversity, and genome-wide selection were evaluated using principal component analysis, ADMIXTURE, nucleotide diversity, linkage disequilibrium, runs of homozygosity, genomic inbreeding (FROH), and selection scans based on FST, &#x3b8;&#x3c0; ratio, and XP-EHH. Population genomic analyses revealed clear differentiation between dromedary and Bactrian camels, whereas Kazakh camel populations exhibited higher nucleotide diversity (&#x3b8;&#x3c0; = 1.307-1.551 &#xd7; 10-3), and lower genomic inbreeding (median FROH: 0.037-0.056) than Arabian populations. Genome-wide selection analyses identified MC4R as the prominent candidate gene in Kazakh dromedaries and RYR1 as a prominent candidate gene in Kazakh Bactrian camels. Functional enrichment analyses highlighted pathways related to energy metabolism, thermogenesis, calcium signaling, skeletal muscle function, mitochondrial activity, and oxidative stress response. These findings provide new insights into genomic variation potentially associated with environmental adaptation in Kazakh camels and offer valuable genomic resources for future conservation, breeding, and evolutionary studies.

MC4R

Ancient Introgression Explains Mitochondrial Genome Capture and Mitonuclear Discordance Among South American Collared Tropidurus Lizards.

Mitonuclear discordance-evolutionary discrepancies between mitochondrial and nuclear DNA phylogenies-can arise from various factors, including introgression, incomplete lineage sorting, recent or ancient demographic fluctuations, sex-biased dispersal asymmetries, among others. Understanding this phenomenon is crucial for accurately reconstructing evolutionary histories, as failing to account for discordance can lead to misinterpretations of species boundaries, phylogenetic relationships, and historical biogeographic patterns. We investigate the evolutionary drivers of mitonuclear discordance in the Tropidurus spinulosus species group, which contains nine species of lizards inhabiting open tropical and subtropical environments in South America. Using a combination of population genetic and phylogenomic approaches applied to mitochondrial and nuclear data, we identified different instances of gene flow that occurred in ancestral lineages of extant species. Our results point to a complex evolutionary history marked by prolonged isolation between species, demographic fluctuations, and potential episodes of secondary contact with genetic admixture. These conditions likely facilitated mitochondrial genome capture while diluting signals of nuclear introgression. Furthermore, we found no strong evidence supporting incomplete lineage sorting or natural selection as primary drivers of the observed mitonuclear discordance. Therefore, the unveiled patterns are most consistent with neutral demographic processes, coupled with ancient mitochondrial introgression, as the main factors underlying the mismatch between nuclear and mitochondrial phylogenies in this system. Future research could further explore the role of other demographic processes, such as asymmetric sex-biased dispersal, in shaping these complex evolutionary patterns.

Animals

Deconstructing empirical fitness seascapes across scales of granularity.

The fitness landscape metaphor remains resonant in evolutionary theory and has facilitated the birth of newer concepts, like the fitness seascape, that consider the role of environmental context in shaping the dynamics of evolution. Since its emergence, the seascape has appeared in numerous studies examining how different and fluctuating environments shape evolutionary outcomes. Despite growing interest, we lack comprehensive examinations of how environmental context shapes features of fitness seascapes. In this study, we address this gap by deconstructing empirical fitness seascapes across scales of granularity: loci, locus interactions (epistasis), alleles, trajectories, and entire seascapes. For each, we examine how environmental context influences qualitative and quantitative aspects of seascapes, and find that they change appreciably, with patterns specific to individual systems of study. We also quantify how much each scale varies across environments, and find that certain scales tend to be more sensitive to context than others. In summary, we reflect on the implications of the seascape metaphor for the incorporation of environmental effects into theoretical population genetics, for understanding how the environment shapes evolution in disease systems, and for contemporary bioengineering efforts.

Genetic Fitness

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals

A Multitrait Locus Regulates Sarbecovirus Pathogenesis.

Infectious diseases have shaped the human population genetic structure, and genetic variation influences the susceptibility to many viral diseases. However, a variety of challenges have made the implementation of traditional human Genome-wide Association Studies (GWAS) approaches to study these infectious outcomes challenging. In contrast, mouse models of infectious diseases provide an experimental control and precision, which facilitates analyses and mechanistic studies of the role of genetic variation on infection. Here we use a genetic mapping cross between two distinct Collaborative Cross mouse strains with respect to severe acute respiratory syndrome coronavirus (SARS-CoV) disease outcomes. We find several loci control differential disease outcome for a variety of traits in the context of SARS-CoV infection. Importantly, we identify a locus on mouse chromosome 9 that shows conserved synteny with a human GWAS locus for SARS-CoV-2 severe disease. We follow-up and confirm a role for this locus, and identify two candidate genes, CCR9 and CXCR6, that both play a key role in regulating the severity of SARS-CoV, SARS-CoV-2, and a distantly related bat sarbecovirus disease outcomes. As such we provide a template for using experimental mouse crosses to identify and characterize multitrait loci that regulate pathogenic infectious outcomes across species. IMPORTANCE Host genetic variation is an important determinant that predicts disease outcomes following infection. In the setting of highly pathogenic coronavirus infections genetic determinants underlying host susceptibility and mortality remain unclear. To elucidate the role of host genetic variation on sarbecovirus pathogenesis and disease outcomes, we utilized the Collaborative Cross (CC) mouse genetic reference population as a model to identify susceptibility alleles to SARS-CoV and SARS-CoV-2 infections. Our findings reveal that a multitrait loci found in chromosome 9 is an important regulator of sarbecovirus pathogenesis in mice. Within this locus, we identified and validated CCR9 and CXCR6 as important regulators of host disease outcomes. Specifically, both CCR9 and CXCR6 are protective against severe SARS-CoV, SARS-CoV-2, and SARS-related HKU3 virus disease in mice. This chromosome 9 multitrait locus may be important to help identify genes that regulate coronavirus disease outcomes in humans.

Animals

Largest-Scale Genomic Resource Reconstructing the Genetic Origin, Population Structure, and Biological Adaptations of the Hui People.

Historical and archaeological records indicate that the Maritime and Land Silk Roads played a pivotal role in facilitating Trans-Eurasian migrations and cultural exchanges. However, the extent to which population movements or the spread of ideas shape Chinese Hui populations remains debated. We present the largest genomic resource to date, including 2,280 Hui individuals sequenced or genotyped from 30 diverse regions, to examine the genetic origins, population structure, and biological adaptations of this underrepresented group in global human genome research. We identified a detailed population structure characterized by five distinct genetic lineages of the Hui, influenced by geography and varying gene flow. The admixture history and demographic events suggest that the northwestern and northern Hui lineages emerged from demic diffusion during the Tang and Yuan Dynasties via the Land Silk Road. In contrast, the southern and island Hui lineages reflect cultural diffusion along the Maritime Silk Road, while the mixed southern-northern lineage likely developed through a combination of demic and cultural diffusion. Our findings support a hybrid model for Hui formation, indicating that both demographic processes and sociocultural transmissions contributed to their population history. We identified east-west highly differentiated variants and pre- and post-admixture adaptations in Hui genomes, demonstrating that admixture-driven adaptive or neutral variants impacted susceptibility to cardiovascular diseases and immune- and diet-related traits. These adaptive signatures include post-admixture signals of SLC24A5 and ECHDC1 in the Hui, as well as pre-admixture signals of the HLA region, BCL2A1, and KCNH8 in the East Asian source. Overall, our study suggests that Han-related genetic components helped the Hui population rapidly adapt to new local environments. Additionally, the frequency spectrum of clinically essential variants differed significantly between Hui and Han individuals, emphasizing the importance of including underrepresented populations in genomic research to promote health equity.

Humans