Search PubMedSearch

SEARCH · Search PubMed

Results for “Population structure metrics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

12 recordsLinked to original sources

Genetic differentiation of the supralittoral gastropod Tectarius striatus (North Atlantic Archipelagos) and development of new microsatellite resources.

Microsatellite markers are invaluable tools for assessing genetic diversity and elucidating population structure across any species. This study reports the development and application of ten novel polymorphic microsatellite loci for Tectarius striatus, a littorinid species native to the shores of Macaronesia, a geographical region that includes the archipelagos of the Azores, Madeira, Selvagens, Canary Islands, and Cabo Verde. These markers together with a portion of the COI gene were used to genotype 65 individuals, using Illumina amplicon -sequencing across five geographically distinct populations. Our analysis shows moderate to high levels of allelic diversity across all populations. Furthermore, microsatellite markers supported genetic structure between a distant population of Cabo Verde Archipelago and the northern Macaronesian archipelagos. Conversely variation of the COI showed high levels of homogeneity across the sampled populations. While the presence of null alleles and moderate levels of missing data at several loci represent challenges to this study, the overall consistency of our results with earlier research underscores the reliability of microsatellite markers for population genetic inference in this marine gastropod. Nevertheless, our findings highlight the need for cautious interpretation of diversity estimates and population structure metrics, particularly when null alleles are frequent, and underscore the value of expanding the panel of available microsatellite markers for Tectarius striatus and related taxa to improve resolution and accuracy for future studies.

Animals

Genomic diversity, inbreeding, and selection signatures in duroc, landrace, and yorkshire pigs from a long-term closed breeding system.

Duroc (DD), Landrace (LL), and Yorkshire (YY) are among the most widely used commercial pig breeds, having undergone intense long-term selection within closed breeding systems. This study presents a comprehensive genomic analysis of genetic diversity, inbreeding patterns, and selection signatures in DD, LL, and YY populations that have been subject to close breeding for over 15 years. Genomic and pedigree data were available for 1,088 animals (DD = 348, LL = 276, YY = 464), genotyped using the GenoBaits® Porcine 100 K SNP panel. Principal component analysis and genetic diversity metrics revealed distinct population structures among the three breeds. Pairwise genetic differentiation supported this pattern, with DD showing the greatest divergence from LL (0.34 ± 0.24) and YY (0.33 ± 0.24), while LL and YY were more closely related (FST = 0.22 ± 0.19). Linkage disequilibrium (LD) analysis further confirmed these differences, as DD exhibited the highest average r² (0.34), followed by LL (0.28) and YY (0.25). Within-breed genetic diversity metrics, including observed heterozygosity (HO: 0.37 in DD, 0.39 in LL, 0.38 in YY), expected heterozygosity (HE: 0.36 in DD, 0.37 in LL, 0.38 in YY), and minor allele frequency (MAF: 0.27 in DD, 0.28 in LL, 0.29 in YY), indicated greater genetic variability in LL and YY compared to DD. Runs of homozygosity (ROH) analyses revealed different patterns of autozygosity, with DD exhibiting more long ROH indicative of recent inbreeding, while YY harbored a higher number of short ROH, suggestive of more ancient demographic events. ROH-based inbreeding coefficients (FROH) consistently exceeded pedigree-based estimates (FPED) across all breeds, highlighting the presence of recent or unrecorded inbreeding that pedigree data may not fully capture. According to Generation Proxy Selection Mapping (GPSM), 17, 1, and 12 significant SNPs were detected in DD, LL, and YY, respectively. Functional annotation of ROH islands and GPSM-significant loci revealed both breed-specific and overlapping QTLs related to traits such as growth, reproduction, and carcass. In general, the findings of this study contribute to a deeper understanding of the genomic consequences of long-term closed breeding and provide reference information to support consideration of breeding strategies that balance continued selection for productivity with the maintenance of genetic diversity in modern commercial pig populations.

Animals

PRISM-G: an interpretable privacy scoring framework for assessing risk in synthetic human genome data.

MOTIVATION: Synthetic genomic data promises broader data access, but unresolved privacy risks remain a major concern. Existing evaluations often rely on similarity-based metrics that measure proximity between real and synthetic genomes, overlooking additional mechanisms through which genomic information may leak. RESULTS: We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genomic data across three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure through rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. By pairing PRISM-G with downstream utility metrics, the framework also enables analysis of privacy-utility trade-offs across generative models. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT solver (Genomator). Our results show that privacy vulnerabilities arise along different axes across models and marker densities, demonstrating that a single similarity-based metric is insufficient to characterize genomic privacy risk. AVAILABILITY AND IMPLEMENTATION: The source code of PRISM-G is available at https://github.com/alejocrojo09/prismg.

Humans

COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs.

Pangenome graphs capture extensive structural diversity, but resolving complex loci from shallow sequencing remains challenging, particularly when samples are of low quality such as in ancient DNA. We introduce COSIGT (COsine SImilarity-based GenoTyper), which assigns diploid genotypes by matching read-depth distributions to haplotype paths via cosine similarity. Because this metric evaluates relative coverage profiles rather than absolute read counts, COSIGT substantially outperforms existing likelihood-based tools at low coverage (1-2X). We demonstrate scalability to thousands of modern and ancient genomes, enabling robust, population-scale analyses of complex variation directly from low-coverage datasets.

Humans

Complement component C4 and neuroimaging in psychiatry: A systematic review.

INTRODUCTION: Genomic, transcriptomic, and proteomic studies suggest that the complement system contributes to the pathophysiology of various psychiatric disorders partly through neurodevelopmental effects linked to C4A protein levels variations. We conducted a systematic review to characterize how brain micro- and macrostructure and connectivity vary with proxies of in vivo brain C4A protein levels in both psychiatric and general-population cohorts. METHODS: We used Medline, Web of Science, and Embase, and included all studies published before April 14, 2025. Inclusion criteria were: (1) inclusion of healthy controls and/or individuals with psychiatric disorders assessed according to recognized diagnostic manuals (DSM or ICD); (2) use of MRI-based neuroimaging; and (3) use of genomic, transcriptomic and/or proteomic approaches as proxies of in vivo brain C4A proteins levels. RESULTS: From 317 identified articles, 11 were included. Associations between C4A levels and brain structure were heterogeneous across regions. Only the mOFC, dlPFC, and entorhinal cortex were implicated in more than one study. Findings for the mOFC and dlPFC varied by the type of metrics and clinical status, whereas higher C4A levels were more consistently associated with smaller entorhinal cortex surface area and cortical thickness in pediatric, middle-aged, and older general-population cohorts. In addition, one study found higher genetically predicted C4A expression to be associated with higher TSPO levels. CONCLUSION: The limited number of available studies and their methodological heterogeneity make synthesis challenging. However, biological hypotheses such as excessive synaptic pruning or broader inflammatory effects on the brain may provide plausible explanatory frameworks for the reported associations.

Humans

Morphological characterization, genetic diversity and population structure of the rice blast pathogen Magnaporthe oryzae in Northeast India.

The blast pathogen, Magnaporthe oryzae, is one of the most destructive fungal pathogens of rice worldwide, yet its morphological features, genetic diversity and population structure in Northeast India remain poorly understood. In this study, twenty‒two M. oryzae isolates collected from eight states of Northeast India were characterized using morphological, molecular, and population genetic analyses. Morphological characterization revealed whitish to greyish‒white mycelia with sparse sporulation and colony diameters ranged from 36 to 90 mm, classifying the isolates into 14 fast and 8 slow‒growing groups. Whole genome sequencing was performed to enable both ITS‒based identification and SSR locus mining from the assembled genomes. Molecular identification using ITS rDNA sequences confirmed all isolates as M. oryzae, with 95.5-100% similarity. Phylogenetic analysis grouped the isolates into two major clades and identified seven ITS sequence types (GenBank Accessions: PX273287-PX273293). Genetic diversity assessed using 30 SSR markers revealed substantial polymorphism, with 1-7 alleles per locus and polymorphism information content (PIC) values ranging from 0.00 to 0.81. Heatmap clustering, dendrogram analysis, and distance metrics consistently identified two major genetic groups, with some isolates forming nearly identical clusters and others showing moderate divergence. Principal Component Analysis (PCA) and Principal Coordinates Analysis (PCoA) accounted for 87.8% of the total variance (PC1 and PC2 accounted for 54.4% and 33.4% respectively of the total variance) and revealed distinct outliers. Analysis of Molecular Variance (AMOVA) attributed 80% of the total genetic variation to differences among populations while only 20% was attributed to within population differences highlighting significant inter‒population divergence and clonal population structure. The study revealed substantial morphological and genetic diversity among M. oryzae populations in Northeast India, underscoring the need for region‒specific disease management strategies.

India

The value of structural variants to conservation genomics in the pangenome era.

Structural variants (SVs) comprise an axis of genetic diversity with strong consequences for phenotype and fitness, making them a potentially important target for conservation genomics. Here, we review how and why SVs can play a role in conservation genomics; the different types of SVs and how they can affect phenotype; and how pangenomes and long-read sequencing are illuminating their evolution in populations, including small populations and those of conservation concern. SVs comprise multinucleotide mutations including insertions, deletions, transpositions, inversions, and other multinucleotide mutations, often overlapping genes and other functional genome regions. As a result, SVs often play important roles in phenotypic evolution and local adaptation and can contribute substantially to genetic load in inbred populations. However, our understanding of the factors influencing SV diversity in populations is still in its infancy and is complicated by the vast range of sizes, effects, and mechanisms of formation of these mutations. We argue that SVs are an important axis of genetic diversity which should be characterized alongside more traditional metrics of genetic diversity in conservation contexts. There are a number of analytical challenges to detecting and studying SVs, but analyses aimed at understanding the role of SVs in inbreeding load and population health are rapidly becoming realizable goals, accelerated by new technologies and analytical approaches. New tools, including population-scale long-read sequencing and pangenome approaches, are beginning to make SVs accessible in ways which can be readily applied in conservation settings.

Genomic Structural Variation

Genetic diversity of Plasmodium falciparum helical interspersed subtelomeric (phistb) gene in Tanzania and neighboring countries.

BACKGROUND: Lysine-rich membrane associated Plasmodium helical interspersed subtelomeric gene (phistb) is a member of the phist family of genes which encodes exported proteins essential for the parasite's survival within infected red blood cells. Recent studies suggest the phistb gene as a promising malaria vaccine candidate, however, its genetic diversity remains understudied. This study assessed the genetic diversity of the phistb gene in regions of varying malaria transmission aiming to generate data and improve our understanding of this promising malaria vaccine candidate gene. METHODS: Genomic data from 1472 Plasmodium falciparum samples from Tanzania, Kenya, Uganda, and Ethiopia were retrieved in variant Calling file format (VCF) format from the MalariaGEN Pf7 database. Variants were filtered to include only biallelic Single Nucleotide Polymorphism (SNPs) with Variant Quality Score Log- Odds (VQSLOD)&#x2009;>&#x2009;1 and "PASS" status. Genetic diversity, differentiation, and selection signatures were analyzed using population genetics metrics. RESULTS: After filtering, 1312 samples were retained. Wright's inbreeding coefficient (Fws) showed that 875 (66.7%) samples had monoclonal infections, with the highest proportion of monoclonal infections in Ethiopia (95.3%), followed by Tanzania (67.2%), Kenya (65.7%), and Uganda (50%). Among the 875 monoclonal samples, 88 haplotypes were identified, with Hap_1 (renamed PF3D7)&#xa0;and Hap_13 comprising 37.9 and 21.5 of the samples, respectively. Nucleotide and haplotype diversity were relatively higher in Kenya with 0.097, and 0.88 respectively, compared to the other study populations. The overall fixation index (Fst) was&#x2009;<&#x2009;0.05, and Principal Component Analysis revealed no clear population sub-structure among countries. Negative Tajima's D values in Tanzania, Kenya, and Ethiopia indicated an excess of low-frequency alleles. CONCLUSION: This study reports low genetic diversity of the phistb gene in the four countries despite varying malaria transmission intensities among them, thus making it a suitable candidate gene for malaria vaccine. Further studies should be conducted to assess individual antibodies recognition of the phistb variants and the ability to elicit cross reactivity to further support its potential as a vaccine candidate.

Plasmodium falciparum

Mitochondrial DNA control-region and coding-region data highlight geographically structured diversity and post-domestication population dynamics in worldwide donkeys.

Donkeys (Equus asinus) have been used extensively in agriculture and transportations since their domestication, ca. 5000-7000 years ago, but the increased mechanization of the last century has largely spoiled their role as burden animals, particularly in developed countries. Consequently, donkey breeds and population sizes have been declining for decades, and the diversity contributed by autochthonous gene pools has been eroded. Here, we examined coding-region data extracted from 164 complete mitogenomes and 1392 donkey mitochondrial DNA (mtDNA) control-region sequences to (i) assess worldwide diversity, (ii) evaluate geographical patterns of variation, and (iii) provide a new nomenclature of mtDNA haplogroups. The topology of the Maximum Parsimony tree confirmed the two previously identified major clades, i.e. Clades 1 and 2, but also highlighted the occurrence of a deep-diverging lineage within Clade 2 that left a marginal trace in modern donkeys. Thanks to the identification of stable and highly diagnostic coding-region mutational motifs, the two lineages were renamed as haplogroup A and haplogroup B, respectively, to harmonize clade nomenclature with the standard currently adopted for other livestock species. Control-region diversity and population expansion metrics varied considerably between geographical areas but confirmed North-eastern Africa as the likely domestication center. The patterns of geographical distribution of variation analyzed through phylogenetic networks and AMOVA confirmed the co-occurrence of both haplogroups in all sampled populations, while differences at the regional level point to the joint effects of demography, past human migrations and trade following the spread of donkeys out of the domestication center. Despite the strong decline that donkey populations have undergone for decades in many areas of the world, the sizeable mtDNA variability we scored, and the possible identification of a new early radiating lineage further stress the need for an extensive and large-scale characterization of donkey nuclear genome diversity to identify hotspots of variation and aid the conservation of local breeds worldwide.

Animals

Multimodal Deep Learning and Foundation Models for Early Detection and Forecasting of Plant Diseases.

Plant diseases destroy 20-40% of global food production annually, posing a critical threat to food security for a projected population of 9.7 billion by 2050. Conventional diagnostic approaches relying on expert visual assessment are slow, costly, and unsuitable for modern agricultural scales. While deep convolutional neural networks demonstrated early promise, single-modality, image-centric systems consistently fail under real-world field conditions characterized by variable lighting, co-occurring infections, and cultivar diversity. This review synthesizes a decade of progress across four interconnected frontiers: the evolution of deep learning architectures for plant disease detection; the adaptation of foundation models including CLIP, SAM, and DINOv2 to agricultural contexts; the development of multimodal fusion frameworks integrating imagery, environmental, genomic, and hyperspectral data; and the transition from static disease diagnosis to descriptive comparison of reported metrics, which suggested that multimodal approaches frequently reported improved diagnostic performance relative to corresponding single-modality baselines, although direct cross-study comparison was limited by methodological heterogeneity. A systematic review following PRISMA guidelines identifies eligible comparative studies. Descriptive comparison of reported performance metrics across these studies indicated that multimodal approaches generally achieved higher accuracy and sensitivity than single-modality models, particularly for pre-symptomatic disease detection. Eight critical research gaps are identified, including the absence of a unified agricultural foundation model and limited climate-aware forecasting under non-stationary climate projections. A structured research agenda is proposed to accelerate translation from laboratory performance to globally equitable, field-deployable crop protection systems.

convolutional neural networks

Age and gender profiles of HIV infection burden and viraemia: novel metrics for HIV epidemic control in African populations with high antiretroviral therapy coverage.

INTRODUCTION: To prioritize and tailor interventions for ending AIDS by 2030 in Africa, it is important to characterize the population groups in which HIV viraemia is concentrating. METHODS: We analysed HIV testing and viral load data collected between 2013-2019 from the open, population-based Rakai Community Cohort Study (RCCS) in Uganda, to estimate HIV seroprevalence and population viral suppression over time by gender, one-year age bands and residence in inland and fishing communities. All estimates were standardized to the underlying source population using census data. We then assessed 95-95-95 targets in their ability to identify the populations in which viraemia concentrates. RESULTS: Following the implementation of Universal Test and Treat, the proportion of individuals with viraemia decreased from 4.9% (4.6%-5.3%) in 2013 to 1.9% (1.7%-2.2%) in 2019 in inland communities and from 19.1% (18.0%-20.4%) in 2013 to 4.7% (4.0%-5.5%) in 2019 in fishing communities. Viraemia did not concentrate in the age and gender groups furthest from achieving 95-95-95 targets. Instead, in both inland and fishing communities, women aged 25-29 and men aged 30-34 were the 5-year age groups that contributed most to population-level viraemia in 2019, despite these groups being close to or had already achieved 95-95-95 targets. CONCLUSIONS: The 95-95-95 targets provide a useful benchmark for monitoring progress towards HIV epidemic control, but do not contextualize underlying population structures and so may direct interventions towards groups that represent a marginal fraction of the population with viraemia.

Universal Test and Treat

Global inequities in hepatitis B and C genomic surveillance revealed through an interactive data integration dashboard.

OBJECTIVES: To assess global disparities in hepatitis B virus (HBV) and hepatitis C virus (HCV) genomic surveillance and to develop an integrated platform that links genomic data with epidemiological burden. STUDY DESIGN: Retrospective observational analysis. METHODS: We reviewed existing viral genomic repositories to identify structural and analytical limitations. Subsequently, we integrated 10&#xa0;996 HBV and 3533 HCV whole-genome sequences (WGS) from public databases with Global Burden of Disease (GBD) estimates to quantify inequities in genomic surveillance across countries and genotypes. Using these data, we developed the open-access Hepatitis Dashboard, incorporating >14&#xa0;000 sequences from 141 countries with GBD metrics to evaluate representativeness and sequencing coverage relative to disease burden. RESULTS: Marked inequities in hepatitis genomic surveillance were identified. Despite increasing HBV- and HCV-associated mortality, virus sequence availability remains geographically and genotypically skewed-dominated by China and the United States, with substantial underrepresentation of HBV genotype E and HCV genotypes 5 and 8. Many high-endemic countries in Africa and the Western Pacific remain severely undersampled. We detected circulating antiviral drug-resistance mutations and developed a burden-adjusted sequencing coverage metric, revealing that several high-burden countries, including China, Nigeria and India, are among the least represented in global genomic datasets. Projections to 2030 indicate that neither HBV nor HCV are currently on track to meet WHO elimination targets. CONCLUSIONS: The Hepatitis Dashboard provides an integrated, continuously updated resource that links genomic and epidemiological data to quantify and visualise global surveillance gaps. This analysis highlights a critical disconnect between sequencing efforts and public health needs, which may limit the effectiveness of surveillance-informed strategies to support progress toward WHO 2030 elimination goals. By enabling burden-adjusted prioritisation and longitudinal tracking of genomic coverage, the platform supports evidence-based sampling strategies, equitable resource allocation, and monitoring of global progress toward hepatitis elimination.

Humans