Search PubMedSearch

SEARCH · Search PubMed

Results for “Genotyping-by-sequencing”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

SNP genotyping in Pseudotsuga menziesii and Pinus radiata using targeted genotyping-by-sequencing (GBS): improved Bayesian SNP calling using a beta-binomial distribution and other optimized input parameters.

BACKGROUND: Single-nucleotide polymorphism markers (SNPs) have important applications in gene conservation, breeding, and fundamental genetics research. Our long-term goal is to develop routine approaches for SNP genotyping in forest trees. Ideally, these approaches would be inexpensive, able to accommodate a wide range of samples and SNPs, available through commercial providers, and produce high-quality SNP data. RESULTS: Using targeted genotyping-by-sequencing (GBS), we developed SNP assays for two highly heterozygous tree species, Douglas-fir (Pseudotsuga menziesii) and radiata pine (Pinus radiata). Using Douglas-fir haploid and diploid data, we optimized Bayesian SNP calling by testing four input parameters: (1) allele and genotype prior probabilities, (2) Rho, the beta-binomial dispersion parameter, (3) estimated read error (BayesReadError), and (4) the logPO cutoff used to filter low confidence SNP calls. logPO is the Bayesian posterior odds ratio for a called SNP. Compared to assuming a binomial distribution of read counts (Rho = 0), the beta-binomial distribution (Rho = 0.33) substantially reduced call error and heterozygote undercalling. Compared to the other Bayesian parameters, genotype priors had little effect on genotyping success. For Douglas-fir, we tested 5,360 SNP assays, and then studied the performance of the best 4,000. For radiata pine, we tested 6,000 SNP assays, and then studied the performance of the best 4,570. In Douglas-fir and radiata pine, our Bayesian approach resulted in median call rates of 95% to 98% for the top-ranked SNPs, with an estimated call error of 1.60% for known homozygous genotypes and 2.27% for known heterozygotes. In radiata pine, median and mean call rates were above 91% for GBS and SNP genotyping using an Axiom fixed genotyping array. Additionally, the median correspondence between the GBS and Axiom genotypes was about 98% overall (mean 96%). CONCLUSIONS: By optimizing Bayesian SNP calling, selecting the best 4-5 K SNPs, and excluding samples with low DNA amounts, we substantially reduced call error and heterozygote undercalling, resulting in SNP genotypes that were nearly identical to genotypes obtained using the Axiom array. Furthermore, genotyping performance should increase even further if our SNP rankings were used to develop less complex probe pools that target fewer SNPs.

Pinus

Exploration of the genetic diversity of Avena Fatua L. (wild oat) through genotyping-by-sequencing and SDS-PAGE.

BACKGROUND: The consumption of oats has rapidly increased due to their exceptional nutritional value. However, concerns over genetic erosion have emerged as oat breeding programs rely on a highly limited genetic pool. This study aimed to expand the genetic diversity pool of oats by collecting wild oat (Avena fatua L.) populations in South Korea and assessing their genetic diversity and seed storage protein patterns. RESULTS: A total of 237 A. fatua individuals were collected in 2022 from eight regions in the southwestern coastal areas of South Korea. Genetic diversity and seed storage protein patterns were analyzed using genotyping-by-sequencing (GBS) and sodium dodecyl sulfate polyacrylamide gel electrophoresis (SDS-PAGE). The GBS analysis identified 20,836 single-nucleotide polymorphisms (SNPs). An analysis of molecular variance (AMOVA) based on regional populations revealed that 40.9% of the genetic variation was attributed to differences among populations, while 59.1% was within populations, indicating high genetic differentiation within regional populations. Subsequent population structure analysis and discriminant analysis of principal components (DAPC) both stated the formation of two distinct genetic groups, with an AMOVA value of 70.9% between the groups, suggesting a high level of genetic variation. Pairwise FST analysis was conducted to compare the genetic differentiation between two populations, revealing that Jindo and Jangheung exhibited the highest level of genetic differentiation (FST = 0.795) among the geographic groups. Seed storage proteins were analyzed using SDS-PAGE, and the patterns were grouped using k-means clustering. A comparison between the groups based on protein patterns and those based on genetic variation revealed no significant correlation. CONCLUSION: This study provides data on the genetic diversity of A. fatua, a wild relative of cultivated oats, aimed at expanding the genetic pool of oats for future breeding programs. These findings are expected to be a foundational resource for oat breeding and genetic improvement efforts.

Genetic Variation

Replacement of chromosome 3D with Thinopyrum chromosome 3St led to increased drought tolerance during the flowering stage in wheat.

The stable 3St(3D) substitution line offers promising genetic potential for improving drought tolerance in wheat during critical reproductive stages. The flowering stage is highly susceptible to drought, which significantly reduces wheat grain yield globally. Low genetic diversity in wheat further limits the discovery of optimal gene variants for breeding climate-resilient varieties. The substitution of chromosome 3D by a group 3 chromosome pair from Thinopyrum intermedium × Th. ponticum artificial hybrid was identified using in situ hybridization and genotyping-by-sequencing. This homoeologous substitution showed good functional compensation for grain yield and fertility, similar to the wheat parents ('Mv9kr1' and 'Mv Karizma') in field and greenhouse trials. The substitution line exhibits a semidwarf phenotype due to the Rht8 and Rht2 dwarfing alleles. Automated shoot phenotyping after a 10-day water withdrawal at flowering revealed efficient water preservation allowing to maintain photosynthetic functions, sustained photosynthetic activity, and less chlorophyll degradation, indicated by Normalized Difference Vegetation Index (NDVI) and modified Normalized Difference Index (mND705) values and moderate level of protective functions shown by the expression of stress-related genes. Compared to the wheat parents, the substitution line developed thicker roots with increased volume under drought, resulting in a lower surface-to-volume ratio. This may enhance water storage efficiency and help reduce yield loss under drought conditions.

Triticum

Breed classification of Lao People's Democratic Republic (Lao PDR) and Thai native chickens using synchrotron radiation-based Fourier transform infrared spectroscopy and genotyping by sequencing.

Lao PDR harbors substantial genetic diversity in native chicken populations, representing an important resource for sustainable production and long-term food security. This study aimed to classify five Lao native chicken breeds-Ou, Black Bone, Horn Chou, Yolk, and Chae-and to discriminate them from a Thai native breed, Leung Hang Khao (LK), using integrative genotype-based approaches. Blood samples were collected from 50 LK and Lao native chickens (32 Ou, 10 Black Bone, 9 Horn Chou, 121 Yolk, and 41 Chae). Genomic DNA was extracted and analyzed using synchrotron radiation-based Fourier-transform infrared (SR-FTIR) spectroscopy to characterize biochemical composition, while genotyping-by-sequencing (GBS) was employed to identify genome-wide single nucleotide polymorphisms (SNPs). SR-FTIR analysis revealed highly significant differences among breeds in nucleotide-associated functional groups, including thymine, adenine, guanine, cytosine, as well as DNA backbone and deoxyribose components (P < 0.001). Multivariate analyses demonstrated that principal component analysis (PCA) of SR-FTIR spectra effectively discriminated chicken breeds, while hierarchical cluster analysis (HCA) further resolved them into two major clusters with distinct sub-clusters, reflecting variation in DNA biochemical composition. In contrast, GBS analysis identified 1484 common SNPs; however, PCA based on SNP data showed limited resolution in clearly separating breeds, despite revealing similar clustering trends. Overall, the results highlight the strong discriminatory power of SR-FTIR spectroscopy for rapid and effective classification of native chicken breeds at the molecular level, outperforming SNP-based differentiation under the current marker density. This study provides novel insights into the application of synchrotron-based spectroscopic techniques in poultry genetics and contributes valuable baseline information for the conservation and utilization of Lao native chicken genetic resources.

Breed classification

Genome-wide association studies reveal genetic variants associated with antineoplastic monoterpenoid indole alkaloid accumulation in Catharanthus roseus.

Catharanthus roseus produces pharmacologically important monoterpenoid indole alkaloids (MIAs), yet their natural accumulation is low, limiting therapeutic exploitation. To dissect the genetic basis of natural variation in MIA accumulation, we integrated phenotypic, chemotypic, and genomic analyses of 93&#xa0;C. roseus accessions sampled from six locations across India, including New Delhi, Lucknow, Jodhpur, Bangalore, and two locations in Gujarat: Navsari and Bardoli. Morphological characterization showed limited differentiation among locations, whereas accessions from Gujarat tended to be taller compared to other locations and more frequently white-flowered. Quantitative HPLC profiling revealed substantial accession- and location-dependent variation in total indole alkaloid levels, with Gujarat accessions showing the highest accumulation, largely driven by vindoline and catharanthine. Genotyping-by-sequencing generated 10,801 high-quality variants comprising 10,087 SNPs and 714 InDels corresponding to an average density of 19.34 variants per Mbp of the genome, revealing three genetic subgroups with overall admixed ancestry and weak geographic stratification. Genome-wide association study (GWAS) using five benchmark models identified 47 variants potentially associated with catharanthine, vindoline, and vinblastine content. These putative candidate loci were located near genes implicated in hormone signaling, mitochondrial function, nitrogen metabolism, and RNA processing, suggesting complex regulatory control of MIA biosynthesis. Notably, two missense variants in a carboxylesterase-like gene were associated with vindoline accumulation, and highly significant intergenic SNP clusters suggested putative regulatory hotspots for vinblastine biosynthesis. These results provide GWAS-based insights into the genetic architecture of MIA metabolism in C. roseus and nominate candidate variants for precision breeding and metabolic engineering to enhance pharmaceutical alkaloid production.

Catharanthus roseus

Genomic selection in timothy (Phleum pratense L.): a comprehensive evaluation of prediction models, multi-trait strategies, and forward validation across Norwegian environments.

This study presents a comprehensive evaluation of genomic selection (GS) in timothy (Phleum pratense L.), comparing nine prediction models across yield and quality traits at two Norwegian locations. Forward validation with independent full-sib (FS2) families revealed a substantial generalization gap, highlighting the need for realistic accuracy assessment in polyploid forage breeding. Timothy (Phleum pratense L.) is the most important forage grass in Northern Europe, yet genomic selection has not been systematically evaluated in this hexaploid species. We assessed 889 FS2-families originating from biparental crosses among 49 cultivars/populations. The FS2-families were genotyped with 30,698 SNP markers derived from genotyping-by-sequencing (GBS) and field tested for three harvest years at a highland and a lowland continental location in Southern Norway. Nine genomic prediction models were compared for six yield traits (dry matter yield per cut and total) and six quality traits (protein, digestibility, and fiber fractions) across three cuts/year. Within-training cross-validation accuracies were moderate to high (mean r = 0.62), with Random Forest and SVR consistently outperforming GBLUP. However, forward validation using 213 independent FS2-families revealed dramatically lower accuracies (mean r = 0.16), with only 16 of 30 trait-dataset combinations reaching statistical significance (p < 0.05). Genomic heritabilities (GREML), estimated across environments, ranged from near zero for the quality traits to 0.55 for the yield traits. Multi-trait models improved accuracy by 3-5% over single-trait approaches, while FS2 families-by-environment interaction models with Random Forest achieved the highest within-training accuracy (mean r = 0.71). Marker density analysis showed accuracy plateauing at approximately 15000 SNPs. Genetic correlations among the yield component traits were estimated by multi-trait REML; correlations among the quality traits could not be estimated reliably because their genomic heritabilities were low. A multi-trait selection index identified top-performing FS2-families for further crossing recommendations. These results provide a benchmark for GS implementation in hexaploid timothy and emphasize that cross-validation substantially overestimates prediction accuracy for truly independent material.

Norway

An Amplicon Panel for High-Throughput and Low-Cost Genotyping of Yesso Scallop Mizuhopecten yessoensis.

The Yesso scallop Mizuhopecten yessoensis was imported from Japan to western Canada in the late 1980s to establish an economically viable scallop aquaculture industry. Since this time, the industry in Canada has operated with existing genetic diversity within the broodstock, which is considerably limited relative to wild populations. The sector has not been able to realise its full potential in part due to idiopathic hatchery failures and farm stock collapses due to disease outbreaks associated with the intracellular bacterial pathogen Francisella halioticida. To support Yesso scallop production and breeding, here we generate a low-density, genotyping-by-sequencing amplicon panel using single nucleotide polymorphism (SNP) markers that are evenly spaced across the M. yessoensis genome and that show high heterozygosity in Canada and Japan. The panel can also exploit the high genetic polymorphism of the M. yessoensis genome, with de novo SNP calling identifying over 2,500 high quality SNPs within the 579 sequenced amplicons. We demonstrate the utility and versatility of this new genotyping tool for breeding applications including parentage assignment, low density family-based genome-wide association study, trait heritability evaluation to determine potential for genomic selection, and species differentiation (against the weathervane scallop Patinopecten caurinus). We did not find any genomic regions significantly associated with F. halioticida resistance but did identify potential for genomic selection. We could separate the two species based on genotypes, and did not see evidence of a past M. yessoensis x P. caurinus hybridization event within the M. yessoensis breeding population at Vancouver Island University. This low-cost genotyping panel is expected to accelerate selective breeding improvements for M. yessoensis in Canada and elsewhere.

Animals

GWAS-based identification of a candidate gene and development of a predictive KASP marker for seed protein and oil contents in soybean.

BACKGROUND: Soybean [Glycine max (L.) Merrill] is one of the most widely cultivated crops worldwide. Its seeds contain about 40% protein and 20% oil, serving as essential nutrient sources for humans. Given the nutritional importance of seed protein and oil, identifying genes that regulate their levels is crucial for improving soybean seed quality. OBJECTIVE: This study aimed to identify genetic factors associated with seed protein and oil content using a genome-wide association study (GWAS). METHODS: Seed protein and oil contents were quantified in 192 soybean mutant accessions in a mutant diversity pool (MDP), and GWAS was conducted using 17,631 SNPs filtered from genotyping-by-sequencing. Expression of a candidate gene was examined across seed developmental stages (R5 to R7), and a significant SNP was converted into a Kompetitive Allele-Specific PCR (KASP) marker for validation. RESULTS: GWAS detected significant SNPs associated with seed protein and oil content. Chr20_7635098 was identified as a nonsynonymous SNP located in the exon of Glyma.20g042400. This gene showed differential expression across seed developmental stages between mutant accessions with contrasting protein and oil contents. The KASP marker for Chr20_7635098 was validated using the MDP and six domestic soybean cultivars showing predictive accuracies of &#x2265;&#x2009;80.50% for protein content and &#x2265;&#x2009;61.18% for oil content. CONCLUSION: Overall, this study identified a candidate gene linked to both seed protein and oil content, providing valuable insights for molecular breeding strategies aimed at efficiently improving these nutritional traits.

Glycine max

A GWAS-derived histone H4 variant linked to ear row number reveals functional insights into the maize ZmHistone gene family.

Ear row number (ERN) is a major yield determinant in maize and a key target for breeding of high-yielding varieties. This study utilized a multi-parent population (MPP) of 780 recombinant inbred lines (RILs) derived from seven inbred lines across three environments. Genotyping-by-sequencing (GBS) of the MPP yielded 638,646 high-quality SNPs. Using genome-wide association study (GWAS), we detected 80 significant SNPs including S2-15316355 and S4-224453431, which were consistently detected in all environments and best linear unbiased prediction (BLUP) analysis. A linkage disequilibrium-defined &#xb1;20&#x202f;kb window around these two lead SNPs contained three positional candidate genes: Zm00001eb072840, Zm00001eb072850 and Zm00001eb202890. Zm00001eb072850 (ZmHistone12), a histone H4 variant, was prioritized for hypothesis-driven follow-up because the lead SNP lies within its coding sequence and the gene is expressed in ear-related tissues. Additionally, we identified 91 ZmHistone genes in the maize genome and described their phylogeny, promoter motif and expression patterns. Public transcriptome and qRT-PCR analysis in seven parental lines provide descriptive evidence of Histone variant genes in maize ear development. These results suggest a potential involvement of chromatin-associated regulation of ERN in maize and provide a foundation for future functional validation.

Ear development

Global genomic population structure of wild and cultivated oat reveals signatures of chromosome rearrangements.

The genus Avena consists of approximately 30 wild and cultivated oat species. Cultivated oat is an important food crop, yet the broader genetic diversity within the Avena gene pool remains underexplored and underexploited. Here, we characterize over 9000 wild and cultivated hexaploid oat accessions of global origin using genotyping-by-sequencing and explore population structure using multidimensional scaling and population-based clustering methods. We also conduct analyses to reveal chromosome regions associated with local adaptation, sometimes resulting from large-scale chromosome rearrangements. We report four distinct genetic populations within the wild species A. sterilis, a distinct population of cultivated A. byzantina, and multiple populations within cultivated A. sativa. Some chromosome regions associated with local adaptation are also associated with confirmed structural rearrangements on chromosomes 1A, 1C, 3C, 4C, and 7D. This work provides evidence suggesting multiple polyploid origins, multiple domestications, and/or reproductive barriers amongst Avena populations caused by differential chromosome structure.

Avena

Role of functional genes for seed vigor related traits through genome-wide association mapping in finger millet (Eleusine coracana L. Gaertn.).

Finger millet (Eleusine coracana (L.) Gaertn.) is a calcium-rich, nutritious and resilient crop that thrives even in harsh environmental conditions. In such ecologies, seed longevity and seedling vigor are crucial for sustainable crop production amid climate change. The current study explores the genetics of accelerated aging on seed longevity traits across 221 diverse accessions of finger millet through genome-wide association approach (GWAS). A significant variation was identified in germination percentage, germination rate indices, mean germination time, seedling vigor indices and dry weight upon aging treatment. GWAS model from 11,832 high-quality SNPs identified through Genotyping-by-Sequencing (GBS) approach produced 491 marker-trait associations (MTAs) for 27 traits, of which 54 were FDR-corrected. A pleiotropic SNP, FM_SNP_9478 identified on chromosome 7B was associated with the traits viz., germination after aging, germination index after aging and their relative measures. Functional annotation revealed DET1 and expansin-A2 influenced seed coat integrity, critical for germination and aging resilience. Probable protein phosphatase 2C3 and piezo-type ion channels contributed to mechanical sensing and stress adaptation in seeds. Beta-amylase and acetyl-CoA carboxylase 2 were identified for seed metabolism and stress response. These insights lay the framework for targeted breeding efforts to improve seed quality and resilience under diverse production conditions.

Eleusine

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria &#xd7; ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria

Scaling up orphan crop research: genebank genetics highlight geographic structure in cultivated cowpea from 10&#x2009;617 global accessions.

Vigna unguiculata (L.) Walp. is a dryland legume crop, providing essential food and nutritional security for millions of people across the semi-arid tropics, in Africa, Asia and Latin America. However, as a typical 'orphan crop', cowpea has long remained underrepresented in global genomic research to support crop improvement. Here, we conducted the largest genetic diversity analysis of cowpea to date, comprising 10&#x2009;617 accessions sourced from seven international collections. Using genotyping-by-sequencing, we characterised the global patterns of genetic diversity, assessed redundancy within and across collections, and examined the geographic structure of the cowpea global allele pool. Our results revealed nine distinct genetic groups with clear geographic associations and fine-scale population differentiation, reflecting dispersal history, regional adaptation and the influence of modern breeding. Duplication across collections was detected, highlighting the need for improved curation and integration of germplasm resources. Landraces from sub-Saharan Africa do not fully capture the genetic diversity present in several other geographic regions, indicating the existence of abundant and untapped genetic resources worldwide. These findings not only provide insights into the genetic structure and evolutionary history of cowpea but also offer a valuable foundation for harnessing global germplasm diversity to enhance breeding potential and accelerate crop improvement.

Vigna

Development of a 10K breeder-friendly SNP chip for faba bean.

INTRODUCTION: Faba bean breeding and genomics have seen steady progress in recent years, supported by genome sequences and high-density genotyping platforms. These tools have been valuable for trait mapping, diversity assessment, and genomic research, but they have limited routine use in breeding programs due to their relatively high cost. Recent progress in establishing an optimized, cost-efficient genotyping-by-sequencing protocol tailored to the large and complex faba bean genome has created the foundation for a more accessible genotyping solution. METHODS: Using this approach, we explored the genetic diversity of faba bean germplasm from various panels, providing a comprehensive representation of the crop's genetic landscape. From this dataset, we identified and selected a high-quality set of informative SNP markers that are evenly distributed across the genome. Building on these resources, we designed a breeder-friendly 10K SNP chip. RESULTS: The 10K SNP chip delivers high accuracy, broad genomic coverage, and affordability. The chip was validated across diverse germplasm panels, demonstrating strong clustering performance, high reproducibility, and applicability to breeding-relevant germplasm. DISCUSSION: This platform offers a cost-effective alternative to higher-density arrays, enabling its integration into genomic selection, marker-assisted breeding, and diversity monitoring, ultimately supporting accelerated genetic gain and the delivery of improved varieties to farmers.

SNP chip

Genomic early growth mechanisms of two endangered Mexican spruces.

This study elucidated the genomic basis of family-level growth variance in the critically endangered endemic Mexican spruces Picea martinezii and P. mexicana by: (i) analyzing family- and population-level variations in seedling basal diameter and height after 12 months of growth under common garden conditions and seed weight as maternal provisioning trait; and (ii) identifying genomic loci (SNPs) associated with these traits. Despite limited sample sizes (77 and 74 families representing all known populations of both species), 32 and 10 outlier SNPs were identified yielding 17 and six annotated candidate genes in P. martinezii and P. mexicana, respectively. These genes showed contrasting multivariate associations suggesting species-specific hypothesized growth strategies at the family level: defense-oriented framework in P. martinezii and plasticity-driven response in P. mexicana. Notably, several candidate genes encode key components of growth hormone pathways, including a gibberellin-regulated protein, a cytokinin hydroxylase and the AP2-like transcription factor ANT, providing valuable insights into how maternal genetic variation corresponds to the hormonal pathways that govern cell proliferation and organ size in the progeny. Integration of these findings with the contrasting demographic histories of both species revealed that population bottlenecks enhance the detectability of growth-associated variants by reducing background genetic variation. These genomic resources provide actionable information for prioritizing conservation measures, implementing assisted gene flow to maintain adaptive potential under climate change and designing future breeding programs. With 80.9-99.6% sequence identity to conserved Picea abies homologs, these findings may extend across the genus.

Picea