Search PubMedSearch

PubMed · 42696558

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Abstract

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Explore related subjects

Keep this discovery

BibTeXRIS

Sebastian Sonnenberg, Thapasya Vijayan, Christina Rupprecht, Yoko Philipina Krenn, Melissa Gruber, Hannah Dorfer, Gerald Kwikiriza, Harald Meimberg, Manuel Curto. 2026. Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.. https://doi.org/10.1111/1755-0998.70198

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

A genome-wide coverage-based pipeline for the identification of host-derived candidate DNA biomarkers from cell-free blood.

We have created a new data-analysis pipeline for the discovery of host-specific candidate DNA biomarkers derived from sequencing data of cell-free blood. Unlike approaches that rely on specific molecular or genetic signatures, our method leverages the coverage distribution of cell-free DNA sequences mapped to a reference genome, applying statistical analyses to identify informative short genomic regions for biomarker discovery. The pipeline is applicable to diverse diseases and can be used to analyze cell-free DNA sequences from plasma or serum to identify candidate biomarkers that are characteristic of disease states in mammals. Core functionalities were developed in Java and integrated with open-source software tools for the preprocessing of raw sequencing data, complemented by Python scripts for the machine-learning analysis and statistical validation. The pipeline is designed for HPC use and users can access the pipeline through a Galaxy workflow, which offers a user-friendly web interface for input selection prior to execution and analysis progress monitoring. Performance tests, carried out using duplicate sets of COVID-19 samples and controls, showed linear scalability of execution time with an increasing dataset size, as well as a substantial reduction in execution time through parallelized computation, whereby each HPC node is used to process the data of one chromosome. Further statistical tests confirmed the quality of the pipeline's results by showing that the set of identified candidate biomarkers remained stable across varying dataset sizes.

Biomarkers

Quo vadis, BGA? A collaborative EDNAP exercise on the challenges and progress in forensic biogeographical ancestry inference.

There is a broad consensus that forensic tests for the prediction of externally visible characteristics (EVC) and analysis of biogeographic ancestry (BGA) of an individual are technically reliable. However, interpretation of the results and population-specific genotype distribution patterns remains challenging. EVC and BGA analyses provide valuable information for population genetics studies and as investigative leads for criminal cases, as well as for historical and contemporary identification tests. However, inaccurate or incorrect predictions, for example, from subjective bias in the interpretations made, have the potential to misdirect police investigations. The legal situation regarding EVC and BGA testing varies by country: ranging from countries where it is explicitly prohibited, to those without specific regulations on biogeographic ancestry prediction, and others that have already enacted laws governing its use. The reluctance to utilize these analyses is not only due to legal restrictions and data protection concerns, but also to initial limited sets of sufficiently comprehensive forensic DNA assays. Forensic BGA marker panels typically contain up to ∼300 SNPs. This relatively small number of genetic markers, along with limited reference population data, complicates the interpretation of results from donors of unknown origin. This paper presents the results of a collaborative EDNAP study, which, for the first time, evaluated the approach to reporting EVC and BGA data between international laboratories. For the study, DNA from nine individuals with self-reported ancestry was collected and analysed using various forensic panels differing in the number and composition of ancestry-informative markers genotyped, comprising: the Precision ID mtDNA Whole Genome Panel, the VISAGE Basic Tool and the VISAGE Enhanced Tool for Appearance and Ancestry Prediction, and the Ion AmpliSeq™ PhenoTrivium Panel. To ensure full data protection, all SNP genotypes and uniparental marker haplotypes obtained were not shared with third parties. Instead, the genetic data were analysed using a range of commonly used population analysis software packages. These analysis outcomes were then distributed to twelve European forensic laboratories (both academic and law enforcement institutions), who were asked to prepare reports based on their interpretation of the phenotypes and ancestry they inferred from the analysis data. A questionnaire sent alongside the genetic information, aimed to evaluate which difficulties were encountered by the participants in processing the BGA analysis data they were given.

Humans

Genetic determinants of gestational diabetes mellitus in thai pregnant women: role of GCKR, CDKAL1, TCF7L2, NEDD1, and CMIP variants.

BACKGROUND: Gestational diabetes mellitus (GDM) has a high global prevalence and arises from complex interactions between genetic predisposition and environmental factors. GDM is associated with metabolic disturbances and chronic low-grade inflammation, both of which contribute to its pathogenesis. This study aimed to investigate the association between GDM and 135 single-nucleotide polymorphisms (SNPs) across 20 genes related to metabolic traits. METHODS: In this case-control study, 152 pregnant women with GDM and 684 pregnant women with normal glucose tolerance (NGT) who underwent antenatal examination at Siriraj Hospital, Bangkok, were enrolled. Clinical data and blood samples were collected from all participants. Genomic DNA was isolated and subjected to whole-genome sequencing using the DNBSEQ-T7RS high-throughput sequencing platform. Genotype analyses were performed using R software, and haplotype analyses were conducted using the online SNPStats software. RESULTS: After adjusting for maternal age and pre-pregnancy body mass index, polymorphisms in TCF7L2 (rs34872471, rs7901695, rs4506565, rs7903146, rs12243326, and rs12255372), NEDD1 (rs10431408, rs11830756, rs249579, rs249585, and rs4762339), CMIP (rs2306115 and rs201681534), CDKAL1 (rs4710942), GCKR (rs2293572 and rs2293571), and GCK (rs5883890) were significantly associated with the risk of GDM. Haplotype analysis demonstrated that the TCF7L2 rs12243326-rs12255372 CA haplotype was associated with a decreased risk of GDM (OR = 0.44, 95% CI: 0.23-0.81), while the NEDD1 rs249579-rs249585-rs4762339 GGT haplotype was associated with an increased risk of GDM (OR = 1.40, 95% CI: 1.08-1.82). CONCLUSIONS: These findings suggest that genetic variations in TCF7L2, NEDD1, CMIP, CDKAL1, GCK, and GCKR contribute to GDM susceptibility in the Thai population.

Humans