Search PubMedSearch

SEARCH · Search PubMed

Results for “high-throughput genotyping”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

An Amplicon Panel for High-Throughput and Low-Cost Genotyping of Yesso Scallop Mizuhopecten yessoensis.

The Yesso scallop Mizuhopecten yessoensis was imported from Japan to western Canada in the late 1980s to establish an economically viable scallop aquaculture industry. Since this time, the industry in Canada has operated with existing genetic diversity within the broodstock, which is considerably limited relative to wild populations. The sector has not been able to realise its full potential in part due to idiopathic hatchery failures and farm stock collapses due to disease outbreaks associated with the intracellular bacterial pathogen Francisella halioticida. To support Yesso scallop production and breeding, here we generate a low-density, genotyping-by-sequencing amplicon panel using single nucleotide polymorphism (SNP) markers that are evenly spaced across the M. yessoensis genome and that show high heterozygosity in Canada and Japan. The panel can also exploit the high genetic polymorphism of the M. yessoensis genome, with de novo SNP calling identifying over 2,500 high quality SNPs within the 579 sequenced amplicons. We demonstrate the utility and versatility of this new genotyping tool for breeding applications including parentage assignment, low density family-based genome-wide association study, trait heritability evaluation to determine potential for genomic selection, and species differentiation (against the weathervane scallop Patinopecten caurinus). We did not find any genomic regions significantly associated with F. halioticida resistance but did identify potential for genomic selection. We could separate the two species based on genotypes, and did not see evidence of a past M. yessoensis x P. caurinus hybridization event within the M. yessoensis breeding population at Vancouver Island University. This low-cost genotyping panel is expected to accelerate selective breeding improvements for M. yessoensis in Canada and elsewhere.

Animals

Development of a 10K breeder-friendly SNP chip for faba bean.

INTRODUCTION: Faba bean breeding and genomics have seen steady progress in recent years, supported by genome sequences and high-density genotyping platforms. These tools have been valuable for trait mapping, diversity assessment, and genomic research, but they have limited routine use in breeding programs due to their relatively high cost. Recent progress in establishing an optimized, cost-efficient genotyping-by-sequencing protocol tailored to the large and complex faba bean genome has created the foundation for a more accessible genotyping solution. METHODS: Using this approach, we explored the genetic diversity of faba bean germplasm from various panels, providing a comprehensive representation of the crop's genetic landscape. From this dataset, we identified and selected a high-quality set of informative SNP markers that are evenly distributed across the genome. Building on these resources, we designed a breeder-friendly 10K SNP chip. RESULTS: The 10K SNP chip delivers high accuracy, broad genomic coverage, and affordability. The chip was validated across diverse germplasm panels, demonstrating strong clustering performance, high reproducibility, and applicability to breeding-relevant germplasm. DISCUSSION: This platform offers a cost-effective alternative to higher-density arrays, enabling its integration into genomic selection, marker-assisted breeding, and diversity monitoring, ultimately supporting accelerated genetic gain and the delivery of improved varieties to farmers.

SNP chip

Development and identification of KASP-SNP markers correlated with Aeromonas hydrophila resistance traits in blunt snout bream (Megalobrama amblycephala).

The blunt snout bream (Megalobrama amblycephala) is an economically important freshwater fish species. However, it is highly susceptible to Aeromonas hydrophila infection, especially in intensive pond aquaculture in China. Molecular marker-assisted selection provides an efficient approach for breeding disease-resistant varieties; however, the key genes or molecular markers linked to A. hydrophila resistance remain scarce in this species. A 436 differential SNP sites with disease-resistant were screened on basis of whole-genome resequencing. Then, a high-throughput genomic KASP genotyping technique was utilized to discover favorable genes and SNP sites associated with A. hydrophila resistance. A total of 46 KASP markers were successfully developed with an accuracy of 92&#xa0;%. These markers were used to genotyping 120 blunt snout bream individuals. Through trait correlation analysis and general linear models (GLM), five SNPs significantly (P&#xa0;<&#xa0;0.05) associated with resistance to A. hydrophila were identified and mapped to five candidate genes (btnl2, cfhr2, slc47a1, neu3, nlrp1). Survival rate of individuals carrying the dominant genotype demonstrated an average survival rate of 81.39&#xa0;%, which represents a 69.35&#xa0;% increase in comparison with that of 48&#xa0;% in total population. This effect was validated in an external population of 100 fish. These findings identify key genetic markers associated with A. hydrophila resistance and provide a direction for elucidating the underlying molecular immune mechanisms, thus establishing a genetic foundation for future breeding strategies.

Cyprinidae

Genetic determinants of gestational diabetes mellitus in thai pregnant women: role of GCKR, CDKAL1, TCF7L2, NEDD1, and CMIP variants.

BACKGROUND: Gestational diabetes mellitus (GDM) has a high global prevalence and arises from complex interactions between genetic predisposition and environmental factors. GDM is associated with metabolic disturbances and chronic low-grade inflammation, both of which contribute to its pathogenesis. This study aimed to investigate the association between GDM and 135 single-nucleotide polymorphisms (SNPs) across 20 genes related to metabolic traits. METHODS: In this case-control study, 152 pregnant women with GDM and 684 pregnant women with normal glucose tolerance (NGT) who underwent antenatal examination at Siriraj Hospital, Bangkok, were enrolled. Clinical data and blood samples were collected from all participants. Genomic DNA was isolated and subjected to whole-genome sequencing using the DNBSEQ-T7RS high-throughput sequencing platform. Genotype analyses were performed using R software, and haplotype analyses were conducted using the online SNPStats software. RESULTS: After adjusting for maternal age and pre-pregnancy body mass index, polymorphisms in TCF7L2 (rs34872471, rs7901695, rs4506565, rs7903146, rs12243326, and rs12255372), NEDD1 (rs10431408, rs11830756, rs249579, rs249585, and rs4762339), CMIP (rs2306115 and rs201681534), CDKAL1 (rs4710942), GCKR (rs2293572 and rs2293571), and GCK (rs5883890) were significantly associated with the risk of GDM. Haplotype analysis demonstrated that the TCF7L2 rs12243326-rs12255372 CA haplotype was associated with a decreased risk of GDM (OR = 0.44, 95% CI: 0.23-0.81), while the NEDD1 rs249579-rs249585-rs4762339 GGT haplotype was associated with an increased risk of GDM (OR = 1.40, 95% CI: 1.08-1.82). CONCLUSIONS: These findings suggest that genetic variations in TCF7L2, NEDD1, CMIP, CDKAL1, GCK, and GCKR contribute to GDM susceptibility in the Thai population.

Humans

Barcoded mutant library enables high-throughput functional genomics in a filamentous fungus.

Advances in sequencing technology enabling rapid and inexpensive whole-genome sequencing highlight how few genes are functionally characterized. This problem is particularly acute in filamentous fungi, where even in the best studied organisms upward of half of genes are poorly characterized or unannotated. High-throughput tools to identify gene function exist for single-celled organisms, like yeast and bacteria. However, filamentous fungi present challenges to high-throughput gene characterization, including low transformation efficiency and multinucleate cells. Filamentous fungi are critical components of nutrient cycling in ecosystems, form symbioses with plants that improve nutrient uptake, and are devastating human, plant, and animal pathogens causing millions of deaths and substantial crop loss each year. Thus, it is critical to overcome challenges to rapid gene characterization in filamentous fungi. We generated a library of hundreds of millions of uniquely barcoded plasmids containing a broad host-range drug resistance marker for ectopic insertion into filamentous fungal genomes by Agrobacterium tumefaciens. We then optimized A. tumefaciens mediated transformation of the biocontrol agent Trichoderma atroviride and made an insertional mutagenesis library containing 83,311 barcoded insertions, disrupting 5,331 of 11,863 predicted genes. This library enables high-throughput screens to rapidly connect genotype to phenotype. Quantifying relative barcode abundance in the pooled library before and after exposure to experimental conditions identified candidate genes and recovered known pathway components in amino acid biosynthetic, fructose utilization, and xylose utilization pathways. This resource establishes a scalable platform for high-throughput functional genomics in filamentous fungi, enabling investigations of fungal biology to improve medical outcomes, biotechnology, and sustainable agriculture.

Genomics

Site-Specific Measurement of Meiotic Crossing-Over Rate with Droplet Digital PCR.

Understanding the frequency and distribution of meiotic crossovers (COs) is critical for both fundamental studies on meiosis and for practical applications in plant breeding, where controlling recombination can accelerate crop improvement. Determining CO rates at specific genomic loci has traditionally relied on labor-intensive methods that require the production and genotyping of large progenies. Here, we present a high-throughput protocol for site-specific quantification of meiotic COs in maize using droplet digital PCR (ddPCR). The method is based on genotyping individual pollen nuclei from hybrid plants to detect recombinant and nonrecombinant alleles at defined chromosomal intervals. By distributing several thousands of pollen nuclei into nanoliter-sized droplets and performing PCR with allele-specific fluorescent probes, this method allows precise quantification of CO frequency with high sensitivity. The protocol provides detailed guidance for nuclei isolation, probe master mix preparation, droplet generation, and data interpretation. This method can be easily adapted for use in other plants.

Journal Article

ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data.

MOTIVATION: Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes. However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive tools for visualization and analysis. Existing workflows often require multiple independent software tools, limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables efficient exploration and sharing of chemical genomics data. RESULTS: We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic profiles, assess gene-gene and condition-condition correlations, perform GO and KEGG enrichment analysis, and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances reproducibility, and promotes collaboration within the research community. AVAILABILITY AND IMPLEMENTATION: ChemGenXplore is freely accessible as a web application at https://chemgenxplore.kaust.edu.sa/. Source code and documentation, including instructions for local installation, are provided on GitHub (https://github.com/Hudaahmadd/ChemGenXplore). A Docker image is also available on DockerHub (https://hub.docker.com/r/hudaahmad/chemgenxplore) to ensure reproducibility and simplify installation.

Software

Next-Generation Sequencing Methods for Sensitive Hepatitis B Viral Genome Analysis: A European Study.

This multicentre study investigated the utility of next-generation sequencing (NGS) to detect and generate hepatitis B virus (HBV) genomes in samples of low viral load (from 0.2 to 6207 IU/mL). 23 HBV DNA-positive plasma samples of genotypes A-E and one HBV-negative control sample were assayed blindly via 9 established NGS methods from 6 European laboratories. Methods included untargeted metagenomics, pre-enrichment by probe-capture followed by Illumina sequencing, and HBV-specific PCR pre-amplification followed by sequencing with Nanopore or Illumina. Full HBV genomes were obtained only from samples with viral loads >&#x2009;1000 IU/mL using probe-capture methods, >&#x2009;200 IU/mL using PCR-Illumina methods, >&#x2009;10 IU/mL using PCR-Nanopore methods, and in no samples using metagenomic methods. Contamination was observed in the negative control and samples with very low viral loads in PCR-based methods. Probe-capture and metagenomic methods detected additional viruses not routinely screened in blood donations, including polyomaviruses and herpesviruses; positive results were confirmed by PCR. In conclusion, NGS may delineate whole-genome sequences at low viral loads if supported by a PCR pre-amplification step. Probe-capture methods also reliably detect HBV without pre-amplification but show limited genome coverage for samples with low viral loads; they may additionally detect a wide range of blood-borne viruses.

Humans

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis

Evaluation of bone preparation approaches using length-based analysis and targeted sequencing for forensic human identification of historic skeletal remains.

Advances in DNA technology have significantly enhanced the forensic community's ability to develop genetic profiles from unidentified human skeletal remains. However, sampling requires mechanical grinding of hard tissues before DNA isolation. This processing can compromise genetic profiles, particularly in aged bones. We compared the industry-standard pulverization method with an alternative powder-free preparation involving prolonged demineralization and subsequent slicing of 19th-century cortical bone. Data from DNA quantification, STR genotyping, and targeted SNP sequencing were used to evaluate powdered samples versus demineralized slices from paired human bones. Average human DNA yields for pulverized samples and demineralized slices were 0.032&#x2009;ng and 0.692&#x2009;ng, respectively. Demineralized slices recovered more amplifiable DNA than traditional homogenization methods (p&#x2009;<&#x2009;0.05). No pulverized samples produced STR profiles, whereas demineralized slices from the same bone samples yielded partial profiles. Samples underwent DNA repair, library preparation, and hybridization capture using the FORensic Capture Enrichment (FORCE) panel. Applying low-coverage (1X) analysis of high-throughput sequencing (HTS) data, demineralized slices outperformed those prepared by traditional pulverization methods (p&#x2009;<&#x2009;0.05) and substantially increased the information recovered compared with conventional STR analysis methods. Based on HTS data from pulverized samples, DNA fragment length ranged from 27 to 95&#x2009;bp, and FORCE SNP recovery was 33.23%. In contrast, for demineralized slices, DNA fragment length ranged from 85 to 114&#x2009;bp, and FORCE SNP recovery was 83.24%. The required reagents and equipment are typically available in forensic labs, and the workflow outlined herein significantly increases the success of DNA recovery from challenging skeletal samples.

Humans

Development of a PCR-based technique for genotyping UGT1A1 gene and distribution of rs3064744 alleles in the Russian population.

BACKGROUND: Accurate determination of tandem thymine-adenine (TA) repeat numbers in the UGT1A1 promoter region (rs3064744) is essential for diagnosing Gilbert's syndrome and personalizing therapy with toxic agents like irinotecan and atazanavir. However, traditional polymerase chain reaction (PCR) assays face severe limitations due to the AT-rich sequence and overlapping melting temperatures (Tm) of the highly homologous 7TA and 8TA alleles. In this context, melting curve analysis (MCA) employing fluorophore-quencher systems has emerged as a promising alternative. The purpose of this study was to develop a novel genotyping approach combining optimized aPCR-MCA analysis with an automated classifier to overcome the limitations posed by the differentiation of highly homologous alleles and to demonstrate its practical application, providing the distribution of rs3064744 genotypes across four regional cohorts of the Russian population. METHODS: A specialized Dual Head 1D-convolutional neural network (1D-CNN) ensemble with Test-Time Augmentation (TTA) was developed. The model was trained and internally validated on 1,620 engineered plasmid samples, and independently evaluated on an external clinical test set of 440 unique patient genomic DNA specimens. Real-time PCR was performed on CFX96 and DTprime platforms. Additionally, population-wide screening was conducted on 997 archival clinical samples from Moscow, Sakha (Yakutia), Dagestan, and Rostov regions. RESULTS: While 5TA and 6TA alleles were easily separated, absolute Tm distributions of 7TA and 8TA alleles overlapped significantly, and non-uniform Tm shifts of 0.8&#xa0;&#xb0;C-1.4&#xa0;&#xb0;C occurred across platforms. Conventional absolute Tm thresholding was therefore inadequate. By assessing relative morphological curve divergence against co-amplified 7TA/7TA and 7TA/8TA reference anchors, the 1D-CNN ensemble neutralized instrument noise. It achieved 100% accuracy on internal validation and 100% concordance (440/440) with clinical reference pyrosequencing. Population screening revealed that Dagestan, Yakutia, and Rostov cohorts closely align with the European population. Rare 5TA and 8TA alleles were detected at low frequencies in Yakutia and Moscow. CONCLUSION: Combining LNA-modified aPCR-MCA with a comparative 1D-CNN model successfully circumvents thermodynamic limitations and eliminates human operator bias. This integrated system offers an accessible, high-throughput, and clinically valid solution for routine UGT1A1 pharmacogenetic testing.

1D-CNN

UPDhmm: detecting uniparental disomy from NGS trio data.

SUMMARY: Uniparental disomies (UPDs) are copy-neutral chromosomal alterations that occur when both copies of a chromosome pair (entire or segmental) come from one parent. UPDs, including isodisomies (identical parental chromosome) and heterodisomies (two different homologs from the same parent), reflect meiotic and/or mitotic aberrations of chromosomal segregation that can be associated with congenital or acquired disease. Despite their relevance, current methods to detect UPDs using sequence data (exomes or genomes) have limited sensitivity for small events, cannot precisely determine the UPD sub-type or coordinates, and perform poorly when including individuals or populations with consanguinity. We present UPDhmm, a novel tool that uses trio-based sequence data (proband and parents) and models inheritance patterns. UPDhmm predicts the most likely inheritance scenario, normal Mendelian inheritance versus UPD event, based on genotype combinations using a Hidden Markov Model (HMM). We validated the method using simulations on exome and genome data from 1000-Genomes projects. UPDhmm overperformed currently available methods in detecting simulated UPD events in both data types. We applied UPDhmm to a collection of nearly 2400 families with a proband with autism spectrum disorder (Simons Simplex Collection Project) and identified UPD events in two affected individuals, one of them previously unreported. These two events, a paternal isodisomy of chr8 and a maternal heterodisomy of chr22, can be genetic causes of the disease, demonstrating the clinical utility of UPDhmm. Thus, UPDhmm can facilitate the incorporation of UPD detection into clinical pipelines of genomic analysis. AVAILABILITY AND IMPLEMENTATION: UPDhmm is implemented in R and is available in the Bioconductor package (version 1.5.0): https://www.bioconductor.org/packages/release/bioc/html/UPDhmm.html. The source code can be found at https://github.com/martasevilla/UPDhmm under the MIT license.

Uniparental Disomy

Columba: fast approximate pattern matching with optimized search schemes.

MOTIVATION: Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity. RESULTS: This paper introduces Columba, a high-performance lossless aligner tailored for Illumina sequencing data. Columba processes single or paired-end reads in FASTQ format and outputs alignments in SAM format. By utilizing advanced search schemes and bit-parallel alignment techniques, Columba achieves exceptional speed. Columba is available in two variants. The first, based on the bidirectional FM-index, prioritizes speed. The second, Columba RLC, uses run-length compression using a bidirectional move structure, significantly reducing memory usage for large, repetitive datasets like pan-genomes. Benchmarks on the human genome, as well as bacterial and human pan-genome datasets, demonstrate that Columba is much faster than existing lossless aligners and even competitive with lossy tools. We integrated Columba into the OptiType HLA genotyping pipeline, where it substantially reduced computational time while maintaining accuracy. These results position Columba as a versatile, state-of-the-art tool for high-sensitivity genomic analyses. AVAILABILITY AND IMPLEMENTATION: The source code of Columba is available at https://github.com/biointec/columba under AGPL license. Scripts to reproduce the benchmarks and analyses are available at https://doi.org/10.5281/zenodo.15849246.

Software