Search PubMedSearch

SEARCH · Search PubMed

Results for “SNP calling”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

16 recordsLinked to original sources

SNP genotyping in Pseudotsuga menziesii and Pinus radiata using targeted genotyping-by-sequencing (GBS): improved Bayesian SNP calling using a beta-binomial distribution and other optimized input parameters.

BACKGROUND: Single-nucleotide polymorphism markers (SNPs) have important applications in gene conservation, breeding, and fundamental genetics research. Our long-term goal is to develop routine approaches for SNP genotyping in forest trees. Ideally, these approaches would be inexpensive, able to accommodate a wide range of samples and SNPs, available through commercial providers, and produce high-quality SNP data. RESULTS: Using targeted genotyping-by-sequencing (GBS), we developed SNP assays for two highly heterozygous tree species, Douglas-fir (Pseudotsuga menziesii) and radiata pine (Pinus radiata). Using Douglas-fir haploid and diploid data, we optimized Bayesian SNP calling by testing four input parameters: (1) allele and genotype prior probabilities, (2) Rho, the beta-binomial dispersion parameter, (3) estimated read error (BayesReadError), and (4) the logPO cutoff used to filter low confidence SNP calls. logPO is the Bayesian posterior odds ratio for a called SNP. Compared to assuming a binomial distribution of read counts (Rho = 0), the beta-binomial distribution (Rho = 0.33) substantially reduced call error and heterozygote undercalling. Compared to the other Bayesian parameters, genotype priors had little effect on genotyping success. For Douglas-fir, we tested 5,360 SNP assays, and then studied the performance of the best 4,000. For radiata pine, we tested 6,000 SNP assays, and then studied the performance of the best 4,570. In Douglas-fir and radiata pine, our Bayesian approach resulted in median call rates of 95% to 98% for the top-ranked SNPs, with an estimated call error of 1.60% for known homozygous genotypes and 2.27% for known heterozygotes. In radiata pine, median and mean call rates were above 91% for GBS and SNP genotyping using an Axiom fixed genotyping array. Additionally, the median correspondence between the GBS and Axiom genotypes was about 98% overall (mean 96%). CONCLUSIONS: By optimizing Bayesian SNP calling, selecting the best 4-5 K SNPs, and excluding samples with low DNA amounts, we substantially reduced call error and heterozygote undercalling, resulting in SNP genotypes that were nearly identical to genotypes obtained using the Axiom array. Furthermore, genotyping performance should increase even further if our SNP rankings were used to develop less complex probe pools that target fewer SNPs.

Pinus

Mining of important genetic loci and evaluation of genetic effects for growth traits in Baicheng You Chicken.

The Baicheng You Chicken is a precious indigenous breed in Xinjiang, China, prized for its strong disease and stress resistance and superior meat quality. However, the lack of scientific breeding and conservation has led to poor production performance, particularly in growth traits. In this study, we collected phenotypic and whole-genome resequencing data from 1,535 18-week-old Baicheng You Chickens (180 males and 1,355 females). After stringent quality control (SNP call rate > 95%, minor allele frequency > 1%), we constructed the breed's first comprehensive SNP-based genome-wide variation map, which comprised 2,020,743 high-quality SNPs across the genome. The filtered SNPs had high mapping quality (99.73% mapped to the bGalGal1.mat.broiler.GRCg7b reference genome, Q30 = 93.26%) and a reasonable Ti/Tv ratio (2.596), guaranteeing the reliability of subsequent analyses. We estimated genetic effects (SNP-based heritability and phenotypic variance explained (PVE) by individual loci) via the restricted maximum likelihood (REML) method, and performed a genome-wide association study (GWAS) using a mixed linear model (MLM) - with sex as a fixed effect and principal components to correct for population stratification - to identify significant loci and their effect sizes (Beta). All eight growth traits showed moderate to high heritability: body weight (BW) had the highest heritability (0.86±0.11), while chest width (CW, 0.41±0.08) and body slanting length (BSL, 0.43±0.09) were the lowest; keel length (KL), chest girth (CG), pelvic width (PW), chest depth (CD) and shank length (SL) had heritabilities of 0.50±0.09, 0.46±0.09, 0.54±0.09, 0.67±0.10 and 0.74±0.10, respectively. GWAS identified 145 significant SNPs, with a maximum Beta value of 0.39 and PVE ranging from 1.25% to 6.25%. We annotated 22 candidate genes, with TAPT1, IGF2BP1, ADGRB3, LDB2, NCAPG and LCORL as key candidates. These quantifiable genetic markers and effect estimates provide direct targets for marker-assisted selection (MAS) and valuable resources for future genomic selection (GS) programs, offering a practical approach to improve the breed's slow growth while preserving its unique meat quality.

Baicheng You Chicken

Reference-Free Variant Calling with Local Graph Construction with ska lo (SKA).

The study of genomic variants is increasingly important for public health surveillance of pathogens. Traditional variant-calling methods from whole-genome sequencing data rely on reference-based alignment, which can introduce biases and require significant computational resources. Alignment- and reference-free approaches offer an alternative by leveraging k-mer-based methods, but existing implementations often suffer from sensitivity limitations, particularly in high mutation density genomic regions. Here, we present ska lo, a graph-based algorithm that aims to identify within-strain variants in pathogen whole-genome sequencing data by traversing a colored De Bruijn graph and building variant groups (i.e. sets of variant combinations). Through in silico benchmarking and real-world dataset analyses, we demonstrate that ska lo achieves high sensitivity in single-nucleotide polymorphism (SNP) calls while also enabling the detection of insertions and deletions, as well as SNP positioning on a reference genome for recombination analyses. These findings highlight ska lo as a simple, fast, and effective tool for pathogen genomic epidemiology, extending the range of reference-free variant-calling approaches. ska lo is freely available as part of the SKA program (https://github.com/bacpop/ska.rust).

Polymorphism, Single Nucleotide

Paralog-aware assembly and filtering strategies reveal minimal nucleotide variation on the macro germline-restricted chromosome of the zebra finch.

The germline-restricted chromosome (GRC) of passerines is a remarkable tissue-specific chromosome that accumulated paralogs of genes from the regular "A chromosomes" over millions of years, often amplified into dozens of gene copies. In addition to its repetitive content, typically uniparental inheritance, and lack of recombination, the GRC resembles non-recombining sex chromosomes and some B chromosomes, for all of which assembly and single-nucleotide polymorphisms (SNPs) calling are difficult. Here, we first show that much of the Australian zebra finch macro-GRC can be assembled using accurate long reads. We then describe a paralog-aware Snakemake pipeline, ParaVar, to map short reads from the GRC to retrieve GRC regions suitable for haplotype-based analysis. ParaVar reliably calls hundreds of SNPs across the GRC, thereby providing an estimate of nucleotide diversity on the highly repetitive zebra finch macro-GRC. Our results show significantly lower nucleotide diversity (20- to 50-fold lower) on the GRC compared to the mitogenome and autosomes, and a strong phylogenetic discordance between the GRC and the mitochondrial genome. Beyond the contribution of background selection, our results suggest that a single GRC haplotype recently spread through the populations while jumping across matrilines via occasional paternal inheritance. We anticipate that our paralog-aware pipeline will be useful for SNP calling and population genetics analyses of repetitive GRCs, sex chromosomes, and B chromosomes.

Animals

An Amplicon Panel for High-Throughput and Low-Cost Genotyping of Yesso Scallop Mizuhopecten yessoensis.

The Yesso scallop Mizuhopecten yessoensis was imported from Japan to western Canada in the late 1980s to establish an economically viable scallop aquaculture industry. Since this time, the industry in Canada has operated with existing genetic diversity within the broodstock, which is considerably limited relative to wild populations. The sector has not been able to realise its full potential in part due to idiopathic hatchery failures and farm stock collapses due to disease outbreaks associated with the intracellular bacterial pathogen Francisella halioticida. To support Yesso scallop production and breeding, here we generate a low-density, genotyping-by-sequencing amplicon panel using single nucleotide polymorphism (SNP) markers that are evenly spaced across the M. yessoensis genome and that show high heterozygosity in Canada and Japan. The panel can also exploit the high genetic polymorphism of the M. yessoensis genome, with de novo SNP calling identifying over 2,500 high quality SNPs within the 579 sequenced amplicons. We demonstrate the utility and versatility of this new genotyping tool for breeding applications including parentage assignment, low density family-based genome-wide association study, trait heritability evaluation to determine potential for genomic selection, and species differentiation (against the weathervane scallop Patinopecten caurinus). We did not find any genomic regions significantly associated with F. halioticida resistance but did identify potential for genomic selection. We could separate the two species based on genotypes, and did not see evidence of a past M. yessoensis x P. caurinus hybridization event within the M. yessoensis breeding population at Vancouver Island University. This low-cost genotyping panel is expected to accelerate selective breeding improvements for M. yessoensis in Canada and elsewhere.

Animals

GENOME TARGETED ENRICHMENT AND SEQUENCING OF HUMAN-INFECTING CRYPTOSPORIDIUM spp.

Cryptosporidium spp. are parasites that cause severe illness in vulnerable human populations. Obtaining pure and sufficient Cryptosporidium DNA from clinical and environmental samples is a challenging task. Oocysts shed in available fecal samples can be limited in quantity, require purification (biased towards dominant strains), and yield limited DNA (<40 fg/oocyst). Here, we use updated genomic sequences from a broad diversity of human-infecting Cryptosporidium species ( C. cuniculus , C. hominis , C. meleagridis , C. parvum , C. tyzzeri , and C. viatorum ) to develop and validate a set of 100,000 RNA baits (CryptoCap_100k) with the aim of enriching Cryptosporidium spp. DNA from varied samples. Compared to unenriched libraries, CryptoCap_100k increases the percentage of reads mapping to target genome sequences, increases the depth and breadth of genome coverage and the reliability of detecting species and mixed infections within a sample, and allows assessment of genetic variation via SNP calling, while decreasing costs.

Journal Article

A Practical Approach to High-Throughput and Accurate Mapping-by-Sequencing in Arabidopsis.

Forward-directed genetic screens are extremely powerful in identifying novel genes involved in a specific biological process, including various chromatin regulatory pathways. However, the traditional ways of genetic mapping are time- and cost-demanding. Recently, the whole process was revolutionized by the development of mapping-by-sequencing (MBS) protocols. In MBS, the causal mutations and their positions within genes are identified directly by whole-genome sequencing and bioinformatics analysis of the bulk of mutant plants selected based on the mutant phenotype from a segregating population. MBS increases precision and economizes the mapping. Here, we describe a general protocol and provide practical tips on how to proceed with the mapping-by-sequencing on the example of Arabidopsis forward-directed genetic screen designed to identify mutants sensitive to a specific type of DNA damage. The described protocol is generally applicable to a wide range of genetic screens in various inbreeding species with a reference genome sequence.

Arabidopsis

Optimizing genetic ancestry adjustment in DNA methylation studies: a comparative analysis of approaches.

BACKGROUND: Genetic ancestry is an important factor to account for in DNA methylation studies because genetic variation influences DNA methylation patterns. One approach uses principal components (PCs) calculated from CpG sites that overlap with common SNPs to adjust for ancestry when genotyping data is not available. However, this method does not remove technical and biological variations, such as sex and age, prior to calculating the PCs. The first PC is therefore often associated with factors other than ancestry. METHODS: We developed and adapted the adapted EpiAnceR+&#x2009;approach, which includes (1) residualizing the CpG data overlapping with common SNPs for control probe PCs, sex, age, and cell type proportions to remove the effects of technical and biological factors, and (2) integrating the residualized data with genotype calls from the SNP probes (commonly referred to as rs probes) present on the arrays, before calculating PCs and evaluated the clustering ability and relationship to genetic ancestry. RESULTS: The PCs generated by EpiAnceR+&#x2009;led to improved clustering for repeated samples from the same individual and stronger associations with genetic ancestry groups predicted from genotype information compared to the original approach. EpiAnceR+&#x2009;also outperformed the use of DNA methylation PCs or surrogate variables for ancestry adjustment. CONCLUSIONS: We show that the EpiAnceR+&#x2009;approach improves the adjustment for genetic ancestry in DNA methylation studies. EpiAnceR+&#x2009;can be integrated into existing R pipelines for commercial methylation arrays, such as 450&#xa0;K, EPIC v1, and EPIC v2. The code is available on GitHub ( https://github.com/KiraHoeffler/EpiAnceR ).

DNA Methylation

Rapid CRISPR-based bovine embryo sexing to streamline genotype-informed cattle breeding.

Cattle in vitro fertilisation and embryo transfer programmes increasingly rely on embryo-level selection to accelerate genetic gain, but current sexing and genotyping workflows can be costly, slow and logistically demanding. This study developed an efficient, low-resource workflow for bovine embryo sexing that combines whole genome amplification (WGA) with recombinase polymerase amplification-CRISPR-Cas12a (RPA-Cas12a). It also assessed whether the same WGA biopsy products could be used for downstream single nucleotide polymorphism (SNP) microarray genotyping. A one-tube RPA-Cas12a assay targeting the bovine Y-chromosome S4 repeat was developed for fluorescence and lateral flow assay (LFA) readouts. Analytical sensitivity was assessed using serially diluted bovine genomic DNA (gDNA), and breed robustness was tested using male and female gDNA from five major beef breeds and Holstein cattle. The workflow was then applied to WGA products from 22 bovine blastocyst biopsies, with sex calls validated against an established real-time PCR melt curve assay and 100K SNP microarray genotyping. The assay detected male bovine gDNA down to 100&#x202f;pg using both fluorescence and LFA readouts, with no signal from female gDNA. Male-specific detection was consistent across all breeds tested. All WGA-RPA-Cas12a sex calls from blastocyst biopsies were concordant with real-time PCR and SNP microarray sex calls, and WGA biopsy products produced genome-wide SNP call rates above 85%. This workflow provides a practical approach for rapid bovine embryo sex triage and could reduce unnecessary cryopreservation and genotyping while improving the efficiency of genotype-informed cattle breeding programmes.

Bovine embryo

A transcriptome-wide approach for rapid pathotype discrimination of Puccinia striiformis f. sp. tritici in north-western India.

Stripe rust of wheat caused by Puccinia striiformis f. sp. tritici (Pst) remains a major constraint to wheat production in India due to the rapid evolution and frequent emergence of virulent pathotypes. Rapid and reliable discrimination of Pst pathotypes is essential for effective resistance deployment and surveillance. In the present study, transcriptome-wide simple sequence repeats (SSRs) and single nucleotide polymorphisms (SNPs) were exploited to develop and validate molecular markers for pathotype-specific detection of Pst pathotypes prevalent in North India (110S119, 238S119, 46S119, 110S84 and 78S84). Microsatellite mining from 6103 core orthologous clusters comprising 51,127 transcripts mined 14,634 SSR loci, from which 93 primer pairs were synthesized. However, only three SSR markers exhibited polymorphism indicating limited discrimination potential of expressed sequence-derived (EST) SSRs for pathotype differentiation. In contrast, SNP discovery through stringent variant calling and filtration yielded 186 pathotype-specific homokaryotic SNPs, of which 56 high-confidence loci were selected for Kompetitive Allele-Specific PCR (KASP) assay development. A total of 48 KASP markers were synthesized and 14 demonstrated clear pathotype- or cluster-specific polymorphism representing substantially higher resolution than SSR markers. The high SNP-to-KASP conversion efficiency (~&#x2009;95%) and reproducible fluorescence-based clustering emphasize the robustness of KASP assay. Comparative evaluation revealed that SNP-based KASP markers provide superior discriminatory capacity for closely related Pst pathotypes and represent a promising complementary molecular approach for rapid identification of predominant Indian Pst pathotypes. The validated marker panel developed in this study can complement conventional virulence phenotyping and field pathogenomics approaches for surveillance of currently known pathotypes, while continued refinement may accommodate future changes in pathogen populations.

India

Use of a dense single nucleotide polymorphism map for in silico mapping in the mouse.

Rapid expansion of available data, both phenotypic and genotypic, for multiple strains of mice has enabled the development of new methods to interrogate the mouse genome for functional genetic perturbations. In silico mapping provides an expedient way to associate the natural diversity of phenotypic traits with ancestrally inherited polymorphisms for the purpose of dissecting genetic traits. In mouse, the current single nucleotide polymorphism (SNP) data have lacked the density across the genome and coverage of enough strains to properly achieve this goal. To remedy this, 470,407 allele calls were produced for 10,990 evenly spaced SNP loci across 48 inbred mouse strains. Use of the SNP set with statistical models that considered unique patterns within blocks of three SNPs as an inferred haplotype could successfully map known single gene traits and a cloned quantitative trait gene. Application of this method to high-density lipoprotein and gallstone phenotypes reproduced previously characterized quantitative trait loci (QTL). The inferred haplotype data also facilitates the refinement of QTL regions such that candidate genes can be more easily identified and characterized as shown for adenylate cyclase 7.

Adenylyl Cyclases

Genomic epidemiology of enteropathogenic Escherichia coli in southwestern Nigeria.

BACKGROUND: Enteropathogenic Escherichia coli (EPEC) are etiological agents of diarrhea. We studied the genetic diversity and virulence factors of EPEC in southwestern Nigeria, where this pathotype is rarely characterized. METHODOLOGY/PRINCIPAL FINDINGS: EPEC isolates (n&#x2009;=&#x2009;96) recovered from recent southwestern Nigeria diarrhea case-control studies were whole genome-sequenced using Illumina technology. Genomes were assembled using SPAdes and quality was evaluated using QUAST. Virulencefinder, Ectyper, and ResFinder were used to identify virulence genes, serotypes, and resistance genes. Multilocus sequence typing was done by STtyping. Single nucleotide polymorphisms (SNPs) were called out of whole genome alignment using SNP-sites and a phylogenetic tree was constructed using IQtree. Thirty-nine of the 96(40.6%) EPEC isolates were from diarrhea cases diarrhea. Nine isolates from diarrhea patients and four from healthy controls were typical EPEC, harboring bundle-forming pilus (bfp) genes whilst the rest were atypical EPEC. There were 15 EPEC-EAEC hybrids. Atypical serotypes O71:H19 (16, 16.6%), O108:H21 (6, 6.3%), O157:H39 (5, 5.2%), and O165:H9 (4, 4.2%) were the most prevalent; only 8 (8.3%) isolates belonged to classical EPEC serovars. The largest, ST517 clade harbored multiple siderophore and serine protease autotransporter genes and included an O71:H19 subclade <10 SNPs apart, representing a likely outbreak involving 15 children, four with diarrhea. Likely outbreaks, of typical O119:H6(ST28) and atypical O127:H29(ST7798) were additionally identified. CONCLUSION/SIGNIFICANCE: EPEC circulating in southwestern Nigeria are diverse and differ substantially from well-characterized lineages seen previously elsewhere. EPEC carriage and outbreaks could be commonplace but are largely undetected, hence, unreported, and require genomic surveillance for identification.

Nigeria

Leveraging ONT move table values for signal aware variant calling.

Oxford Nanopore Technologies (ONT) sequencing enables long-range haplotype phasing and contiguous genome assembly but still exhibits elevated error rates that challenge small variant calling, particularly for insertions and deletions (Indels). While raw electrical signals contain rich information, existing signal-aware methods require computationally intensive processing of large signal files. Here, we present Clair3 v2, a method that leverages the ONT move table-a lightweight byproduct of basecalling that maps signal events to nucleotide positions-to improve variant calling accuracy. Clair3 v2 builds upon Clair3 and integrates signal-level dwelling time to significantly enhance variant calling performance. We also propose a genome position based circular buffer to incorporate dwelling time with minimal computational overhead. Benchmarking across six Genome in a Bottle samples demonstrates substantial improvements in variant calling accuracy. With HAC basecalling, Clair3 v2 achieves a mean SNP F1-score of 97.69% at 10 &#xd7; depth (compared to 96.45% for baseline Clair3), and Indel F1 scores improved from 64.27% to 76.70%, while gains persisted at higher depths. The benefits were most pronounced for longer Indels and in complex genomic regions, where Indel F1 scores in long homopolymer regions improved from 14.3% to 45.2%. Benchmark results across various basecalling modes, samples, and coverage settings outperformed Clair3 baselines and other methods, including DeepVariant and Dorado Variant, and demonstrate the significant benefits of Clair3 v2. Furthermore, Clair3 v2 incurs negligible runtime compared to standard Clair3, making it practical for routine use.

Sequence Analysis, DNA

Accelerated long-read variant calling with Clair3 for whole-genome sequencing.

SUMMARY: The rapid growth of genomic data and increasing adoption of long-read sequencing technologies have rendered variant calling one of the most computationally demanding tasks in genomic analysis. Although deep learning-based methods currently outperform conventional approaches in distinguishing true variants from complex sequencing noise, they impose prohibitive computational and time requirements. To address this limitation, we present a computational framework based on Clair3 that integrates parallelized feature generation, enhanced variant phasing, in-memory read haplotagging, and GPU-accelerated neural network inference to accelerate variant calling. By dynamically optimizing the use of both GPU and CPU resources, our method achieves substantial runtime improvements without compromising accuracy. We evaluated our framework across a range of sequencing depths, diverse samples, and multiple hardware configurations. Our results demonstrate that the optimized pipeline completes variant calling for a 30&#xd7; whole-genome sequence in 12-20&#x2009;minutes using standard computational resources (32 CPU threads and one NVIDIA GPU), and in 12-15&#x2009;minutes on an Apple Mac Studio (32 threads), which is &#x223c;10-20-fold speedup compared with its initial release. In addition to exceptional efficiency, our method maintains state-of-the-art accuracy, achieving SNP F1-scores of 99.32% and 99.70% on 30&#xd7; ONT and PacBio GIAB HG003 datasets, respectively. This work introduces a rapid, accurate, and scalable variant calling framework that effectively supports large-cohort genomic studies and time-sensitive clinical applications. AVAILABILITY AND IMPLEMENTATION: The accelerated implementation of Clair3 is open source and available at: https://github.com/HKU-BAL/Clair3/tree/gpu.

Whole Genome Sequencing

Copy number variant scan in more than four thousand Holstein cows bred in Lombardy, Italy.

Copy Number Variants (CNV) are modifications affecting the genome sequence of DNA, for instance, they can be duplications or deletions of a considerable number of base pairs (i.e., greater than 1000 bp and up to millions of bp). Their impact on the variation of the phenotypic traits has been widely demonstrated. In addition, CNVs are a class of markers useful to identify the genetic biodiversity among populations related to adaptation to the environment. The aim of this study was to detect CNVs in more than four thousand Holstein cows, using information derived by a genotyping done with the GGP (GeneSeek Genomic Profiler) bovine 100K SNP chip. To detect CNV the SVS 8.9 software was used, then CNV regions (CNVRs) were detected. A total of 123,814 CNVs (4,150 non redundant) were called and aggregated into 1,397 CNVRs. The PCA results obtained using the CNVs information, showed that there is some variability among animals. For many genes annotated within the CNVRs, the role in immune response is well known, as well as their association with important and economic traits object of selection in Holstein, such as milk production and quality, udder conformation and body morphology. Comparison with reference revealed unique CNVRs of the Holstein breed, and others in common with Jersey and Brown. The information regarding CNVs represents a valuable resource to understand how this class of markers may improve the accuracy in prediction of genomic value, nowadays solely based on SNPs markers.

Cattle

Single nucleotide polymorphism-based validation of exonic splicing enhancers.

Because deleterious alleles arising from mutation are filtered by natural selection, mutations that create such alleles will be underrepresented in the set of common genetic variation existing in a population at any given time. Here, we describe an approach based on this idea called VERIFY (variant elimination reinforces functionality), which can be used to assess the extent of natural selection acting on an oligonucleotide motif or set of motifs predicted to have biological activity. As an application of this approach, we analyzed a set of 238 hexanucleotides previously predicted to have exonic splicing enhancer (ESE) activity in human exons using the relative enhancer and silencer classification by unanimous enrichment (RESCUE)-ESE method. Aligning the single nucleotide polymorphisms (SNPs) from the public human SNP database to the chimpanzee genome allowed inference of the direction of the mutations that created present-day SNPs. Analyzing the set of SNPs that overlap RESCUE-ESE hexamers, we conclude that nearly one-fifth of the mutations that disrupt predicted ESEs have been eliminated by natural selection (odds ratio = 0.82 +/- 0.05). This selection is strongest for the predicted ESEs that are located near splice sites. Our results demonstrate a novel approach for quantifying the extent of natural selection acting on candidate functional motifs and also suggest certain features of mutations/SNPs, such as proximity to the splice site and disruption or alteration of predicted ESEs, that should be useful in identifying variants that might cause a biological phenotype.

Alleles