Search PubMedSearch

SEARCH · Search PubMed

Results for “single-nucleotide polymorphism”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Testing for Genetic Interactions in Complex Disease With Distance Correlation.

Understanding epistasis (genetic interaction) may shed some light on the genomic basis of common diseases, including disorders of maximum interest due to their high socioeconomic burden, like schizophrenia. Distance correlation is an association measure that characterizes general statistical independence between random variables, not only the linear one. Here, we propose distance correlation as a novel tool for the detection of epistasis from case-control data of single-nucleotide polymorphisms. On the methodological side, we highlight the derivation of the explicit asymptotic null distribution of the test statistic. We show that this is the only way to obtain enough computational speed for the method to be used in practice, in a scenario where the resampling techniques found in the literature are impractical. Our simulations show satisfactory calibration of significance, as well as comparable or better power than existing methodology. We conclude with the application of our technique to a schizophrenia genetics dataset, obtaining biologically sound insights.

Epistasis, Genetic

Panmixia in a Widespread Butterfly: High Dispersal and Ecological Generalism Buffer Against Landscape Fragmentation.

Habitat fragmentation is widely expected to reduce population connectivity and increase genetic differentiation, although the strength of these effects depends on species-specific traits such as dispersal ability. Here, we investigated the population genetic structure of the cosmopolitan butterfly, Pieris rapae L. (Lepidoptera: Pieridae), across western Germany using genome-wide single-nucleotide polymorphism (SNP) data. To analyze the effects of landscape structure on genetic connectivity, we applied a paired study design comprising four landscape pairs, each consisting of a highly intensified, modern agricultural landscape and a more heterogeneous, traditional landscape. Our results revealed no evidence of genetic differentiation. Pairwise FST values were close to zero; we detected no isolation by distance, and clustering analyses supported a single genetic population. No meaningful associations between genetic variation and environmental variables were detected, with landscape effects explaining less than 0.4% of genomic variation. Consequently, we found no evidence for stronger genetic structuring in modern compared to more connected traditional landscapes. Our results suggest that extensive habitat fragmentation does not necessarily translate into reduced genetic connectivity in highly mobile, generalist species. In P. rapae , high dispersal ability and ecological generalism appear to buffer against the genetic consequences of landscape modification, resulting in panmictic population structure even across strongly contrasting agricultural landscapes.

Pieris rapae

Genome-wide scans reveal candidate genes associated with wing morph differentiation in Tetrix japonica.

Wing dimorphism is an important dispersal-related trait in insects, but its genomic basis remains poorly understood in pygmy grasshoppers. Here, we integrated genome-wide single-nucleotide polymorphism (SNP) analyses, population structure inference, selection scans, and functional annotation to investigate genomic differentiation between long- and short-winged Tetrix japonica. Principal component analysis (PCA), ADMIXTURE, and phylogenetic analyses revealed weak genome-wide separation between morphs, indicating differentiation on a largely shared genetic background. Genome-wide scans based on the fixation index (FST), nucleotide diversity ratios, and Tajima's D, using 50-kb non-overlapping windows and empirical top-5% outlier thresholds, identified multiple candidate regions across seven chromosomes. The broader long- and short-winged candidate sets spanned 9.35 Mb and 9.37 Mb and directly overlapped 82 and 77 genes, respectively. Candidate genes were associated with signaling/hormone regulation, membrane transport, metabolism, cytoskeletal organization, extracellular matrix structure, and development. Short-winged candidate genes were significantly enriched for ABC-type transporter activity and ATP hydrolysis activity. Because all individuals originated from a single laboratory-maintained population with weak genome-wide structure, these regions should be regarded as candidate loci from a screening-stage analysis that require validation in independent populations and by functional assays, rather than as confirmed targets of selection.

Animals

Genome-wide characterization of heat shock protein genes reveals thermal stress-responsive candidates in Litopenaeus vannamei.

Heat shock proteins (HSPs) are conserved molecular chaperones involved in protein folding, refolding, aggregation prevention, and degradation of damaged proteins. However, the genomic organization and thermal responsiveness of HSP genes in the Pacific white shrimp (Litopenaeus vannamei) remain incompletely understood. Here, we performed a genome-wide analysis of the HSP gene family and examined its phylogenetic relationships, structural features, duplication patterns, sequence variation, interaction networks, and transcriptional responses to acute heat stress. A total of 34 HSP genes were identified and classified into the HSP90, HSP70, HSP40/DNAJ, HSP60, and small HSP families. Phylogenetic, motif, gene structure, synteny, and subcellular localization analyses revealed evolutionary conservation and structural diversification among family members. Three duplicated gene pairs were identified, comprising two segmental duplications and one tandem duplication. All pairs exhibited Ka/Ks ratios below 1, consistent with purifying selection of varying strength. Sequence analysis identified 295 nonsynonymous single-nucleotide polymorphisms, of which 12 were consistently predicted to be deleterious by multiple algorithms. Protein-protein interaction analysis indicated enrichment of protein-folding and cellular stress-response functions. RT-qPCR analysis showed significant induction of HSPA4, HSP90AA1, TRAP1, BiP, and DNAJA1 after 6, 12, and 24 h of exposure to 34 °C, whereas DNAJC3 was significantly induced only at 12 h. All six genes reached their highest transcript abundance at 12 h. These findings may provide a genomic framework for HSP genes in L. vannamei and identify candidate genes and variants associated with thermal stress responses.

Animals

Genetic determinants of gestational diabetes mellitus in thai pregnant women: role of GCKR, CDKAL1, TCF7L2, NEDD1, and CMIP variants.

BACKGROUND: Gestational diabetes mellitus (GDM) has a high global prevalence and arises from complex interactions between genetic predisposition and environmental factors. GDM is associated with metabolic disturbances and chronic low-grade inflammation, both of which contribute to its pathogenesis. This study aimed to investigate the association between GDM and 135 single-nucleotide polymorphisms (SNPs) across 20 genes related to metabolic traits. METHODS: In this case-control study, 152 pregnant women with GDM and 684 pregnant women with normal glucose tolerance (NGT) who underwent antenatal examination at Siriraj Hospital, Bangkok, were enrolled. Clinical data and blood samples were collected from all participants. Genomic DNA was isolated and subjected to whole-genome sequencing using the DNBSEQ-T7RS high-throughput sequencing platform. Genotype analyses were performed using R software, and haplotype analyses were conducted using the online SNPStats software. RESULTS: After adjusting for maternal age and pre-pregnancy body mass index, polymorphisms in TCF7L2 (rs34872471, rs7901695, rs4506565, rs7903146, rs12243326, and rs12255372), NEDD1 (rs10431408, rs11830756, rs249579, rs249585, and rs4762339), CMIP (rs2306115 and rs201681534), CDKAL1 (rs4710942), GCKR (rs2293572 and rs2293571), and GCK (rs5883890) were significantly associated with the risk of GDM. Haplotype analysis demonstrated that the TCF7L2 rs12243326-rs12255372 CA haplotype was associated with a decreased risk of GDM (OR = 0.44, 95% CI: 0.23-0.81), while the NEDD1 rs249579-rs249585-rs4762339 GGT haplotype was associated with an increased risk of GDM (OR = 1.40, 95% CI: 1.08-1.82). CONCLUSIONS: These findings suggest that genetic variations in TCF7L2, NEDD1, CMIP, CDKAL1, GCK, and GCKR contribute to GDM susceptibility in the Thai population.

Humans

Dissemination of blaKPC-3-harbouring Klebsiella pneumoniae across ST48 and ST628 in multiple healthcare facilities in the Republic of Korea.

Klebsiella pneumoniae carbapenemase-3 (KPC-3) remains rare in South Korea, where KPC-2 is the dominant carbapenemase, making the repeated detection of a concentrated blaKPC-3 signal over five years notable. We performed genomic analyses of blaKPC-3-harbouring K. pneumoniae from a regional healthcare network. Two chromosomally distinct lineages with concordant capsule loci (ST628/KL15 and ST48/KL62) presented multidrug-resistant phenotypes, and the virulence-associated loci were confined to ST48. Single-nucleotide polymorphism (SNP) analyses revealed near-clonal relatedness within lineages, with 0-38 pairwise SNPs among ST628 isolates and 8 SNPs between the two ST48 isolates. Core-genome multilocus sequence typing (cgMLST) supported this structure, as ST628 isolates were assigned to complex type 19149 with 0-7 allelic differences, and ST48 isolates were assigned to complex type 19150 with 5 allelic differences. These patterns support vertical spread via clonal expansion across multiple facilities. Despite substantial chromosomal separation, most isolates carried the same IncFII(K) plasmid backbone and blaKPC-3, and they were nearly indistinguishable from a plasmid previously reported in South Korea. One isolate carried blaKPC-3 on a distinct multireplicon IncFIB(K)/IncFII(K) plasmid, indicating that the signal was not confined to a single plasmid backbone. In both plasmids, blaKPC-3 was embedded within Tn4401b. These findings indicate that a rare blaKPC-3 genotype can persist regionally through sustained clonal dissemination and that cross-lineage linkage is compatible with past horizontal transfer involving a conserved plasmid. These findings underscore the need for subtype-resolved, regionally coordinated genomic surveillance in connected healthcare networks to detect uncommon carbapenemase variants early.

Klebsiella pneumoniae

Scalable medium-density genotyping platforms for cultivar identification, pedigree authentication, marker-assisted and genomic selection, and other applications in strawberry.

A broad spectrum of high-density genotyping approaches, including single-nucleotide polymorphism (SNP) arrays, genotyping-by-sequencing, and whole-genome reduced-representation sequencing, have been shown to perform well in strawberry (Fragaria × ananassa), despite the inherent complexity of the octoploid genome. While these approaches are effective, their routine deployment in breeding programs can be constrained by cost, computational requirements, and workflow complexity. In parallel, many breeding programs continue to rely on locus-specific assays for marker-assisted selection, resulting in fragmented and inefficient genotyping strategies. Here, we describe medium-density amplicon-based genotyping platforms for strawberry designed to provide cost-effective, turnkey solutions that integrate markers used for marker-assisted selection with genome-wide markers suitable for genomic prediction in a single laboratory assay. These platforms were developed by targeting 1,650 or 4,811 target SNPs via amplicon sequencing, and are interoperable with existing high-density genotyping resources, including a widely used 50K SNP array, thereby facilitating data integration across platforms. We benchmarked their performance relative to the 50K SNP array across breeding-relevant applications, including identity and purity testing, pedigree authentication, marker-assisted selection, and genomic selection, and further evaluated the feasibility of genotype imputation to enhance genome-wide information content. Across analyses, the 1,650- and 4,811-amplicon platforms produced results comparable to higher-density platforms while substantially reducing genotyping cost and analytical overhead. This work demonstrates that targeted amplicon-based genotyping can support efficient, scalable, and integrated genome-informed breeding, enabling the routine application of both marker-assisted and genomic selection within strawberry breeding workflows. Open-source R workflows are provided to support streamlined analyses in breeding contexts.

Fragaria

Identification and characterization of PsFwC9 conferring Fusarium wilt resistance in pea.

Pea (Pisum sativum L.) is one of the most important edible legumes in China, with both planting area and total yield ranking among the highest in the world. Fusarium wilt, caused by Fusarium oxysporum f. sp. pisi (Fop), is a severe factor limiting pea production. The deployment of resistant pea cultivars is the most effective and sustainable strategy for controlling this disease. In the present study, a novel resistance gene PsFwC9, conferring resistance to Fop race 5, was identified in the resistant pure line Chengwan 9-8 (CW9-8), and its candidate gene Psat4g213640 was characterized and functionally validated to be associated with disease resistance. Genetic analysis of the F₂ population derived from the cross between the resistant parent CW9-8 and the susceptible parent Chengwan 9-1 (CW9-1) revealed that PsFwC9 was controlled by a single dominant gene. Based on whole-genome resequencing, bulked segregant analysis sequencing (BSA-seq) and fine mapping, PsFwC9 was localized to an 817.06-kb region on chromosome 4 (i.e. linkage group IV, chr4LG4), flanked by KASP markers A016508 and A016511, and co-segregated with four markers. Haplotype analysis revealed that only the marker A016615 was significantly associated with Fusarium wilt resistance, and this marker was designated as a diagnostic marker for PsFwC9. Marker A016615 was located at 425 699 725 bp on chr4LG4, corresponding to the 277 bp within Psat4g213640, where a 'A/G' single-nucleotide polymorphism caused an amino acid substitution leading to an alteration in protein structure; therefore, Psat4g213640 was identified as the PsFwC9 candidate gene. Quantitative real-time PCR analysis showed no significant difference in the expression levels of Psat4g213640 between CW9-8 and CW9-1. Overexpression of the candidate gene Psat4g213640CW9-8 in the hairy root system significantly enhanced the resistance of CW9-1 to Fusarium wilt, whereas RNA interference-mediated silencing of Psat4g213640CW9-8 reduced the resistance of CW9-8, indicating that Psat4g213640CW9-8 played a crucial role in pea resistance to Fusarium wilt. In addition, subcellular localization showed that the protein encoded by Psat4g213640 was targeted to the endoplasmic reticulum. Collectively, these findings not only enriched the gene resources for disease resistance in pea and provided an important foundation for elucidating the molecular mechanism of PsFwC9-mediated resistance, but also provided important technical support for the practical application of molecular breeding for disease resistance in pea.

Journal Article

Genomic Regions Associated with Resistance to Soybean Cyst Nematode (Heterodera glycines Ichinohe) Population HG Type 1.2.5.7 in Dry Beans (Phaseolus vulgaris L.).

North Dakota, the largest dry bean (Phaseolus vulgaris L.) producing state in the U.S., faces an emerging production threat caused by the soybean cyst nematode (SCN; Heterodera glycines Ichinohe, 1952). Host resistance is an effective management strategy, yet resistance to the virulent SCN population HG type 1.2.5.7 has not been genetically characterized in dry beans. In this study, 170 dry bean genotypes (113 breeding lines/cultivars and 57 germplasm accessions) were evaluated for response to HG type 1.2.5.7 under controlled conditions using female index (FI) as the resistance phenotype. FI values ranged from 4.1% to 78.1%, with one genotype (PI 313733) classified as resistant, 35 moderately resistant, 104 moderately susceptible, and 30 susceptible. Genome-wide association analysis using 2,044 single-nucleotide polymorphism (SNP) markers from the 3.8K Bean Panel chip and the BLINK model identified four significant marker-trait associations on chromosomes Pv02, Pv05, Pv07, and Pv11. Linkage disequilibrium-defined candidate intervals spanned 108 kb (Pv02), 1.50 Mb (Pv05), 798 kb (Pv07), and 1.45 Mb (Pv11), collectively containing 126 annotated genes: 20 on Pv02, 39 on Pv05, 35 on Pv07, and 32 on Pv11. The intervals contained putative genes annotated for signaling and transcriptional regulation, cell wall and carbohydrate metabolism, transport, and secondary metabolism. Together, these findings indicate that the response to HG type 1.2.5.7 in dry bean is quantitative and associated with multiple genomic regions. The identified intervals provide candidate targets for independent validation, fine mapping, functional analysis, and future marker development to support breeding for SCN resistance.

Disease Resistance

Comparative Genome-Wide Association Studies of Metabolites and Grain-Related Traits in Common Wheat.

The metabolome is highly diverse and the closest layer to phenotype; therefore, it is commonly regarded as a bridge between the genome and phenome in plants. Here, we performed large-scale metabolome analysis using liquid chromatography-tandem mass spectrometry (LC-MS/MS) and 33 grain-related traits in a diverse panel of natural accessions and a recombinant inbred line (RIL) population. We identified a new network of 2286 associations between 947 metabolites and 33 grain-related traits. Systematic integration of metabolic genome-wide association study (mGWAS) and metabolic quantitative trait locus (mQTL) analyses identified 33 566 significant single-nucleotide polymorphisms (SNPs) and 3128 mQTL. Thirteen annotated metabolites co-localized within a physical interval on 7A. Integration of metabolite-based and phenotype-based GWAS and QTL revealed an overlapped region for gibberellin A4 (GA4) content and grain roundness on 4A. Phenotyping of an ethyl methanesulfonate (EMS)-induced mutant confirmed the role of TaSDR in regulating GA4 content and grain morphology. These findings provide novel insights into the metabolic pathways influencing key grain-related traits and advance our understanding of the complex molecular mechanisms regulating grain metabolites and phenotypes in wheat. The identified metabolic markers and candidate genes provide valuable targets for molecular breeding programs aimed at improving wheat yield and quality.

QTL

Identification of OsCsLF6 Gene Responsible for Rice Seed Submergence Germination Through Genome-Wide Association Analysis.

Flooding stress is a primary environmental barrier that severely limits the widespread adoption of direct-seeded rice systems. Under submerged conditions, rapid coleoptile elongation serves as a vital morphological strategy that facilitates anaerobic germination and successful seedling establishment, yet its underlying molecular mechanisms remain poorly understood. Through a genome-wide association study, we identified a critical locus governing anaerobic coleoptile elongation, in which OsCsLF6, encoding a mixed-linkage glucan (MLG) synthase, was characterized as the causal gene. Genetic and biochemical analyses demonstrated that OsCsLF6 positively regulated coleoptile elongation by directly mediating MLG deposition into the primary cell wall. Mechanistically, we identified OsERF74, an AP2/ERF transcription factor, as an upstream master repressor that directly binds to a conserved core cis-element within the OsCsLF6 promoter. Under submergence, OsERF74 deficiency (oserf74 mutants) completely releases this transcriptional suppression, triggering a substantial upregulation of OsCsLF6 expression and subsequent hyper-accumulation of cell wall MLG. In contrast, constitutive overexpression of OsERF74 persistently blocks MLG biosynthesis. Crucially, a natural single-nucleotide polymorphism located within the OsERF74 binding element in the promoter defines two distinct haplotypes. The elite haplotype (Hap1) effectively disrupts OsERF74 binding affinity, which in turn attenuates transcriptional repression and sustains high OsCsLF6 expression, ultimately driving accelerated MLG synthesis and coleoptile elongation. Our findings establish a condition-specific OsERF74-OsCsLF6 regulatory module that serves as a central biochemical hub orchestrating cell wall remodelling during anaerobic germination. This module thus represents a promising molecular target and elite genetic resource for molecular breeding of flood-tolerant and direct-seeded rice varieties.

OsCsLF6

Whole-Exome and Whole-Genome Sequencing of Candidate Pharmacogenomic and Schizophrenia-Related Genes in Sudanese Families with Schizophrenia.

BACKGROUND: Schizophrenia is considered a neuro-developmental disorder leading to disastrous lifelong disability of the patients and their families. There is a lack of data regarding pharmacogenomics of schizophrenia in Sudan. This study aimed to identify different genes affecting the treatment outcomes in Sudanese patients with schizophrenia. METHODS: A case-control study was conducted on seven families having more than one member diagnosed with schizophrenia. This was a small exploratory family-based sequencing study involving 18 affected individuals and 8 controls from seven families. Ethical clearance and informed consent were obtained. Demographic data were collected using a standardized data collection sheet. DNA was extracted from blood samples collected from patients and control groups. Then, whole-exome and genome sequencing were performed. Sixty-six genes associated with schizophrenia, treatment, and treatment resistance were selected from the variant calling file. Variants showing single-nucleotide polymorphisms (SNPs) were identified. These variants were then classified based on their impact on the protein-coding sequence into high- and moderate-impact. Moreover, indel mutations were also identified. RESULTS: Twelve variants of seven genes (COMT, FMO1, LPL, CYP2E1, ABCC1, GRM3, CYP2C9) were identified as genes with impact and potential association with schizophrenia (p-value=0.006632). Forty-three genes had a moderate impact, and they showed a potential association with schizophrenia (p-value=0.0004436). Two variants were indel mutations (CYP2D6, DTNBP1) and showed association with schizophrenia (p-value=0.004741). The p-values were generated from different databases. CONCLUSION: This exploratory family-based sequencing study identified several potentially relevant pharmacogenomic and schizophrenia-associated variants in Sudanese families, warranting validation in larger and ethnically diverse cohorts.

antipsychotics

Genetic associations in sepsis and ARDS.

Critical illness syndromes, such as sepsis and acute respiratory distress syndrome (ARDS), are characterized by substantial clinical heterogeneity and remain major causes of morbidity and mortality worldwide. Increasing evidence suggests that genetic variation contributes to susceptibility, disease severity, and clinical outcomes in critically ill patients. However, the molecular mechanisms linking genetic predisposition to the pathophysiology of sepsis and ARDS remain incompletely understood. In this review, we evaluated genetic associations reported in sepsis and ARDS, including 13 genome-wide studies identifying 19 unique single-nucleotide polymorphisms (SNPs) across 17 distinct genomic loci, as well as 21 meta-analyses of candidate-gene studies identifying 21 SNPs across 16 genes. The identified variants were primarily associated with pathways involved in pathogen recognition, immune and inflammatory signaling, leukocyte recruitment, and endothelial dysfunction. Collectively, these findings support a polygenic basis for susceptibility to critical illness and highlight several biologically relevant pathways that may contribute to sepsis and ARDS pathogenesis. Improved understanding of the functional consequences of these variants may facilitate the identification of potential therapeutic targets and support the development of precision-guided approaches to critical care.

ARDS

Genetic diversity, phylogenetic relationships, and marker development between Hydrangea serrata and H. macrophylla based on plastome and 45S nrDNA.

Ornamental hydrangeas (genus Hydrangea) are cultivated worldwide for their diverse flower colors and attractive morphology. Here, we assembled the complete plastid genome (plastome) and 45S nuclear ribosomal DNA (45S nrDNA) sequences of 22 individuals representing H. serrata, H. macrophylla, and related species (H. arborescens, H. paniculata, H. petiolaris, and H. hydrangeoides). The plastomes contained up to 2,344 single-nucleotide polymorphisms (SNPs) and 367 insertions/deletions (InDels) within the genus, whereas the assembled 45S nrDNA sequences showed 119 SNPs and 10 InDels. Phylogenetic analyses based on plastome and 45S nrDNA sequences clearly separated H. serrata and H. macrophylla from the other Hydrangea species. In the plastome-based tree, H. petiolaris was placed in the same clade as H. arborescens, whereas in the 45S nrDNA-based tree it showed a close relationship to H. hydrangeoides. The H. serrata and H. macrophylla samples were not always separated according to their species boundaries, as observed in samples Hse8-Hse12. Notably, one H. serrata sample (Hse8), collected from a wild mountainous region of Japan, exhibited a closer genetic relationship to H. macrophylla samples, indicating that cultivated hydrangeas may have originated from a specific wild lineage of H. serrata adapted to mountainous habitats. Using plastome-derived molecular markers, 66 Hydrangea samples were further classified into five groups, with Group II comprising both cultivated H. macrophylla and a subset of wild H. serrata samples, suggesting a close genetic affinity between this group and the ancestral gene pool of cultivated H. macrophylla. Based on these genomic resources, eight plastome-derived molecular markers were developed to differentiate cultivated hydrangeas from wild genotypes and to assess genetic diversity within H. serrata and H. macrophylla, providing practical tools for germplasm identification, breeding, and genetic resource management of Hydrangea species.

hydrangea

TET2 promotes monocyte inflammatory activation in asthma via ALKBH5-m6A regulation and PI3K signaling: evidence from m6A-SNP and single-cell analyses.

Asthma is a complex inflammatory airway disease with strong genetic determinants, yet the functional relevance of most asthma-associated non-coding variants remains unclear. Emerging evidence suggests that N6-methyladenosine (m6A) modification may serve as a critical epitranscriptomic link between genetic variation and immune regulation. In this study, we aimed to systematically identify functionally relevant m6A-regulated genes in asthma by integrating large-scale GWAS data, m6A-SNP annotations, and single-cell transcriptomic analyses, and to investigate their roles in monocyte-driven airway inflammation. We identified TET2 as a key m6A-regulated gene associated with both asthma and lung function, which was selectively upregulated in monocytes during asthma and accompanied by activation of inflammatory and PI3K signaling pathways. Mechanistic experiments further demonstrated that inflammatory stimulation induced ALKBH5 expression, reduced m6A modification of TET2 mRNA, and increased TET2 protein levels, thereby promoting PI3K/AKT signaling and pro-inflammatory cytokine production, whereas inhibition of TET2 or ALKBH5 attenuated these effects. Collectively, these findings demonstrate that ALKBH5-mediated m6A regulation of TET2 enhances PI3K/AKT signaling in monocytes, thereby promoting inflammatory responses in asthma. Our study establishes TET2 as a key m6A-regulated gene linking genetic susceptibility to monocyte-driven inflammation, and highlights the ALKBH5-m6A-TET2 axis as a potential therapeutic target for modulating aberrant immune responses in asthma.

Humans

Genome-wide annotation of human multi-nucleotide variants reveals widespread functional differences from single nucleotide variants.

Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.

Humans

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans