Search PubMedSearch

SEARCH · Search PubMed

Results for “SV”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans

Molecular differences between young and mature stria vascularis from organotypic explants and transcriptomics.

The stria vascularis (SV) is an essential component of the inner ear that regulates the ionic environment required for hearing. SV degeneration disrupts cochlear homeostasis, leading to irreversible hearing loss, yet a comprehensive understanding of the SV, and consequently therapeutic availability for SV degeneration, is lacking. We developed a whole-tissue explant model from neonatal and mature mice to create a platform for advancing SV research. We validated our model by demonstrating that the proliferative behavior of the SV in&#xa0;vitro mimics SV in&#xa0;vivo. We also provided evidence for pharmacological experimentation by investigating the role of Wnt/&#x3b2;-catenin signaling in SV proliferation. Finally, we performed single-cell RNA sequencing from in&#xa0;vivo neonatal and mature mouse SV and surrounding tissue and revealed key genes and pathways that may play a role in SV proliferation and maintenance. Together, our results contribute new insights into investigating biological solutions for SV-associated hearing loss.

Biochemistry

Sawfish: improving long-read structural variant discovery and genotyping with local haplotype modeling.

MOTIVATION: Structural variants (SVs) play an important role in evolutionary and functional genomics but are challenging to characterize. High-accuracy, long-read sequencing can substantially improve SV characterization when coupled with effective calling methods. While state-of-the-art long-read SV callers are highly accurate, further improvements are achievable by systematically modeling local haplotypes during SV discovery and genotyping. RESULTS: We describe sawfish, an SV caller for mapped high-quality long reads incorporating systematic SV haplotype modeling to improve accuracy and resolution. Assessment against the draft Genome in a Bottle (GIAB) SV benchmark from the T2T-HG002-Q100 diploid assembly shows that sawfish has the highest accuracy among state-of-the-art long-read SV callers across every tested SV size group. Additionally, sawfish maintains the highest accuracy at every tested depth level from 10- to 32-fold coverage, such that other callers required at least 30-fold coverage to match sawfish accuracy at 15-fold coverage. Sawfish also shows the highest accuracy in the GIAB challenging medically relevant genes benchmark, demonstrating improvements in both comprehensive and medically relevant contexts.When joint-genotyping seven samples from CEPH-1463, sawfish has over 9000 more pedigree-concordant calls than other state-of-the-art SV callers, with the highest proportion of concordant SVs (81%). Sawfish's quality model enables selection for an even higher proportion of concordant SVs (88%), while still calling nearly 5000 more pedigree-concordant SVs than other callers. These results demonstrate that sawfish improves on the state-of-the-art for long-read SV calling accuracy across both individual and joint-sample analyses. AVAILABILITY AND IMPLEMENTATION: Sawfish source code, pre-compiled Linux binaries, and documentation are released on GitHub: https://github.com/PacificBiosciences/sawfish.

Haplotypes

insilicoSV: a flexible grammar-based framework for structural variant simulation and placement.

SUMMARY: Structural variants (SVs) are key drivers of genetic variation and disease in the genome. Their discovery remains challenging, however, in large part due to the scarcity of validated SV callsets and comprehensive benchmarks, which are essential for method development and evaluation. The growing number of data-driven learning-based approaches for SV discovery, in particular, requires large, diverse, and well-balanced training datasets to achieve reliable performance. To address this need, SV simulation has served as a key tool for assessing method performance and training SV models. However, existing SV simulators only support a fixed and limited set of SV classes and do not provide fine-grained control over the placement of SVs within specific contexts of the genome. Here we present insilicoSV, a versatile framework for SV simulation, which models SVs using a simple and flexible grammar, allowing users to easily define standard and custom arbitrary genome rearrangements, as well as encode genome placement constraints. This design allows insilicoSV to naturally support new and bespoke SV types, such as the complex rearrangements of cancer genomes. In addition to grammar-based modeling, insilicoSV provides built-in support for 26 predefined SV types, placement of user-provided SVs, small variant simulation, streamlined workflows for the simulation of genome evolution and genome mixtures, read simulation, alignment, and visualization. These features enable the creation of comprehensive genomic datasets for a variety of downstream applications, such as in-depth benchmarking of alignment and variant calling methods, as well as training of data-driven learning-based approaches for SV detection. AVAILABILITY AND IMPLEMENTATION: insilicoSV is available under the MIT license at https://github.com/PopicLab/insilicoSV and https://doi.org/10.5281/zenodo.17402009.

Software

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software

Panorama of Chromosomal Instability in Lung Cancer.

Lung cancer is a highly heterogeneous disease primarily driven by tobacco smoking. About 20% of lung cancers occur among patients who have never smoked (LCINS) with differences in patient ancestry, sex, tumor histology, and clinical features. Our understanding of chromosomal instability in lung cancer, especially LCINS, is still limited. Here, we perform a comprehensive study of 182,429 somatic structural variations (SVs) detected in 1,209 whole-genome sequenced lung cancers, of which 864 LCINS. SVs are more abundant in tumors from patients who have smoked (LCSS); however, they are more complex and play more important roles in tumorigenesis in LCINS. EGFR mutations and KRAS mutations profoundly and independently shape the SV landscape. EGFR-mutant tumors have higher SV burden and more cancer-driving SVs. In contrast, KRAS mutations are associated with lower SV burden and less driver SVs. We decompose 16 SV signatures for both complex and simple SVs that likely represent divergent molecular mechanisms. The SV breakpoints have distinct distributions across the genome depending on the signatures due to mutagenic mechanisms and positive selection. Many established cancer-driving genes are recurrently rearranged by multiple SV signatures suggesting functional convergence of these genome instability mechanisms.

Journal Article

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

Comprehensive benchmarking of somatic structural variant detection at ultra-low allele fractions.

Postzygotic mosaicism gives rise to somatic structural variants (SVs) at ultra-low variant allele fractions (VAFs), which pose challenges for detection due to the high-coverage sequencing required and noise introduced by sequencing artifacts. Although somatic SV detection has been extensively studied in cancer, these studies are not directly applicable to the study of tissue mosaicism, as they rely on matched normals, target higher VAF ranges, and are enriched for different types of SVs. We present comprehensive benchmark data and best practices for non-cancer somatic SV detection. We created a synthetic mosaic sample by combining six HapMap individuals at varying proportions, generating allele fractions as low as 0.25%. This sample was sequenced to ~2,300x total coverage using Illumina, PacBio, and Nanopore technologies across multiple sequencing centers. A high-confidence benchmark SV set containing over 21,000 pseudo-somatic insertions and deletions &#x2265;50bp was derived from haplotype-resolved assemblies. We evaluated 12 SV discovery pipelines and identified caller-specific strengths and sequencing platform-specific shortcomings. We find that short read-based approaches show reduced recall for insertions and repeat-associated SVs, whereas long-read sequencing achieves high accuracy throughout the genome, increasing linearly with coverage. The best algorithm's sensitivity exceeded 80% for VAFs &#x2265;4% and 15% for VAFs of 0.5-1% with 60x coverage. The publicly available benchmarking data and comparative analysis of current methods provide a foundation for robust discovery of SV mosaicism in non-cancer tissues..

Journal Article

Severus detects somatic structural variation and complex rearrangements in cancer genomes using long-read sequencing.

For the detection of somatic structural variation (SV) in cancer genomes, long-read sequencing is advantageous over short-read sequencing with respect to mappability and variant phasing. However, most current long-read SV detection methods are not developed for the analysis of tumor genomes characterized by complex rearrangements and heterogeneity. Here, we present Severus, a breakpoint graph-based algorithm for somatic SV calling from long-read cancer sequencing. Severus works with matching normal samples, supports unbalanced cancer karyotypes, can characterize complex multibreak SV patterns and produces haplotype-specific calls. On a comprehensive multitechnology cell line panel, Severus consistently outperforms other long-read and short-read methods in terms of SV detection F1 score (harmonic mean of the precision and recall). We also illustrate that compared to long-read methods, short-read sequencing systematically misses certain classes of somatic SVs, such as insertions or clustered rearrangements. We apply Severus to several clinical cases of pediatric leukemia/lymphoma, revealing clinically relevant cryptic rearrangements missed by standard genomic panels.

Humans

GiGCN: a network-based framework for uncovering synthetic lethal and viable genetic interactions.

Genetic interactions (GIs) underpin the functional connectivity of genes and pathways, and are important for dissecting genotype-phenotype relationships and identifying therapeutic targets for diseases. However, the scale of the human genome restricts systematic experimental interrogation of GIs. Existing computational tools focus on predicting synthetic lethality (SL) and synthetic viability (SV), the two primary forms of GIs, yet their accuracy and biological interpretability are compromised by inadequate modeling of the molecular mechanisms behind positive and negative interactions, as well as the limitation of negative samples. To overcome these challenges, we developed Genetic Interaction Graph Convolutional Network (GiGCN), a signed network modeling framework for the joint identification of gene pairs with SL and SV. We built a high-confidence signed genetic network by integrating verified GIs, and non-interacting gene pairs, together with gene semantic similarity derived from biological processes. By leveraging disentangled subspace decomposition, this framework separately models distinct functional dimensions within gene networks, enabling robust representation of context-dependent regulatory relationships and accurate discrimination of SL and SV events. Benchmark experiments demonstrate that GiGCN outperforms state-of-the-art approaches (area under receiver operating-characteristic curve: 0.978, and area under precision-recall curve: 0.944). Further analyses reveal biologically meaningful insights, including known and novel SL interactions centered on the oncogene MYC Proto-Oncogene (MYC), as well as SV interactions linked to autophagy and mitophagy pathways. This study provides a robust and interpretable network-based strategy for systematically exploring GIs. The GiGCN framework not only improves the precision of SL and SV prediction, but also offers mechanistic insights into gene functional relationships, thereby supporting the discovery of actionable therapeutic targets for cancer and other human diseases.

Humans

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics

The landscape of structural variation in pediatric cancer.

Structural variants (SVs) account for over 60% of the driver variants in pediatric cancer, and in many cases act as the cancer initiating event. To study SVs from a pan-cancer perspective, we analyzed 1,616 pediatric cancer genomes in 16 major cancer types of hematological malignancies (n = 908), brain tumors (n = 183), and solid tumors (n = 525) and compared their profiles to those of 2,203 adult cancers. The SV burden varied ~100-fold across pediatric cancer types and demonstrated an 8- to 16-fold reduction compared to adult brain and solid tumors but was comparable in pediatric versus adult hematological malignancies. Recurrent SV hotspots occurred uniquely in pediatric acute lymphoblastic leukemias (ALLs) in proximity to RAG-mediated recombination signal sequences (RSS) and disrupted multiple immune-related loci as well as 69 genes, which often involved cryptic RSS sites. By contrast, such hotspots affected only immune-related loci but not driver genes in adult lymphoid cancers. Eight SV signatures extracted from the cohort had varying distributions across cancer types, with clustered translocations reflecting templated insertions in osteosarcoma, and medium-sized deletions (10 kb to 1 Mb) enriched in cancers with RAG-mediated deletions. Intra-patient evolutionary analysis in 13 patients with multiple spatiotemporally distinct samples revealed that RAG-mediated recombination in leukemia and complex rearrangements in solid tumors occurred both early in disease initiation and continuously during later diversification, contributing to clonal heterogeneity. Finally, we found that both driver genes and fragile sites were the two genomic regions most frequently disrupted by SVs. The unique and diverse SV landscapes that emerged from this comprehensive analysis expand the scope of RSS-mediated mutagenesis in pediatric ALL and will be a valuable resource for guiding future functional studies and the design of clinical genomic testing in pediatric cancer.

Journal Article

Refining the genetic diagnostic puzzle: A case report on a Chinese ARPKD patient with a reciprocal balanced translocation and c.2507&#x2009;T&#x2009;>&#x2009;C (p.V836A) in PKHD1.

INTRODUCTION: Autosomal recessive polycystic kidney disease (ARPKD) ranks among the most severe chronic kidney diseases (CKD). Its primary cause is variants in the Polycystic Kidney and Hepatic Disease 1 gene (PKHD1). The clinical spectrum of ARPKD varies widely, ranging from mild late-onset symptoms to severe perinatal mortality. However, achieving an early genetic diagnosis in ARPKD patients before clinical symptoms appear proves challenging. CASE PRESENTATION: This case is a 4-year-old boy who experienced a convulsion characterized by a generalized tonic attack lasting approximately 3-5 minutes and later sought treatment to our hospital. However, routine abdominal ultrasound examination accidentally detected that he had diffuse liver lesions, splenomegaly, and bilateral renal enlargement with renal pelvis dilation. Given the uncertainty regarding the underlying cause of the patient's structural abnormalities and convulsions, karyotyping, whole exome sequencing (WES), structural variant analysis (SV analysis) of whole genome sequencing (WGS) were recommended. The result of SV analysis revealed that he has an RBT impacting PKHD1 and the precise location of breakpoints was confirmed through Long-Range Polymerase Chain Reaction (LR-PCR). However, WES did not screen out pathogenic variants initially, the WES data was reviewed subsequently based on SV analysis results. CONCLUSION: We identified an infrequent variant combination, c.2507T>C (p.V836A) in PKHD1 and an RBT with broken PKHD1, which extends the genetic spectrum of ARPKD, and provide a basis for further genetic counselling to the family.

Humans

Longitudinal ctDNA tracking in early and recurrent breast cancer using an ultrasensitive structural variant-based assay: an extended analysis from the TRACER study.

BACKGROUND: Detection of circulating tumor DNA (ctDNA) following curative-intent therapy is prognostic of disease recurrence in early-stage breast cancer (EBC). An ultrasensitive structural variant (SV)-based ctDNA assay was evaluated previously in a 100-patient EBC cohort treated with neoadjuvant therapy, demonstrating high sensitivity, specificity, and a long lead-time to relapse. The stability of primary tumor-specific SVs at and after metastatic recurrence and their utility for longer-term ctDNA monitoring had not been established. PATIENTS AND METHODS: An updated retrospective analysis of ctDNA dynamics was conducted in an expanded cohort of 121 patients with EBC treated with neoadjuvant therapy. Plasma samples were collected at key clinical timepoints and serially in several patients who experienced metastatic recurrence. Clinical variables were abstracted from medical records. Associations between ctDNA detection, dynamics, and clinical outcomes were evaluated in the early-stage and metastatic settings. RESULTS: Thirty of 121 patients experienced clinical recurrence (28 distant, 2 local) over a median follow-up of 4.2 years (range 0.5-8.8; 25 ctDNA evaluable with adjuvant timepoints). All patients with detectable ctDNA in the adjuvant setting developed metastatic recurrence (22/22). Median lead time from ctDNA detection to metastatic recurrence was 346 days (range 0-1937). Among recurrent cases, 79% of primary tumor-specific SVs (n = 17 patients, tumor fraction &#x2265;0.1%) remained detectable in plasma [range 7% (1/14 SV)-100% (15/15); median: 92%]. ctDNA dynamics in the recurrent metastatic setting demonstrated a strong relationship with radiographic outcomes in evaluable patients (n = 9). CONCLUSION: This SV-based digital PCR assay provided ultrasensitive ctDNA detection in an expanded EBC cohort, maintaining 100% positive predictive value for metastatic recurrence. In patients with recurrence, ctDNA dynamics were concordant with radiographic outcomes. Prospective studies evaluating the clinical utility of longitudinal ctDNA monitoring are warranted.

MRD

Elucidating the evolution of meat quality, water distribution, microstructure, and protein structure during sous-vide and micro-pressure cooking.

This study investigated the evolution of eating quality (colour, texture and volatile flavour compounds), water status, microstructure and protein structure of pork meat under different cooking methods. The methods analysed included traditional cooking (TC: 10, 20, 30 and 40&#xa0;min, 100&#xa0;&#xb0;C), sous-vide cooking (SV: 1, 2, 3 and 4&#xa0;h, 60&#xa0;&#xb0;C) and micro-pressure cooking (MC: 10, 20, 30 and 40&#xa0;min, 120&#xa0;&#xb0;C). Across the three cooking processes, as cooking time increased, cooking loss, lightness, yellowness, P23, &#x3b2;-sheet, random coil and surface hydrophobicity of the meat samples increased. By contrast, redness, P22, hydrogen proton density, esters content, &#x3b1;-helix, &#x3b2;-turn and sulfhydryl group content decreased. Moreover, the Warner-Bratzler shear force (WBSF), adhesiveness, hardness, springiness, gumminess, chewiness, alcohols, aldehydes, ketones and fluorescence intensity of the meat samples, initially increased and then decreased as cooking progressed. SV resulted in higher water-holding capacity (WHC), improved redness and increased alcohol and ester levels, whereas MC produced softer meat and greater water mobility. Furthermore, MC enhanced the degree of microstructural damage and protein structural unfolding in the meat. MC requires less time to achieve textures and flavours similar to those obtained using the TC and SV methods. Thus, MC is an efficient cooking method for the catering industry to obtain desired meat quality rapidly.

Cooking

Allelic variation and light-responsive regulation of FaMYB10-2 underlie tissue-specific anthocyanin accumulation in strawberry.

Anthocyanins critically determine fruit color, nutrition, and stress resilience in cultivated strawberry (Fragaria &#xd7; ananassa), directly influencing consumer preference. Despite complex genetic and environmental regulation of their biosynthesis, the basis for tissue-specific pigmentation, notably the widespread occurrence of red skin and pale flesh, remains poorly understood. We integrated genomic, transcriptomic, and functional analyses across 200 cultivars to dissect receptacle pigmentation regulation. Approaches included FaMYB10-2 allele mining, promoter structural variant (SV) identification, expression profiling, regulatory interaction assays, and characterization of upstream light-responsive factors. FaMYB10-2 was identified as the key R2R3-MYB regulator of fruit anthocyanin biosynthesis. Alleles FaMYB10-2.2 and FaMYB10-2.3 encode truncated proteins retaining bHLH-binding capacity but lacking activation domains, functioning as dominant-negative repressors. A promoter SV 986&#x2005;bp upstream of FaMYB10-2 was associated with reduced pale fruit due to cis-regulatory divergence. The SV (Alt) allele is prevalent in Asian cultivars, while the Ref allele is enriched in Western germplasm. Crucially, a light-responsive FaHYH-FaWRKY71 cascade activates FaMYB10-2 and structural genes haplotype-dependently, compensating for weak MYB activity in the skin. Our findings reveal a multilayered regulatory system integrating allelic variation, cis-regulatory divergence, and environmental signals, advancing anthocyanin understanding and providing engineering targets for polyploid crop color improvement.

Fragaria

Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci.

The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. The point of template switching in 4 samples was shown to be a segment of &#x223c;2.2-5.5 kb of 100% nucleotide similarity within inverted repeat pairs. These data provide experimental evidence that inverted low-copy repeats act as recombinant substrates. This type of CGR can result in multiple conformers generating diverse SV haplotypes in susceptible dosage-sensitive loci.

Humans