Search PubMedSearch

SEARCH · Search PubMed

Results for “Structural variants”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Molecular residual disease assessment in colorectal and bladder cancer by somatic structural variant analysis of cell-free DNA whole-genome sequencing data.

BACKGROUND: Whole-genome sequencing (WGS)-based methods for circulating tumor DNA (ctDNA) detection typically rely on tumor-informed identification of somatic single nucleotide variants (SNVs). Somatic structural variants (SVs) are another type of cancer-specific genomic alteration, which owing to their larger genomic footprint and unique breakpoint junctions, are easier to distinguish from sequencing noise than SNVs. They are, however, rarely used for ctDNA detection because of (1) artifacts from WGS procedures that SV callers may falsely interpret as genuine SVs. This makes it difficult to establish high-confidence SV catalogos from short-read tumor WGS and can cause false-positive ctDNA detections. (2) Lack of robust strategies to quantify SV-supporting reads in plasma WGS. To address these barriers and enable integration of SV biomarkers into WGS-based ctDNA detection, we present a bioinformatic framework for algorithmic curation of somatic SV calls from fresh-frozen and formalin-fixed paraffin-embedded (FFPE) tumors, coupled with a novel approach for sensitive, accurate mapping and quantification of SV breakpoint-supporting reads in plasma WGS. METHODS: Tumor, normal and plasma WGS data from 144 patients with stage III colorectal cancer was used to establish the bioinformatic framework. This included ~30x WGS data from 1564 serially collected plasma samples. The framework was validated using tumor/normal/plasma WGS data from 32 patients with muscle-invasive bladder cancer. SV-based ctDNA detection was benchmarked against previously published SNV-based ctDNA results for the same samples. RESULTS: After curation of SV calls and quantification in plasma WGS, our SV-based approach enabled robust ctDNA detection with overall specificity exceeding 99% in plasma samples. Furthermore, we observed strong concordance (Pearson&#x2019;s r&#x2009;>&#x2009;0.93, p&#x2009;<&#x2009;2.2&#x2009;&#xd7;&#x2009;10&#x2212; 16) between ctDNA-positive samples identified by our SV-based method and previous SNV-based analyses, validating the reliability of our approach. Finally, we demonstrated application of the method in an independent bladder cancer cohort, highlighting its generalizability and potential clinical use. CONCLUSIONS: We provide a bioinformatic framework that establishes somatic SVs as ultra-specific biomarkers for WGS-based, tumor-informed ctDNA detection. The approach delivers specific detection even when the SV catalogos are established from FFPE samples. The SV framework can stand alone or enhance SNV-based analysis pipelines.

Humans

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

insilicoSV: a flexible grammar-based framework for structural variant simulation and placement.

SUMMARY: Structural variants (SVs) are key drivers of genetic variation and disease in the genome. Their discovery remains challenging, however, in large part due to the scarcity of validated SV callsets and comprehensive benchmarks, which are essential for method development and evaluation. The growing number of data-driven learning-based approaches for SV discovery, in particular, requires large, diverse, and well-balanced training datasets to achieve reliable performance. To address this need, SV simulation has served as a key tool for assessing method performance and training SV models. However, existing SV simulators only support a fixed and limited set of SV classes and do not provide fine-grained control over the placement of SVs within specific contexts of the genome. Here we present insilicoSV, a versatile framework for SV simulation, which models SVs using a simple and flexible grammar, allowing users to easily define standard and custom arbitrary genome rearrangements, as well as encode genome placement constraints. This design allows insilicoSV to naturally support new and bespoke SV types, such as the complex rearrangements of cancer genomes. In addition to grammar-based modeling, insilicoSV provides built-in support for 26 predefined SV types, placement of user-provided SVs, small variant simulation, streamlined workflows for the simulation of genome evolution and genome mixtures, read simulation, alignment, and visualization. These features enable the creation of comprehensive genomic datasets for a variety of downstream applications, such as in-depth benchmarking of alignment and variant calling methods, as well as training of data-driven learning-based approaches for SV detection. AVAILABILITY AND IMPLEMENTATION: insilicoSV is available under the MIT license at https://github.com/PopicLab/insilicoSV and https://doi.org/10.5281/zenodo.17402009.

Software

OctopuSV and TentacleSV: a one-stop toolkit for multi-sample, cross-platform structural variant comparison and analysis.

MOTIVATION: Structural variants (SVs) influence gene regulation, disease progression, and diagnostics, yet integrating SV calls across platforms remains difficult due to inconsistent annotations, limited merging flexibility, and fragmented workflows. Ambiguous breakend (BND) annotations, which comprise many variant calls, are often discarded or misclassified, hindering variant characterization. Existing tools lack advanced merging operations essential for precise identification of disease-specific or somatic variants across samples or patient groups. Additionally, current SV analysis pipelines require extensive manual intervention and complex parameter tuning, compromising reproducibility and scalability. Addressing these gaps is crucial for improving the accuracy, interpretability, and clinical utility of SV analyses. RESULTS: We developed OctopuSV and TentacleSV to address these long-standing challenges in SV analysis. OctopuSV features a specialized BND correction module that converts ambiguous BND annotations into canonical SV types, recovering important variants that are often overlooked by existing tools. Additionally, it provides advanced set operations (difference, complement, custom-defined) that enable sophisticated variant filtering without programming expertise, critical for identifying tumor-specific SVs or variants unique to specific sample groups. TentacleSV completes our solution by automating the entire SV analysis process from raw sequencing data to high-confidence callsets, ensuring consistency and reproducibility across projects. Benchmarking across short-read and long-read platforms showed superior F1 score, complete SV type consistency compared to existing tools. Our framework enables experimental biologists and clinical researchers to perform sophisticated analyses ranging from cancer subtype-specific SV identification to multi-sample comparative studies without requiring specialized programming skills. AVAILABILITY AND IMPLEMENTATION: All codes are available at https://github.com/ylab-hi/OctopuSV; https://github.com/ylab-hi/TentacleSV.

Software

The human IG heavy chain constant gene locus is enriched for large structural variants and coding polymorphisms that vary among human populations.

The human immunoglobulin heavy chain constant (IGHC) domain of antibodies (Ab) is responsible for effector functions critical to immunity. This domain is encoded by genes in the IGHC locus, where descriptions of genomic diversity remain incomplete. We utilized long-read sequencing to build an IGHC haplotype/variant catalog from 105 individuals of diverse ancestry. We discovered uncharacterized single nucleotide variants (SNV) and large structural variants (SVs, n=7), representing new genes and alleles enriched for non-synonymous substitutions, highlighting potential functional effects. Of the 221 identified IGHC alleles, 192 were novel. SNV, SV, and gene allele/genotype frequencies revealed population differentiation, including (i) hundreds of SNVs in African and East Asian populations exceeding a fixation index (FST) of 0.3, and (ii) an IGHG4 haplotype carrying coding variants uniquely enriched in Asian populations. Our results illuminate missing signatures of IGHC diversity and establish a new foundation for investigating IGHC germline variation in Ab function and disease.

Journal Article

Non-coding single-nucleotide and structural variants affecting the EYS putative promoter cause autosomal recessive retinitis pigmentosa.

PURPOSE: Variants in untranslated genomic regions are difficult to identify as pathogenic but are capable of causing disease by interfering with gene expression. This study aimed to characterize the effect of variants identified in the 5'-untranslated region of EYS in patients with autosomal recessive retinitis pigmentosa (RP). METHODS: Variant screening included gene panels, Sanger, exome, and genome sequencing. Functional validation included an electrophoretic mobility shift assay and various luciferase assays. RESULTS: Patients with RP from 6 EYS biallelic Arab-Muslim families harbored a 5' noncoding EYS variant, c.-453G>T, and 4 harbored a structural variant affecting the 5' noncoding exons. Electrophoretic mobility shift assay analysis revealed an effect on binding of transcription factors for c.-453G>T and a neighboring variant c.-454G>T. Dual luciferase assays using overexpression of various transcription factors showed distinct effects on expression. c.-453G>T was associated with higher luciferase expression with CRX overexpression and c.-454G>C with OTX2 overexpression. In addition, the 2 variants were found to influence translation by affecting upstream initiation codons. Interestingly, visual function of EYS RP patients who harbor c.-453G>T are better than those with biallelic null EYS variants. CONCLUSION: Our analysis revealed both single-nucleotide and structural variants in the EYS promoter as the cause of autosomal recessive RP. These variants may affect EYS expression via a dual mechanism by altering transcription factor binding affinity at the EYS promoter and by affecting upstream open reading frames.

Humans

SynFlow: an interactive online genome structural variant viewer.

MOTIVATION: Structural variations (SVs), including inversions, translocations (TRAs), duplications, and large insertions or deletions, are key drivers of genome evolution and phenotypic diversity. With the increasing number of high-quality, chromosome-scale genome assemblies, the ability to detect and interpret SVs has become a crucial aspect of modern genomics. While SV detection has advanced, most visualization methods produce static plots that fall short when researchers, particularly in comparative genomics, need to interactively explore large datasets, zoom into specific genomic regions, or dynamically filter structural events in real time. RESULTS: To address this gap, we introduce SynFlow, a lightweight, web-based interactive application specifically designed for exploring and visualizing SVs identified by SyRI. We demonstrate that SynFlow can reproduce complex static synteny plots published in literature, but transforms them into dynamic, shareable visualizations that support real-time filtering, reordering, and deep exploration of specific SVs, including TRAs. SynFlow is available as a web server and offers multiple entry points: browsing precomputed datasets (e.g. banana and grapevine genomes), uploading user-provided SyRI outputs, or running an integrated workflow to produce and visualize SVs on the fly. AVAILABILITY AND IMPLEMENTATION: https://synflow.southgreen.fr; source code https://github.com/SouthGreenPlatform/synflow; preprocessing Snakemake workflow https://gitlab.cirad.fr/agap/cluster/snakemake/synflow.

Software

Human serum amyloid A (SAA): biosynthesis and postsynthetic processing of preSAA and structural variants defined by complementary DNA.

To study structural variants of human serum amyloid A (SAA), an apoprotein of high-density lipoprotein, complementary DNA clones were isolated from a human liver library with the use of two synthetic oligonucleotide mixtures containing sequences that could code for residues 33-38 and 90-95 of the protein sequence. The SAA-specific cDNA clone (pA1) contains the nucleotide sequence coding for the mature SAA and 10 amino acids of the 18-residue signal peptide. It also includes a 70 nucleotide long 3'-untranslated region and approximately 120 bases of the poly(A) tail. The derived amino acid sequence of pA1 is identical with the alpha form of apoSAA1. A fragment of pA1 containing the conserved (residues 33-38) region of SAA also hybridized with RNA from human acute phase liver and acute phase stimulated, but not unstimulated, mouse and rabbit liver. In contrast, a fragment corresponding to the variable region hybridized to a much greater extent with human than with rabbit or murine RNA. Human acute phase liver SAA mRNA (approximately 600 nucleotides in length) directs synthesis of preSAA (Mr 14 000) in a cell-free translating system. In a Xenopus oocyte translation system preSAA is synthesized and processed to the mature Mr 12 000 product. The complete 18 amino acid signal peptide sequence of preSAA was derived from sequencing cDNA synthesized by "primer extension" from the region of SAA mRNA corresponding to the amino terminus of the mature product. Two other SAA-specific cDNA clones (pA6 and pA10) differed from pA1 in that they lack the internal PstI restriction enzyme site spanning residues 54-56 of pA1.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

The value of structural variants to conservation genomics in the pangenome era.

Structural variants (SVs) comprise an axis of genetic diversity with strong consequences for phenotype and fitness, making them a potentially important target for conservation genomics. Here, we review how and why SVs can play a role in conservation genomics; the different types of SVs and how they can affect phenotype; and how pangenomes and long-read sequencing are illuminating their evolution in populations, including small populations and those of conservation concern. SVs comprise multinucleotide mutations including insertions, deletions, transpositions, inversions, and other multinucleotide mutations, often overlapping genes and other functional genome regions. As a result, SVs often play important roles in phenotypic evolution and local adaptation and can contribute substantially to genetic load in inbred populations. However, our understanding of the factors influencing SV diversity in populations is still in its infancy and is complicated by the vast range of sizes, effects, and mechanisms of formation of these mutations. We argue that SVs are an important axis of genetic diversity which should be characterized alongside more traditional metrics of genetic diversity in conservation contexts. There are a number of analytical challenges to detecting and studying SVs, but analyses aimed at understanding the role of SVs in inbreeding load and population health are rapidly becoming realizable goals, accelerated by new technologies and analytical approaches. New tools, including population-scale long-read sequencing and pangenome approaches, are beginning to make SVs accessible in ways which can be readily applied in conservation settings.

Genomic Structural Variation

Sawfish: improving long-read structural variant discovery and genotyping with local haplotype modeling.

MOTIVATION: Structural variants (SVs) play an important role in evolutionary and functional genomics but are challenging to characterize. High-accuracy, long-read sequencing can substantially improve SV characterization when coupled with effective calling methods. While state-of-the-art long-read SV callers are highly accurate, further improvements are achievable by systematically modeling local haplotypes during SV discovery and genotyping. RESULTS: We describe sawfish, an SV caller for mapped high-quality long reads incorporating systematic SV haplotype modeling to improve accuracy and resolution. Assessment against the draft Genome in a Bottle (GIAB) SV benchmark from the T2T-HG002-Q100 diploid assembly shows that sawfish has the highest accuracy among state-of-the-art long-read SV callers across every tested SV size group. Additionally, sawfish maintains the highest accuracy at every tested depth level from 10- to 32-fold coverage, such that other callers required at least 30-fold coverage to match sawfish accuracy at 15-fold coverage. Sawfish also shows the highest accuracy in the GIAB challenging medically relevant genes benchmark, demonstrating improvements in both comprehensive and medically relevant contexts.When joint-genotyping seven samples from CEPH-1463, sawfish has over 9000 more pedigree-concordant calls than other state-of-the-art SV callers, with the highest proportion of concordant SVs (81%). Sawfish's quality model enables selection for an even higher proportion of concordant SVs (88%), while still calling nearly 5000 more pedigree-concordant SVs than other callers. These results demonstrate that sawfish improves on the state-of-the-art for long-read SV calling accuracy across both individual and joint-sample analyses. AVAILABILITY AND IMPLEMENTATION: Sawfish source code, pre-compiled Linux binaries, and documentation are released on GitHub: https://github.com/PacificBiosciences/sawfish.

Haplotypes

Variability of human rRNA genes: inheritance and nonrandom chromosomal distribution of structural variants of nontranscribed spacer sequences.

Human rRNA genes contain variable regions, one of which is located in nontranscribed spacers (NTSs) closely downstream from the 3'-end of the transcribed region. This polymorphism may be detected by means of blot hybridization analysis as a set of distinct restriction fragments corresponding to this part of the rRNA genes. We have analyzed DNA of 51 individuals and found eight structural NTS variants of this region; two of these were common to all individuals analyzed, and six others were found in different combinations and with different frequencies. The copy number of each variant also differed but was not less than 10-20 copies per cell. The analysis of DNA isolated from leukocytes of the members of 11 families indicated that some of the structural variants (of the NTS region) are inherited as a single Mendelian locus. We propose that rRNA genes that belong to one particular structural variant form clusters on separate chromosomes. To test this proposition, we developed a combined method, including AgNO3-staining of chromosomes, in situ hybridization, and DNA analysis with methylation-sensitive restrictases, and used it for study of persons who had methylated rRNA genes located on AgNO3-negative nucleolar organizers. It was found that in three of four cases methylated genes really belonged to one structural variant. This approach may be used for detailed localization of separate classes of NTS structural variants of human rRNA genes.

Blotting, Southern

cDNA clones encoding murine IgE-binding factors represent multiple structural variants of intracisternal A-particle genes.

Previously [Moore, K. W., Jardieu, P., Mietz, J. A., Trounstine, M. L., Kuff, E. L., Ishizaka, K. & Martens, C. L. (1986) J. Immunol. 136, 4283-4290], we examined a T-hybridoma-derived cDNA clone, 8.3, that encodes a biologically active murine IgE-binding factor (IgE-BF), and we showed that it was a variant member of the endogenous retroviral gene family related to mouse intracisternal A particles (IAPs). We have now characterized four more IgE-BF cDNA clones by heteroduplex and restriction enzyme analysis and found that they all represent different structural variants of the full-size IAP genomic element. In clones 8.3 and 10.2, which have been fully sequenced, the open reading frames span deletions 3.4 and 1.9 kilobases (kb) long, respectively, and specify different gag-pol fusion polypeptides. Clone 9.5 contains a 2.1-kb deletion entirely within the pol region. Two other clones (4.2 and 11.7) contain no internal deletion and may represent truncated cDNA copies of full-size (7.2 kb) IAP gene transcripts. Structural variants very similar to clone 10.2 are common in the mouse genome, and clone 9.5 is also probably not a unique gene form. The sequences of clones 8.3 and 10.2 are different in detail, but each is closely homologous to a randomly cloned mouse genomic IAP element throughout the gag-related portions of their open reading frames. Antibodies against two oligopeptides specified by the sequence of clone 8.3 immunoprecipitated IAP-related proteins from mouse neuroblastoma and myeloma cells, confirming that the IgE-BF produced by this clone shares sequence with expressed IAP elements in different cell types. Thus, information related to the IgE-BF is an integral part of the murine IAP retrotransposon gag gene.

Amino Acid Sequence

Complex de novo structural variants are an underestimated cause of rare disorders.

Complex de novo structural variants (dnSVs) are crucial genetic factors in rare disorders, yet their prevalence and characteristics in rare disorders remain poorly understood. Here, we conduct a comprehensive analysis of whole-genome sequencing data of 12,568 families, including 13,698 offspring with rare diseases, obtained as part of the UK 100,000 Genomes Project. We identify 1,870 dnSVs, constituting the largest dnSV dataset reported to date. Complex dnSVs (n&#x2009;=&#x2009;158; 8.4%) emerge as the third most common type of SV, following simple deletions and duplications. We classify 65% of these complex dnSVs into 11 subtypes. Among probands with dnSVs (n&#x2009;=&#x2009;1,696), 9% exhibit exon-disrupting pathogenic dnSVs associated with the probands' phenotype. Notably, 12% of exon-disrupting pathogenic dnSVs and 22% of de novo deletions or duplications previously identified by array-based or whole-exome sequencing methods are found to be complex dnSVs. We also find distinct genomic properties of de novo deletions depending on the parent of origin. This study highlights the importance of complex dnSVs in the cause of rare disorders and demonstrates the necessity of specific genomic analysis to avoid overlooking these variants.

Humans

Structural variants of human T200 glycoprotein (leukocyte-common antigen).

Structural variation in the primary structure of human T200 glycoprotein has been detected. Three cDNA variants have been characterized each of which encode T200 molecules that differ in size as a result of sequence differences in their amino-terminal regions. The largest form of the molecule is distinguished from the smallest by an insert of 161 amino acids, after the first eight amino-terminal residues. The other variant has an insert at the same location of 47 amino acids identical to residues 75-121 in the larger insert. Both extra domains are rich in serine and threonine residues and are likely to display multiple O-linked oligosaccharides. These structural variants which probably arise by cell-type-specific alternative splicing provide a molecular basis for the previously observed structural and antigenic heterogeneity of T200 glycoprotein. In addition to the variable amino-terminal region, the external domain of human T200 glycoprotein consists of a second cysteine-rich region of about 400 amino acids, a single transmembrane-spanning region and a large cytoplasmic domain of 707 amino acids shared by all of the structural variants and highly conserved between species. The gene encoding human T200 is located on the long arm of chromosome 1.

Amino Acid Sequence

needLR: long-read structural variant annotation with population-scale frequency estimation.

SUMMARY: We present needLR, a structural variant (SV) annotation tool that can be used for filtering and prioritization of candidate pathogenic SVs from long-read sequencing data using population allele frequencies, annotations for genomic context, and gene-phenotype associations. When using population data from 500 presumably healthy individuals to evaluate nine test cases with known pathogenic SVs, needLR assigned allele frequencies to over 97.5% of all detected SVs and reduced the average number of novel genic SVs to 121 per case while retaining all known pathogenic variants. AVAILABILITY AND IMPLEMENTATION: needLR is implemented in bash with dependencies including Truvari v4.2.2, BEDTools v2.31.1, and BCFtools v1.19. Source code, documentation, and pre-computed population allele frequency data are freely available at https://github.com/jgust1/needLR under an MIT license and archived on Zenodo at https://zenodo.org/records/19463479.

Software

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58&#x2009;706 SVs in a study sample of 11&#x2009;556 CAD cases and 42&#x2009;907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci.

The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. The point of template switching in 4 samples was shown to be a segment of &#x223c;2.2-5.5 kb of 100% nucleotide similarity within inverted repeat pairs. These data provide experimental evidence that inverted low-copy repeats act as recombinant substrates. This type of CGR can result in multiple conformers generating diverse SV haplotypes in susceptible dosage-sensitive loci.

Humans