Search PubMedSearch

SEARCH · Search PubMed

Results for “Structural variants”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Variability of human rRNA genes: inheritance and nonrandom chromosomal distribution of structural variants of nontranscribed spacer sequences.

Human rRNA genes contain variable regions, one of which is located in nontranscribed spacers (NTSs) closely downstream from the 3'-end of the transcribed region. This polymorphism may be detected by means of blot hybridization analysis as a set of distinct restriction fragments corresponding to this part of the rRNA genes. We have analyzed DNA of 51 individuals and found eight structural NTS variants of this region; two of these were common to all individuals analyzed, and six others were found in different combinations and with different frequencies. The copy number of each variant also differed but was not less than 10-20 copies per cell. The analysis of DNA isolated from leukocytes of the members of 11 families indicated that some of the structural variants (of the NTS region) are inherited as a single Mendelian locus. We propose that rRNA genes that belong to one particular structural variant form clusters on separate chromosomes. To test this proposition, we developed a combined method, including AgNO3-staining of chromosomes, in situ hybridization, and DNA analysis with methylation-sensitive restrictases, and used it for study of persons who had methylated rRNA genes located on AgNO3-negative nucleolar organizers. It was found that in three of four cases methylated genes really belonged to one structural variant. This approach may be used for detailed localization of separate classes of NTS structural variants of human rRNA genes.

Blotting, Southern

cDNA clones encoding murine IgE-binding factors represent multiple structural variants of intracisternal A-particle genes.

Previously [Moore, K. W., Jardieu, P., Mietz, J. A., Trounstine, M. L., Kuff, E. L., Ishizaka, K. & Martens, C. L. (1986) J. Immunol. 136, 4283-4290], we examined a T-hybridoma-derived cDNA clone, 8.3, that encodes a biologically active murine IgE-binding factor (IgE-BF), and we showed that it was a variant member of the endogenous retroviral gene family related to mouse intracisternal A particles (IAPs). We have now characterized four more IgE-BF cDNA clones by heteroduplex and restriction enzyme analysis and found that they all represent different structural variants of the full-size IAP genomic element. In clones 8.3 and 10.2, which have been fully sequenced, the open reading frames span deletions 3.4 and 1.9 kilobases (kb) long, respectively, and specify different gag-pol fusion polypeptides. Clone 9.5 contains a 2.1-kb deletion entirely within the pol region. Two other clones (4.2 and 11.7) contain no internal deletion and may represent truncated cDNA copies of full-size (7.2 kb) IAP gene transcripts. Structural variants very similar to clone 10.2 are common in the mouse genome, and clone 9.5 is also probably not a unique gene form. The sequences of clones 8.3 and 10.2 are different in detail, but each is closely homologous to a randomly cloned mouse genomic IAP element throughout the gag-related portions of their open reading frames. Antibodies against two oligopeptides specified by the sequence of clone 8.3 immunoprecipitated IAP-related proteins from mouse neuroblastoma and myeloma cells, confirming that the IgE-BF produced by this clone shares sequence with expressed IAP elements in different cell types. Thus, information related to the IgE-BF is an integral part of the murine IAP retrotransposon gag gene.

Amino Acid Sequence

Structures of five sulfated hexasaccharides prepared from porcine intestinal heparin using bacterial heparinase. Structural variants with apparent biosynthetic precursor-product relationships for the antithrombin III-binding site.

Porcine intestinal heparin was extensively digested with Flavobacterium heparinase and size-fractionated by gel chromatography. Subfractionation of the hexasaccharide fraction by anion exchange high pressure liquid chromatography yielded 10 fractions. Six contained oligosaccharides derived from the repeating disaccharide region, whereas four contained glycoserines from the glycosaminoglycan-protein linkage region. The latter structures were reported recently (Sugahara, K., Tsuda, H., Yoshida, K., Yamada, S., de Beer, T., and Vliegenthart, J.F.G. (1995) J. Biol. Chem. 270, 22914-22923). In this study, the structures of one tetra- and five hexasaccharides from the repeat region were determined by chemical and enzymatic analyses as well as 500-MHz 1H NMR spectroscopy. The tetrasaccharide has the hexasulfated structure typical of heparin. The five hexa- or heptasulfated hexasaccharides share the common core pentasulfated structure delta HexA(2S) alpha 1-4GlcN-(NS, 6S) alpha 1-4IdoA alpha/GlcA beta 1-4GlcN(6S) alpha 1-4GlcA beta 1-4GlcN (NS) with one or two additional sulfate groups (delta HexA, GlcN, IdoA, and GlcA represent 4-deoxy-alpha-L-threo-hex-4-enepyranosyluronic acid, D-glucosamine, L-iduronic acid, and D-glucuronic acid, whereas 2S, 6S and NS stand for 2-O-, 6-O-, and 2-N-sulfate, respectively). Three components have the following hitherto unreported structures: delta HexA(2S) alpha 1-4GlcN(NS, 6S) alpha 1-4GlcA beta 1-4GlcN(NS, 6S) alpha 1-4GlcA beta 1-4GlcN(NS,6S), delta HexA(2S) alpha 1-4GlcN(NS, 6S) alpha 1-4IdoA alpha 1-4GlcNAc(6S)-alpha 1-4GlcA beta 1-4GlcN(NS, 3S), and delta HexA(2S) alpha 1-4GlcN-(NS,6S) alpha 1-4IdoA (2S) alpha 1-4GlcNAc(6S) alpha 1-4GlcA beta 1-4GlcN(NS, 6S). Two of the five hexasaccharides are structural variants derived from the antithrombin III-binding sites containing 3-O-sulfated GlcN at the reducing termini with or without a 6-O-sulfate group on the reducing N,3-disulfated GlcN residue. Another contains the structure identical to that of the above heptasulfated antithrombin III-binding site fragment but lacks the 3-O-sulfate group and therefore is a pro-form for the binding site. Another has an extra sulfate group on the internal IdoA residue of this pro-form and therefore can be considered to have diverged from the binding site in the biosynthetic pathway. Thus, the isolated hexasacharides in this study include the three overlapping pairs of structural variants with an apparent biosynthetic precursor-product relationship, which may reflect biosynthetic regulatory mechanisms of the binding site.

Animals

Complex de novo structural variants are an underestimated cause of rare disorders.

Complex de novo structural variants (dnSVs) are crucial genetic factors in rare disorders, yet their prevalence and characteristics in rare disorders remain poorly understood. Here, we conduct a comprehensive analysis of whole-genome sequencing data of 12,568 families, including 13,698 offspring with rare diseases, obtained as part of the UK 100,000 Genomes Project. We identify 1,870 dnSVs, constituting the largest dnSV dataset reported to date. Complex dnSVs (n = 158; 8.4%) emerge as the third most common type of SV, following simple deletions and duplications. We classify 65% of these complex dnSVs into 11 subtypes. Among probands with dnSVs (n = 1,696), 9% exhibit exon-disrupting pathogenic dnSVs associated with the probands' phenotype. Notably, 12% of exon-disrupting pathogenic dnSVs and 22% of de novo deletions or duplications previously identified by array-based or whole-exome sequencing methods are found to be complex dnSVs. We also find distinct genomic properties of de novo deletions depending on the parent of origin. This study highlights the importance of complex dnSVs in the cause of rare disorders and demonstrates the necessity of specific genomic analysis to avoid overlooking these variants.

Humans

High-resolution separation of amyloid beta-peptides: structural variants present in Alzheimer's disease amyloid.

In Alzheimer's disease (AD), one of the cardinal neuropathological signs is deposition of amyloid, primarily consisting of the amyloid beta-peptide (Abeta). Structural variants of AD-associated Abeta peptides have been difficult to purify by high-resolution chromatographic techniques. We therefore developed a novel chromatographic protocol, enabling high-resolution reverse-phase liquid chromatography (RPLC) purification of Abeta variants displaying very small structural differences. By using a combination of size-exclusion chromatography and the novel RPLC protocol, Abeta peptides extracted from AD amyloid were purified and subsequently characterized. Structural analysis by microsequencing and electrospray-ionization mass spectrometry revealed that the RPLC system resolved a complex mixture of Abeta variants terminating at either residue 40 or 42. Abeta variants differing by as little as one amino acid residue could be purified rapidly to apparent homogeneity. The resolution of the system was further illustrated by its ability to separate the structural isomers of Abeta1-40. The present chromatography system might provide further insight into the role of N-terminally and posttranslationally modified Abeta variants, because each variant can now be studied individually.

Alzheimer Disease

Structural variants of human T200 glycoprotein (leukocyte-common antigen).

Structural variation in the primary structure of human T200 glycoprotein has been detected. Three cDNA variants have been characterized each of which encode T200 molecules that differ in size as a result of sequence differences in their amino-terminal regions. The largest form of the molecule is distinguished from the smallest by an insert of 161 amino acids, after the first eight amino-terminal residues. The other variant has an insert at the same location of 47 amino acids identical to residues 75-121 in the larger insert. Both extra domains are rich in serine and threonine residues and are likely to display multiple O-linked oligosaccharides. These structural variants which probably arise by cell-type-specific alternative splicing provide a molecular basis for the previously observed structural and antigenic heterogeneity of T200 glycoprotein. In addition to the variable amino-terminal region, the external domain of human T200 glycoprotein consists of a second cysteine-rich region of about 400 amino acids, a single transmembrane-spanning region and a large cytoplasmic domain of 707 amino acids shared by all of the structural variants and highly conserved between species. The gene encoding human T200 is located on the long arm of chromosome 1.

Amino Acid Sequence

needLR: long-read structural variant annotation with population-scale frequency estimation.

SUMMARY: We present needLR, a structural variant (SV) annotation tool that can be used for filtering and prioritization of candidate pathogenic SVs from long-read sequencing data using population allele frequencies, annotations for genomic context, and gene-phenotype associations. When using population data from 500 presumably healthy individuals to evaluate nine test cases with known pathogenic SVs, needLR assigned allele frequencies to over 97.5% of all detected SVs and reduced the average number of novel genic SVs to 121 per case while retaining all known pathogenic variants. AVAILABILITY AND IMPLEMENTATION: needLR is implemented in bash with dependencies including Truvari v4.2.2, BEDTools v2.31.1, and BCFtools v1.19. Source code, documentation, and pre-computed population allele frequency data are freely available at https://github.com/jgust1/needLR under an MIT license and archived on Zenodo at https://zenodo.org/records/19463479.

Software

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58 706 SVs in a study sample of 11 556 CAD cases and 42 907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci.

The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. The point of template switching in 4 samples was shown to be a segment of ∼2.2-5.5 kb of 100% nucleotide similarity within inverted repeat pairs. These data provide experimental evidence that inverted low-copy repeats act as recombinant substrates. This type of CGR can result in multiple conformers generating diverse SV haplotypes in susceptible dosage-sensitive loci.

Humans

The activity of 25 paroxetine/femoxetine structure variants in various reactions, assumed to be important for the effect of antidepressants.

Structure-activity relationships for 25 structural variants around the 5-hydroxytryptamine (5-HT) uptake inhibitors paroxetine and femoxetine have been investigated. Three parameters related to the 5-HT system were investigated: (i) The inhibition of [3H]5-HT uptake into rat brain synaptosomes, (ii) the inhibition of [3H]paroxetine binding to rat neuronal membranes and (iii) the effect of the compounds on the affinity of [3H]imipramine for the human platelet membrane binding site, measured as the dissociation rate of the [3H]imipramine human platelet membrane binding site complex. A highly significant correlation was found for 5-HT uptake inhibition and inhibition of [3H]paroxetine binding for the different substances, indicating that the two parameters are closely connected. However the slope of the regression line was only 0.6 and not 1.0; this may indicate that [3H]paroxetine binding is necessary, but not sufficient for 5-HT uptake inhibition. No correlation was found between the inhibition of [3H]paroxetine binding and the affinity of the compounds for the [3H]imipramine binding site complex. The two binding sites are therefore probably situated on different parts of the 5-HT transport system, the [3H]paroxetine binding site being part of the 5-HT transport mechanism whereas the [3H]imipramine binding site may represent a site modulating the activity of, and affinity for, 5-HT in the 5-HT transport mechanism. Structure-activity relationships among the substances showed that stereochemical changes from (-)- to (+)-trans changed the activity towards both 5-HT uptake inhibition and [3H]paroxetine displacement for most of the (-)-/(+)-pairs. The substitution of -H with -F or -CH3 also affected the activity.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals

Structural variants of the alpha-fetoprotein gene in different inbred strains of rat.

Two structural variants of the rat alpha-fetoprotein (AFP) gene have been detected in different inbred strains of rats by EcoRI or HindIII restriction enzyme cleavage of cellular DNA, agarose gel electrophoresis and Southern blot hybridization using 32P-labeled cloned rat AFP cDNA probes. The type I AFP gene variant is characteristic of the Sprague-Dawley strain, and type II is found in Buffalo rats. These variants appear to represent two different allelic forms of the rat AFP gene since they are inherited in a normal Mendelian fashion when Sprague-Dawley and Buffalo rats are crossed. The mapping results suggest that the two allelic variants differ from each other by multiple cleavage site variations (base pair substitutions) and by an insertion or deletion of DNA sequences. An extensive sequence variation appears to exist between the two forms of the rat AFP gene; we have estimated that as much as 2.7% of the nucleotides in this region vary between the two alleles.

Animals

Comprehensive benchmarking of somatic structural variant detection at ultra-low allele fractions.

Postzygotic mosaicism gives rise to somatic structural variants (SVs) at ultra-low variant allele fractions (VAFs), which pose challenges for detection due to the high-coverage sequencing required and noise introduced by sequencing artifacts. Although somatic SV detection has been extensively studied in cancer, these studies are not directly applicable to the study of tissue mosaicism, as they rely on matched normals, target higher VAF ranges, and are enriched for different types of SVs. We present comprehensive benchmark data and best practices for non-cancer somatic SV detection. We created a synthetic mosaic sample by combining six HapMap individuals at varying proportions, generating allele fractions as low as 0.25%. This sample was sequenced to ~2,300x total coverage using Illumina, PacBio, and Nanopore technologies across multiple sequencing centers. A high-confidence benchmark SV set containing over 21,000 pseudo-somatic insertions and deletions ≥50bp was derived from haplotype-resolved assemblies. We evaluated 12 SV discovery pipelines and identified caller-specific strengths and sequencing platform-specific shortcomings. We find that short read-based approaches show reduced recall for insertions and repeat-associated SVs, whereas long-read sequencing achieves high accuracy throughout the genome, increasing linearly with coverage. The best algorithm's sensitivity exceeded 80% for VAFs ≥4% and 15% for VAFs of 0.5-1% with 60x coverage. The publicly available benchmarking data and comparative analysis of current methods provide a foundation for robust discovery of SV mosaicism in non-cancer tissues..

Journal Article

Improving long-read somatic structural variant calling with pangenome and de novo personal genome assembly.

Accurate detection of mosaic and somatic structural variants (SVs) provides early diagnostic and therapeutic evidence for cancers. While long-read whole-genome sequencing leads to more accurate SV detection than short read sequencing, existing long-read SV callers only look at alignment against a single reference genome and are susceptible to systematic false discovery caused by germline differences between the individual genome and the reference genome. Here we develop a new SV filtering method that jointly considers the alignment against a pangenome and the de novo assembly of the germline genome. It dramatically reduces false positive mosaic and somatic SVs in cancer cell lines with little loss in sensitivity for existing long read SV callers. Our study highlights the essential need for pangenome or personal genome assembly to integrate SV calls for both SV discoveries and clinical diagnostics.

Journal Article

Structural variants linked to Alzheimer's disease and other common age-related clinical and neuropathologic traits.

BACKGROUND: Alzheimer's disease (AD) is a complex neurodegenerative disorder with substantial genetic influence. While genome-wide association studies (GWAS) have identified numerous risk loci for late-onset AD (LOAD), the functional mechanisms underlying most of these associations remain unresolved. Large genomic rearrangements, known as structural variants (SVs), represent a promising avenue for elucidating such mechanisms within some of these loci. METHODS: By leveraging data from two ongoing cohort studies of aging and dementia, the Religious Orders Study and Rush Memory and Aging Project (ROS/MAP), we performed genome-wide association analysis testing 20,205 common SVs from 1088 participants with whole genome sequencing (WGS) data. A range of Alzheimer's disease and other common age-related clinical and neuropathologic traits were examined. RESULTS: First, we mapped SVs across 81 AD risk loci and discovered 22 SVs in linkage disequilibrium (LD) with GWAS lead variants and directly associated with the phenotypes tested. The strongest association was a deletion of an Alu element in the 3'UTR of the TMEM106B gene, in high LD with the respective AD GWAS locus and associated with multiple AD and AD-related disorders (ADRD) phenotypes, including tangles density, TDP-43, and cognitive resilience. The deletion of this element was also linked to lower TMEM106B protein abundance. We also found a 22-kb deletion associated with depression in ROS/MAP and bearing similar association patterns as GWAS SNPs at the IQCK locus. In addition, we leveraged our catalog of SV-GWAS to replicate and characterize independent findings in SV-based GWAS for AD and five other neurodegenerative diseases. Among these findings, we highlight the replication of genome-wide significant SVs for progressive supranuclear palsy (PSP), including markers for the 17q21.31 MAPT locus inversion and a 1483-bp deletion at the CYP2A13 locus, along with other suggestive associations, such as a 994-bp duplication in the LMNTD1 locus, suggestively linked to AD and a 3958-bp deletion at the DOCK5 locus linked to Lewy body disease (LBD) (P = 3.36 × 10-4). CONCLUSIONS: While still limited in sample size, this study highlights the utility of including analysis of SVs for elucidating mechanisms underlying GWAS loci and provides a valuable resource for the characterization of the effects of SVs in neurodegenerative disease pathogenesis.

Humans

Regulation of cellular interactions with laminin by integrin cytoplasmic domains: the A and B structural variants of the alpha 6 beta 1 integrin differentially modulate the adhesive strength, morphology, and migration of macrophages.

Several integrin alpha subunits have structural variants that are identical in their extracellular and transmembrane domains but that differ in their cytoplasmic domains. The functional significance of these variants, however, is unknown. In the present study, we examined the possibility that the A and B variants of the alpha 6 beta 1 integrin laminin receptor differ in function. For this purpose, we expressed the alpha 6A and alpha 6B cDNAs, as well as a truncated alpha 6 cDNA (alpha 6-delta CYT) in which the cytoplasmic domain sequence was deleted after the GFFKR pentapeptide, in P388D1 cells, an alpha 6 deficient macrophage cell line. Populations of stable alpha 6A, alpha 6B, and alpha 6-delta CYT transfectants that expressed equivalent levels of cell surface alpha 6 were obtained by fluorescence-activated cell sorter and shown to form heterodimers with endogenous beta 1 subunits. Upon attachment to laminin, the alpha 6A transfectants extended numerous pseudopodia. In contrast, the alpha 6B transfectants remained rounded and extended few processes. The transfectants were also examined for their ability to migrate toward a laminin substratum using Transwell chambers. The alpha 6A transfectants were three- to fourfold more migratory than the alpha 6B transfectants. The alpha 6-delta CYT transfectants did not attach to laminin in normal culture medium, but they did attach in the presence of Mn2+. The alpha 6-delta CYT transfectants migrated to a lesser extent than either the alpha 6A or alpha 6B transfectants in the presence of Mn2+. The alpha 6 transfectants differed significantly in the concentration of substratum bound laminin required for half-maximal adhesion in the presence of Mn2+:alpha 6A (2.1 micrograms/ml), alpha 6B (6.3 micrograms/ml), and alpha 6-delta CYT (8.8 micrograms/ml). Divalent cation titration studies revealed that these transfectants also differed significantly in both the [Ca2+] and [Mn2+] required to obtain half-maximal adhesion to laminin. These data demonstrate that the A and B variants of the alpha 6 cytoplasmic domain can differentially modulate the function of the alpha 6 beta 1 extracellular domain.

Animals

Chromosome-scale genome remodeling in tumor evolution: Copy number alterations and structural variants as two sides of the same coin.

Chromosome-scale genomic rearrangements are a dominant force in tumor evolution. Copy-number alterations (CNAs) and structural variants (SVs) constitute two complementary axes of this process. Although detection technologies now deliver near-comprehensive catalogs, technical resolution has outpaced conceptual integration. In this review, we frame CNAs and SVs as inextricable facets of chromosomal aberrations. They reshape cancer genomes through altered gene dosage and three-dimensional regulatory rewiring. CNAs quantify the gene-dosage imbalance, yet arise through mechanistically distinct routes. Segmental CNAs typically require chromosomal breakage, and therefore often coincide with SV junctions. By contrast, whole-chromosome aneuploidy and whole-genome doubling (WGD) primarily reflect mitotic or cytokinetic failure and can occur without local breakpoints, while nevertheless reshaping the karyotypic landscape and seeding subsequent structural complexity. SVs, in turn, range from unbalanced events that alter copy number to ostensibly balanced exchanges that predominantly rewire regulatory architecture. Despite their diverse and sometimes catastrophic architectures, SVs are ultimately rooted in double-strand break formation and error-prone resolution. By integrating CNAs and SVs within a unified mechanistic and functional framework, we aim to convert catalogs into concepts and distill the organizing principles that govern tumor genome evolution.

Humans

[Structural variants in hemoglobin occurring in the Czech Republic].

The authors present a review of clinical and laboratory findings of seven in the Czech Republic hitherto diagnosed structural haemoglobin variants. Unstable variants are found most frequently: Hb-Köln, Hb-St. Louis, Hb-Nottingham, Hb-E and Hb-Hradec Králové. The variant Hb-Hradec Králové (Hb-HK) or alpha 2 beta 2 115 (G17) Ala-Asp was newly detected. The great instability of Hb-HK chains makes classical diagnosis of Hb-pathy impossible. It was possible to identify it only at a molecular genetic level. A manifestation of Hb-HK instability is also the thalassaemic feature of the disease and the formation of Heinz bodies from free chains. The only representative of haemoglobins with a high oxygen affinity identified in this country was newly detected. It was given the name Hb-Olomouc or alpha 2 beta 2 86 (F2) Ala-Asp. This haemoglobin variant leads to erythrocytosis in father and son and the same clinical manifestations were recently described also in Japan. The last structural variant of haemoglobin found in this country is Hb-M Milwaukee or alpha 2 beta 2 67 (E11) Val-Glu which in our patients is manifested rather by haemolysis with formation of Heinz bodies than classical cyanosis. The cause of instability of Hb-M in our patients is not known. Hb-S was not diagnosed so far in the Czech Republic.

Adolescent

De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.

Three-dimensional genome organization plays a critical role in gene regulation, and disruptions can lead to developmental disorders by altering the contact between genes and their distal regulatory elements. Structural variants (SVs) can disturb local genome organization, such as the merging of topologically associating domains upon boundary deletion. Testing large numbers of SVs experimentally for their effects on chromatin structure and gene expression is time and cost prohibitive. To address this, we propose a computational approach to predict SV impacts on genome folding, which can help prioritize causal hypotheses for functional testing. We develop a weighted scoring method that measures chromatin contact changes specifically affecting regions of interest, such as regulatory elements or promoters, and implement it in the SuPreMo-Akita software. With this tool, we rank hundreds of de novo SVs (dnSVs) from autism spectrum disorder (ASD) individuals and their unaffected siblings based on predicted disruptions to nearby neuronal regulatory interactions. This reveals that putative cis-regulatory element interactions (CREints) are more disrupted by dnSVs from ASD probands versus unaffected siblings. We prioritize candidate variants that disrupt ASD CREints and validate our top-ranked locus using isogenic excitatory neurons with and without the dnSV, confirming accurate predictions of disrupted chromatin contacts. This study suggests that disrupted genome folding is a potential genetic mechanism in a subset of ASD cases and provides a general strategy for prioritizing variants predicted to disrupt regulatory interactions across tissues.

Humans