Search PubMedSearch

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Structural genome variation drives adaptation of the xylose-fermenting yeast Scheffersomyces stipitis to lignocellulosic hydrolysates.

Second-generation (2G) bioethanol from lignocellulosic feedstocks is a sustainable alternative to fossil fuels. However, its production is constrained by the poor performance of industrial microbes in hydrolysates that are generated during biomass pretreatment. Scheffersomyces stipitis is a native xylose fermenting yeast and a promising platform for 2G bioethanol production, and adaptive evolution under hydrolysate stress has yielded strains with enhanced performance. However, the chromosomal basis of this adaptation is unknown. Here, we demonstrate that chromosome scale structural variation, rather than point mutations, underlies the improved phenotype of the evolved strains. By integrating long- and short-read genome sequencing, we identify two major chromosomal rearrangements in the top performing isolate: a reciprocal translocation between chromosomes 1 and 2 that disrupts the NUDIX hydrolase gene YSA1, and the formation of a mitotically stable 175 kb minichromosome derived from chromosome 5. Functional analyses show that disruption of YSA1 enhances xylose utilisation and ethanol yield, while the minichromosome contributes to improved performance in hydrolysate conditions. These findings provide direct evidence that balanced rearrangements and minichromosome formation can be selected during prolonged stress and can generate adaptive phenotypes. Taken together, our study establishes genome reorganisation as a key driver of adaptation in S. stipitis.

Xylose

Generating three-dimensional genome structures with a variational quantum algorithm.

Chromosome conformation capture experiments have revealed the underlying spatial interactions that govern three-dimensional (3D) genome organization and topology. Detecting 3D contacts between genomic loci considerably enhances our understanding of fundamental regulatory processes. Modeling 3D structures from experimental contact matrices can further contextualize the relationship between 3D genome organization and regulation. While classical algorithms have been successful in reconstructing genomic conformations, we investigate the prospect of quantum computation to aid in modeling the conformational space. In this context, we propose a novel variational quantum algorithm (VQA) to model the distribution of 3D genomic structures from experimental contact data. Through rigorous evaluations, we demonstrate the capability of our algorithm to sample ensembles of viable 3D conformations that agree well with experimental and simulated contact data. Furthermore, we extend our methodology to model the conformational space of a single cell or a population of cells. In the advent of sufficient quantum utility, the insights gained from this study can serve as a foundation for investigating high-resolution, large-scale ensembles of genomic conformations through generative VQAs.

Algorithms

Chromatin accessibility analysis reveals functional cis-regulatory regions related to fruit development and domestication in tomato.

Non-coding DNA sequences harbor vast regulatory programs that ensure the precise spatiotemporal control of gene expression, which is essential for proper plant development and trait formation. Chromatin accessibility analysis could identify functional DNA regions within the extensive non-coding sequences and infer regulatory elements, serving as a crucial approach to unravel the mysteries of non-coding DNA sequences. Tomato fruit, a fleshy organ, provides a special system for studying fruit development and trait formation. However, the role of cis-accessible chromatin regions (cis-ACRs) during tomato fruit development, particularly in comparison with protein-coding DNA sequences, remains poorly understood. Here, we used ATAC-seq to define the landscape of cis-ACRs during fruit development and domestication in tomato. Temporal differential analysis revealed the dynamic opening and closing of cis-ACRs during fruit development. Comparative analysis of cis-ACRs between cultivated and wild tomatoes highlighted their significant contributions to fruit domestication. Combining analysis with genomic structural variations (SVs) suggested that SVs are likely a key factor in the formation of specific accessible cis-ACRs in cultivated tomatoes. Moreover, using gene editing, we identified a functional cis-ACR within the intron of the MBP3 gene that regulates fruit development and size traits. Overall, our findings provide a comprehensive perspective on the roles of cis-ACRs in tomato fruit development and domestication.

Solanum lycopersicum

Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci.

The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. The point of template switching in 4 samples was shown to be a segment of ∼2.2-5.5 kb of 100% nucleotide similarity within inverted repeat pairs. These data provide experimental evidence that inverted low-copy repeats act as recombinant substrates. This type of CGR can result in multiple conformers generating diverse SV haplotypes in susceptible dosage-sensitive loci.

Humans

Sawfish: improving long-read structural variant discovery and genotyping with local haplotype modeling.

MOTIVATION: Structural variants (SVs) play an important role in evolutionary and functional genomics but are challenging to characterize. High-accuracy, long-read sequencing can substantially improve SV characterization when coupled with effective calling methods. While state-of-the-art long-read SV callers are highly accurate, further improvements are achievable by systematically modeling local haplotypes during SV discovery and genotyping. RESULTS: We describe sawfish, an SV caller for mapped high-quality long reads incorporating systematic SV haplotype modeling to improve accuracy and resolution. Assessment against the draft Genome in a Bottle (GIAB) SV benchmark from the T2T-HG002-Q100 diploid assembly shows that sawfish has the highest accuracy among state-of-the-art long-read SV callers across every tested SV size group. Additionally, sawfish maintains the highest accuracy at every tested depth level from 10- to 32-fold coverage, such that other callers required at least 30-fold coverage to match sawfish accuracy at 15-fold coverage. Sawfish also shows the highest accuracy in the GIAB challenging medically relevant genes benchmark, demonstrating improvements in both comprehensive and medically relevant contexts.When joint-genotyping seven samples from CEPH-1463, sawfish has over 9000 more pedigree-concordant calls than other state-of-the-art SV callers, with the highest proportion of concordant SVs (81%). Sawfish's quality model enables selection for an even higher proportion of concordant SVs (88%), while still calling nearly 5000 more pedigree-concordant SVs than other callers. These results demonstrate that sawfish improves on the state-of-the-art for long-read SV calling accuracy across both individual and joint-sample analyses. AVAILABILITY AND IMPLEMENTATION: Sawfish source code, pre-compiled Linux binaries, and documentation are released on GitHub: https://github.com/PacificBiosciences/sawfish.

Haplotypes

needLR: long-read structural variant annotation with population-scale frequency estimation.

SUMMARY: We present needLR, a structural variant (SV) annotation tool that can be used for filtering and prioritization of candidate pathogenic SVs from long-read sequencing data using population allele frequencies, annotations for genomic context, and gene-phenotype associations. When using population data from 500 presumably healthy individuals to evaluate nine test cases with known pathogenic SVs, needLR assigned allele frequencies to over 97.5% of all detected SVs and reduced the average number of novel genic SVs to 121 per case while retaining all known pathogenic variants. AVAILABILITY AND IMPLEMENTATION: needLR is implemented in bash with dependencies including Truvari v4.2.2, BEDTools v2.31.1, and BCFtools v1.19. Source code, documentation, and pre-computed population allele frequency data are freely available at https://github.com/jgust1/needLR under an MIT license and archived on Zenodo at https://zenodo.org/records/19463479.

Software

The value of structural variants to conservation genomics in the pangenome era.

Structural variants (SVs) comprise an axis of genetic diversity with strong consequences for phenotype and fitness, making them a potentially important target for conservation genomics. Here, we review how and why SVs can play a role in conservation genomics; the different types of SVs and how they can affect phenotype; and how pangenomes and long-read sequencing are illuminating their evolution in populations, including small populations and those of conservation concern. SVs comprise multinucleotide mutations including insertions, deletions, transpositions, inversions, and other multinucleotide mutations, often overlapping genes and other functional genome regions. As a result, SVs often play important roles in phenotypic evolution and local adaptation and can contribute substantially to genetic load in inbred populations. However, our understanding of the factors influencing SV diversity in populations is still in its infancy and is complicated by the vast range of sizes, effects, and mechanisms of formation of these mutations. We argue that SVs are an important axis of genetic diversity which should be characterized alongside more traditional metrics of genetic diversity in conservation contexts. There are a number of analytical challenges to detecting and studying SVs, but analyses aimed at understanding the role of SVs in inbreeding load and population health are rapidly becoming realizable goals, accelerated by new technologies and analytical approaches. New tools, including population-scale long-read sequencing and pangenome approaches, are beginning to make SVs accessible in ways which can be readily applied in conservation settings.

Genomic Structural Variation

Temperature and Pressure Shaped the Evolution of Antifreeze Proteins in Polar and Deep Sea Zoarcoid Fishes.

Antifreeze proteins (AFPs) have enabled teleost fishes to repeatedly colonize polar seas. Four AFP types have convergently evolved in several fish lineages. AFPs inhibit ice crystal growth and lower tissue freezing point. In lineages with AFPs, species inhabiting colder environments may possess more AFP copies. Elucidating how differences in AFP copy number evolve is challenging due to the genes' tandem array structure and consequently poor resolution of these repetitive regions. Here, we explore the evolution of type III AFPs (AFP III) in the globally distributed suborder Zoarcoidei, leveraging six new long-read genome assemblies. Zoarcoidei has fewer genomic resources relative to other polar fish clades while it is one of the few groups of fishes adapted to both the Arctic and Southern Oceans. Combining these new assemblies with additional long-read genomes available for Zoarcoidei, we conducted a comprehensive phylogenetic test of AFP III evolution and modeled the effects of thermal habitat and depth on AFP III gene family evolution. We confirm a single origin of AFP III via neofunctionalization of the enzyme sialic acid synthase B. We also show that AFP copy number increased under low temperature but decreased with depth, potentially because pressure lowers freezing point. Associations between the environment and AFP III copy number were driven by duplications of paralogs that were translocated out of the ancestral locus at which AFP III arose. Our results reveal novel environmental effects on AFP evolution and demonstrate the value of high-quality genomic resources for studying how structural genomic variation shapes convergent adaptation.

Animals

Mechanistic Perspectives From Genomics and Pangenomics of Medicinal and Aromatic Plants: Linking Genome Architecture to Phytochemical Diversity.

Medicinal and aromatic plants (MAPs) produce a remarkable diversity of specialized metabolites with significant pharmaceutical, nutraceutical, and industrial value. Although advances in long-read sequencing, chromosome-scale genome assembly, and pangenomics have greatly expanded genomic resources, the mechanistic links between genome architecture and phytochemical diversity remain incompletely understood. The present review synthesizes current evidence describing how structural genomic variation may contribute to phytochemical diversity, while acknowledging that many proposed genome-to-metabolite relationships require further experimental validation. Examples illustrate how genome architecture is associated with specialized-metabolite biosynthesis through multiple regulatory processes. However, the strength of supporting evidence varies considerably among MAP species. Moreover, relatively few genome-to-metabolite relationships have been confirmed through direct functional validation. We further discuss how pangenomics, multiomics integration, genome editing, synthetic biology, and artificial intelligence support the discovery, validation, and engineering of specialized metabolic pathways. Casual conclusions are evaluated according to the strength of available evidence, highlighting where causal relationships have been experimentally established and where conclusions remain primarily association-based. Overall, this review provides an integrated conceptual and evidence-based perspective summarizing proposed relationships between genome architecture and phytochemical diversity and outlines future priorities for functional genomics, precision breeding, metabolic engineering, and sustainable utilization of MAPs.

artificial intelligence

Analysis of deep-resequencing data of 984 soybean accessions reveals structural variations underlying agronomic traits.

Genomic structural variants (SVs) are major sources of genetic variation and have profound impacts on phenotypic traits. However, their functional effects remain largely unexplored in soybean. Here, we resequence 940 soybean accessions. Together with 44 publicly available datasets, we identify 602,281 SVs. Using a graph-based genome, we detect an additional 58,760 presence/absence variations (PAVs) that broadly affect gene expression. Population genomic analyses reveal that SVs serve as a core driving force for soybean domestication and improvement. Integrating SVs with QTLs for oil and protein content, and performing GWAS on 27 traits, we identify key functional SVs. These include transposable element insertions altering seed coat color, multiple insertions within a cytochrome P450 gene modifying flower and hypocotyl color, and a GmMATE1 deletion enhancing seed size. Together, our study establishes a comprehensive SV map of soybean, offering a valuable resource for dissecting the genetic basis of complex traits to accelerate molecular breeding.

Glycine max

Chromosome-Scale Genome of Zoonotic Eyeworm Thelazia callipaeda from China.

Thelazia callipaeda is a vector-borne zoonotic eyeworm infecting companion animals, wildlife, and humans, but chromosome-scale genomic resources from Chinese clinical material remain limited. We generated a genome supported by Pacific Biosciences (PacBio) high-fidelity (HiFi) sequencing and high-throughput chromosome conformation capture (Hi-C) from 100 adult worms recovered from naturally infected dogs in Beijing and compared its chromosome-scale organization with Portuguese assembly GCA_965194785.1. The final assembly spans 119.53 megabases (Mb) and comprises 115 top-level sequences, including four pseudomolecules totaling 91.26 Mb (76.34%) and 111 unanchored sequences. Genome-mode Benchmarking Universal Single-Copy Orthologs (BUSCO) analysis recovered 98.5% complete chromadorean orthologues, and the representative 11,788-protein gene set recovered 92.6%. Sequence-level alignment resolved Chinese chromosomes 1-4 (chr1-chr4) to Portuguese chr1, chrX, chr3, and chr2, respectively, with retained alignments covering 95.9-99.2% of each Chinese pseudomolecule and estimated sequence identities of 99.75-99.91%. Strong chromosome-scale collinearity was accompanied by localized reverse-collinear regions, including 0.243 Mb and 0.115 Mb intervals on chr2-chrX and chr3-chr3. The anchored sequences contained 96.7% of predicted genes and were substantially more gene-dense than the unanchored sequences. These results establish a clinically sourced Chinese chromosome-scale reference and provide a validated framework for future individual-worm, population-genomic, structural-variation, and comparative genomic studies of this parasite.

Hi-C

SVbyEye: a visual tool to characterize structural variation among whole-genome assemblies.

MOTIVATION: We are now in the era of being able to routinely generate highly contiguous (near telomere-to-telomere) genome assemblies of human and nonhuman species. Complex structural variation and regions of rapid evolutionary turnover are being discovered for the first time. Thus, efficient and informative visualization tools are needed to evaluate and directly observe structural differences between two or more genomes. RESULTS: We developed SVbyEye, an open-source R package to visualize and annotate sequence-to-sequence alignments along with various functionalities to process these alignments. The tool facilitates the characterization of complex structural variants in the context of sequence homology helping resolve the mechanisms underlying their formation. AVAILABILITY AND IMPLEMENTATION: SVbyEye is available on GitHub (https://github.com/daewoooo/SVbyEye) and via Zenodo (https://doi.org/10.5281/zenodo.15303553).

Software

Severus detects somatic structural variation and complex rearrangements in cancer genomes using long-read sequencing.

For the detection of somatic structural variation (SV) in cancer genomes, long-read sequencing is advantageous over short-read sequencing with respect to mappability and variant phasing. However, most current long-read SV detection methods are not developed for the analysis of tumor genomes characterized by complex rearrangements and heterogeneity. Here, we present Severus, a breakpoint graph-based algorithm for somatic SV calling from long-read cancer sequencing. Severus works with matching normal samples, supports unbalanced cancer karyotypes, can characterize complex multibreak SV patterns and produces haplotype-specific calls. On a comprehensive multitechnology cell line panel, Severus consistently outperforms other long-read and short-read methods in terms of SV detection F1 score (harmonic mean of the precision and recall). We also illustrate that compared to long-read methods, short-read sequencing systematically misses certain classes of somatic SVs, such as insertions or clustered rearrangements. We apply Severus to several clinical cases of pediatric leukemia/lymphoma, revealing clinically relevant cryptic rearrangements missed by standard genomic panels.

Humans

CHITRA: an interactive visualization tool for comparative genomic rearrangement analysis.

MOTIVATION: The increasing availability of chromosome-scale genome assemblies has fuelled a renewed interest in studying chromosomal evolution and rearrangements. Synteny visualization plays a critical role in understanding genome organization, structural variations, and evolutionary relationships. However, existing tools often have steep learning curves, produce static plots, or are limited in their ability to analyse multiple genomes simultaneously. There is a growing need for an intuitive and interactive visualization tool that can effectively explore syntenic relationships and chromosomal rearrangements. RESULTS: Here, we present CHITRA, a web-based interactive tool designed to visualize synteny blocks, chromosomal rearrangements, and breakpoints in both linear and circular styles. CHITRA-enables real-time exploration of genome structural variations with an intuitive graphical interface, customizable visualization options, and high-resolution export capabilities for publication-ready figures. The tool supports chromosome- and scaffold-level assemblies and allows users to filter, highlight, and interactively examine syntenic relationships. AVAILABILITY AND IMPLEMENTATION: CHITRA is freely available at https://chitra.bioinformaticsonline.com/, with comprehensive documentation at https://chitra.bioinformaticsonline.com/docs. The source code is open-source and accessible on GitHub at https://github.com/pranjalpruthi/CHITRA.

Journal Article

Unveiling the Genetic Landscape of Coronary Artery Disease Through Common and Rare Structural Variants.

BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58 706 SVs in a study sample of 11 556 CAD cases and 42 907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.

Humans

Pan-genome-based resequencing of 2,320 accessions reveals structural variations and accelerates breeding advances in cultivated peanut.

The cultivated peanut is a crucial global legume crop that is essential for food security and nutrition, particularly in developing regions. However, its limited genetic variation hampers breeding progress and yield improvement. Here we constructed a graph-based pan-genome for peanut, incorporating 14 genomes that represent all 6 peanut varieties. Using this pan-genome, we genotyped 2,320 accessions, covering 88.03% of ICRISAT and 59.21% of USDA core germplasm, enriching valuable resources for genomic studies and breeding. We cataloged genomic structural variations and investigated the role of homoeologous exchanges in population divergence. Through our pan-genome approach, we overcame the challenges of genotyping posed by homoeologous exchanges and identified key genes associated with flowering and dwarfism in peanut. By integrating superior haplotypes and germplasm resources guided by the pan-genome, we further developed high-yield dwarf lines. This work provides essential genomic resources to accelerate functional gene discovery and modern peanut breeding.

Journal Article

Chromosome-scale genome remodeling in tumor evolution: Copy number alterations and structural variants as two sides of the same coin.

Chromosome-scale genomic rearrangements are a dominant force in tumor evolution. Copy-number alterations (CNAs) and structural variants (SVs) constitute two complementary axes of this process. Although detection technologies now deliver near-comprehensive catalogs, technical resolution has outpaced conceptual integration. In this review, we frame CNAs and SVs as inextricable facets of chromosomal aberrations. They reshape cancer genomes through altered gene dosage and three-dimensional regulatory rewiring. CNAs quantify the gene-dosage imbalance, yet arise through mechanistically distinct routes. Segmental CNAs typically require chromosomal breakage, and therefore often coincide with SV junctions. By contrast, whole-chromosome aneuploidy and whole-genome doubling (WGD) primarily reflect mitotic or cytokinetic failure and can occur without local breakpoints, while nevertheless reshaping the karyotypic landscape and seeding subsequent structural complexity. SVs, in turn, range from unbalanced events that alter copy number to ostensibly balanced exchanges that predominantly rewire regulatory architecture. Despite their diverse and sometimes catastrophic architectures, SVs are ultimately rooted in double-strand break formation and error-prone resolution. By integrating CNAs and SVs within a unified mechanistic and functional framework, we aim to convert catalogs into concepts and distill the organizing principles that govern tumor genome evolution.

Humans

Comparative genomics reveals lineage-associated structural variation and diversification in a barley fungal pathogen.

Leaf rust, caused by Puccinia hordei, is a major barley disease worldwide. Despite repeated shifts in virulence, contrasting reproductive histories, and emerging fungicide insensitivity, the genomic basis of its diversification and adaptation remains poorly understood. In this study, we generated haplotype-resolved, chromosome-level genome assemblies for two isolates with contrasting virulence and analyzed 41 Australian isolates collected over 54 yr (1966-2020), integrating comparative and population genomics, mating-type gene phylogenies, chromosome-specific k-mer profiling, genome-wide copy-number variation (CNV) analysis, and gene-expression analysis. We identified a structurally dynamic chromosome characterized by repeat-associated rearrangements, structural variation, and lineage-associated CNV, representing the first evidence in a rust fungus of chromosome-scale structural diversification of this extent. Population analyses distinguished clonally expanded lineages from recombination-associated lineages, with mating-type gene phylogenies providing further support for lineage differentiation. More recently collected isolates showed increased duplication-associated variation, and CNV boundaries were associated with structural-variant breakpoints. We also identified lineage-associated amplification of Cyp51, with increased copy number associated with higher transcript abundance, supporting a potential role in fungicide adaptation. Overall, our findings highlight structural variation, contrasting reproductive histories, and lineage-associated CNV as important contributors to diversification in P. hordei, providing insights for future rust pathogen surveillance and management strategies.

Cyp51 gene