Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Dual-dimensional profiling of host genomic variations and HPV integration in PD-L1-stratified cervical cancer via Oxford Nanopore Technology.

BACKGROUND: The integration of human papillomavirus (HPV) DNA into the host genome is a key step in the development of HPV-associated cervical cancer (CC). However, the genomic characteristics of host genomic variations and HPV integration within the context of programmed death-ligand 1 (PD-L1) expression stratification have not been systematically investigated. METHODS: Whole-genome sequencing was performed using Oxford Nanopore Technology (ONT) on six samples (three from the high PD-L1 expression group and three from the low PD-L1 expression group). The characteristics of host genomic variations under different PD-L1 expression stratifications were explored, including structural variations (SV), copy number variations (CNV), single nucleotide polymorphisms (SNP), and insertion-deletions (Indel). Subsequently, the distribution features of HPV integration sites were analyzed, different integration types were identified, and pathway analysis was conducted. RESULTS: Whole-genome SV analysis revealed that the total number of SVs and the composition of mutation types were similar between the high and low PD-L1 expression groups, with insertions (INS) and deletions (DEL) predominating in both. These variations were primarily enriched in intergenic regions and introns. In the low PD-L1 expression group, integration events were observed at multiple chromosomal loci, with the most frequent integration occurring in the KLF5 gene region on chromosome 13. No frequently integrated loci were identified in the high PD-L1 expression group. Additionally, four distinct HPV integration breakpoint patterns were preliminarily identified and analyzed. CONCLUSION: PD-L1 expression stratification did not significantly alter the overall genomic instability of the host. However, differences were observed in the distribution patterns of HPV integration sites. These findings provide new insights into the genomic heterogeneity of CC under different PD-L1 expression backgrounds and may lay the groundwork for future research exploring stratified immunotherapy based on HPV integration features.

Humans

Optical mapping in Black genomes: Distinct LCR22 structures and 22q11.2 deletion syndrome mechanisms.

PURPOSE: The genomic architecture of 22q11.2 deletion syndrome (22q11.2DS) has primarily been studied in White populations, despite evidence suggesting a lower prevalence in Black individuals. This study aims to improve our understanding of the population-specific organization of 22q11.2 genomic structures. METHODS: Optical mapping data from 106 genomes, representing various Black and White individuals, were analyzed to assess the structure and variation of the 22q11.2 low copy repeats (LCR22s). RESULTS: Extensive variability in copy-number and orientation of LCR22 elements was observed between Black and White genomes. Several novel copy-number variants and haplotype configurations were identified, some being private or more prevalent within specific groups. Notably, copy-number variants diversity was particularly striking among Black genomes. Comparisons of Black and White families with de novo 22q11.2DS probands revealed unique nonallelic homologous recombination scenarios, with Black families exhibiting recombination patterns that are not previously observed. CONCLUSION: Perhaps the unique and highly variable LCR22 haplotype configurations in Black individuals contribute to the lower observed prevalence of 22q11.2DS by inhibiting the likelihood of nonallelic homologous recombination, the mechanism that leads to the syndrome.

Humans

Primulina pan-genome reveals differential gene retention following whole-genome duplications and provides insights into edaphic specialization.

Primulina, a genus of >200 species specialized to extreme soils, provides a model for edaphic adaptation. We assemble seven genomes and construct a pan-genome spanning nine species from karst, Danxia, and acidic soils. Comparative analyses reveal that karst-adapted species have smaller genomes. Two lineage-specific whole-genome duplications (WGDs) exhibit biased duplicate loss in large gene families but preferential retention of transcription factors, indicating combined adaptive and nonadaptive forces. Pan-genome analyses identify ion channel and transporter genes enriched in variant hotspots and under positive selection in karst lineages. Candidate genes for drought and salt stress tolerance include ABC transporters and ion channels. Notably, an ABC transporter shows positive selection in karst species and unique structural variation in non-karst species. Together, our findings show that genome downsizing, biased post-WGD retention, and evolution of ion-transport pathways shape adaptation to extreme soils. The Primulina pan-genome provides a resource for dissecting mechanisms underlying edaphic specialization.

Gene Duplication

Decoding missense variants pleiotropy in the immune GPCR P2RY8.

G protein-coupled receptors (GPCRs) form the largest family of cell surface receptors and remain a central focus in pharmacology and drug discovery. Despite extensive structural and pharmacological studies, the functional impact of missense variation across GPCRs remains poorly understood, particularly for receptors involved in immune regulation. In this issue of Cell Genomics, LaFlam et al.1 systematically map P2RY8 variant functions using deep mutational scanning (DMS) combined with structural biology approaches, revealing pleiotropy and mechanisms linking GPCR variation to B cell confinement and lymphoma.

Humans

Lineage-associated small inversions disrupt dosT, dnaE2, and a promoter-adjacent region in some Mycobacterium tuberculosis isolates.

UNLABELLED: Large molecular inversions in the genome of Mycobacterium tuberculosis (Mtb) due to factors like the presence of insertion sequences and transposases are widely known. However, smaller inversions within coding sequences and non-coding control elements are rarely reported. The present study aims to identify inversions and their potential impact on Mtb biology in a lineage-specific manner. Structural variants (SVs) could only be detected by long reads. For this, we simulated long reads by de novo assembling the short-read sequencing data sets and subsequently aligned representative strains from each lineage using the Progressive Mauve algorithm. Independently, long-read sequencing from the Pacific Biosciences platform was acquired and analyzed using the structural variant identification method. Variants were merged, and Fisher's exact test was carried out to identify the inversion association with lineages. To visualize deoxyribonucleic acid (DNA) features, the DNA-features-viewer tool was used. Simulated reads from short-read sequencing gave indications of lineage (L)-specific inversions. The long-read sequencing approach led to the identification of seven unique inversions: two positively associated with L1, one positively associated with L3, two negatively associated with L4, and two positively associated with L3 but negatively associated with L4 (P < 0.05). The inversions encompassed primarily non-essential genes like sdaA, dosT, Rv2026c, dnaE2, Rv1341, Rv1342, and lprD. An interesting inversion was observed in the upstream control element of purB and Rv0776c. The study sheds light on small inversions that may be causing alterations in expression, formation of fusion genes, and nonsense mutations that may have a role in lineage-specific phenotypic changes. IMPORTANCE: The role of mutations like SNPs and INDELs and their association with drug resistance is well known in Mycobacterium tuberculosis (Mtb). However, structural variations, especially inversions, are largely overlooked and unreported. In this paper, publicly available whole-genome sequencing datasets from Illumina and Pacific Biosciences-Oxford Nanopore Technologies platform have been used to detect inversions and report seven unreported Mtb lineage-specific small inversions.

Mycobacterium tuberculosis

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common &#x223c;4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

A chromosome-scale genome of Capsicum pubescens provides insights into candidate terpene-associated gene clusters and pan variation of terpene synthases.

A chromosome-scale genome of Capsicum pubescens and comparative pan-TPS analysis support structural characterization and gene-level prioritization of a chromosome-9 terpene-associated candidate locus in this accession. Capsicum pubescens is one of the five domesticated Capsicum species, mainly cultivated in mid- to high-elevation regions of the Americas. Despite its distinctive morphology and fruit traits, genomic resources for C. pubescens remain less developed than those for the widely cultivated C. annuum. Here, we assembled a chromosome-scale reference genome for accession HNUCP0001, spanning 3.70&#xa0;Gb with a scaffold N50 of 278.01&#xa0;Mb. Comparative genomics revealed 679 significantly expanded gene families enriched in sesquiterpenoid and triterpenoid biosynthesis. Genome-wide biosynthetic gene-cluster mining identified multiple terpene-associated candidate loci, which were subsequently prioritized using genome-derived structural criteria and Capsicum pubescens-specific expression evidence. Subsequently, we curated the terpene synthase (TPS) repertoire and, across 16 Capsicum genomes, resolved 36 TPS orthogroups with pronounced presence/absence variation, highlighting dynamic lineage-specific diversification. Together, these analyses establish HNUCP0001 as an accession-specific genomic resource and provide a comparative framework for prioritizing terpene-associated TPS genes and candidate BGCs in Capsicum. These candidate loci, together with accession-level transcriptomic and metabolomic evidence, offer testable hypotheses for future functional studies of specialized terpenoid metabolism in C. pubescens.

Alkyl and Aryl Transferases

Integrated multi-omics analyses provide new insights into genomic variation landscape and regulatory network candidate genes associated with walnut endocarp.

Persian walnut (Juglans regia) is an economically important nut oil tree; the fruit has a hard endocarp/shell to protect seeds, thus playing a key role in its evolution, and the shell thickness is an important trait for walnut breeding. However, the genomic landscape and the gene regulatory networks associated with walnut shell development remain to be systematically elucidated. Here, we report a high-quality genome assembly of the walnut cultivar 'Xiangling' and construct a graphic structure pan-genome of eight Juglans species to reveal the genetic variations at the genome level. We re-sequence 285 accessions to characterize the genomic variation landscape. Through genome-wide association studies (GWAS), we identified 19 loci associated with more than 268 loci that underwent selection during walnut domestication and improvement. Multi-omics analyses, including transcriptomics, metabolomics, DNA methylation, and spatial transcriptomics across eleven developmental stages, revealed several candidate genes related to secondary cell biosynthesis and lignin accumulation. This integrated multi-omics approach revealed several candidate genes associated with secondary cell biosynthesis and lignin accumulation, such as UGP, MYB308, MYB83, NAC043, NAC073, CCoAOMT1, CCoAOMT7, CHS2, CESA7, LAC7, COBL4, and IRX12. Overexpression of JrUGP and JrMYB308 in Arabidopsis thaliana confirmed their roles in lignin biosynthesis and cell wall thickening. Consequently, our comprehensive multi-omics findings offer novel insights into walnut genetic variation and network regulation of endocarp development and shell thickness, which enable further genome-informed breeding strategies for walnut cultivar improvement.

Juglans

Polygenic and monogenic adaptation drive evolutionary rescue at different magnitudes of environmental change.

Understanding the genetic basis of rapid adaptation is key to predicting species' evolutionary responses to environmental change. However, it is still debatable whether many small-effect mutations or a few large-effect mutations underlie rapid adaptation, and how this knowledge can predict population survival or extinction. To address this question, we performed a series of ecologically grounded forward-in-time genetic simulations to study rapid adaptation and extinction with increasing magnitudes of environmental change. These simulations were seeded with genomic variation of the plant Arabidopsis thaliana to have a realistic genomic structure, with one (monogenic) to 1,000 (polygenic) variants with varying heritabilities contributing to an environmental adaptive trait. Our results revealed two distinct scenarios of rapid adaptation and population rescue. Under small-to-moderate environmental shifts, high polygenic traits increased evolutionary rescue probability. Under extreme environmental shifts, high polygenic traits lead predictably to extinction, yet monogenic traits sometimes produce one-off winning adaptive genotypes. We interpret our rapid evolutionary rescue findings in terms of the fundamental theorem of natural selection, where trait polygenicity shapes the distribution of genetic variance in fitness across replicates and, in turn, the probability of population survival, with polygenic architectures producing more stable and predictable fitness variance and monogenic architectures generating highly skewed and variable outcomes. These results highlight the insights genomics gives us into the (un)predictability of species' evolutionary responses to global change, with management implications for assisted adaptation and conservation.

Arabidopsis

Evolution and domestication-trait associations of ultra-long centromere haplotypes in pepper plants.

Centromeric and pericentromeric regions of most eukaryotic genomes are highly repetitive and strongly recombination-suppressed, confounding efforts to resolve genetic variation, population structure and phenotypic associations. Pepper (Capsicum annuum) centromeres are nearly devoid of satellite repeats, facilitating assembly and population-level comparison of centromeric regions. Here we integrate 9 near-complete genome assemblies, CENH3 ChIP-seq profiles from 26 diverse accessions, and resequencing and phenotypic data from ~400 cultivated and wild accessions to investigate population-level diversity and phenotypic relevance of pepper peri/centromeric regions. Functional centromere positions are largely fixed on 8 of 12 chromosomes, whereas the remaining 4 carry distinct centromeric epialleles shaped mainly by centromere repositioning and pericentromeric inversions. Pepper centromeres are embedded within ultra-long centromere-spanning haplotype (cenhap) blocks, ranging from 29.8 to 112.9&#x2009;Mb and collectively covering 23.96% of the genome; each block contains only 1-4 major haplotypes. Some cenhaps may act as supergene-like units and are strongly associated with fruit traits, probably because recombination-suppressed intervals harbour multiple fruit-related genes, including OFP and F-box genes. F2 segregation assays further reveal transmission distortion of chromosomes carrying alternative cenhaps. Together, these findings highlight peri/centromeric regions as underrecognized reservoirs of agronomically important variation.

Centromere

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics

A family of retrotransposons and associated genomic variation in wheat.

A family of related retroelements was characterized in the genomes of some Graminease species. The structure of these retroelements indicates that they are retrotransposons containing reading frames with sequence similarity to the polyproteins of copia and Ty. This family of retroelements (termed WIS-2) occurs in the genomes of barley, wheat, rye, oats, and Aegilops species. Ongoing genomic variation both within individual plants of a wheat variety and within and between varieties of wheat is associated with some members of the WIS-2 family.

Amino Acid Sequence

Integrating Optical Genome Mapping into the Genetic Diagnostic Algorithm: Clinical Utility in Unresolved Autosomal Recessive Disorders from a Large Cohort.

INTRODUCTION: The identification of precise genetic etiologies is indispensable for the clinical management of monogenic disorders. However, conventional diagnostic methods and exome sequencing (ES) frequently fail to identify complex structural variations (SVs), leaving the genetic basis unexplained in approximately 30-60% of suspected cases. Optical genome mapping (OGM) emerges as a high-resolution technology capable of detecting cryptic SVs inaccessible to standard methodologies. METHODS: In this study, we evaluated the clinical utility of integrating OGM into the diagnostic algorithm for unresolved monogenic diseases. Following negative or inconclusive results from standard ES pipelines, OGM was applied to a targeted subset of patients (n = 7) selected from a comprehensive clinical cohort of 1,257 individuals with suspected genetic disorders. RESULTS: The integration of OGM identified candidate SVs that may represent the second allelic alteration in two distinct cases; however, confirmation through parental segregation analysis remains pending. Specifically, OGM identified an intronic insertion in the TTLL5 gene and a deletion in a putative regulatory region approximately 400 kb upstream of the NMNAT1 gene, both of which were missed by prior diagnostic testing. CONCLUSION: Our findings suggest that OGM has potential value in investigating the missing heritability of autosomal recessive disorders. By detecting candidate SVs invisible to conventional methods, OGM may warrant consideration as a complementary diagnostic approach following inconclusive ES; however, larger cohorts and confirmatory functional studies are needed to establish its clinical utility.

Autosomal recessive disorders

Comparative genomics reveals population structure and functional differentiation in Limosilactobacillus fermentum.

Limosilactobacillus fermentum is a widely distributed lactic acid bacterium frequently detected in fermented foods and host-associated microbiota, yet its global genomic diversity and functional variability remain insufficiently characterized. Here, we performed a large-scale comparative genomic analysis of 336 high-quality L. fermentum genomes curated from public databases. Species identity was validated using average nucleotide identity (ANI), and population structure was examined using pairwise ANI comparisons together with Mash-based phylogenetic reconstruction. Clustering at &#x2265;&#x2009;99% ANI resolved the dataset into 15 genomic clusters, with four dominant lineages comprising the majority of genomes. Pangenome reconstruction identified 5,853 gene clusters, including 1,325 core genes (22.6%) and a large accessory component dominated by low-frequency genes. Heap's law modeling (&#x3bb;&#x2009;=&#x2009;0.19) indicated a weakly open pangenome, suggesting ongoing gene acquisition as additional genomes are sampled. Functional annotation revealed that core genes were primarily associated with essential cellular processes, whereas accessory genes were enriched in carbohydrate metabolism, membrane-associated functions, and defense-related systems. Variation in carbohydrate-active enzymes (CAZymes), transport systems, and stress-response genes was observed across lineages, indicating strain-level functional diversity. Although genomes from human and food sources were broadly distributed across phylogenetic lineages, multivariate analysis showed that gene-content variation was more strongly associated with genomic lineage than with isolation source. These results provide a population genomic framework for understanding genomic diversity and functional potential in L. fermentum.

Phylogeny

PangyPlot: multi-scale interactive visualization of pangenome variation graphs.

SUMMARY: Pangenome variation graphs integrate multiple samples into a unified representation, mitigating the reference bias inherent to linear genomes. However, these graphs can be large and structurally complex. Existing visualization tools are each confined to a fixed scale of resolution, requiring researchers to switch between multiple tools to examine variation at different levels of detail. PangyPlot is an interactive pangenome browser designed for multi-scale exploration of reference variation graphs from full chromosome to nucleotide-level sequence segments. PangyPlot anchors navigation to linear reference coordinates, organizes variation into hierarchical bubble structures, and uses a force-directed layout engine for automatic node arrangement. AVAILABILITY AND IMPLEMENTATION: An instance preloaded with data is available at https://pangyplot.research.sickkids.ca. Source code and documentation are openly available at https://github.com/strug-hub/pangyplot under the MIT License.

Software

Giant G+C% mosaic structures of the human genome found by arrangement of GenBank human DNA sequences according to genetic positions.

To determine the overall variation in the G+C% distribution over long ranges of the human genome, DNA sequences of human genes, which were closely linked genetically or physically, were surveyed from the GenBank Data Bank. A total of 72 sequences longer than 2 kb, which were mutually linked within 500 kb, were identified. The sequences belonged to 17 linkage groups and were ordered in each group according to their genetic positions. Analyses of the G+C% distribution along the ordered sequences showed that sequences within each group almost always had similar G+C% levels, but those belonging to different groups often had different levels. Similar analyses of more distantly linked sequences (e.g., greater than 10 Mb) showed mosaic structures of G+C% distribution. These findings are consistent with predictions made from the "isochore" structures found by CsCl equilibrium centrifugation, in that the structures having homogeneous base compositions stretched over at least several hundred kilobases. A possible boundary of the giant G+C% mosaic structures was identified between X-linked G6PD and F8C.

Base Composition

Substantial non-homologous recombination and structural variation results from Brassica AABC and CCAB hybrid meiosis.

Meiotic crossovers contribute to genetic diversity and play a crucial role in homologous chromosome segregation. Non-homologous crossovers in Brassica, involving the exchange of genetic material between genomes, can be valuable for transferring novel traits or characteristics between Brassica species. However, there are a limited number of studies that specifically investigate crossover frequencies in populations of interspecific hybrids. We investigated the distribution and frequency of homologous crossover events, as well as non-homologous recombination and structural variation, in hybrids between B. juncea (AABB)&#x2009;&#xd7;&#x2009;B. napus (AACC) (resulting in AABC hybrids; 5 genotypes) and B. napus (AACC)&#x2009;&#xd7;&#x2009;B. carinata (BBCC) (resulting in CCAB hybrids; 4 genotypes). The analysis was performed on individuals derived from microspore culture of both unreduced and reduced gametes produced by the AABC and CCAB hybrids. All AABC and almost all CCAB unreduced gamete-derived individuals and most AABC and CCAB reduced gamete-derived individuals showed copy number variation indicative of non-homologous (A-C) recombination. Additionally, a higher frequency of homologous crossovers, also in centromeric and pericentromic regions, was observed in the diploid genomes of the AABC and CCAB hybrids. Overall, these hybrid types show high frequencies of A-C introgressions, which may be useful in B. juncea or B. carinata introgression breeding, and this increased recombination frequency may help break up existing linkage disequilibrium blocks in the Brassica A and C genomes.

Meiosis

Optical genome mapping enhanced by refined variant interpretation in pediatric acute lymphoblastic leukemia.

Reliable detection of structural variants (SVs) and copy number variations (CNVs) is crucial in the contemporary diagnostics of pediatric B-cell acute lymphoblastic leukemia (B-ALL). However, limitations of commonly used conventional and molecular cytogenetic methods may hinder the accurate genetic characterization of patients. Optical genome mapping (OGM) offers a reliable alternative by enabling high-resolution, genome-wide detection of CNVs and SVs. Chromosomal aberrations were screened using OGM in 51 children with B-ALL. The results were compared with those of karyotyping, fluorescence in situ hybridization (FISH), digital multiplex ligation-dependent probe amplification (digitalMLPA), and targeted RNA sequencing (RNA-seq). OGM data showed high congruency with karyotyping and FISH findings, detecting clinically relevant variants beyond G-banding results and unraveling a complex KMT2A fusion undetected by FISH. Gene fusions involved in complex ETV6::RUNX1 translocations, but not detected by RNA-seq, were confirmed using FISH. Normalization of OGM copy number values with DNA-index-improved concordance with FISH-derived copy numbers in near-tri/tetraploid cases. In the peripheral regions of OGM variants (fringe-zones), a novel evaluation strategy called 'FriZone' was applied, which significantly improved the concordance between OGM and digitalMLPA. In addition, a co-segregation analysis revealed strong associations between ETV6::RUNX1 fusion and deletions of ETV6, RAG2, and NR3C2. OGM uncovered complex rearrangements undetected by widely used methods in 15% of cases, improving genetic classification and risk stratification in 10% of the patients. The FriZone analysis and normalization by DNA-index provide a refined, more accurate approach to OGM variant interpretation, facilitating the efficient application of OGM in clinical diagnostics. &#xa9; 2026 The Author(s). The Journal of Pathology published by John Wiley & Sons Ltd on behalf of The Pathological Society of Great Britain and Ireland.

Humans