Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Gene conversion confined to a direct repeat of the acceptor splice site generates allelic diversity at human glycophorin (GYP) locus.

The glycophorin locus (GYP) on the long arm of chromosome 4 encodes antigens of the MNSs blood group system and displays considerable allelic variation among human populations. The genomic structure and organization of a variant glycophorin allele specifying a novel Miltenberger (Mi)-related phenotype, MiX, were examined. This variant probably arose from a gene conversion event involving a direct repeat of the acceptor splice site. Southern blot analysis indicated that MiX gene derived its 5' and 3' portions from glycophorin B or delta gene but its internal part from glycophorin A or alpha gene. Genomic sequences encompassing the rearranged regions of the MiX gene were amplified by single copy polymerase chain reaction. Direct DNA sequencing showed that during the formation of MiX gene, a short stretch of alpha exon III with a donor splice site has replaced a silent sequence in the delta gene containing a cryptic acceptor splice site. The upstream delta-alpha breakpoint is flanked by the direct repeats of the acceptor splice site, whereas the down-stream alpha-delta breakpoint is located in the adjacent intron. This segmental transfer produced a new composite exon whose expression not only transactivated a portion of silent sequence but also created intraexon and interexon hybrid junctions that characterize the antigenic specificities of MiX glycophorin. The identification of MiX as yet another delta-alpha-delta hybrid different from MiIII and MiVI in gene conversion sites suggests that shuffling of expressed and unexpressed sequences through particular genomic DNA motifs has been an important mechanism for shaping the antigenic diversity of MNSs blood group system during evolution.

Alleles

The cold case of state transition 7 (stt7) mutants of Chlamydomonas reinhardtii, solved by whole-genome sequencing.

The process of State Transitions (ST) corresponds to an STT7 kinase-driven redistribution of the transmembrane LHCII antenna proteins between Photosystem II (PSII) and Photosystem I (PSI), which results from changes in their phosphorylation state. For the past two decades, two LHCII-kinase mutants, stt7-1 and stt7-9, have been instrumental in the study of STs in Chlamydomonas reinhardtii, the former being a null mutant for the kinase but quasi-sterile in crosses, while the latter, although fertile, has a leaky phenotype. Using long-read sequencing, this study further characterized the genetic lesions of the stt7 mutant strains through whole-genome reconstruction and de novo chromosome assembly. In addition, two new stt7 null mutants were generated, one derived by crosses from the original stt7-1 and one obtained by Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated protein 9 (Cas9) technology. This work provides a comprehensive genomic characterization of the original stt7-1 null mutant, revealing extensive chromosomal rearrangements and high levels of aneuploidy, associated with increased cell size and meiotic dysfunction. Reassessment of their physiology and genetic backgrounds highlights the need for caution in interpreting genetic information. We thus produced more reliable null mutants for the LHCII-kinase, amenable to genetic crosses for the study of STs in a variety of genetic backgrounds.

Chlamydomonas reinhardtii

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis

Complete plastid genome of Iris orchioides and comparative analysis with 19 Iris plastomes.

Iris is a cosmopolitan genus comprising approximately 280 species distributed throughout the Northern Hemisphere. Although Iris is the most diverse group in the Iridaceae, the number of taxa is debatable owing to various taxonomic issues. Plastid genomes have been widely used for phylogenetic research in plants; however, only limited number of plastid DNA markers are available for phylogenetic study of the Iris. To understand the genomic features of plastids within the genus, including its structural and genetic variation, we newly sequenced and analyzed the complete plastid genome of I. orchioides and compared it with those of 19 other Iris taxa. Potential plastid markers for phylogenetic research were identified by computing the sequence divergence and phylogenetic informativeness. We then tested the utility of the markers with the phylogenies inferred from the markers and whole-plastome data. The average size of the plastid genome was 152,926 bp, and the overall genomic content and organization were nearly identical among the 20 Iris taxa, except for minor variations in the inverted repeats. We identified 10 highly informative regions (matK, ndhF, rpoC2, ycf1, ycf2, rps15-ycf, rpoB-trnC, petA-psbJ, ndhG-ndhI and psbK-trnQ) and inferred a phylogeny from each region individually, as well as from their concatenated data. Remarkably, the phylogeny reconstructed from the concatenated data comprising three selected regions (rpoC2, ycf1 and ycf2) exhibited the highest congruence with the phylogeny derived from the entire plastome dataset. The result suggests that this subset of data could serve as a viable alternative to the complete plastome data, especially for molecular diagnoses among closely related Iris taxa, and at a lower cost.

Iris Plant

Diversity of ribosomes at the level of rRNA variation associated with human health and disease.

With hundreds of copies of rDNA, it is unknown whether they possess sequence variations that form different types of ribosomes. Here, we developed an algorithm for long-read variant calling, termed RGA, which revealed that variations in human rDNA loci are predominantly insertion-deletion (indel) variants. We developed full-length rRNA sequencing (RIBO-RT) and in situ sequencing (SWITCH-seq), which showed that translating ribosomes possess variation in rRNA. Over 1,000 variants are lowly expressed. However, tens of variants are abundant and form distinct rRNA subtypes with different structures near indels as revealed by long-read rRNA structure probing coupled to dimethyl sulfate sequencing. rRNA subtypes show differential expression in endoderm/ectoderm-derived tissues, and in cancer, low-abundance rRNA variants can become highly expressed. Together, this study identifies the diversity of ribosomes at the level of rRNA variants, their chromosomal location, and unique structure as well as the association of ribosome variation with tissue-specific biology and cancer.

Humans

Ancient DNA and Human Physiology.

Ancient DNA (aDNA) enables the reconstruction of chronologically sampled genomes from ancient humans, animals, plants, pathogens, and microorganisms, as well as environmental DNA, providing a record of biological changes through time. Improvements in short and degraded DNA extraction methods and low-cost sequencing now enable the generation of broad, cross-regional datasets that expand evolutionary analyses from past population demography to biological mechanisms. By tracking temporal shifts of allele frequencies, integrating functional genomics resources (e.g., gene expression, chromatin structure variation), modeling population demography to separate selection from genetic drift, and aligning genetic changes with archaeological, cultural, and climatic data, aDNA has the potential to link sequence variation to physiological function within their temporal and environmental contexts. In this review, we summarize illustrative case studies from aDNA research spanning complex traits, dietary adaptations, and responses to pathogens and other environmental changes, showing how human biology has evolved under multiple selective pressures through time. These dated signals help triage experimental work and expose mechanisms that are rare or absent in living cohorts. Although some challenges remain, such as geographic and temporal sampling disparities, limitations in data resolution and variant detection, and genotype-phenotype uncertainties, rapid methodological progress and stronger ethical frameworks are expanding what can be inferred, making aDNA a promising tool for refining physiological pathways, their timing, and their drivers.

Humans

Polyploidy-mediated variations in glutamate receptor proteins linked to Fusarium wilt resistance in upland cotton.

Cotton production in the US faces a serious threat from Fusarium oxysporum f. sp. vasinfectum race 4 (FOV4), a soil-borne fungus causing Fusarium wilt by infecting the roots and vascular system of susceptible cotton, leading to rapid wilting and death. Here, we investigate genetic mechanisms of resistance to FOV4 in the highly resistant upland cotton genotype "U1" using an early-generation segregating biparental population ("U1" × "CSX8308") with comprehensive genomic resources. Reference-grade genomic assemblies of the parents revealed minor structural variations between "U1" haplotypes, a high degree of collinearity at chromosome synteny and micro-synteny levels, and significant divergence from "CSX8308" with 8.9 million SNPs. QTL analysis identified significant markers on chromosomes D03 and A02 linked to reduced Fusarium wilt severity. Within these regions, two glutamate-receptor-like (GLR) genes showed structural variation and overlapped between translocated segments on A02 and D03, suggesting a rare but important reinforcing effect of parallel evolution between susceptible and resistant genotypes. Transcriptome profiles of "U1" under FOV4 infection reveal activation of calcium-binding proteins and transcription factors regulating plant hormones (ethylene, abscisic acid, jasmonic acid, and salicylic acid), along with enzymes involved in cell wall remodeling and phytoalexin production. Advancing cotton improvement depends on incorporating durable genetic disease resistance into high-yielding, high-quality cultivars.

Fusarium

Extensive longevity and DNA virus-driven adaptation in nearctic Myotis bats.

The genus Myotis is one of the largest clades of bats, and exhibits some of the most extreme variation in lifespans among mammals alongside unique adaptations to viral tolerance and immune defense. To study the evolution of longevity-associated traits and infectious disease, we generated cell lines and near-complete genome assemblies for 8 closely related species of Myotis. Using genome-wide screens of positive selection, analyses of structural variation, and functional experiments in primary cells, we identify new patterns of adaptation contributing to longevity, cancer resistance, and viral interactions in bats. We show that the recurrent evolution of longevity seen in Myotis leads to some of the highest predicted increases in cancer risk across mammals and demonstrate a unique DNA damage response in primary cells of the long-lived M. lucifugus. We also find evidence of abundant adaptation in response to DNA viruses - but not RNA viruses - in Myotis and other bats in sharp contrast with other mammals, potentially contributing to the role of bats as reservoirs of zoonoses. Together, our results demonstrate how genomics and primary cells derived from diverse taxa uncover the molecular bases of extreme adaptations in non-model organisms.

Aging

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals

Plastid genome evolution and phylogenomics with broad taxon sampling: insights into intrafamilial classification of Hamamelidaceae.

Hamamelidaceae, within the order Saxifragales, comprises 27 genera and approximately 120 species. The family has a pantropical and temperate distribution across the Americas, Asia, Africa, and Australia. Previous molecular investigations, constrained by limited taxon sampling and inadequate genetic markers, supported a five-subfamily classification system. However, these studies predominantly focused on Asian taxa, resulting in poor resolution of the evolutionary relationships among American, African, and Australian genera. To address these sampling gaps, we employed near-complete generic sampling (26 of 27 genera) to investigate plastome architecture, structural variation, and phylogenetic relationships. We newly sequenced and assembled 15 plastid genomes representing geographically and taxonomically underrepresented genera and analyzed them alongside 59 publicly available plastomes retrieved from GenBank. Plastid genomes exhibited conserved quadripartite architecture with sizes ranging from 158, 076 bp to 160, 814 bp, minimal structural variation, consistent GC content (37.7-38.2%), and identical gene order. Inverted repeat (IR) regions had limited size variation (26, 211-26, 429 bp). Simple sequence repeat (SSR) distribution (2, 219 loci) showed no clear correlation with the genus-level phylogenetic relationships. We identified ten hypervariable regions, including coding sequences (accD, ycf1, clpP, ndhF, and rpl22) and intergenic spacers (rpl33-rps18, the trnG-UCC intron, trnH-GUG-psbA, accD-psaI, and petA-psbJ), as promising candidate regions for future applications in species delimitation and phylogenetic studies. Phylogenetic analyses revealed largely congruent topologies across datasets and methods, providing improved resolution and strong support for most subfamilial and tribal relationships compared with previous studies. This study highlights the utility of plastid genome data for resolving deep-level phylogenetic relationships within Hamamelidaceae. The genome architecture reflects the high conservation of plastid genomes, while the identified mutation hotspots represent potential resources for future taxonomic and phylogenetic studies. Our results support the existing subfamily classification while improving geographical coverage and generic representation, providing a robust framework for future taxonomic and evolutionary studies of this globally distributed and taxonomically complex family.

Hamamelidaceae

Comparisons Between Large-Scale Genomic Variants and SNPs in Driving Population Divergence and Local Adaptation.

Genomic variations, such as indels (2-49 bp) and structural variants (SVs, ≥50 bp), are larger-scale mutations than single nucleotide polymorphisms (SNPs) and can substantially impact evolutionary processes, including speciation, adaptation, and phenotypes. Despite their functional importance, integrative population genetic analyses that jointly consider genome-wide SNPs, indels, and SVs remain under-explored. The ground tit (Pseudopodoces humilis), an endemic species to the Qinghai-Tibet Plateau (QTP), exhibits divergence across distinct glacial refugia, accompanied by habitat and morphological divergence, making it an excellent example for investigating how different types of genomic variants contribute to population divergence and local adaptation. Here, by retrieving 81 whole-genome sequence data, over 13 million SNPs, 2 million indels, and 22,101 SVs were identified. Variants were unevenly distributed across the genome, characterized by distinct hotspot regions. Indels and SVs revealed four genetic clusters consistent with previous SNP-based results, thereby validating the reliability of our variant datasets. FST and genotype-environment association (GEA) analyses independently revealed numerous candidate indels and SVs; each showed minimal overlap with previously identified SNPs, and were enriched in similar functional pathways such as signal transduction, skeletal muscle development, water transport, DNA repair, reproduction, nervous system development, and immunity. Collectively, our results demonstrated that indels and SVs could capture additional signatures besides SNPs. Furthermore, similar but distinct gene functions among different types of genomic variants collectively and complementarily drive genomic divergence across environmental gradients in such a high-elevation endemic species, underscoring its evolutionary relevance in local adaptation.

indels

Dissecting genetic architecture and improving machine learning‑based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC > 0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

A haploid wild yeast resource for exploring the natural ecology of Saccharomyces cerevisiae.

Saccharomyces cerevisiae occurs predominantly in the diploid state in nature, limiting genetic analyses of wild populations. Here, we establish a haploid collection from 32 Taiwanese S. cerevisiae isolates through targeted HO disruption. This resource spans predomesticated Asian wild lineages and enables the investigation of reproductive isolation and ecological trait variation. Although all pairwise hybridizations formed zygotes, many yielded reduced spore viability, revealing strong postzygotic barriers. Genome analyses associated reduced hybrid fertility with lineage-specific structural variation, including elevated levels of intra-chromosomal inversions in H413-8/TW1 and inter-chromosomal rearrangements in PD35A/CHN-V, rather than sequence divergence alone. Phenotyping revealed ecological differentiation, with TW1 favoring cooler growth and a natural hybrid exhibiting heterosis with expanded thermotolerance. Most wild strains grew poorly on maltose, whereas anthropogenic strains displayed enhanced utilization linked to MAL + regulatory alleles and maltose-specific transporters. Together, this haploid collection links structural variation and metabolic divergence to ecological and reproductive differentiation in wild S. cerevisiae.

Saccharomyces cerevisiae

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial

STK11 Mutations and Deletions Define an Aggressive Molecular Subgroup of Cervical Adenocarcinoma.

Cervical adenocarcinoma accounts for 15%-20% of cervical cancers and is associated with poorer survival and reduced response to screening and immunotherapy compared with squamous cell carcinoma (SCC). The genomic drivers underlying this molecular subgroup remain incompletely characterized. Whole-exome sequencing was performed on 302 invasive cervical cancers from Guatemala and Venezuela. Structural variation analysis was conducted using SNP-array and whole-genome sequencing data. Findings were replicated in more than 4600 additional cervical cancer samples from TCGA, AACR Project GENIE, MSKCC, and Caris datasets. TP53 mutations were more frequent in adenocarcinoma than SCC, particularly in HPV-negative tumors. STK11 alterations, including mutations and focal deletions, were significantly enriched in HPV-positive adenocarcinomas compared with SCC and affected 23% of adenocarcinomas overall. Whole-genome analyses identified recurrent focal deletions, inversions, chromosomal rearrangements, and breakage-fusion-bridge events involving chromosome 19p and STK11 that were not detected by exome sequencing alone. STK11 alterations were associated with younger age at diagnosis, poorer overall survival, and inferior outcomes following immune checkpoint inhibitor (ICI) therapy. STK11 alterations significantly co-occurred with YAP1 amplification but were largely mutually exclusive with PIK3CA mutation. Cervical adenocarcinomas also demonstrated significantly lower CD274 (PD-L1) expression than SCC. STK11 alterations define a distinct molecular subgroup of cervical adenocarcinoma characterized by structural disruption of chromosome 19p, younger age at onset, and poorer clinical outcomes. These findings have implications for molecular classification and future targeted therapeutic approaches in cervical cancer.

Humans

Panmixia in a Widespread Butterfly: High Dispersal and Ecological Generalism Buffer Against Landscape Fragmentation.

Habitat fragmentation is widely expected to reduce population connectivity and increase genetic differentiation, although the strength of these effects depends on species-specific traits such as dispersal ability. Here, we investigated the population genetic structure of the cosmopolitan butterfly, Pieris rapae L. (Lepidoptera: Pieridae), across western Germany using genome-wide single-nucleotide polymorphism (SNP) data. To analyze the effects of landscape structure on genetic connectivity, we applied a paired study design comprising four landscape pairs, each consisting of a highly intensified, modern agricultural landscape and a more heterogeneous, traditional landscape. Our results revealed no evidence of genetic differentiation. Pairwise FST values were close to zero; we detected no isolation by distance, and clustering analyses supported a single genetic population. No meaningful associations between genetic variation and environmental variables were detected, with landscape effects explaining less than 0.4% of genomic variation. Consequently, we found no evidence for stronger genetic structuring in modern compared to more connected traditional landscapes. Our results suggest that extensive habitat fragmentation does not necessarily translate into reduced genetic connectivity in highly mobile, generalist species. In P. rapae , high dispersal ability and ecological generalism appear to buffer against the genetic consequences of landscape modification, resulting in panmictic population structure even across strongly contrasting agricultural landscapes.

Pieris rapae

Panorama of Chromosomal Instability in Lung Cancer.

Lung cancer is a highly heterogeneous disease primarily driven by tobacco smoking. About 20% of lung cancers occur among patients who have never smoked (LCINS) with differences in patient ancestry, sex, tumor histology, and clinical features. Our understanding of chromosomal instability in lung cancer, especially LCINS, is still limited. Here, we perform a comprehensive study of 182,429 somatic structural variations (SVs) detected in 1,209 whole-genome sequenced lung cancers, of which 864 LCINS. SVs are more abundant in tumors from patients who have smoked (LCSS); however, they are more complex and play more important roles in tumorigenesis in LCINS. EGFR mutations and KRAS mutations profoundly and independently shape the SV landscape. EGFR-mutant tumors have higher SV burden and more cancer-driving SVs. In contrast, KRAS mutations are associated with lower SV burden and less driver SVs. We decompose 16 SV signatures for both complex and simple SVs that likely represent divergent molecular mechanisms. The SV breakpoints have distinct distributions across the genome depending on the signatures due to mutagenic mechanisms and positive selection. Many established cancer-driving genes are recurrently rearranged by multiple SV signatures suggesting functional convergence of these genome instability mechanisms.

Journal Article

Dual-dimensional profiling of host genomic variations and HPV integration in PD-L1-stratified cervical cancer via Oxford Nanopore Technology.

BACKGROUND: The integration of human papillomavirus (HPV) DNA into the host genome is a key step in the development of HPV-associated cervical cancer (CC). However, the genomic characteristics of host genomic variations and HPV integration within the context of programmed death-ligand 1 (PD-L1) expression stratification have not been systematically investigated. METHODS: Whole-genome sequencing was performed using Oxford Nanopore Technology (ONT) on six samples (three from the high PD-L1 expression group and three from the low PD-L1 expression group). The characteristics of host genomic variations under different PD-L1 expression stratifications were explored, including structural variations (SV), copy number variations (CNV), single nucleotide polymorphisms (SNP), and insertion-deletions (Indel). Subsequently, the distribution features of HPV integration sites were analyzed, different integration types were identified, and pathway analysis was conducted. RESULTS: Whole-genome SV analysis revealed that the total number of SVs and the composition of mutation types were similar between the high and low PD-L1 expression groups, with insertions (INS) and deletions (DEL) predominating in both. These variations were primarily enriched in intergenic regions and introns. In the low PD-L1 expression group, integration events were observed at multiple chromosomal loci, with the most frequent integration occurring in the KLF5 gene region on chromosome 13. No frequently integrated loci were identified in the high PD-L1 expression group. Additionally, four distinct HPV integration breakpoint patterns were preliminarily identified and analyzed. CONCLUSION: PD-L1 expression stratification did not significantly alter the overall genomic instability of the host. However, differences were observed in the distribution patterns of HPV integration sites. These findings provide new insights into the genomic heterogeneity of CC under different PD-L1 expression backgrounds and may lay the groundwork for future research exploring stratified immunotherapy based on HPV integration features.

Humans