Search PubMedSearch

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics

A family of retrotransposons and associated genomic variation in wheat.

A family of related retroelements was characterized in the genomes of some Graminease species. The structure of these retroelements indicates that they are retrotransposons containing reading frames with sequence similarity to the polyproteins of copia and Ty. This family of retroelements (termed WIS-2) occurs in the genomes of barley, wheat, rye, oats, and Aegilops species. Ongoing genomic variation both within individual plants of a wheat variety and within and between varieties of wheat is associated with some members of the WIS-2 family.

Amino Acid Sequence

Polymorphic variations in the ori sequences from the mitochondrial genomes of different wild-type yeast strains.

We determined the restriction maps and primary structures of two as yet poorly characterized regions of the mitochondrial genomes of different wild-type strains of Saccharomyces cerevisiae. These regions respectively comprised the ori1 sequence and the newly identified ori8 sequence. Ori1 and ori8, together with their flanking sequences, exhibit a large polymorphism, resulting from specific variations due to insertions or deletions of optional GC clusters at different locations. The mechanisms underlying such sequence rearrangements are discussed.

Base Sequence

Integrating Optical Genome Mapping into the Genetic Diagnostic Algorithm: Clinical Utility in Unresolved Autosomal Recessive Disorders from a Large Cohort.

INTRODUCTION: The identification of precise genetic etiologies is indispensable for the clinical management of monogenic disorders. However, conventional diagnostic methods and exome sequencing (ES) frequently fail to identify complex structural variations (SVs), leaving the genetic basis unexplained in approximately 30-60% of suspected cases. Optical genome mapping (OGM) emerges as a high-resolution technology capable of detecting cryptic SVs inaccessible to standard methodologies. METHODS: In this study, we evaluated the clinical utility of integrating OGM into the diagnostic algorithm for unresolved monogenic diseases. Following negative or inconclusive results from standard ES pipelines, OGM was applied to a targeted subset of patients (n = 7) selected from a comprehensive clinical cohort of 1,257 individuals with suspected genetic disorders. RESULTS: The integration of OGM identified candidate SVs that may represent the second allelic alteration in two distinct cases; however, confirmation through parental segregation analysis remains pending. Specifically, OGM identified an intronic insertion in the TTLL5 gene and a deletion in a putative regulatory region approximately 400 kb upstream of the NMNAT1 gene, both of which were missed by prior diagnostic testing. CONCLUSION: Our findings suggest that OGM has potential value in investigating the missing heritability of autosomal recessive disorders. By detecting candidate SVs invisible to conventional methods, OGM may warrant consideration as a complementary diagnostic approach following inconclusive ES; however, larger cohorts and confirmatory functional studies are needed to establish its clinical utility.

Autosomal recessive disorders

Comparative genomics reveals population structure and functional differentiation in Limosilactobacillus fermentum.

Limosilactobacillus fermentum is a widely distributed lactic acid bacterium frequently detected in fermented foods and host-associated microbiota, yet its global genomic diversity and functional variability remain insufficiently characterized. Here, we performed a large-scale comparative genomic analysis of 336 high-quality L. fermentum genomes curated from public databases. Species identity was validated using average nucleotide identity (ANI), and population structure was examined using pairwise ANI comparisons together with Mash-based phylogenetic reconstruction. Clustering at ≥ 99% ANI resolved the dataset into 15 genomic clusters, with four dominant lineages comprising the majority of genomes. Pangenome reconstruction identified 5,853 gene clusters, including 1,325 core genes (22.6%) and a large accessory component dominated by low-frequency genes. Heap's law modeling (λ = 0.19) indicated a weakly open pangenome, suggesting ongoing gene acquisition as additional genomes are sampled. Functional annotation revealed that core genes were primarily associated with essential cellular processes, whereas accessory genes were enriched in carbohydrate metabolism, membrane-associated functions, and defense-related systems. Variation in carbohydrate-active enzymes (CAZymes), transport systems, and stress-response genes was observed across lineages, indicating strain-level functional diversity. Although genomes from human and food sources were broadly distributed across phylogenetic lineages, multivariate analysis showed that gene-content variation was more strongly associated with genomic lineage than with isolation source. These results provide a population genomic framework for understanding genomic diversity and functional potential in L. fermentum.

Phylogeny

A variation in the structure of the protein-coding region of the human p53 gene.

An extensive analysis of genomic DNA preparations from a number of normal and malignant tissues revealed BglII site polymorphism of the human p53 gene. Approximately 10% of p53 gene alleles were found to contain an additional BglII site localized in a region of intron I. This allelic form of p53 gene was also responsible for p53 protein having altered electrophoretic mobility. Molecular cloning and sequencing of both the alleles of p53 gene revealed a base-pair change in codon 72 causing arginine----proline substitution in the allele with the additional BglII site. Both variants of the p53 gene may occur in homozygous state and are therefore functional.

Amino Acid Sequence

PangyPlot: multi-scale interactive visualization of pangenome variation graphs.

SUMMARY: Pangenome variation graphs integrate multiple samples into a unified representation, mitigating the reference bias inherent to linear genomes. However, these graphs can be large and structurally complex. Existing visualization tools are each confined to a fixed scale of resolution, requiring researchers to switch between multiple tools to examine variation at different levels of detail. PangyPlot is an interactive pangenome browser designed for multi-scale exploration of reference variation graphs from full chromosome to nucleotide-level sequence segments. PangyPlot anchors navigation to linear reference coordinates, organizes variation into hierarchical bubble structures, and uses a force-directed layout engine for automatic node arrangement. AVAILABILITY AND IMPLEMENTATION: An instance preloaded with data is available at https://pangyplot.research.sickkids.ca. Source code and documentation are openly available at https://github.com/strug-hub/pangyplot under the MIT License.

Software

Giant G+C% mosaic structures of the human genome found by arrangement of GenBank human DNA sequences according to genetic positions.

To determine the overall variation in the G+C% distribution over long ranges of the human genome, DNA sequences of human genes, which were closely linked genetically or physically, were surveyed from the GenBank Data Bank. A total of 72 sequences longer than 2 kb, which were mutually linked within 500 kb, were identified. The sequences belonged to 17 linkage groups and were ordered in each group according to their genetic positions. Analyses of the G+C% distribution along the ordered sequences showed that sequences within each group almost always had similar G+C% levels, but those belonging to different groups often had different levels. Similar analyses of more distantly linked sequences (e.g., greater than 10 Mb) showed mosaic structures of G+C% distribution. These findings are consistent with predictions made from the "isochore" structures found by CsCl equilibrium centrifugation, in that the structures having homogeneous base compositions stretched over at least several hundred kilobases. A possible boundary of the giant G+C% mosaic structures was identified between X-linked G6PD and F8C.

Base Composition

Substantial non-homologous recombination and structural variation results from Brassica AABC and CCAB hybrid meiosis.

Meiotic crossovers contribute to genetic diversity and play a crucial role in homologous chromosome segregation. Non-homologous crossovers in Brassica, involving the exchange of genetic material between genomes, can be valuable for transferring novel traits or characteristics between Brassica species. However, there are a limited number of studies that specifically investigate crossover frequencies in populations of interspecific hybrids. We investigated the distribution and frequency of homologous crossover events, as well as non-homologous recombination and structural variation, in hybrids between B. juncea (AABB) × B. napus (AACC) (resulting in AABC hybrids; 5 genotypes) and B. napus (AACC) × B. carinata (BBCC) (resulting in CCAB hybrids; 4 genotypes). The analysis was performed on individuals derived from microspore culture of both unreduced and reduced gametes produced by the AABC and CCAB hybrids. All AABC and almost all CCAB unreduced gamete-derived individuals and most AABC and CCAB reduced gamete-derived individuals showed copy number variation indicative of non-homologous (A-C) recombination. Additionally, a higher frequency of homologous crossovers, also in centromeric and pericentromic regions, was observed in the diploid genomes of the AABC and CCAB hybrids. Overall, these hybrid types show high frequencies of A-C introgressions, which may be useful in B. juncea or B. carinata introgression breeding, and this increased recombination frequency may help break up existing linkage disequilibrium blocks in the Brassica A and C genomes.

Meiosis

Genome size variation in North American minnows (Cyprinidae). II. Variation among 20 species.

Genome sizes (nuclear DNA contents) from 200 individuals representing 20 species of North American cyprinid fishes (minnows) were examined spectrophotometrically. The distributions of DNA values of individuals within populations of the 20 species were essentially continuous and normal; the distribution of DNA values among species was continuous and overlapping. These observations suggest that changes in DNA quantity in cyprinids are small in amount, involve both gains and losses of DNA, and are cumulative and independent in effect. Significant heterogeneity in mean genome size occurs both between individuals within populations of species and among species. The former averages maximally around 6% of the cyprinid genome and is nearly the same as the amount of DNA theoretically needed for the entire cyprinid structural gene component. The majority of the DNA content variation among the 20 species is distributed above the level of individuals within populations. Comparisons of average genome size difference or distance between individuals drawn from different levels of taxonomic organization indicate that considerably greater divergence in genome size has occurred in the extremely speciose cyprinid genus Notropis as compared with other North American cyprinid genera. This may suggest that genome size change is concentrated in speciation episodes. Finally, no associations were found between interspecific variation in genome size and five life-history characters. This suggests that much of the variation in genome size within and among the 20 species may be phenotypically inconsequential.

Animals

Optical genome mapping enhanced by refined variant interpretation in pediatric acute lymphoblastic leukemia.

Reliable detection of structural variants (SVs) and copy number variations (CNVs) is crucial in the contemporary diagnostics of pediatric B-cell acute lymphoblastic leukemia (B-ALL). However, limitations of commonly used conventional and molecular cytogenetic methods may hinder the accurate genetic characterization of patients. Optical genome mapping (OGM) offers a reliable alternative by enabling high-resolution, genome-wide detection of CNVs and SVs. Chromosomal aberrations were screened using OGM in 51 children with B-ALL. The results were compared with those of karyotyping, fluorescence in situ hybridization (FISH), digital multiplex ligation-dependent probe amplification (digitalMLPA), and targeted RNA sequencing (RNA-seq). OGM data showed high congruency with karyotyping and FISH findings, detecting clinically relevant variants beyond G-banding results and unraveling a complex KMT2A fusion undetected by FISH. Gene fusions involved in complex ETV6::RUNX1 translocations, but not detected by RNA-seq, were confirmed using FISH. Normalization of OGM copy number values with DNA-index-improved concordance with FISH-derived copy numbers in near-tri/tetraploid cases. In the peripheral regions of OGM variants (fringe-zones), a novel evaluation strategy called 'FriZone' was applied, which significantly improved the concordance between OGM and digitalMLPA. In addition, a co-segregation analysis revealed strong associations between ETV6::RUNX1 fusion and deletions of ETV6, RAG2, and NR3C2. OGM uncovered complex rearrangements undetected by widely used methods in 15% of cases, improving genetic classification and risk stratification in 10% of the patients. The FriZone analysis and normalization by DNA-index provide a refined, more accurate approach to OGM variant interpretation, facilitating the efficient application of OGM in clinical diagnostics. © 2026 The Author(s). The Journal of Pathology published by John Wiley & Sons Ltd on behalf of The Pathological Society of Great Britain and Ireland.

Humans

Comparison of the 5' and 3' untranslated genomic regions of virulent and attenuated foot-and-mouth disease viruses (strains O1 Campos and C3 Resende).

The complete 5' and 3' non-coding regions of two attenuated South American foot-and-mouth disease virus (FMDV) vaccine strains, O1C-O/E and C3R-O/E, and their corresponding virulent parental strains, O1 Campos and C3 Resende, have been cloned from polymerase chain reaction-amplified primary cDNA. Differences observed in the derived nucleotide sequences between attenuated and virulent viruses seem not to affect regulatory signal structures, supporting the theory that genetic variations, primarily in the 3' halves of the viral genomes, contribute to the attenuation phenotype of the vaccine strains. In addition, this is the first report on the complete sequence of the 5' untranslated region of a C-type aphthovirus. Approximately 10% of the nucleotides differ from the corresponding known sequences of serotypes A or O.

Aphthovirus

Population history rather than tree age contributes to the evolutionary importance of ancient trees in an endangered conifer.

Ancient trees are in global decline and face increasing conservation challenges. Their exceptional longevity has fostered the view that they are genetic reservoirs, yet whether old age is synonymous with unique genetic variation remains unclear. Here we assembled a ~8-Gb chromosome-level reference genome for the critically endangered conifer Glyptostrobus pensilis, now largely restricted to southern China with scattered populations in Vietnam and Laos, and resequenced 147 individuals, including 64 ancient (>100 years old and persisting in human-dominated landscapes), 33 wild and 50 recently cultivated individuals. Ancient individuals comprised both likely natural relics and historically introduced individuals and formed two deeply divergent lineages and one ancestral-admixed group, each with distinct demographic histories of prolonged contraction and genomic erosion. Lineage identity explained more variation in genome-wide diversity, inbreeding and genetic load than the three conservation types, despite broad differences in age structure. Rare-allele analyses revealed pronounced heterogeneity among ancient trees: only relic and ancestral-origin individuals from high-diversity lineages contributed substantial unique variation, much of which is poorly represented in wild and cultivated populations. Together, our findings suggest that ancient trees are not uniformly genetically irreplaceable and that, at least in this conifer, evolutionary importance is shaped more strongly by population history than by age alone.

Endangered Species

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins

Optical genome mapping improves clinical interpretation of constitutional copy-number gains and reduces their VUS burden.

PURPOSE: Genomic structure of copy-number gains is critical for their clinical interpretation but cannot be determined by chromosomal microarray (CMA) analysis, which does not provide information about chromosomal location and orientation of multiplied regions. We thus hypothesized that in CMA testing gains have higher probability than losses to be classified as variants of uncertain significance (VUS) and that structural information from optical genome mapping (OGM) may improve their interpretation. METHODS: Using a χ2 test, we assessed the association between classification of copy-number variants as VUS and their type (gains vs losses) in a cohort of 4073 CMA cases. Thirty-three VUS gains involving disease-associated genes were characterized by OGM to evaluate if OGM data enable their more conclusive clinical interpretation. RESULTS: The proportion of variants reported as VUS compared with likely pathogenic/pathogenic was significantly higher for gains than losses, confirming their increased VUS burden. OGM successfully determined genomic structure for all 33 copy-number gains, showing that 26 of 33 were tandem duplications and 7 of 33 were complex rearrangements. Structural information facilitated clinical interpretation in majority of the cases; it supported benign nature for 27 of 33 gains and was inconclusive or supported pathogenic role for 6 of 33. An estimated 20% of reported VUS gains would not have been reportable if we had OGM data. CONCLUSION: We illustrate a specific advantage of OGM compared with CMA: in addition to detecting both copy-number variants and balanced rearrangements, OGM improves clinical interpretation of copy-number gains by providing structural information and is thus expected to significantly decrease their VUS burden.

Humans

Involvement of cross-genus phages in bacterial resistance to chlorine disinfection.

Chlorine disinfection resistance in pathogenic microorganisms poses severe environmental concerns and public health risks. While phages play critical roles in host adaptation to environmental stress, how poly-host phages contribute to bacterial resistance to chlorine disinfectants remains poorly understood. Here, we investigated shifts in the population dynamics, transcriptional profiles, and function potentials of cross-genus phage-bacterial communities under exposure to chlorine disinfectants in a continuously operated anaerobic-anoxic-oxic system over a 92-day period, using integrated metagenomic and metatranscriptomic approaches. In the presence and absence of chlorine disinfectants, the genomic abundance and diversity of phage and bacterial communities showed similar variation trends, and the community structures of both exhibited clear differences. A strong significant positive correlation was observed between phage and bacterial diversity under chlorine exposure (R&#x202f;=&#x202f;0.975, p&#x202f;=&#x202f;0.00,057), whereas no significant correlation was detected in the absence of chlorine disinfection (R&#x202f;=&#x202f;-0.314, p&#x202f;=&#x202f;0.613), suggesting that chlorine disinfectants may enhance phage-bacteria interactions. Host-associated phages exhibited high consistency with their corresponding putative hosts in terms of genomic abundance (M2&#x202f;=&#x202f;0.0945, p&#x202f;=&#x202f;0.001) and transcript abundance (M2&#x202f;=&#x202f;0.3668, p&#x202f;=&#x202f;0.001), and they were also significantly correlated with cross-genus phages in both genomic abundance (R&#x202f;=&#x202f;0.97, p&#x202f;<&#x202f;2.2e-16) and transcript abundance (R&#x202f;=&#x202f;0.83, p&#x202f;<&#x202f;2.2e-16), which collectively suggests the critical role of cross-genus phages in the resistance of microbial communities to chlorine disinfectants. Bipartite association network analysis shows that cross-genus phages carry highly homologous genes to their putative hosts and may be involved in the horizontal transfer of these genes among bacteria. These homologous genes are involved in DNA repair, redox balance regulation, environmental stress adaptation and efflux pump functions, suggesting a synergistic role between cross-genus phages and their putative hosts in chlorine resistance. Our findings reveal that cross-genus phages can contribute to the resistance of bacterial communities to chlorine disinfectants, providing the theoretical foundation for evaluating the role of poly-host phages in microbial communities.

Chlorine resistance

De novo genome assemblies of threatened Asian hornbills (Bucerotidae) reveal declining population trajectories during the late Pleistocene.

BACKGROUND: Asian hornbills are flagship species of the wet tropics that face significant threats from hunting, habitat loss, and fragmentation. Despite being conservation flagships, whole genome information is available for only two of the 32 Asian hornbill species. In this study, we provide the first de novo genome assemblies for four hornbill species (Bucerotidae) in Asia. METHODS: We used a combination of long-read and short-read sequencing data to assemble and annotate de novo hybrid genomes of four species of hornbills. We also assembled and compared mitochondrial genomes of these species. Using a comparative genomics approach, we performed orthology assignment and gene evolution analyses to identify unique gene families in Asian hornbills, gene families that showed significant expansion, their functions and structural variation. Furthermore, using the Pairwise Sequentially Markov Coalescent (PSMC) method, we reconstructed demographic histories of hornbill species to examine changes in their population trajectories in the past. RESULTS: We present hybrid genome assemblies for Great Hornbill (B. bicornis - GH), Rufous-necked Hornbill (A. nipalensis- RNH), Malabar Pied Hornbill (A. coronatus- MPH) and Wreathed Hornbill (R. undulatus- WH). The genome sizes of these hornbills range from 1.1 Gb to 1.3 Gb, with over 95.9% completeness and gene prediction BUSCO. We reported 10,525 orthogroups shared among four Asian hornbill species and identified significant expansion in gene families associated with structural keratin development in Asian hornbills compared to their ancestors. We also provide annotated mitogenomes for each of these species. Furthermore, we found that the WH, a more abundant, widely distributed, and migratory species, showed a higher Ne than the other three hornbill species. However, an overall decline in Ne for all species was recorded during the Pleistocene climatic fluctuations. CONCLUSIONS: We present the first-ever, high-quality reference genomes for the threatened hornbill species from Asia. Hornbills have shown significant expansion in genes involved in structural keratin development. Our results indicate that Pleistocene climatic fluctuations have led to dramatic population declines in all four species. We believe that this study provides robust genomic resources to support future comparative and conservation genomics efforts for hornbills.

Animals