Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Repeated sequences in bacterial chromosomes and plasmids: a glimpse from sequenced genomes.

To gain insight into the extent of exact DNA repeats in sequenced bacterial genomes and their plasmids, we analyzed the collection of completely sequenced bacterial genomes available at GenBank using the program Miropeats. This program draws graphical representations of exact DNA repeats in whole genomes. In this work, we present maps showing the extent and type (inverted or direct) of exact DNA repeats longer than 300 bp for the whole collection. These repeats may participate in a variety of events relevant for bacterial genome plasticity, such as amplifications, deletions, inversions, and translocations (via homologous recombination), as well as transposition. Additionally, we review recent data showing that high-frequency architectural variations in genomic structure occur at both the interspecies and interstrain levels.

Bacteria↗

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans↗

Complex de novo structural variants are an underestimated cause of rare disorders.

Complex de novo structural variants (dnSVs) are crucial genetic factors in rare disorders, yet their prevalence and characteristics in rare disorders remain poorly understood. Here, we conduct a comprehensive analysis of whole-genome sequencing data of 12,568 families, including 13,698 offspring with rare diseases, obtained as part of the UK 100,000 Genomes Project. We identify 1,870 dnSVs, constituting the largest dnSV dataset reported to date. Complex dnSVs (n = 158; 8.4%) emerge as the third most common type of SV, following simple deletions and duplications. We classify 65% of these complex dnSVs into 11 subtypes. Among probands with dnSVs (n = 1,696), 9% exhibit exon-disrupting pathogenic dnSVs associated with the probands' phenotype. Notably, 12% of exon-disrupting pathogenic dnSVs and 22% of de novo deletions or duplications previously identified by array-based or whole-exome sequencing methods are found to be complex dnSVs. We also find distinct genomic properties of de novo deletions depending on the parent of origin. This study highlights the importance of complex dnSVs in the cause of rare disorders and demonstrates the necessity of specific genomic analysis to avoid overlooking these variants.

Humans↗

De novo structural variants in autism spectrum disorder disrupt distal regulatory interactions of neuronal genes.

Three-dimensional genome organization plays a critical role in gene regulation, and disruptions can lead to developmental disorders by altering the contact between genes and their distal regulatory elements. Structural variants (SVs) can disturb local genome organization, such as the merging of topologically associating domains upon boundary deletion. Testing large numbers of SVs experimentally for their effects on chromatin structure and gene expression is time and cost prohibitive. To address this, we propose a computational approach to predict SV impacts on genome folding, which can help prioritize causal hypotheses for functional testing. We develop a weighted scoring method that measures chromatin contact changes specifically affecting regions of interest, such as regulatory elements or promoters, and implement it in the SuPreMo-Akita software. With this tool, we rank hundreds of de novo SVs (dnSVs) from autism spectrum disorder (ASD) individuals and their unaffected siblings based on predicted disruptions to nearby neuronal regulatory interactions. This reveals that putative cis-regulatory element interactions (CREints) are more disrupted by dnSVs from ASD probands versus unaffected siblings. We prioritize candidate variants that disrupt ASD CREints and validate our top-ranked locus using isogenic excitatory neurons with and without the dnSV, confirming accurate predictions of disrupted chromatin contacts. This study suggests that disrupted genome folding is a potential genetic mechanism in a subset of ASD cases and provides a general strategy for prioritizing variants predicted to disrupt regulatory interactions across tissues.

Humans↗

Aplf/Dna2 variants drive chromosomal fission and accelerate speciation in zokors.

Chromosomal fissions and fusions are common, yet the molecular mechanisms and implications in speciation remain poorly understood. Here, we confirm a fission event in one zokor species through multiple-omics and functional analyses. We traced this event to a mutation in a splicing enhancer of the DNA repair gene Aplf in the fission-bearing species, which caused exon skipping and produced a truncated protein that disrupted DNA repair. An intronic deletion in Dna2, known to facilitate neo-telomere formation when knocked out, reduced gene activity. These variants collectively drove chromosomal fission in this zokor species. The newly formed chromosome became fixed due to carrying essential genes and strong selective pressure. While geographic isolation likely initiated the divergence of this species and the sister one, the fission event and associated decline at the chromosome level in gene flow probably exacerbated the speciation process. Our work elucidates the genetic basis of chromosomal fission and underscores its role in speciation dynamics.

Multiomics↗

The mosaic structure of variation in the laboratory mouse genome.

Most inbred laboratory mouse strains are known to have originated from a mixed but limited founder population in a few laboratories. However, the effect of this breeding history on patterns of genetic variation among these strains and the implications for their use are not well understood. Here we present an analysis of the fine structure of variation in the mouse genome, using single nucleotide polymorphisms (SNPs). When the recently assembled genome sequence from the C57BL/6J strain is aligned with sample sequence from other strains, we observe long segments of either extremely high (approximately 40 SNPs per 10 kb) or extremely low (approximately 0.5 SNPs per 10 kb) polymorphism rates. In all strain-to-strain comparisons examined, only one-third of the genome falls into long regions (averaging >1 Mb) of a high SNP rate, consistent with estimated divergence rates between Mus musculus domesticus and either M. m. musculus or M. m. castaneus. These data suggest that the genomes of these inbred strains are mosaics with the vast majority of segments derived from domesticus and musculus sources. These observations have important implications for the design and interpretation of positional cloning experiments.

Albinism↗

Sequence variation in 5' termini of rubella virus genomes: changes affecting structure of the 5' proximal stem-loop.

Variation within a 523 nucleotide region proximal to the 5' terminus of seven rubella virus strains has been analysed. Compared to the Therien strain twenty sites of nucleotide variation have been identified, three of which are in the 5' untranslated region. Individual strains have between three and nine nucleotide differences, only three of which result in amino acid substitutions. TO-336 has a serine for threonine at amino acid (aa) 42 and CM arginine for histidine at aa 159. RA27/3 has arginine for lysine at aa 3 and serine for threonine at aa 42. Nucleotide differences which affect a stem-loop structure reported to be important for binding of host cell proteins have been identified.

Amino Acid Substitution↗

Intraspecific variation in gamma-radiation resistance and genomic structure in the filamentous fungus Alternaria alternata: a case study of strains inhabiting Chernobyl reactor no. 4.

This is probably the first report on intraspecific variation in radiation resistance for filamentous fungi. It was revealed that natural ("field") strains of the filamentous fungus Alternaria alternata are extremely variable in response to gamma-irradiation ranging from supersensitive to highly resistant to radiation. At the same time nearly all strains originating from the highly radiation-polluted reactor of the Chernobyl (Ukraine) Nuclear Power Plant possessed high radiation resistance. The genome structure of strains studied by universally primed polymerase chain reaction (UP-PCR) was found to be well conserved in "reactor" but not in "control" strains. The "reactor" strains appear to be genetically adapted to this high radiation habitat by means of selection, thus providing a natural source of genetically homogeneous fungal lineages.

Alternaria↗

Genome screen for quantitative trait loci underlying normal variation in femoral structure.

Femoral structure contributes to bone strength at the proximal femur and predicts hip fracture risk independently of bone mass. Quantitative components of femoral structure are highly heritable traits. To identify genetic loci underlying variation in these structural phenotypes, we conducted an autosomal genome screen in 309 white sister pairs. Seven structural variables were measured from femoral radiographs and used in multipoint sib-pair linkage analyses. Three chromosomal regions were identified with significant evidence of linkage (log10 of the odds ratio [LOD] > 3.6) to at least one femoral structure phenotype. The maximum LOD score of 4.3 was obtained for femur neck axis length on chromosome 5q. Evidence of linkage to chromosome 4q was found with both femur neck axis length (LOD = 3.9) and midfemur width (LOD = 3.5). Significant evidence of linkage also was found to chromosome 17q, with a LOD score of 3.6 for femur head width. Two additional chromosomal regions 3q and 19p gave suggestive (LOD > 2.2) evidence of linkage with at least two of the structure phenotypes. Chromosome 3 showed evidence of linkage with pelvic axis length (LOD = 3.1), midfemur width (LOD = 2.8), and femur head width (LOD = 2.3), spanning a broad (60 cm) region of chromosome 3q. Linkage to chromosome 19 was supported by two phenotypes, femur neck axis length (LOD = 2.8) and femur head width (LOD = 2.8). This study is the first genome screen for loci underlying variation in femoral structure and represents an important step toward identifying genes contributing to the risk of osteoporotic hip fracture in the general population.

Adult↗

Beyond HLA: the significance of genomic variation for allogeneic hematopoietic stem cell transplantation.

The last 2 years have seen much excitement in the field of genetics with the identification of a formerly unappreciated level of "structural variation" within the normal human genome. Genetic structural variants include deletions, duplications, and inversions in addition to the recently discovered, copy number variants. Single nucleotide polymorphisms are the most extensively evaluated variant within the genome to date. Combining our knowledge from these studies with our rapidly accumulating understanding of structural variants, it is apparent that the extent of genetic dissimilarity between any 2 individuals is considerable and much greater than that which was previously recognized. Clearly, this more diverse view of the genome has significant implications for allogeneic hematopoietic stem cell transplantation, not least in the generation of transplant antigens but also in terms of individual susceptibility to transplant-related toxicities. With advances in DNA sequencing technology we now have the capacity to perform genome-wide analysis in a high throughput fashion, permitting a detailed genetic analysis of patient and donor prior to transplantation. Understanding the significance of this additional genetic information and applying it in a clinically meaningful way will be one of the challenges faced by transplant clinicians in the future.

Genetic Predisposition to Disease↗

The human growth hormone gene locus: structure, evolution, and allelic variations.

Genomic clones containing the closely related genes for human growth hormone (hGH) and chorionic somatomammotropin (hCS) were obtained from genomic bacteriophage lambda and cosmid libraries. The entire GH/CS chromosomal locus was reconstructed utilizing overlapping restriction fragments characterized from the isolated clones. The hGH/hCS locus contains two GH genes and three CS genes spanning 48 kb of DNA in the order: 5'-(hGH-1/hCS-5/hCS-1/hGH-2/hCS-2)-3', confirming analysis of cosmid clones obtained from a different human library (Barsh et al., 1983). To complete the characterization of the hCS genes, the nucleotide sequence of the hCS-5 gene was determined. Sequence analysis revealed a mutation of the 5' splice site at the exon II-intron B boundary, suggesting that the hCS-5 gene is a pseudogene. The nucleotide sequence of an allelic variant of the hCS-2 gene was determined and found to contain a single amino acid substitution and the deletion of a single codon. The hGH/hCS gene locus was further characterized by the localization of at least 27 Alu-type repetitive sequences and identification of three unique sequences in the vicinity of several hGH and hCS genes which define the probable breakpoints of the evolutionary duplication units. These data, combined with the nucleotide sequences of all five GH and CS genes, indicate that the hGH/hCS gene locus has evolved by duplication mechanisms. Evidence for the occurrence of at least one gene conversion event involving the hCS-1 gene precursor and the hCS-2 gene was found, indicating that the hGH/hCS gene locus has evolved by concerted mechanisms. The structure of the hCS genes is discussed in light of recent studies of CS genes from other mammalian species.

Alleles↗

Comparative Analysis of Mammalian Adaptive Immune Loci Revealed Spectacular Divergence and Common Genetic Patterns.

Adaptive immune responses are mediated by the production of adaptive immune receptors, antibodies, and T-cell receptors, which bind antigens, thus causing their neutralization. Unlike other proteins, adaptive immune receptors are not fully encoded in the germline genome and result from a complex of somatic processes collectively called V(D)J recombination affecting germline immunoglobulin (IG) and T-cell receptor (TR) loci consisting of template genes. While various existing studies report extreme diversity of antibodies and T-cell receptors, little is known about the diversity of germline IG and TR loci. To overcome this gap, the first comparative analysis of full-length sequences of IG/TR loci across 46 mammalian species from 13 taxonomic orders was performed. First, germline gene counts were shown to correlate in immunoglobulin heavy chain immunoglobulin heavy chain (IGH)/immunoglobulin lambda (IGL) loci and T-cell receptor alpha (TRA)/T-cell receptor beta (TRB) and anticorrelate in immunoglobulin kappa (IGK)/IGL, possibly indicating coevolution between corresponding chains. Second, structures of IG/TR loci were analyzed, and it was shown that IG/TR loci formed by long arrays of high multiplicity repeats are more common for species that have experienced population bottlenecks. Finally, haplotypes of IG/TR loci with little or no sequence similarity within a species were found, suggesting that they may have a limited potential for homologous recombination. These results demonstrate that IG/TR loci are rapidly evolving genomic regions whose structural variation is shaped by the population history of the species and open new perspectives for immunogenomics studies.

Animals↗

PARTAGE: Parallel analysis of replication timing and gene expression.

The human genome is partitioned into functional compartments that replicate at specific times during the S-phase. This temporal program, referred to as replication timing (RT), is co-regulated with the 3D genome organization, is cell type-specific, and changes during development in coordination with gene expression. Moreover, RT alterations are linked to abnormal gene expression, genome instability, and structural variation in multiple diseases, including cancer. However, mechanistic links between RT, large-scale 3D genome architecture, and transcriptional regulation remain poorly understood. A major limitation is that current approaches require the separate profiling of RT and transcriptomes from independent batches of samples, obscuring the complex co-regulation between the epigenome and transcriptome. Here, we developed PARTAGE, a multiomics approach that enables joint profiling of copy number variation (CNV), RT, and gene expression from the same sample, providing a more accurate integrative view of the complex relationships between RT and gene regulation.

Journal Article↗

Global variation in G+C content along vertebrate genome DNA. Possible correlation with chromosome band structures.

The global, rather than local, variation in G+C content along the nuclear DNA sequences of various organisms was studied using GenBank sequence data. When long DNA sequences of the genomes of Escherichia coli and Saccharomyces cerevisiae were examined, the levels of their G+C content (G+C%) were found to be within a narrow range around that of the whole genome. The G+C% levels for sequences of vertebrate genomes, however, were found to cover a wide range, showing that their genome is a mosaic of sequences with different G+C% levels, in each of which the sequence is fairly homogeneous in its G+C% for a very long distance. Through surveying a human genetic map and GenBank DNA sequences, the global variations in G+C% along the human genome DNA were found to be correlated with chromosome band structures.

Animals↗

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗

TAR cloning: insights into gene function, long-range haplotypes and genome structure and evolution.

The structural and functional analysis of mammalian genomes would benefit from the ability to isolate from multiple DNA samples any targeted chromosomal segment that is the size of an average human gene. A cloning technique that is based on transformation-associated recombination (TAR) in the yeast Saccharomyces cerevisiae satisfies this need. It is a unique tool to selectively recover chromosome segments that are up to 250 kb in length from complex genomes. In addition, TAR cloning can be used to characterize gene function and genome variation, including polymorphic structural rearrangements, mutations and the evolution of gene families, and for long-range haplotyping.

Animals↗

The fine-scale structure of recombination rate variation in the human genome.

The nature and scale of recombination rate variation are largely unknown for most species. In humans, pedigree analysis has documented variation at the chromosomal level, and sperm studies have identified specific hotspots in which crossing-over events cluster. To address whether this picture is representative of the genome as a whole, we have developed and validated a method for estimating recombination rates from patterns of genetic variation. From extensive single-nucleotide polymorphism surveys in European and African populations, we find evidence for extreme local rate variation spanning four orders in magnitude, in which 50% of all recombination events take place in less than 10% of the sequence. We demonstrate that recombination hotspots are a ubiquitous feature of the human genome, occurring on average every 200 kilobases or less, but recombination occurs preferentially outside genes.

Base Composition↗

Comparisons of genetic variability and genome structure among mosquito strains selected for refractoriness to a malaria parasite.

Restriction fragment length polymorphism (RFLP) markers were used to evaluate Aedes aegypti genome structure and genetic variability within and between substrains selected for different levels of refractoriness to the malaria parasite, Plasmodium gallinaceum. The MOYO-R substrain was previously selected for complete refractoriness and the MOYO-IS substrain for intermediate susceptibility from the Moyo-In-Dry (MOYO) strain by selective inbreeding (F = 0.5). Eighteen mapped RFLP markers were used to provide coverage of the mosquito genome. The two substrains showed reduced genetic diversity compared with the MOYO strain, including significant reductions in mean heterozygosity, number of alleles per locus, and proportion of polymorphic loci. Genetic differentiation between the two substrains was statistically significant, as reflected by differences in allele frequencies. Significant pairwise linkage disequillbrium among the RFLP loci was detected in all three strains, most evidently in the MOYO strain. This is surprising because the RFLP loci examined are separated by large map distances, and therefore linkage disequilibrium should decay to zero after many generations of laboratory culture. Our hypothesis to explain this phenomena is that lack of recombination, or low recombination rates in some regions of the A. aegypti genome, is a result of chromosome inversions. Finally, we used graphical genotyping, wherein whole genome genotypic information for individual mosquitoes is represented in a simple graphic format, to illustrate genome structure and allelic variation within and among the mosquito strains. Our analysis revealed an apparent chromosomal deletion on chromosome 3 for some individuals in the MOYO strain and MOYO-IS substrain.

Aedes↗