Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Molecular evolution of a tandemly repeated trnF(GAA) gene in the chloroplast genomes of Microseris (Asteraceae) and the use of structural mutations in phylogenetic analyses.

We sequenced the first ca. 900 bp of the 5'-trnL(UAA)-trnV(UAC)/ndhJ region of the chloroplast DNA of different Microseris accessions in order to resolve homoplasious length variation detected in the trnL(UAA)-trnF(GAA) region. We found two to four tandemly repeated trnF genes in the species of Microseris (Asteraceae, Lactuceae) and two in their sister genus Uropappus. Sequences indicated nonhomologous transitions between two, three, and four trnF genes in different Microseris taxa. Independent origins of similar trnF copy numbers were inferred from a chloroplast phylogeny of Microseris. The taxa involved grow on separate continents, supporting parallel origins of similar length variants. The changes in trnF copy numbers were best explained by interchromosomal recombination with unequal crossing over. The 5' copies of the repeats showed the highest sequence conservation, suggesting that these copies are likely to be functional trnF genes, whereas the other ones probably represent pseudogenes. Our results show that length polymorphisms accumulate once a duplicated sequence has become incorporated. Due to parallel gains of similar trnF copy numbers, homoplasious length variation was introduced into the data matrix. The data demonstrate that length polymorphisms cannot be used as indicators for phylogenetic distance unless they can be analyzed at the sequence level.

Base Sequence↗

Intron-genome size relationship on a large evolutionary scale.

The intron-genome size relationship was studied across a wide evolutionary range (from slime mold and yeast to human and maize), as well as the relationship between genome size and the ratio of intervening/coding sequence size. The average intron size is scaled to genome size with a slope of about one-fourth for the log-transformed values; i.e., on the global scale its increase in evolution is lower than the increase in genome size by four orders of magnitude. There are exceptions to the general trend. In baker's yeast introns are extraordinarily long for its genome size. Tetrapods also have longer introns than expected for their genome sizes. In teleost fish the mean intron size does not differ significantly, notwithstanding the differences in genome size. In contrast to previous reports, avian introns were not found to be significantly shorter than introns of mammals, although avian genomes are smaller than genomes of mammals on average by about a factor of 2.5. The extra-/intragenic ratio of noncoding DNA can be higher in fungi than in animals, notwithstanding the smaller fungal genomes. In vertebrates and invertebrates taken separately, this ratio is increasing as the increase in genome size. Two hypotheses are proposed to explain the variation in the extra-/intragenic ratio of noncoding DNA in organisms with similar numbers of genes: transition (dynamic) and equilibrium (static). According to the transition model, this variation arises with the rapid shift of genome size because the bulk of extragenic DNA can be changed more rapidly than the finely interspersed intron sequences. The equilibrium model assumes that this variation is a result of selective adjustment of genome size with constraints imposed on the intron size due to its putative link to chromatin structure (and constraints of the splicing machinery).

Animals↗

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans↗

The relationship between HP1 and S2 bacteriophages of Haemophilus influenzae.

Comparison of the nucleotide sequences of the left arms of two Haemophilus influenzae phages, S2 and HP1 is presented. They exhibit a characteristic mosaic pattern of homologous and non-homologous regions. The homology extends over the attP site and int, orf 5 to 9, rep and the 3' part of cI genes. Two major non-homologous regions were detected. One is found between the int and cI genes; the other spans the region of promoters and the cox gene. Variations in the region of the promotors which is involved in the choice between a lysogenic and a lytic pathway and some divergences in the cI coding sequences are probably responsible for the observed immunity differences between the two phages. Distinctions in the distribution of consensus sequences for an integration host factor (IHF) and integrase-binding sites and promoters are described. These data offer an explanation of the relationship between three types of S2/HP1 phages. It allows in turn a final settlement of the nomenclature variation in the literature. The results presented, which are similar to those obtained for other phage groups, suggest that the mosaic structure of phage genomes is a normal outcome of phage divergence.

Amino Acid Sequence↗

A multipartite mitochondrial genome in the potato cyst nematode Globodera pallida.

The mitochondrial genome (mtDNA) of the plant parasitic nematode Globodera pallida exists as a population of small, circular DNAs that, taken individually, are of insufficient length to encode the typical metazoan mitochondrial gene complement. As far as we are aware, this unusual structural organization is unique among higher metazoans, although interesting comparisons can be made with the multipartite mitochondrial genome organizations of plants and fungi. The variation in frequency between populations displayed by some components of the mtDNA is likely to have major implications for the way in which mtDNA can be used in population and evolutionary genetic studies of G. pallida.

Animals↗

A Multitrait Locus Regulates Sarbecovirus Pathogenesis.

Infectious diseases have shaped the human population genetic structure, and genetic variation influences the susceptibility to many viral diseases. However, a variety of challenges have made the implementation of traditional human Genome-wide Association Studies (GWAS) approaches to study these infectious outcomes challenging. In contrast, mouse models of infectious diseases provide an experimental control and precision, which facilitates analyses and mechanistic studies of the role of genetic variation on infection. Here we use a genetic mapping cross between two distinct Collaborative Cross mouse strains with respect to severe acute respiratory syndrome coronavirus (SARS-CoV) disease outcomes. We find several loci control differential disease outcome for a variety of traits in the context of SARS-CoV infection. Importantly, we identify a locus on mouse chromosome 9 that shows conserved synteny with a human GWAS locus for SARS-CoV-2 severe disease. We follow-up and confirm a role for this locus, and identify two candidate genes, CCR9 and CXCR6, that both play a key role in regulating the severity of SARS-CoV, SARS-CoV-2, and a distantly related bat sarbecovirus disease outcomes. As such we provide a template for using experimental mouse crosses to identify and characterize multitrait loci that regulate pathogenic infectious outcomes across species. IMPORTANCE Host genetic variation is an important determinant that predicts disease outcomes following infection. In the setting of highly pathogenic coronavirus infections genetic determinants underlying host susceptibility and mortality remain unclear. To elucidate the role of host genetic variation on sarbecovirus pathogenesis and disease outcomes, we utilized the Collaborative Cross (CC) mouse genetic reference population as a model to identify susceptibility alleles to SARS-CoV and SARS-CoV-2 infections. Our findings reveal that a multitrait loci found in chromosome 9 is an important regulator of sarbecovirus pathogenesis in mice. Within this locus, we identified and validated CCR9 and CXCR6 as important regulators of host disease outcomes. Specifically, both CCR9 and CXCR6 are protective against severe SARS-CoV, SARS-CoV-2, and SARS-related HKU3 virus disease in mice. This chromosome 9 multitrait locus may be important to help identify genes that regulate coronavirus disease outcomes in humans.

Animals↗

Multi-omics Investigations of Immune Microenvironment of Human Colorectal Cancer.

BACKGROUND/AIM: Colorectal cancer (CRC) remains a leading cause of cancer-related morbidity and mortality worldwide. Although immunotherapy has improved outcomes for a subset of patients, its limited efficacy in many cases highlights the need for a more comprehensive understanding of the CRC immune microenvironment. This study aimed to characterize the molecular landscape of the CRC immune microenvironment using an integrated multi-omics approach and to identify candidate regulatory molecules associated with immune remodelling. MATERIALS AND METHODS: We integrated structural variation, DNA methylation, chromatin accessibility, proteomic, and phosphoproteomic data generated from an in-house CRC cohort with transcriptomic data from The Cancer Genome Atlas (TCGA). Analyses focused on 1,539 immune-related genes (IRGs) associated with CD4+ T cells, B cells, and natural killer (NK) cells. Multi-layered genomic and proteomic analyses were performed to identify altered immune-related pathways, hub genes, candidate transcription factors, and upstream kinases. RESULTS: Higher infiltration of CD4+ T cells, B cells, and NK cells was associated with CRC. IRGs exhibited widespread alterations across genomic, epigenomic, transcriptomic, proteomic, and phosphoproteomic levels. IL10, LEP, ITGAM, and EGFR emerged as candidate hub genes. EGFR phosphorylation at S991 and T693 was significantly decreased in CRC. STAT2 and HSF1 were identified as candidate upstream transcription factors, while CDK2 emerged as a candidate upstream kinase associated with immune infiltration and immune checkpoint expression. CONCLUSION: This study provides a systematic multi-omics characterization of immune microenvironment remodelling in CRC and identifies candidate molecular regulators that may serve as potential targets for future immunotherapy research.

Humans↗

Lymphadenopathy-associated virus: from molecular biology to pathogenicity.

Recent data indicate that the lymphadenopathy-associated virus (LAV) is morphologically similar to animal lentiviruses, such as equine infectious anemia and visna viruses. This finding, together with the cross-reactivity of the core proteins of LAV with those of the equine infectious anemia virus and a similarity in genome structure and biological properties, allows LAV to be placed in the retroviral subfamily of Lentivirinae. Molecular data indicate a high degree of genetic variation of the virus, especially in the envelope gene, which have important implications for the origin of the virus (the T4 lymphotropism may be a recently acquired property) and for future immunization. Another problem is the role of viral infection in the induction of irreversible immunodeficiency. This syndrome occurs in a minority of infected persons, who generally have in common a past of antigenic stimulation and of immune depression before LAV infection.

Acquired Immunodeficiency Syndrome↗

Complicated organization of a single repeated DNA sequence in the chicken genome is revealed by cloning.

The structural organization of a family of repeated DNA sequences in the chicken genome has been determined by hybridization of a cloned repeated DNA sequence to Southern blots of total DNA. The length of the cloned DNA fragment is 3600 nucleotide pairs. This fragment consists principally, if not entirely, of a single repeated DNA sequence occurring only once within the cloned fragment. In the chicken genome, the family of repeated DNA sequences homologous to the cloned sequence has a limited number of alternative forms. Some of the restriction fragments of total DNA to which the cloned sequence hybridizes correspond to those expected from the location of restriction endonuclease cleavage sites within the cloned sequence. There are also a limited number of other genomic restriction fragments, each present in multiple copies, to which the cloned sequence hybridizes but which do not relate in any obvious way to the length of the cloned sequence. These various restriction fragments differ from one another in that they appear to be present in unequal amounts in total DNA, and many of them do not contain the entire cloned sequence. This study provides some new information about the structure of repeated DNA sequences in the chicken genome. The copies of a repeated DNA sequence may differ from one another both by minor variations in nucleotide sequence (divergence) and in more substantial ways as would be expected to arise from processes such as insertion, deletion, and translocation. In addition to this description of a single cloned repeated DNA sequence from the chicken genome, this paper reports the cloning of more than 100 different restriction fragments of chicken DNA, each of which contains one or more repeated DNA sequences.

Animals↗

Genomic signatures of host-range divergence in the generalist Beauveria bassiana and the specialist Beauveria brongniartii.

Entomopathogenic fungi of the genus Beauveria are widely used biological control agents that infect diverse insect hosts and can also associate with plants as rhizosphere colonizers and endophytes. Within this genus, Beauveria bassiana is a cosmopolitan generalist, whereas Beauveria brongniartii exhibits a narrower host range, primarily targeting soil-dwelling coleopteran larvae with limited evidence of plant colonization. To explore genomic differentiation associated with this ecological divergence, the commercially exploited B. brongniartii strain BIPESCO2 and B. bassiana ATHUM 4946 were sequenced using Oxford Nanopore technology, followed by comparative genomic analyses across multiple strains. Orthology identified species-specific gene families, although overall genome architecture and core gene content were highly conserved. The CAZyme repertoires were nearly identical, indicating retention of a versatile enzymatic toolkit supporting plant association, saprotrophy, and insect pathogenicity. In contrast, biosynthetic gene clusters displayed substantial variation, including structural remodeling of Beauveria-specific virulence-associated clusters and expansion of type I polyketide synthase clusters in B. brongniartii. Effector prediction revealed a conserved core of largely uncharacterized proteins alongside species-specific orthogroups enriched in adhesion-, immunity-, and cuticle-interaction domains. Together, these findings indicate that host-range divergence in Beauveria is associated with compartmentalized genomic differentiation, particularly in secondary metabolism and a limited subset of lineage-specific virulence factors, rather than in the conserved core infection machinery.

Beauveria↗

A highly conserved sequence in H1 histone genes as an oligonucleotide hybridization probe: isolation and sequence of a duck H1 gene.

A 3.5-kb HindIII fragment of a histone gene cluster was isolated from a recombinant phage out of a duck genomic library. This DNA contains a duck H1 gene and its flanking sequences. The hybridization probe, which was used to screen for the H1 gene, had been designed on the basis of a comparative analysis of available H1 gene and protein data. Most H1 histones contain repeated motifs in their C-terminal domain, and these form part of an octapeptide (ser pro lys lys ala lys lys pro) that is highly conserved in many H1 histone proteins. A comparison of the duck H1 described here with two different published chicken H1 histone sequences reveals conservative amino acid exchanges at 22 (of 217 and 218, respectively) positions. The homology is maintained at the flanking sequences, and includes the putative H1 histone gene-specific signal structures and the established 3' stem and loop structures and the CAAGA box. The duck H1 gene and its flanking sequence have been found in identical arrangements in two recombinant bacteriophages, but minor sequence variations and genomic Southern blotting after HindIII digestion suggest that we have either isolated alleles of this genome segment or that the gene described may occur twice per haploid duck genome.

Amino Acid Sequence↗

Targeted long-read genomic and epigenomic profiling enhances timely comprehensive variant discovery in hypotonia and muscle weakness.

BACKGROUND: Identifying the genetic basis of hypotonia and muscle weakness is critical for patient management and family counseling. However, diagnosis is often hindered by diverse genomic alterations, including repeat expansions, structural variants (SVs), and methylation defects. Standard-of-care testing, largely based on short-read sequencing, is limited in its ability to detect this heterogeneous variation landscape, leaving many patients undiagnosed or requiring lengthy sequential testing. Long-read sequencing represents a promising solution. However, its application as a first-tier diagnostic assay for hypotonia remains unexplored. METHODS: We retrospectively analyzed 227 patients with hypotonia to assess diagnostic yield, time-to-diagnosis, and costs associated with standard-of-care testing. A long-read whole-genome sequencing (LR-WGS) workflow with targeted analysis of hypotonia-associated genes was developed to detect and prioritize pathogenic SNVs, SVs, and CNVs, repeat expansions, and methylation changes at key disease loci. The workflow was validated in a reference-positive cohort with known diagnoses (n = 15) and applied to an unsolved cohort (n = 14). Variant interpretation followed ACMG guidelines and was confirmed with orthogonal methods. RESULTS: Standard-of-care testing achieved a diagnostic yield of 42% with an average time-to-diagnosis of 68.7 days; however, 30% of diagnosed patients experienced significant delays (average 169 days) due to sequential testing. The LR-WGS based approach identified all known pathogenic variants in the positive cohort, including SMN1 deletions, methylation defects at 15q11.2/Prader-Willi locus, FMR1 repeat expansions, and sequence and copy-number variants in > 100 genes underlying myopathies and muscular dystrophies. The targeted long-read pipeline reduced prioritized variant calls by 97.9-99.9% and, in the unsolved cohort, yielded one definitive diagnosis (de novo COL6A3 deletion) and one possible diagnosis (aberrant methylation and copy number at POMK), for an additional 14% yield. Among patients diagnosed after sequential testing (n = 29), LR-WGS is expected to reduce time-to-diagnosis by ~ 85% and decrease cumulative diagnostic delays, with projected healthcare cost savings of $396,000-439,000. Across the entire 227 patient cohort, LR-WGS is anticipated to reduce testing costs by 6.5%, yielding an average savings of $105 per patient. CONCLUSIONS: LR-WGS enables comprehensive discovery of genomic and epigenomic variants in hypotonia and muscle weakness, improving diagnostic yield, shortening diagnostic timelines, and reducing costs compared with current standard-of-care testing.

Humans↗

Charting host structural variations in cervical cancer by long-read sequencing pinpoints a functional deletion in PIAS1.

Host structural variations (SVs) are critical in cancer development but their landscape and interaction with HPV integration in cervical carcinogenesis remain unclear. In this study, we performed Nanopore long-read sequencing on five HPV-positive cervical cancer tissues and two cell lines to profile host SVs. We identified thousands of SVs and statistically demonstrated their significant enrichment in genomic windows ±25 to ±50 kb from HPV integration sites. Cross-sample analysis revealed 60 shared SVs, including a recurrent deletion within the PIAS1 gene. Multi-omics integration (Hi-C, H3K27ac ChIP-seq, and TCGA data) showed that this deletion is associated with reduced PIAS1 expression, disruption of local topologically associating domains, advanced pathological tumor stage, and poorer overall survival. Functional assays confirmed that PIAS1 deficiency inhibits cervical cancer cell proliferation and migration. Our findings identify a PIAS1 deletion as a candidate driver event, and underscore the pivotal role of host genomic instability in HPV-associated oncogenesis.

Cervical cancer↗

Immunoglobulin heavy chain gene organization and complexity in the skate, Raja erinacea.

Immunoglobulin heavy chain genes from Raja erinacea have been isolated by cross hybridization with probes derived from the immunoglobulin genes of Heterodontus francisci (horned shark), a representative of a different elasmobranch order. Heavy chain variable (VH), diversity (DH) and joining (JH) segments are linked closely to constant region (CH) exons, as has been described in another elasmobranch. The nucleotide sequence homology of VH gene segments within Raja and between different elasmobranch species is high, suggesting that members of this phylogenetic subclass may share one VH family. The organization of immunoglobulin genes segments is diverse; both VD-J and VD-DJ joined genes have been detected in the genome of non-lymphoid cells. JH segment sequence diversity is high, in contrast to that seen in a related elasmobranch. These data suggest that the clustered V-D-J-C form of immunoglobulin heavy chain organization, including germline joined components, may occur in all subclasses of elasmobranchs. While variation in VH gene structure is limited, gene organization appears to be diverse.

Amino Acid Sequence↗

Mini-chromosomes in Fusarium sporotrichioides are mosaics of dispersed repeats and unique sequences.

Variations in trichothecene patterns of 26 Fusarium sporotrichioides isolates from different plant and geographic origins showed no correlation with electrophoretic karyotype polymorphisms. When intact chromosomes were examined, interisolate karyotype differences were observed only in the mini-chromosome range. Further polymorphisms were revealed in Notl-digested samples. By summing the Notl fragments the average genome size of F. sporotrichioides was estimated to be 20.4 Mb. Mini-chromosomes shared common sequences with the larger ones; however, clones (RMS-1 and RMS-2) specific to these structures have also been found. These clones contained no coding region and no promising similarities were observed when they were compared to sequences held at GenBank. Mini-chromosomes in F. sporotrichioides constitute a mosaic composed of dispersed repeats and unique sequences. This mosaic structure was maintained in all noninterbreeding, genetically isolated strains examined.

Chromosomes, Fungal↗

Plant molecular diversity and applications to genomics.

Surveys of nucleotide diversity are beginning to show how genomes have been shaped by evolution. Nucleotide diversity is also being used to discover the function of genes through the mapping of quantitative trait loci (QTL) in structured populations, the positional cloning of strong QTL, and association mapping.

Chromosome Mapping↗

Cross-kingdom genomic variation in chicken gut microbiomes: insights from China's diverse local breeds.

BACKGROUND: The gut microbiome possesses substantial genetic diversity that supports microbial adaptation, but the genomic variation patterns across its prokaryotic and viral populations remain incompletely characterized. RESULTS: Through integrated metagenomic and metatranscriptomic analysis of ten indigenous chicken breeds from China, we recovered 1527 representative prokaryotic MAGs, 37,555 representative DNA viral contigs, and 1867 representative RNA viral contigs (primarily comprising Bacillota/Bacteroidota, Uroviricota, and Lenarviricota/Pisuviricota, respectively). By integrating complementary short-read and long-read metagenomics with metatranscriptomics, we identified structural variants (SVs) and single-nucleotide variants (SNVs) in these cross-kingdom genomes. Positive SV-SNV density correlations occurred consistently across all microbial groups, indicating coordinated mutational processes. DNA viruses exhibited the highest variant prevalence (86.9% SNVs, 47.7% SVs), with temperate phages accumulating significantly more variants than virulent phages. Functionally, prokaryotic variants accumulated in carbohydrate metabolism and amino acid metabolism, while viral variants demonstrated broad metabolic hijacking. Horizontal gene transfer (HGT) was characterized by a strong virus-associated signature (69.40% of 536 events) and marked by an asymmetric pattern, with phage-to-bacteria (P-to-B) flow alone constituting 37.50% of all events. Random forest analysis revealed a strong bidirectional predictive relationship between SV and SNV densities across prokaryotic, DNA viral, and RNA viral populations, suggesting coupled genomic instability. Niche breadth emerged as a major driver of SNVs across kingdoms and was positively correlated with variant density. In prokaryotes, HGT events significantly shaped variant patterns. For viruses, genomic GC content was an important factor and consistently showed a negative correlation with SNV density in both DNA and RNA viruses. CONCLUSIONS: These findings demonstrate that coordinated mutational processes and kingdom-specific intrinsic factors drive genomic variation, with viruses serving as key genetic exchange vectors in chicken gut ecosystems. Video Abstract.

Animals↗