Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Haplotype block structure is conserved across mammals.

Genetic variation in genomes is organized in haplotype blocks, and species-specific block structure is defined by differential contribution of population history effects in combination with mutation and recombination events. Haplotype maps characterize the common patterns of linkage disequilibrium in populations and have important applications in the design and interpretation of genetic experiments. Although evolutionary processes are known to drive the selection of individual polymorphisms, their effect on haplotype block structure dynamics has not been shown. Here, we present a high-resolution haplotype map for a 5-megabase genomic region in the rat and compare it with the orthologous human and mouse segments. Although the size and fine structure of haplotype blocks are species dependent, there is a significant interspecies overlap in structure and a tendency for blocks to encompass complete genes. Extending these findings to the complete human genome using haplotype map phase I data reveals that linkage disequilibrium values are significantly higher for equally spaced positions in genic regions, including promoters, as compared to intergenic regions, indicating that a selective mechanism exists to maintain combinations of alleles within potentially interacting coding and regulatory regions. Although this characteristic may complicate the identification of causal polymorphisms underlying phenotypic traits, conservation of haplotype structure may be employed for the identification and characterization of functionally important genomic regions.

Animals↗

Whole-Genome Sequencing Reveals Population Structure, Genetic Diversity, and Selection Signatures in Kazakh Dromedary and Bactrian Camels.

Understanding the genomic basis of environmental adaptation is essential for the conservation and genetic improvement of domestic camels. In this study, we investigated the population structure, genetic diversity, and genomic variation potentially associated with environmental adaptation of Kazakh dromedary and Bactrian camels using whole-genome sequencing. Whole-genome sequencing data were generated for Kazakh camels (15 dromedaries and 16 Bactrian camels) and integrated with 131 publicly available genomes representing camel populations from the Arabian Peninsula, Iran, Xinjiang, Inner Mongolia, and Mongolian wild camels. Population structure, genetic diversity, and genome-wide selection were evaluated using principal component analysis, ADMIXTURE, nucleotide diversity, linkage disequilibrium, runs of homozygosity, genomic inbreeding (FROH), and selection scans based on FST, θπ ratio, and XP-EHH. Population genomic analyses revealed clear differentiation between dromedary and Bactrian camels, whereas Kazakh camel populations exhibited higher nucleotide diversity (θπ = 1.307-1.551 × 10-3), and lower genomic inbreeding (median FROH: 0.037-0.056) than Arabian populations. Genome-wide selection analyses identified MC4R as the prominent candidate gene in Kazakh dromedaries and RYR1 as a prominent candidate gene in Kazakh Bactrian camels. Functional enrichment analyses highlighted pathways related to energy metabolism, thermogenesis, calcium signaling, skeletal muscle function, mitochondrial activity, and oxidative stress response. These findings provide new insights into genomic variation potentially associated with environmental adaptation in Kazakh camels and offer valuable genomic resources for future conservation, breeding, and evolutionary studies.

MC4R↗

Genetic variability and characterization of non-structural region 5 of hepatitis C virus genome from Chinese patients.

Sequence variation in the putative non-structural region 5b (NS5b) of hepatitis C virus (HCV) was analyzed in China. Complementary DNA fragments from sera of 49 Chinese patients were amplified by polymerase chain reaction (PCR) and the products were cloned and sequenced. Based on the comparison in NS5b of 33 clones of genotype 1b and 16 clones of genotype 2a, Chinese isolates of HCV belong to the same subtype as HCV-J, and HC-J6 from Japan. There does exist, however, some heterogeneity in the primary structure of the nucleotide acid. Higher homology was found among Chinese isolates than among Chinese isolates and Japanese isolates. Furthermore, among Chinese isolates, we found some conserved nucleotide acid positions different from those of Japanese isolates. Comparison of average homology among the 33 clones of genotype 1b and the 16 clones of genotype 2a indicated that the average homology among genotype 2a was lower than that among genotype 1b. In addition, a deletion of three nucleotide acids and a frame-shift, resulting in the introduction of an in-frame stop codon, were first observed in the NS5b region. These results indicated geographical differences in the distribution of individual HCV isolates, and the existence of a local variant in the same subtype. Our findings also suggested the need for further study on the sequence of genotype 2a, to improve diagnosis and help to advance the development of a vaccine.

Adult↗

How homologous recombination generates a mutable genome.

Recombination and mutation have traditionally been regarded as independent evolutionary processes: the latter generates variation, which the former reshuffles. Recent studies, however, have suggested that allelic recombination influences the underlying mutation rate, as high mutation rates are inferred in regions of high recombination. Furthermore, recombination between duplicated sequences introduces structural variation into the human genome and facilitates the formation of clustered gene families. Comparisons of whole-genome sequences reveal the expansion of gene family clusters to be an important mode of genome evolution. The negative aspect of this genomic dynamism is the contribution of these rearrangements to genetic diseases.

Genetics, Medical↗

Genome rearrangements, deletions, and amplifications in the natural population of Bartonella henselae.

Cats are the natural host for Bartonella henselae, an opportunistic human pathogen and the agent of cat scratch disease. Here, we have analyzed the natural variation in gene content and genome structure of 38 Bartonella henselae strains isolated from cats and humans by comparative genome hybridizations to microarrays and probe hybridizations to pulsed-field gel electrophoresis (PFGE) blots. The variation in gene content was modest and confined to the prophage and the genomic islands, whereas the PFGE analyses indicated extensive rearrangements across the terminus of replication with breakpoints in areas of the genomic islands. We observed no difference in gene content or structure between feline and human strains. Rather, the results suggest multiple sources of human infection from feline B. henselae strains of diverse genotypes. Additionally, the microarray hybridizations revealed DNA amplification in some strains in the so-called chromosome II-like region. The amplified segments were centered at a position corresponding to a putative phage replication initiation site and increased in size with the duration of cultivation. We hypothesize that the variable gene pool in the B. henselae population plays an important role in the establishment of long-term persistent infection in the natural host by promoting antigenic variation and escape from the host immune response.

Animals↗

Fine scale structural variants distinguish the genomes of Drosophila melanogaster and D. pseudoobscura.

BACKGROUND: A primary objective of comparative genomics is to identify genomic elements of functional significance that contribute to phenotypic diversity. Complex changes in genome structure (insertions, duplications, rearrangements, translocations) may be widespread, and have important effects on organismal diversity. Any survey of genomic variation is incomplete without an assessment of structural changes. RESULTS: We re-examine the genome sequences of the diverged species Drosophila melanogaster and D. pseudoobscura to identify fine-scale structural features that distinguish the genomes. We detect 95 large insertion/deletion events that occur within the introns of orthologous gene pairs, the majority of which represent insertion of transposable elements. We also identify 143 microinversions below 5 kb in size. These microinversions reside within introns or just upstream or downstream of genes, and invert conserved DNA sequence. The sequence conservation within microinversions suggests they may be enriched for functional genetic elements, and their position with respect to known genes implicates them in the regulation of gene expression. Although we found a distinct pattern of GC content across microinversions, this was indistinguishable from the pattern observed across blocks of conserved non-coding sequence. CONCLUSION: Drosophila has long been known as a genus harboring a variety of large inversions that disrupt chromosome colinearity. Here we demonstrate that microinversions, many of which are below 1 kb in length, located in/near genes may also be an important source of genetic variation in Drosophila. Further examination of other Drosophila genome sequences will likely identify an array of novel microinversion events.

Animals↗

Structure-function relationship studies in human cholinesterases reveal genomic origins for individual variations in cholinergic drug responses.

1. Due to their involvement in the termination of neurotransmission at cholinergic synapses and neuromuscular junctions, cholinesterases are the target proteins for numerous drugs of neuro-psychopharmacology importance. 2. In order to perform structure-function relationship studies on human cholinesterases with respect to such drugs, a set of expression vectors was engineered, all of which include cloned cDNA inserts encoding various forms of human acetyl- and butyrylcholinesterase. These vectors were designed to be transcribed in vitro into their corresponding mRNA products which, when microinjected into Xenopus oocytes, are efficiently translated to yield their catalytically active enzymes, each with its distinct substrate specificity and sensitivity to selective inhibitors. 3. A fully automated microtiter plate assay for evaluating the inhibition of said enzymes by tested cholinergic drugs and/or poisons has been developed, in conjunction with computerized data analysis, which offers prediction of such inhibition data on the authentic human enzymes and their natural or mutagenized variants. 4. Thus, it was found that asp70-->gly substitution renders butyrylcholinesterase succinylcholine insensitive and resistant to oxime reactivation while ser 425-->Pro with gly70 gives rise to the "atypical" butyrylcholinesterase phenotype, abolishing dibucaine binding. 5. Furthermore, differences in cholinesterase affinities to physostigmine, ecothiophate and bambuterol were shown in these natural variants. 6. Definition of key residues important for drug interactions may initiate rational design of more specific cholinesterase inhibitors, with fewer side effects. This, in turn, offers therapeutic potential in the treatment of clinical syndromes such as Alzheimer's and Parkinson's disease, glaucoma and myasthenia gravis.

Animals↗

Structural divergence of chromosomal segments that arose from successive duplication events in the Arabidopsis genome.

Using the extensive segmental duplications of the Arabidopsis thaliana genome, a comparative study of homoeologous segments occurring in chromosomes 1, 2, 4 and 5 was performed. The gene-by-gene BLASTP approach was applied to identify duplicated genes in homoeologues. The levels of synonymous substitutions between duplicated coding sequences suggest that these regions were formed by at least two rounds of duplications. Moreover, remnants of even more ancient duplication events were recognised by a whole-genome study. We describe a subchromosomal organisation of genes, including the tandemly repeated genes, and the distribution of transposable elements (TEs). In certain cases, evidence of the possible mechanisms of structural rearrangements within the segments could be found. We provide a probable scenario of the rearrangements that took place during the evolution of the homoeologous regions. Furthermore, on the basis of the comparative analysis of the chromosomal segments in the Columbia and Landsberg erecta accessions, an additional structural variation in the A.thaliana genome is described. Analysis of the segments, spanning 7 Mb or 5.6% of the genome, permitted us to propose a model of evolution at the subchromosomal level.

Arabidopsis↗

Long-read sequencing to interrogate strain-level variation among adherent-invasive Escherichia coli isolated from human intestinal tissue.

Adherent-invasive Escherichia coli (AIEC) is a pathovar linked to inflammatory bowel diseases (IBD), especially Crohn's disease, and colorectal cancer. AIEC are genetically diverse, and in the absence of a universal molecular signature, are defined by in vitro functional attributes. The relative ability of difference AIEC strains to colonize, persist, and induce inflammation in an IBD-susceptible host is unresolved. To evaluate strain-level variation among tissue-associated E. coli in the intestines, we develop a long-read sequencing approach to identify AIEC by strain that excludes host DNA. We use this approach to distinguish genetically similar strains and assess their fitness in colonizing the intestine. Here we have assembled complete genomes using long-read nanopore sequencing for a model AIEC strain, NC101, and seven strains isolated from the intestinal mucosa of Crohn's disease and non-Crohn's tissues. We show these strains can colonize the intestine of IBD susceptible mice and induce inflammatory cytokines from cultured macrophages. We demonstrate that these strains can be quantified and distinguished in the presence of 99.5% mammalian DNA and from within a fecal population. Analysis of global genomic structure and specific sequence variation within the ribosomal RNA operon provides a framework for efficiently tracking strain-level variation of closely-related E. coli and likely other commensal/pathogenic bacteria impacting intestinal inflammation in experimental settings and IBD patients.

Animals↗

Protein superfamily evolution and the last universal common ancestor (LUCA).

By exploiting three-dimensional structure comparison, which is more sensitive than conventional sequence-based methods for detecting remote homology, we have identified a set of 140 ancestral protein domains using very restrictive criteria to minimize the potential error introduced by horizontal gene transfer. These domains are highly likely to have been present in the Last Universal Common Ancestor (LUCA) based on their universality in almost all of 114 completed prokaryotic (Bacteria and Archaea) and eukaryotic genomes. Functional analysis of these ancestral domains reveals a genetically complex LUCA with practically all the essential functional systems present in extant organisms, supporting the theory that life achieved its modern cellular status much before the main kingdom separation (Doolittle 2000). In addition, we have calculated different estimations of the genetic and functional versatility of all the superfamilies and functional groups in the prokaryote subsample. These estimations reveal that some ancestral superfamilies have been more versatile than others during evolution allowing more genetic and functional variation. Furthermore, the differences in genetic versatility between protein families are more attributable to their functional nature rather than the time that they have been evolving. These differences in tolerance to mutation suggest that some protein families have eroded their phylogenetic signal faster than others, hiding in many cases, their ancestral origin and suggesting that the calculation of 140 ancestral domains is probably an underestimate.

Evolution, Molecular↗

Whole-genome sequencing implicates rare, low-frequency and structural non-coding variation at the SCN5A locus in Brugada syndrome.

Brugada syndrome (BrS) is an inherited cardiac condition characterized by a hallmark ECG pattern and an increased risk of sudden cardiac death. Central to the aetiology of BrS, the SCN5A region harbours both common non-coding risk variants and rare coding variants that are causative in approximately 20% of patients. However, rare non-coding genetic variation in this region remains largely unexplored. Here, we used whole-genome sequencing (WGS) of 752 European-ancestry BrS cases and 1,827 ancestry-matched controls to identify BrS-associated rare non-coding genetic variation at the SCN5A locus. Sliding-window and cis-regulatory element (CRE)-based rare-variant aggregate testing implicated three conserved CREs, including a dense aggregation of case singleton variants within a 178 bp enhancer in intron 17 of SCN5A which replicated in an independent BrS cohort. Prioritised BrS-associated rare and low-frequency non-coding variants within these elements were predicted to alter cardiac transcription factor motifs, and altered CRE activity in hiPSC-CM luciferase assays or were associated with BrS-relevant ECG endophenotypes in the UK Biobank. Single-variant analysis across the region identified a Bonferroni-significant five-fold case-enriched low-frequency variant within a known CRE in intron 1 of SCN5A, which replicated, was associated with slower cardiac conduction in the UK Biobank and accounted for part of the BrS GWAS signal at this locus. Structural variant analyses identified a 10.5 kb deletion upstream of SCN5A in a BrS case that encompassed a cardiac CRE and reduced sodium current density in a hiPSC-CM model, as well as a 6 kb BrS-enriched retrotransposon insertion in SCN5A that appeared to underlie part of the GWAS signal in this region. Together, these findings implicate rare and low-frequency non-coding variation at the SCN5A locus in BrS susceptibility and demonstrate the value of targeted WGS analysis of key disease loci.

Journal Article↗

In silico prediction of the impact of genomic variations in the small conductance calcium activated potassium channel SK3 structure and function.

The small-conductance calcium-activated potassium channel SK3, encoded by the KCNN3 gene, plays a critical role in regulating dopaminergic neuron (DN) firing patterns by modulating after hyperpolarization currents. SK3 dysfunction has been implicated in neuropsychiatric and neurodegenerative disorders. We analyzed structural and functional consequences of KCNN3 splicing and genetic variation. Alternative splicing variants of the KCNN3 gene were retrieved from the Ensembl database and aligned using T-Coffee, manually inspected and curated. Protein domains were identified with Pfam 35.0, SMART 9.0, and InterPro 98.0, and visualized. An AlphaFold2 model of SK3 full-length protein (UniProt: Q9UGI6) used as reference and structural models of its splicing variants were predicted with ColabFold. Functional domains (S1-S6 transmembrane helices, H5 pore loop, and calmodulin-binding) were defined and superimposed onto the AlphaFold2 reference. Domain integrity was assessed based on completeness of all expected residue indices within each functional region. SNPs and CNVs across all coding KCNN3 splicing variants were analyzed, classified, and filtered to isolate pathogenic variants prioritizing non-synonymous amino acid substitutions. Differential variant impacts across splicing isoforms were assessed by mapping variant positions to individual transcript protein sequences and used to predict functional consequences. Two long and two short splicing variants are known. Short variants lack the motif required for potassium channels. Pathogenic variants result from missense mutations resulting in amino acid substitutions. In all cases, the consequential effects depend on the specific location and role of the amino acid being changed.

SK3 channels↗

Gene-history correlation and population structure.

Correlation of gene histories in the human genome determines the patterns of genetic variation (haplotype structure) and is crucial to understanding genetic factors in common diseases. We derive closed analytical expressions for the correlation of gene histories in established demographic models for genetic evolution and show how to extend the analysis to more realistic (but more complicated) models of demographic structure. We identify two contributions to the correlation of gene histories in divergent populations: linkage disequilibrium, and differences in the demographic history of individuals in the sample. These two factors contribute to correlations at different length scales: the former at small, and the latter at large scales. We show that recent mixing events in divergent populations limit the range of correlations and compare our findings to empirical results on the correlation of gene histories in the human genome.

Genetic Variation↗

The evolution of isochores.

One of the most striking features of mammalian chromosomes is the variation in G+C content that occurs over scales of hundreds of kilobases to megabases, the so-called 'isochore' structure of the human genome. This variation in base composition affects both coding and non-coding sequences and seems to reflect a fundamental level of genome organization. However, although we have known about isochores for over 25 years, we still have a poor understanding of why they exist. In this article, we review the current evidence for the three main hypotheses.

Chromosomes, Human↗

Predicted stem-loop structures and variation in nucleotide sequence of 3' noncoding regions among animal calicivirus genomes.

Caliciviruses are nonenveloped with a polyadenylated genome of approximately 7.6 kb and a single capsid protein. The "RNA Fold" computer program was used to analyze 3'-terminal noncoding sequences of five feline calicivirus (FCV), rabbit hemorrhagic disease virus (RHDV), and two San Miguel sea lion virus (SMSV) isolates. The FCV 3'-terminal sequences are 40-46 nucleotides in length and 72-91% similar. The FCV sequences were predicted to contain two possible duplex structures and one stem-loop structure with free energies of -2.1 to -18.2 kcal/mole. The RHDV genomic 3'-terminal RNA sequences are 54 nucleotides in length and share 49% sequence similarity to homologous regions of the FCV genome. The RHDV sequence was predicted to form two duplex structures in the 3'-terminal noncoding region with a single stem-loop structure, resembling that of FCV. In contrast, the SMSV 1 and 4 genomic 3'-terminal noncoding sequences were 185 and 182 nucleotides in length, respectively. Ten possible duplex structures were predicted with an average structural free energy of -35 kcal/mole. Sequence similarity between the two SMSV isolates was 75%. Furthermore, extensive cloverleaflike structures are predicted in the 3' noncoding region of the SMSV genome, in contrast to the predicted single stem-loop structures of FCV or RHDV.

Base Sequence↗

Molecular characterization of a short interspersed repetitive element from tobacco that exhibits sequence homology to specific tRNAs.

We have characterized a family of tRNA-derived short interspersed repetitive elements (SINEs) in the tobacco genome. Members of this family of SINEs, designated TS, have a composite structure and include a region structurally similar to a rabbit tRNA(Lys), a tRNA-unrelated region, and a TTG repeat of variable length at the 3' end. Southern blot hybridization, together with a search of the GenBank data base, showed that various plants belonging to the families Solanaceae and Convolvulaceae contain sequences homologous to the TS family in the introns and flanking regions of many genes, whereas Arabidopsis in the family Cruciferae and several species of monocoytledonous plants do not. The TS family is widely involved in structural and genetic variations in the genomes of many plants that belong to the order Tubiflorae. All of nine sequences identified in a data base search are truncated at their 5' regions and lack the tRNA-related region of the TS family. We characterized the entire sequence of the members of the TS family and found that this family can be categorized as a member of a group of SINEs with a tRNA(Lys)-like structure, as can several animal SINEs. The TS family can be divided into two major subfamilies by analysis of diagnostic positions, and one of the subfamilies is clearly younger than the other. Amplification of many copies of the full sequence of the younger subfamily occurred during the recent evolution of the tobacco lineage. We also discuss mechanisms that could be involved in the generation of SINEs in animals and also in plants.

Base Sequence↗

The Rise of Plant Pan-Genomes: From Genome Variation to Predictive Breeding.

Plant pan-genomics is entering a new phase beyond genome variation discovery, requiring a shift from cataloguing genomic diversity toward understanding how variation generates biological function and breeding value. Here, we propose that the future of plant pan-genomics will be shaped by three conceptual transitions. First, structural variation (SV), presence-absence variation (PAV), and haplotype diversity should be interpreted not merely as genomic differences, but as regulatory components that influence gene networks, chromatin organization, and complex traits. Second, the expansion from species-level pan-genomes to genus-level super pan-genomes provides an evolutionary framework for uncovering adaptive genetic modules preserved in wild relatives and overlooked during domestication. Third, integrating pan-genomes with pan-omics, three-dimensional genome analyses, and artificial intelligence will enable the transformation of genomic variation into predictive models for crop improvement. We further propose that the ultimate value of pan-genomes lies not in generating increasingly complete genome collections, but in establishing a mechanistic bridge between genome diversity, biological function, and breeding decisions. This transition will move crop improvement from empirical selection toward rational genome design, where evolutionary diversity can be systematically interpreted, predicted, and engineered.

Journal Article↗

Characterization of the mitochondrial genome of Diphyllobothrium latum (Cestoda: Pseudophyllidea) - implications for the phylogeny of eucestodes.

The complete nucleotide sequence of the mitochondrial genome was determined for the fish tapeworm Diphyllobothrium latum. This genome is 13,608 bp in length and encodes 12 protein-coding genes (but lacks the atp8), 22 transfer RNA (tRNA) and 2 ribosomal RNA (rRNA) genes, corresponding to the gene complement found thus far in other flatworm mitochondrial (mt) DNAs. The gene arrangement of this pseudophyllidean cestode is the same as the 6 cyclophyllidean cestodes characterized to date, with only minor variation in structure among these other genomes; the relative position of trnS2 and trnL1 is switched in Hymenolepis diminuta. Phylogenetic analyses of the concatenated amino acid sequences for 12 protein-coding genes of all complete cestode mtDNAs confirmed taxonomic and previous phylogenetic assessments, with D. latum being a sister taxon to the cyclophyllideans. High nodal support and phylogenetic congruence between different methods suggest that mt genomes may be of utility in resolving ordinal relationships within the cestodes. All species of Diphyllobothrium infect fish-eating vertebrates, and D. latum commonly infects humans through the ingestion of raw, poorly cooked or pickled fish. The complete mitochondrial genome provides a wealth of genetic markers which could be useful for identifying different life-cycle stages and for investigating their population genetics, ecology and epidemiology.

Animals↗