Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Graph-based pan-genome reveals structural and functional diversity across oil palm domestication gradients.

BACKGROUND: Oil palm (Elaeis guineensis Jacq.), the world's most land-efficient oil crop, underpins global vegetable oil supply yet faces mounting constraints from limited expansion, climate stress, and disease pressure. These challenges highlight the urgent need for genomic resources that capture species-wide diversity to support sustainable improvement. While recent reference assemblies have advanced trait discovery, single linear genomes fail to represent the full spectrum of structural and gene-content variation, limiting resolution of agronomic alleles. RESULTS: Here, we constructed a graph-based pan-genome from 30 diverse oil palm assemblies representing wild, semi-domesticated, and commercial accessions. We characterized structural variants, gene presence-absence variation, and copy-number gains, with focusing on functional stratification and resistance gene dynamics. The graph-based pan-genome revealed extensive structural and gene-content variation, including a large conserved core, complemented by shell and unique fractions enriched or biased toward regulatory, stress-responsive, and defense-related functions. Structural variation and duplication-derived copy-number gains contributed substantially to gene-content diversity, with semi-domesticated accessions exhibiting the greatest variability. Resistance gene repertoires showed contrasting patterns: receptor-like kinases remained comparatively stable, whereas the CNL subclass of NLR genes contributed disproportionately to shell-genome variation and duplication-associated turnover. CONCLUSIONS: This graph-based pan-genome provides a curated multi-assembly reference and comparative framework for oil palm genomics. By capturing structural variants, gene-content variations, copy-number gains, and resistance gene dynamics across domestication gradients, it establishes a foundation for future pan-GWAS analysis, functional genomics, and molecular breeding strategies aimed at improving resilience and productivity in this globally important crop.

Arecaceae↗

Genome-wide SNP data reveal geographic structure and landscape-associated genomic differentiation in a widespread lizard in arid Eastern Central Asia.

Arid landscapes provide important systems for examining how geographic structure and environmental heterogeneity shape genomic differentiation. In topographically complex desert regions, however, it remains challenging to determine whether population structure primarily reflects landscape resistance, geographic distance, or contemporary environmental variation. Here, we use genome-wide SNP data to investigate population structure, phylogenetic relationships, historical gene flow, demographic history, and landscape correlates of genomic differentiation in the variegated racerunner (Eremias vermiculata), a widespread lacertid lizard across arid Eastern Central Asia. Analyses of 164 individuals recovered six geographically structured nuclear clusters associated with major desert basins and mountain-bounded regions. Nuclear phylogenies resolved two broad regional clades corresponding to northeastern and southwestern parts of the species' range, while PCA and ADMIXTURE analyses recovered six finer-scale genetic clusters. Mitochondrial phylogenies, based on combined NCBI-derived Cyt b and COI sequences from the same individuals, recovered four deeper maternal lineages. These patterns indicate overall phylogeographic agreement between nuclear and mitochondrial datasets, with genome-wide SNPs providing finer-scale resolution of population structure. Demographic reconstructions further uncovered regionally heterogeneous Late Pleistocene histories among clusters, including signals of expansion, stability, and decline. Landscape genomic analyses revealed that genomic differentiation is primarily associated with landscape resistance, particularly elevation and land cover, as well as geographic distance, whereas contemporary environmental variables explained comparatively little variation after controlling for spatial structure. Together, our results suggest that genomic differentiation in E. vermiculata reflects the interplay of persistent landscape configuration, historical connectivity, and region-specific demographic histories across arid Eastern Central Asia. More broadly, this study highlights the value of integrating phylogeographic and landscape genomic approaches for understanding population differentiation and evolutionary history in topographically heterogeneous desert ecosystems.

Arid Eastern Central Asia↗

The fine-scale structure of recombination rate variation in the human genome.

The nature and scale of recombination rate variation are largely unknown for most species. In humans, pedigree analysis has documented variation at the chromosomal level, and sperm studies have identified specific hotspots in which crossing-over events cluster. To address whether this picture is representative of the genome as a whole, we have developed and validated a method for estimating recombination rates from patterns of genetic variation. From extensive single-nucleotide polymorphism surveys in European and African populations, we find evidence for extreme local rate variation spanning four orders in magnitude, in which 50% of all recombination events take place in less than 10% of the sequence. We demonstrate that recombination hotspots are a ubiquitous feature of the human genome, occurring on average every 200 kilobases or less, but recombination occurs preferentially outside genes.

Base Composition↗

Genetic Differentiation is Constrained to Chromosomal Inversions and Putative Centromeres in Locally Adapted Populations With Higher Gene Flow.

The impact of genome structure on adaptation is a growing focus in evolutionary biology, revealing an important role for structural variation and recombination landscapes in shaping genetic diversity across genomes and among populations. This is particularly relevant when local adaptation occurs despite gene flow, where clustering of differentiated loci can maintain locally adapted variants by reducing recombination between them. However, the limited genomic resources for nonmodel species, including reference genomes and recombination maps, have constrained our understanding of these patterns. In this study, we leverage the Atlantic silverside-a nonmodel fish with extensive local adaptation across a steep latitudinal gradient-as an ideal system to explore how genome structure influences adaptation under varying levels of gene flow, using a newly available reference genome and multiple recombination maps. Analyzing 168 genomes from four populations, we found a continuum of genome-wide differentiation increasing from south to north, reflecting higher connectivity among southern populations and reduced gene flow at northern latitudes. With increasing gene flow, the number and clustering of FST outlier loci also increased, with differentiated loci found exclusively within large haploblocks harboring inversions and smaller peaks overlapping putative centromeric regions. Notably, sequence divergence was only evident in inversions, supporting their role in adaptive divergence with gene flow, whereas centromeric regions appeared differentiated because of low recombination and diversity, with no indication of elevated divergence. Our results support the hypothesis that clustered genomic architectures evolve with high gene flow and enhance our understanding of how inversions and centromeres are linked to different evolutionary processes.

Gene Flow↗

Comparisons of genetic variability and genome structure among mosquito strains selected for refractoriness to a malaria parasite.

Restriction fragment length polymorphism (RFLP) markers were used to evaluate Aedes aegypti genome structure and genetic variability within and between substrains selected for different levels of refractoriness to the malaria parasite, Plasmodium gallinaceum. The MOYO-R substrain was previously selected for complete refractoriness and the MOYO-IS substrain for intermediate susceptibility from the Moyo-In-Dry (MOYO) strain by selective inbreeding (F = 0.5). Eighteen mapped RFLP markers were used to provide coverage of the mosquito genome. The two substrains showed reduced genetic diversity compared with the MOYO strain, including significant reductions in mean heterozygosity, number of alleles per locus, and proportion of polymorphic loci. Genetic differentiation between the two substrains was statistically significant, as reflected by differences in allele frequencies. Significant pairwise linkage disequillbrium among the RFLP loci was detected in all three strains, most evidently in the MOYO strain. This is surprising because the RFLP loci examined are separated by large map distances, and therefore linkage disequilibrium should decay to zero after many generations of laboratory culture. Our hypothesis to explain this phenomena is that lack of recombination, or low recombination rates in some regions of the A. aegypti genome, is a result of chromosome inversions. Finally, we used graphical genotyping, wherein whole genome genotypic information for individual mosquitoes is represented in a simple graphic format, to illustrate genome structure and allelic variation within and among the mosquito strains. Our analysis revealed an apparent chromosomal deletion on chromosome 3 for some individuals in the MOYO strain and MOYO-IS substrain.

Aedes↗

Evolutionary instability of operon structures disclosed by sequence comparisons of complete microbial genomes.

Gene orders have been shown to be generally unstable by comprehensive analyses in several complete genomes. In this study, we examined instability of genome structures within operons, where functionally related genes are clustered. We compared gene orders of known operons obtained from Escherichia coli and Bacillus subtilis with corresponding those of operons in 11 complete genome sequences. We found that in many cases, gene orders within operons could be shuffled frequently during evolution, although several operon structures, such as ribosomal protein operons, were well conserved. This suggests that shuffling of a genome structure is virtually neutral in long-term evolution. Moreover, degrees of instability of the operon structures depended on the genomes examined. Variation in degrees of instability of the genome structures was likely to be related to differences in amounts of insertion sequences. Effects on transcription regulation are also discussed in association with operon destruction.

Bacillus subtilis↗

Rapid emergence of novel antigenic and genetic variants of equine infectious anemia virus during persistent infection.

Previous results from our laboratory have demonstrated that equine infectious anemia virus displays structural variations in its surface glycoproteins and RNA genome during passage and chronic infections in experimentally infected Shetland ponies (Montelaro et al., J. Biol. Chem. 259:10539-10544, 1984; Payne et al., J. Gen. Virol. 65:1395-1399, 1984). The present study was undertaken to obtain an antigenic and biochemical characterization of equine infectious anemia virus isolates recovered from an experimentally infected pony during sequential disease episodes, each separated by intervals of only 4 to 8 weeks. The virus isolates could be distinguished antigenically by neutralization assays with serum from the infected pony and by Western blot analysis with a monoclonal antibody against the major surface glycoprotein gp90, thus demonstrating that novel antigenic variants of equine infectious anemia virus predominate during each clinical episode. The respective virion glycoproteins displayed different electrophoretic mobilities on sodium dodecyl sulfate-polyacrylamide gels, indicating structural variation. Tryptic peptide and glycopeptide maps of the viral proteins of each virus isolate revealed biochemical alterations involving amino acid sequence and glycosylation patterns in the virion surface glycoproteins gp90 and gp45. In contrast, no structural variation was observed in the internal viral proteins pp15, p26, and p9 from any of the four virus isolates. Oligonucleotide mapping experiments revealed similar but unique RNase T1-resistant oligonucleotide fingerprints of the RNA genomes of each of the virus isolates. Localization of altered oligonucleotides for one virus isolate placed two of three unique oligonucleotides within the predicted env gene region of the genome, perhaps correlating with the structural variation observed in the envelope glycoproteins. Thus these results support the concept that equine infectious anemia virus is indeed capable of relatively rapid genomic variations during replication, some of which result in altered glycoprotein structures and antigenic variants which are responsible for the unique periodic disease nature observed in persistently infected animals. The findings of envelope specific differences in isolates of visna virus and of human T-cell lymphotropic virus III (acquired immune deficiency syndrome-related virus) suggest that this variation may be a common characteristic of the subfamily Lentivirinae.

Animals↗

Identification of novel non-autonomous CemaT transposable elements and evidence of their mobility within the C. elegans genome.

We describe here two new transposable elements, CemaT4 and CemaT5, that were identified within the sequenced genome of Caenorhabditis elegans using homology based searches. Five variants of CemaT4 were found, all non-autonomous and sharing 26 bp inverted terminal repeats (ITRs) and segments (152-367 bp) of sequence with similarity to the CemaT1 transposon of C. elegans. Sixteen copies of a short, 30 bp repetitive sequence, comprised entirely of an inverted repeat of the first 15 bp of CemaT4's ITR, were also found, each flanked by TA dinucleotide duplications, which are hallmarks of target site duplications of mariner-Tc transposon transpositions. The CemaT5 transposable element had no similarity to maT elements, except for sharing identical ITR sequences with CemaT3. We provide evidence that CemaT5 and CemaT3 are capable of excising from the C. elegans genome, despite neither transposon being capable of encoding a functional transposase enzyme. Presumably, these two transposons are cross-mobilised by an autonomous transposon that recognises their shared ITRs. The excisions of these and other non-autonomous elements may provide opportunities for abortive gap repair to create internal deletions and/or insert novel sequence within these transposons. The influence of non-autonomous element mobility and structural diversity on genome variation is discussed.

Animals↗

HCV genotypes--role in pathogenesis of disease and response to therapy.

Hepatitis C virus (HCV) shows considerable variation in its genomic structure, allowing classification into six main genotypes. Epidemiological studies have shown marked differences in genotype distribution by geographical region, and between patient groups. Improved understanding of the rate of nucleotide sequence mutation in HCV has allowed the approximate time of divergence of major genotypes to be estimated, and the origin and spread of the present epidemic of hepatitis C to be better defined. Improved methods of genotype definition over the last few years have enabled the importance of genotype in the progression of HCV-related disease and response to anti-viral therapy to be studied. Present data strongly indicates that HCV genotype is an important determinant of response to treatment, but the effect of genotype on disease progression has been harder to clarify. This is largely due to the absence of model systems of HCV infection, the epidemiological differences in patient groups infected with the different genotypes, and the lack of good prospective longitudinal clinical data. As a result of advances in methodology, and recent results of large clinical trials of combination therapy, a knowledge of HCV genotype is now central to the clinician in the management of patients with chronic hepatitis C.

Antiviral Agents↗

Globin gene structure and the nature of mutation.

In this chapter I have tried to relate the salient features of globin gene structure, genomic organization, normal variation, and mutations affecting gene expression. The lessons learned from the normal beta-globin gene cluster and mutations producing beta-thalassemia should be highly applicable to studies of other inherited diseases. As more and more gene probes become available, which have relevance to the study of human disease, the striking extent of genetic heterogeneity producing single gene disorders of man will be illuminated.

Gene Expression Regulation↗

Gene structure and the nature of mutation.

In this paper we have tried to relate the salient features of gene structure, genomic organization, normal variation, and mutations affecting gene expression. The lessons learned from the normal beta-globin gene cluster and mutations producing beta-thalassemia should be greatly applicable to studies of other inherited diseases. As more and more gene probes become available which have relevance to the study of human disease, the striking extent of genetic heterogeneity producing single gene disorders of man will be illuminated.

Animals↗

Genome assembly comparison identifies structural variants in the human genome.

Numerous types of DNA variation exist, ranging from SNPs to larger structural alterations such as copy number variants (CNVs) and inversions. Alignment of DNA sequence from different sources has been used to identify SNPs and intermediate-sized variants (ISVs). However, only a small proportion of total heterogeneity is characterized, and little is known of the characteristics of most smaller-sized (<50 kb) variants. Here we show that genome assembly comparison is a robust approach for identification of all classes of genetic variation. Through comparison of two human assemblies (Celera's R27c compilation and the Build 35 reference sequence), we identified megabases of sequence (in the form of 13,534 putative non-SNP events) that were absent, inverted or polymorphic in one assembly. Database comparison and laboratory experimentation further demonstrated overlap or validation for 240 variable regions and confirmed >1.5 million SNPs. Some differences were simple insertions and deletions, but in regions containing CNVs, segmental duplication and repetitive DNA, they were more complex. Our results uncover substantial undescribed variation in humans, highlighting the need for comprehensive annotation strategies to fully interpret genome scanning and personalized sequencing projects.

Base Sequence↗

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99&#xd7;) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including&#x2009;~&#x2009;17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8&#x2009;&#xb1;&#x2009;8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (&#x3c0;&#x2009;=&#x2009;0.00267), followed by lowland (&#x3c0;&#x2009;=&#x2009;0.00233), whereas highland chickens showed the lowest diversity (&#x3c0;&#x2009;=&#x2009;0.00203) and elevated genomic inbreeding (FROH and FHOM &#x2248; 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals↗

Importance of genetic drift during Pleistocene divergence as revealed by analyses of genomic variation.

Determining what factors affect the structuring of genetic variation is key to deciphering the relative roles of different evolutionary processes in species differentiation. Such information is especially critical to understanding how the frequent shifts and fragmentation of species distributions during the Pleistocene translates into species differences, and why the effect of such rapid climate change on patterns of species diversity varies among taxa. Studies of mitochondrial DNA (mtDNA) have detected significant population structure in many species, including those directly impacted by the glacial cycles. Yet, understanding the ultimate consequence of such structure, as it relates to how species divergence occurs, requires demonstration that such patterns are also shared with genomic patterns of differentiation. Here we present analyses of amplified fragment length polymorphisms (AFLPs) in the montane grasshopper Melanoplus oregonensis to assess the evolutionary significance of past demographic events and associated drift-induced divergence as inferred from mtDNA. As an inhabitant of the sky islands of the northern Rocky Mountains, this species was subject to repeated and frequent shifts in species distribution in response to the many glacial cycles. Nevertheless, significant genetic structuring of M. oregonensis is evident at two different geographic and temporal scales: recent divergence associated with the recolonization of the montane meadows in individual sky islands, as well as older divergence associated with displacements into regional glacial refugia. The genomic analyses indicate that drift-induced divergence, despite the lack of long-standing geographic barriers, has significantly contributed to species divergence during the Pleistocene. Moreover, the finding that divergence associated with past demographic events involves the repartitioning of ancestral variation without significant reductions of genomic diversity has intriguing implications - namely, the further amplification of drift-induced divergence by selection.

Animals↗

Haplotype block structure is conserved across mammals.

Genetic variation in genomes is organized in haplotype blocks, and species-specific block structure is defined by differential contribution of population history effects in combination with mutation and recombination events. Haplotype maps characterize the common patterns of linkage disequilibrium in populations and have important applications in the design and interpretation of genetic experiments. Although evolutionary processes are known to drive the selection of individual polymorphisms, their effect on haplotype block structure dynamics has not been shown. Here, we present a high-resolution haplotype map for a 5-megabase genomic region in the rat and compare it with the orthologous human and mouse segments. Although the size and fine structure of haplotype blocks are species dependent, there is a significant interspecies overlap in structure and a tendency for blocks to encompass complete genes. Extending these findings to the complete human genome using haplotype map phase I data reveals that linkage disequilibrium values are significantly higher for equally spaced positions in genic regions, including promoters, as compared to intergenic regions, indicating that a selective mechanism exists to maintain combinations of alleles within potentially interacting coding and regulatory regions. Although this characteristic may complicate the identification of causal polymorphisms underlying phenotypic traits, conservation of haplotype structure may be employed for the identification and characterization of functionally important genomic regions.

Animals↗

Whole-Genome Sequencing Reveals Population Structure, Genetic Diversity, and Selection Signatures in Kazakh Dromedary and Bactrian Camels.

Understanding the genomic basis of environmental adaptation is essential for the conservation and genetic improvement of domestic camels. In this study, we investigated the population structure, genetic diversity, and genomic variation potentially associated with environmental adaptation of Kazakh dromedary and Bactrian camels using whole-genome sequencing. Whole-genome sequencing data were generated for Kazakh camels (15 dromedaries and 16 Bactrian camels) and integrated with 131 publicly available genomes representing camel populations from the Arabian Peninsula, Iran, Xinjiang, Inner Mongolia, and Mongolian wild camels. Population structure, genetic diversity, and genome-wide selection were evaluated using principal component analysis, ADMIXTURE, nucleotide diversity, linkage disequilibrium, runs of homozygosity, genomic inbreeding (FROH), and selection scans based on FST, &#x3b8;&#x3c0; ratio, and XP-EHH. Population genomic analyses revealed clear differentiation between dromedary and Bactrian camels, whereas Kazakh camel populations exhibited higher nucleotide diversity (&#x3b8;&#x3c0; = 1.307-1.551 &#xd7; 10-3), and lower genomic inbreeding (median FROH: 0.037-0.056) than Arabian populations. Genome-wide selection analyses identified MC4R as the prominent candidate gene in Kazakh dromedaries and RYR1 as a prominent candidate gene in Kazakh Bactrian camels. Functional enrichment analyses highlighted pathways related to energy metabolism, thermogenesis, calcium signaling, skeletal muscle function, mitochondrial activity, and oxidative stress response. These findings provide new insights into genomic variation potentially associated with environmental adaptation in Kazakh camels and offer valuable genomic resources for future conservation, breeding, and evolutionary studies.

MC4R↗

Genetic variability and characterization of non-structural region 5 of hepatitis C virus genome from Chinese patients.

Sequence variation in the putative non-structural region 5b (NS5b) of hepatitis C virus (HCV) was analyzed in China. Complementary DNA fragments from sera of 49 Chinese patients were amplified by polymerase chain reaction (PCR) and the products were cloned and sequenced. Based on the comparison in NS5b of 33 clones of genotype 1b and 16 clones of genotype 2a, Chinese isolates of HCV belong to the same subtype as HCV-J, and HC-J6 from Japan. There does exist, however, some heterogeneity in the primary structure of the nucleotide acid. Higher homology was found among Chinese isolates than among Chinese isolates and Japanese isolates. Furthermore, among Chinese isolates, we found some conserved nucleotide acid positions different from those of Japanese isolates. Comparison of average homology among the 33 clones of genotype 1b and the 16 clones of genotype 2a indicated that the average homology among genotype 2a was lower than that among genotype 1b. In addition, a deletion of three nucleotide acids and a frame-shift, resulting in the introduction of an in-frame stop codon, were first observed in the NS5b region. These results indicated geographical differences in the distribution of individual HCV isolates, and the existence of a local variant in the same subtype. Our findings also suggested the need for further study on the sequence of genotype 2a, to improve diagnosis and help to advance the development of a vaccine.

Adult↗

How homologous recombination generates a mutable genome.

Recombination and mutation have traditionally been regarded as independent evolutionary processes: the latter generates variation, which the former reshuffles. Recent studies, however, have suggested that allelic recombination influences the underlying mutation rate, as high mutation rates are inferred in regions of high recombination. Furthermore, recombination between duplicated sequences introduces structural variation into the human genome and facilitates the formation of clustered gene families. Comparisons of whole-genome sequences reveal the expansion of gene family clusters to be an important mode of genome evolution. The negative aspect of this genomic dynamism is the contribution of these rearrangements to genetic diseases.

Genetics, Medical↗