Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

The Indian Genome Variation database (IGVdb): a project overview.

Indian population, comprising of more than a billion people, consists of 4693 communities with several thousands of endogamous groups, 325 functioning languages and 25 scripts. To address the questions related to ethnic diversity, migrations, founder populations, predisposition to complex disorders or pharmacogenomics, one needs to understand the diversity and relatedness at the genetic level in such a diverse population. In this backdrop, six constituent laboratories of the Council of Scientific and Industrial Research (CSIR), with funding from the Government of India, initiated a network program on predictive medicine using repeats and single nucleotide polymorphisms. The Indian Genome Variation (IGV) consortium aims to provide data on validated SNPs and repeats, both novel and reported, along with gene duplications, in over a thousand genes, in 15,000 individuals drawn from Indian subpopulations. These genes have been selected on the basis of their relevance as functional and positional candidates in many common diseases including genes relevant to pharmacogenomics. This is the first large-scale comprehensive study of the structure of the Indian population with wide-reaching implications. A comprehensive platform for Indian Genome Variation (IGV) data management, analysis and creation of IGVdb portal has also been developed. The samples are being collected following ethical guidelines of Indian Council of Medical Research (ICMR) and Department of Biotechnology (DBT), India. This paper reveals the structure of the IGV project highlighting its various aspects like genesis, objectives, strategies for selection of genes, identification of the Indian subpopulations, collection of samples and discovery and validation of genetic markers, data analysis and monitoring as well as the project's data release policy.

Databases, Genetic↗

Integrated multi-omics analyses provide new insights into genomic variation landscape and regulatory network candidate genes associated with walnut endocarp.

Persian walnut (Juglans regia) is an economically important nut oil tree; the fruit has a hard endocarp/shell to protect seeds, thus playing a key role in its evolution, and the shell thickness is an important trait for walnut breeding. However, the genomic landscape and the gene regulatory networks associated with walnut shell development remain to be systematically elucidated. Here, we report a high-quality genome assembly of the walnut cultivar 'Xiangling' and construct a graphic structure pan-genome of eight Juglans species to reveal the genetic variations at the genome level. We re-sequence 285 accessions to characterize the genomic variation landscape. Through genome-wide association studies (GWAS), we identified 19 loci associated with more than 268 loci that underwent selection during walnut domestication and improvement. Multi-omics analyses, including transcriptomics, metabolomics, DNA methylation, and spatial transcriptomics across eleven developmental stages, revealed several candidate genes related to secondary cell biosynthesis and lignin accumulation. This integrated multi-omics approach revealed several candidate genes associated with secondary cell biosynthesis and lignin accumulation, such as UGP, MYB308, MYB83, NAC043, NAC073, CCoAOMT1, CCoAOMT7, CHS2, CESA7, LAC7, COBL4, and IRX12. Overexpression of JrUGP and JrMYB308 in Arabidopsis thaliana confirmed their roles in lignin biosynthesis and cell wall thickening. Consequently, our comprehensive multi-omics findings offer novel insights into walnut genetic variation and network regulation of endocarp development and shell thickness, which enable further genome-informed breeding strategies for walnut cultivar improvement.

Juglans↗

Polygenic and monogenic adaptation drive evolutionary rescue at different magnitudes of environmental change.

Understanding the genetic basis of rapid adaptation is key to predicting species' evolutionary responses to environmental change. However, it is still debatable whether many small-effect mutations or a few large-effect mutations underlie rapid adaptation, and how this knowledge can predict population survival or extinction. To address this question, we performed a series of ecologically grounded forward-in-time genetic simulations to study rapid adaptation and extinction with increasing magnitudes of environmental change. These simulations were seeded with genomic variation of the plant Arabidopsis thaliana to have a realistic genomic structure, with one (monogenic) to 1,000 (polygenic) variants with varying heritabilities contributing to an environmental adaptive trait. Our results revealed two distinct scenarios of rapid adaptation and population rescue. Under small-to-moderate environmental shifts, high polygenic traits increased evolutionary rescue probability. Under extreme environmental shifts, high polygenic traits lead predictably to extinction, yet monogenic traits sometimes produce one-off winning adaptive genotypes. We interpret our rapid evolutionary rescue findings in terms of the fundamental theorem of natural selection, where trait polygenicity shapes the distribution of genetic variance in fitness across replicates and, in turn, the probability of population survival, with polygenic architectures producing more stable and predictable fitness variance and monogenic architectures generating highly skewed and variable outcomes. These results highlight the insights genomics gives us into the (un)predictability of species' evolutionary responses to global change, with management implications for assisted adaptation and conservation.

Arabidopsis↗

Conservation and variation of gene regulation in embryonic stem cells assessed by comparative genomics.

We have examined the gene structure and regulatory regions of octamer-binding transcription factor 3/4 (Oct 3/4), sex determining region Y box 2 (Sox2), signal transducer and activator of transcription 3 (Stat3), embryonal stem cell-specific gene 1 (ESG), Nanog homeobox (Nanog), and several other genes highly expressed in embryonic stem (ES) cells across different species. Our analysis showed that ES cell-expressed Ras (ERAS) was orthologous to a human pseudogene Harvey Ras (HRASP) and that the promoter and other regulatory sequences were highly divergent. No ortholog of (ES) cell-derived homeobox containing gene (Ehox) could be identified in human, and the closest paralogs PEPP gene subfamily 1 (PEPP1), PEPP2, and extraembryonic, spermatogenesis, homeobox 1 (Esx1) were not expressed by ES cells and shared little homology. The Sox2 promoter was the most conserved across species and the Oct3/4 promoter region showed significant homology particularly in the distal enhancer active in ES cells. Analysis suggested common and divergent pathways of regulation. Conserved Oct3/4 and Sox2 co-binding domains were identified in most ES expressed genes, highlighting the importance of this transcriptional pathway. Conserved fibroblast growth factor response element sites were identified in regulatory regions, suggesting a potential parallel pathway for regulation by FGFs. A central role of Stat3 activation in self-renewal and in a regulatory feedback loop was suggested by the identification of the conserved binding sites in most pathways. Although most pathways were evolutionarily conserved, promoters and genomic structure of the leukemia inhibitory factor (LIF) pathway components were divergent, likely explaining the differential requirement of LIF for human and rodent cells. Our analysis further suggested that the Nanog regulatory pathway was relatively independent of the LIF/Oct pathway and may interact with the Nodal/transforming growth factor-beta pathway. These results provide a framework for examining the current reported differences between rodent and human ES cells and define targets for future perturbation studies.

Amino Acid Sequence↗

Heavy chain joining region segments of the channel catfish. Genomic organization and phylogenetic implications.

The JH locus of the channel catfish has been characterized to determine the organization and structural diversity of JH segments. These analyses indicate that there are a total of nine JH segments tightly clustered within a region spanning about 2.2 kb. The JH locus is closely linked to the CH 1 domain of the expressed catfish H chain; the distance between the CH proximal JH segment (JH9) and the CH 1 domain is about 1.8 kb. Each JH segment has an upstream recombination sequence, which includes a T-rich nonamer, a 22- to 24-bp spacer, and a phylogenetically conserved heptamer. Each JH segment also has an open reading frame that encodes the conserved framework region 4 tryptophan (Trp-103) and terminates with a RNA donor splice site. The catfish JH locus contains an internal repetitive sequence region characterized by a short (183-188 bp) repeat that occurs sequentially five times. Strong sequence homology as well as the unified length of the repeated sequences indicate that JH segments JH3-JH7 probably arose as the result of a series of homologous but unequal crossover events. Sequence alignments of the duplicated JH segments indicates that there is diversity within the 5-11 nucleotides located immediately downstream from the heptamer, an observation which indicates that closely related JH segments can serve to enhance CDR3 diversity in the expressed H chain. Comparisons of the genomic JH sequences with different cDNA clones indicate that each JH segment is probably functional and that junctional diversity serves an important role in the generation of CDR3 diversity. In addition, single base differences observed in comparisons of JH-encoded regions indicate that there is probably somatic mutation or allelic variation of genomic JH segments. These studies suggest that the characteristic structure and organizational pattern of JH segments in higher vertebrates may have evolved early in vertebrate phylogeny at the level of the bony fish.

Amino Acid Sequence↗

Study of correlations in DNA sequences.

We present a method for unified statistical analysis of short and long range correlations between various nucleotides in genomic DNA strands. The approach is based on the mutual study of Fourier structure factor spectra and pair correlation functions. The analysis of cross correlations in the different ranges of structural spectra permits identification of the main sources of correlations, namely, the coherent point mutations, coincident periodicities or large scale density variations. The technique for assessment of structural coupling between various genes in the genome is also described.

Animals↗

KCFtools: rapid alignment-free method for introgression screening and GWAS using k-mer profiles.

MOTIVATION: In the era of multiple genome references, researchers often align sequencing reads against distinct assemblies or even multiple references simultaneously. This enables applications such as the detection of introgressed segments or highly variable genomic regions, which are especially prevalent in large-genome crop species such as lettuce or wheat. However, these applications come at the cost of increased computational burden, inconsistencies in mapping methods, and reduced reproducibility across studies. To address these limitations, we developed KCFtools, a Java-based toolkit that identifies the presence and absence of k-mers in nonoverlapping genomic or transcriptomic windows by comparing query and reference genomes. This alignment-free approach enables the efficient computation of an identity score for each window, thereby facilitating robust detection of introgressed or variable regions across genomes. RESULTS: We systematically evaluated the performance and accuracy of the k-mer-based method implemented in KCFtools, benchmarking it against conventional single nucleotide variation-based introgression detection pipelines. Our results demonstrate that KCFtools effectively captures introgressed segments and structurally diverse regions, even in species with fragmented or highly divergent reference genomes. In addition, we extended KCFtools to generate genotype matrices from k-mer variation tables. These matrices are compatible with genome-wide association studies software and allow the identification of loci associated with phenotypic traits. We showcase the utility of this approach by detecting known and novel associations for downy mildew resistance in lettuce, underscoring the pipeline's potential for high-resolution, reference-agnostic population genetic analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/sivasubramanics/kcftools.

Software↗

Evolution and domestication-trait associations of ultra-long centromere haplotypes in pepper plants.

Centromeric and pericentromeric regions of most eukaryotic genomes are highly repetitive and strongly recombination-suppressed, confounding efforts to resolve genetic variation, population structure and phenotypic associations. Pepper (Capsicum annuum) centromeres are nearly devoid of satellite repeats, facilitating assembly and population-level comparison of centromeric regions. Here we integrate 9 near-complete genome assemblies, CENH3 ChIP-seq profiles from 26 diverse accessions, and resequencing and phenotypic data from ~400 cultivated and wild accessions to investigate population-level diversity and phenotypic relevance of pepper peri/centromeric regions. Functional centromere positions are largely fixed on 8 of 12 chromosomes, whereas the remaining 4 carry distinct centromeric epialleles shaped mainly by centromere repositioning and pericentromeric inversions. Pepper centromeres are embedded within ultra-long centromere-spanning haplotype (cenhap) blocks, ranging from 29.8 to 112.9 Mb and collectively covering 23.96% of the genome; each block contains only 1-4 major haplotypes. Some cenhaps may act as supergene-like units and are strongly associated with fruit traits, probably because recombination-suppressed intervals harbour multiple fruit-related genes, including OFP and F-box genes. F2 segregation assays further reveal transmission distortion of chromosomes carrying alternative cenhaps. Together, these findings highlight peri/centromeric regions as underrecognized reservoirs of agronomically important variation.

Centromere↗

Ultrastructure meets reproductive success: performance of a sphecid wasp is correlated with the fine structure of the flight-muscle mitochondria.

Organisms show a remarkable inter-individual variation in reproductive success. The proximate causes of this variation are not well understood. We hypothesized that the ultrastructure of costly or complex tissues or organelles might affect reproductive performance. We tested this hypothesis in females of a sphecid wasp, the European beewolf, Philanthus triangulum (Hymenoptera, Sphecidae), that show considerable variation in reproductive success. The most critical component of reproduction in beewolf females is flying with paralysed honeybees, which more than double their weight. Because of the high energetic requirements for flight, we predicted that the ultrastructure of the flight-muscle mitochondria might influence female success. We determined the density of mitochondria and the density of the inner mitochondrial membranes (DIMM) of the flight muscles as well as age, body size and fat content. Only DIMM had a significant influence on female reproductive success, which might be mediated by an elevated adenosine triphosphate (ATP) supply. The variation in DIMM might result from differences in larval provisions or from an accumulation of mutations in the mitochondrial genome. Our results support the hypothesis that the organization of complex structures contributes to inter-individual variation in reproductive success.

Adenosine Triphosphate↗

Inference of Genetic Structure and the Process of Population Formation in Nepalese Native Goats Using Uniparental and Genome-Wide Markers.

Nepal is a small, landlocked country with marked elevational variation from the Terai plains to the Himalayas. Here, four indigenous goat populations (Chyangra, Sinhal, Khari, and Terai) are raised at different elevations. This study aimed to clarify the genetic structure of these populations and how they are formed and propagated across the Himalayan region. We analyzed 136 Nepalese goats using mitochondrial (mt) DNA D-loop and sex-determining region Y (SRY) 3'-untranslated region (UTR) sequences, as well as 50 K SNP array data. The mtDNA haplogroups D (0.162) and G (0.03) were detected only in Chyangra, whereas haplogroup B was predominant in Sinhal (0.42), followed by Khari (0.260). Regarding SRY haplotypes, Y2B was detected in all populations, whereas Y1AB (0.42) was found only in Chyangra. Genome-wide SNP analysis showed that Chyangra was genetically related to Tibetan and Central Asian goats, while Terai resembled South Asian goats. Interestingly, Sinhal formed a distinct cluster, whereas Khari exhibited an admixed genetic structure. These findings suggest that Nepalese goats originate from at least three ancestral lineages and that an additional migration route may have existed through the southern Himalayas.

50K SNP↗

A comparison of chromosomal and allozymal variation across a narrow hybrid zone in the grasshopper Caledia captiva.

A hydrid zone between the Moreton and Torresian taxa of the grasshopper Caledia captiva in S.E. Queensland has been characterised in terms of allozyme and chromosome variation within the same individuals.--On chromosomal criteria (pericentric rearrangements), the zone is asymmetrical with evidence of high levels of introgression of Torresian chromosomes into the Moreton taxon. This is apparent from the analysis of two independent transects across the hydrid zone. Major changes in chromosomal frequency occur over distances of less than 0.5 km. and the level of introgression differs between the two transects, with much higher levels in the northern Moreton populations, characterised by an acrocentric X-chromosome, when compared with the southern metacentric-X Moreton populations. Chromosome analysis of samples taken from the same transect over two years has revealed no major changes in the structure of the zone. Moreover, a Moreton population located only 0.5 km. from the null point was found to be stable over 6 generations with evidence for a new balanced genome having originated following the differential incorportation of Torresian chromosomes.--Contrary to the chromosomal situation, the same hybrid zone was found to be symmetrical with respect to allozyme variation with evidence of movement of diagnostic alleles in both directions across the zone. The alleles are independent and not tightly linked to any of the pericentric rearrangements. Thus these 5 alleles are acting as markers of the background genome and reveal the relatively free movement of genes which are located outside the pericentric rearrangements.--It is proposed that the hybrid zone in Caledia captiva is unstable and is moving slowly in a westerly direction into the Torresian territory. This is due to the ability of the Moreton taxon to incorporate more readily into its genome those Torresian chromosomes or chromosome segments which increase the fitness of the Moreton taxon. On chromosomal criteria, the Torresian taxon does not share the same capacity.--It is suggested that, so long as the two taxa retain their ability to hybridise with subsequent asymmetrical introgression, the zone will continue to move westwards and eventually lead to the selective incorporation of the Torresian genome into the Moreton taxon. This will result in a polymorphic situation with clinal variation in chromosomal frequencies. The structure of the zone is dependent upon a fine balance between genomic reorganisation in recombinant genotypes and the relative dispersal capacities of the two hydridising taxa.

Animals↗

High expression of UDP-N-acetylglucosamine: beta-D mannoside beta-1,4-N-acetylglucosaminyltransferase III (GnT-III) in chronic myelogenous leukemia in blast crisis.

The activity and mRNA expression of UDP-N-acetylglucosamine: beta-D mannoside beta-1,4-N-acetylglucosaminyl transferase III (GnT-III: EC 2.4.1.144) were investigated in hematological malignancies. GnT-III activity was elevated in patients with chronic myelogenous leukemia (CML) in blast crisis and patients with multiple myeloma (MM), as compared to normal healthy subjects and patients with other hematological malignancies including CML in chronic phase. The GnT-III transcript was the same size in leukemic cells from various hematological diseases and cell lines, while expression of the transcript was not found to correlate significantly with enzyme activity, implying that post-translational modification might regulate the activity of GnT-III. Southern-blot analysis showed no significant variation in the structure and position of the GnT-III genome, indicating that the gene is present as a single copy without isoforms. Furthermore, analyses by immunoprecipitation and Western blot revealed that high GnT-III activity in KU812 cell, a CML cell line, resulted in an increase in E4-PHA binding to CD45, a major surface glycoprotein of the leukocyte, indicating that more bisecting GlcNAc was added to CD45 catalyzed by elevated GnT-III.

Acetylglucosamine↗

Innovation from reduction: gene loss, domain loss and sequence divergence in genome evolution.

Analyses of genome sequences have revealed a surprisingly variable distribution of genes, reflecting the generation of novel genes, lateral gene transfer and gene loss. The impact of gene loss on organisms has been difficult to examine, but the loss of protein coding genes, the loss of domains within proteins and the divergence of genes have made surprising contributions to the differences among organisms. This paper reviews surveys of gene loss and divergence in fungal and archaeal genomes that indicate suites of functionally related genes tend to undergo loss and divergence. Instances of fungal gene loss highlighted here suggest that specific cellular systems have changed, such as Ca 2+ biology in Saccharomyces cerevisiae and peroxisome function in Schizosaccharomyces pombe. Analyses of loss and divergence can provide specific predictions regarding protein-protein interactions, and the relationship between networks of protein interactions and loss may form a part of a parametric model of genome evolution.

Chromosome Mapping↗

Chromosomal inversion polymorphism leads to extensive genetic structure: a multilocus survey in Drosophila subobscura.

The adaptive character of inversion polymorphism in Drosophila subobscura is well established. The O(ST) and O(3+4) chromosomal arrangements of this species differ by two overlapping inversions that arose independently on O(3) chromosomes. Nucleotide variation in eight gene regions distributed along inversion O(3) was analyzed in 14 O(ST) and 14 O(3+4) lines. Levels of variation within arrangements were quite similar along the inversion. In addition, we detected (i) extensive genetic differentiation between arrangements in all regions, regardless of their distance to the inversion breakpoints; (ii) strong association between nucleotide variants and chromosomal arrangements; and (iii) high levels of linkage disequilibrium in intralocus and also in interlocus comparisons, extending over distances as great as approximately 4 Mb. These results are not consistent with the higher genetic exchange between chromosomal arrangements expected in the central part of an inversion from double-crossover events. Hence, double crossovers were not produced or, alternatively, recombinant chromosomes were eliminated by natural selection to maintain coadapted gene complexes. If the strong genetic differentiation detected along O(3) extends to other inversions, nucleotide variation would be highly structured not only in D. subobscura, but also in the genome of other species with a rich chromosomal polymorphism.

Animals↗

Structural variation in the Waxy gene and differentiation in foxtail millet [Setaria italica (L.) P. Beauv.]: implications for multiple origins of the waxy phenotype.

The origin and evolution of the waxy type of foxtail millet [Setaria italica (L.) P. Beauv] were studied by analyzing structural variation in the Waxy gene. Initially, the Waxy gene was amplified by RT-PCR, RACE and genomic PCR from a non-waxy strain to determine the structure of the wild-type gene. Secondly, we screened by PCR for polymorphisms at the Waxy locus in 79 strains with various waxy phenotypes. We then carried out genomic Southern analysis on 67 strains and identified seven RFLP classes which were designated as types I-VII. RFLP type was correlated with phenotype, such that types I and II corresponded to non-waxy, types III and VI to low-amylose, and types IV, V and VII to waxy phenotypes. The differences between RFLP types could be attributed to insertions in the Waxy gene. Types II and VI were caused by the insertion of a Tourist element into intron 1 and a SINE-like sequence into intron 12, respectively. Types III, IV, V and VII were characterized by the insertion of large sequences into the Waxy gene that may alter the expression of the gene. Thus, multiple, independent insertions in the Waxy gene appear to have caused the loss-of-function waxy phenotypes. Furthermore, the geographical distributions of the three RFLP types associated with the waxy phenotype (types IV, V and VII) were distinct, with type IV being found mainly in Taiwan and Japan, type V in Korea, and type VII in Myanmar. These results indicate a polyphyletic origin for the waxy phenotype in landraces of foxtail millet.

Base Sequence↗

A family of retrotransposons and associated genomic variation in wheat.

A family of related retroelements was characterized in the genomes of some Graminease species. The structure of these retroelements indicates that they are retrotransposons containing reading frames with sequence similarity to the polyproteins of copia and Ty. This family of retroelements (termed WIS-2) occurs in the genomes of barley, wheat, rye, oats, and Aegilops species. Ongoing genomic variation both within individual plants of a wheat variety and within and between varieties of wheat is associated with some members of the WIS-2 family.

Amino Acid Sequence↗

Subpathways of nucleotide excision repair and their regulation.

Nucleotide excision repair provides an important cellular defense against a large variety of structurally unrelated DNA alterations. Most of these alterations, if unrepaired, may contribute to mutagenesis, oncogenesis, and developmental abnormalities, as well as cellular lethality. There are two subpathways of nucleotide excision repair; global genomic repair (GGR) and transcription coupled repair (TCR), that is selective for the transcribed DNA strand in expressed genes. Some of the proteins involved in the recognition of DNA damage (including RNA polymerase) are also responsive to natural variations in the secondary structural features of DNA. Gratuitous repair events in undamaged DNA might then contribute to genomic instability. However, damage recognition enzymes for GGR are normally maintained at very low levels unless the cells are genomically stressed. GGR is controlled through the SOS stress response in E. coli and through the activated p53 tumor suppressor in human cells. These inducible responses in human cells are important, as they have been shown to operate upon chemical carcinogen DNA damage at levels to which humans are environmentally exposed. Interestingly, most rodent tissues are deficient in the p53-dependent GGR pathway. Since rodents are used as surrogates for environmental cancer risk assessment, it is essential that we understand how they differ from humans with respect to DNA repair and oncogenic responses to environmental genotoxins. In the case of terminally differentiated mammalian cells, a new paradigm has appeared in which GGR is attenuated but both strands of expressed genes are repaired efficiently.

Animals↗

Distribution and sequence analysis of a family of type ill-dependent effectors correlate with the phylogeny of Ralstonia solanacearum strains.

In Ralstonia solanacearum, we previously have reported on the characterization of popP1 and popP2 genes. These genes encode type III-dependent pathogenicity effectors related to the large family of AvrRxv/YopJ cysteine proteases that are shared among pathogens of plants and animals. In this study, we identify a third gene, named popP3, that is inactivated in the genome sequence of strain GMI1000 by insertion of a copy of the insertion sequence ISRso13. The three popP genes are localized on two large chromosomal pathogenicity islands, with popP1 and popP2 being present on the same island. Phylogenic analysis demonstrated that the PopP2 and PopP3 proteins are clearly distinct from other effectors of this family previously characterized in plant and animal pathogens. Analysis of the distribution and allelic variations of the three genes in 30 strains representative of the biodiversity of R. solanacearum established that popP genes are distributed widely among strains from two of the three phyla previously defined on the basis of the structure of the core genome. Sequencing of the popP genes from the different strains revealed limited allelic variations at the three loci but did not show evidence of recombination between the popP genes. Limited allelic variation together with occurrence of insertion sequences within or in the close vicinity of popP genes and the presence of gene duplications in these pathogenicity islands suggest that genomic rearrangements might be a major evolutionary driving force controlling evolution of the genes encoded in these regions. The implications of these observations in terms of bacterial evolution, gene acquisition, and horizontal gene transfers are discussed.

Alleles↗