Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Recombination drives the evolution of GC-content in the human genome.

Unraveling the evolutionary forces responsible for variations of neutral substitution patterns among taxa or along genomes is a major issue in the identification of functional sequence features. Mammalian genomes show large-scale regional variations of GC-content (the isochores), but the substitution processes at the origin of this structure are poorly understood. We have analyzed the pattern of neutral substitutions in 14.3 Mb of primate noncoding regions. We show that the GC-content toward which sequences are evolving is strongly correlated (r(2) = 0.61, P </= 2 10(-16)) with the rate of crossovers (notably in females). This demonstrates that recombination drives the evolution of base composition in human (probably via the process of biased gene conversion). The present substitution patterns are very different from what they had been in the past, resulting in a major modification of the isochore structure of our genome. This non-equilibrium situation suggests that changes of recombination rates occur relatively frequently during evolution, possibly as a consequence of karyotype rearrangements. These results have important implications for understanding the spatial and temporal variations of substitution processes in a broad range of sexual organisms, and for detecting the hallmarks of natural selection in DNA sequences.

Animals↗

Population history rather than tree age contributes to the evolutionary importance of ancient trees in an endangered conifer.

Ancient trees are in global decline and face increasing conservation challenges. Their exceptional longevity has fostered the view that they are genetic reservoirs, yet whether old age is synonymous with unique genetic variation remains unclear. Here we assembled a ~8-Gb chromosome-level reference genome for the critically endangered conifer Glyptostrobus pensilis, now largely restricted to southern China with scattered populations in Vietnam and Laos, and resequenced 147 individuals, including 64 ancient (>100&#x2009;years old and persisting in human-dominated landscapes), 33 wild and 50 recently cultivated individuals. Ancient individuals comprised both likely natural relics and historically introduced individuals and formed two deeply divergent lineages and one ancestral-admixed group, each with distinct demographic histories of prolonged contraction and genomic erosion. Lineage identity explained more variation in genome-wide diversity, inbreeding and genetic load than the three conservation types, despite broad differences in age structure. Rare-allele analyses revealed pronounced heterogeneity among ancient trees: only relic and ancestral-origin individuals from high-diversity lineages contributed substantial unique variation, much of which is poorly represented in wild and cultivated populations. Together, our findings suggest that ancient trees are not uniformly genetically irreplaceable and that, at least in this conifer, evolutionary importance is shaped more strongly by population history than by age alone.

Endangered Species↗

Human calcium/calmodulin-dependent protein kinase II gamma gene (CAMK2G): cloning, genomic structure and detection of variants in subjects with type II diabetes.

AIMS/HYPOTHESIS: Ca(2+)/calmodulin-dependent protein kinase II, is expressed in the pancreatic beta cells and is activated by glucose and other secretagogues in a manner correlating with insulin secretion. The activation of Ca(2+)/calmodulin-dependent protein kinase II mediates some of the actions of Ca(2+) on the exocytosis of insulin. We therefore investigated the gene encoding the gamma isoform ( CAMK2G) which has been shown to be expressed in human beta cells as a candidate gene for Type II (non-insulin-dependent) diabetes mellitus. METHODS: Human CAMK2G was cloned from a total human P1 artificial chromosome library using a partial Ca(2+)/calmodulin-dependent protein kinase gamma(E) cDNA probe. Positive PAC clones were localised to chromosome 10q22 by fluorescence in situ hybridisation. To obtain structural information and the sequences of the exon-intron boundaries, the published genomic structures of the rat and mouse genes allowed the putative exon-intron boundaries of human CAMK2G to be amplified by vectorette polymerase chain reaction and sequenced. Sequence variants in each exon were identified using single stranded conformational polymorphism analysis. RESULTS: The human CAMK2G gene comprises 22 exons which range in size between 43 to 230 bp. Screening of the exons and exon-intron boundaries identified two single nucleotide polymorphisms. These did not show association with diabetes in 122 patients and 144 control subjects. CONCLUSIONS/INTERPRETATION: We have identified the genomic structure of CAMK2G to enable further study of this potential candidate gene. Variation in this gene is not strongly associated with diabetes in Caucasians in the United Kingdom. We have identified two single nucleotide polymorphisms which, with appropriately large case control studies, can be used to assess the role of CAMK2G in the susceptibility to Type II diabetes.

Alleles↗

Complete nucleotide sequences of all three poliovirus serotype genomes. Implication for genetic relationship, gene function and antigenic determinants.

The complete nucleotide sequences of the genomes of the type 2 ( P712 , Ch, 2ab ) and type 3 (Leon 12a1b ) poliovirus vaccine strains were determined. Comparison of the sequences with the previously established genome sequence of type 1 (LS-c, 2ab ) poliovirus vaccine strain revealed that 71% of the nucleotides in the genome RNAs were common, that the 5' and 3' termini of the genomes were highly homologous, and that more than 80% of the nucleotide differences in the coding region occurred in the third letter position of in-phase codons, resulting in a low frequency of amino acid difference. These results strongly suggested that the serotypes of poliovirus derived from a common prototype. A comparison of the amino acid sequences predicted from the genome sequences showed highest variation in the capsid protein region, whereas non-structural proteins are highly conserved. Initiation of polyprotein synthesis occurs in all three strains more than 740 nucleotides downstream from the 5' end. An analysis of the non-coding region suggests that small peptides that could potentially originate from this region are conserved. The amino acid sequences immediately surrounding the cleavage signals, however, show a higher than average degree of variation. The analysis of the amino acid sequences of the capsid protein VP1 of all serotypes has led to the prediction of potential antigenic sites on the virion involved in neutralization.

Amino Acid Sequence↗

S100 proteins in mouse and man: from evolution to function and pathology (including an update of the nomenclature).

The S100 protein family is the largest subgroup within the superfamily of proteins carrying the Ca2+-binding EF-hand motif. Despite their small molecular size and their conserved functional domain of two distinct EF-hands, S100 proteins developed a plethora of tissue-specific intra- and extracellular functions. Accordingly, various diseases such as cardiomyopathies, neurodegenerative and inflammatory disorders, and cancer are associated with altered S100 protein levels. Here, we review the different S100 protein functions and related diseases from an evolutionary point of view. We analyzed the structural variations, which are the basis of functional diversification, as well as the genomic organization of the S100 family in human and compared it with the S100 repertoires in mouse and rat. S100 genes and proteins are highly conserved between the different mammalian species. Moreover, we identified evolutionary related subgroups of S100 proteins within the three species, which share functional similarity and form subclusters on the genomic level. The available S100-specific mouse models are summarized and the consequences of our results are discussed with regard to the use of genetically engineered mice as human disease models. An update of the S100 nomenclature is included, because some of the recently identified S100 genes and pseudogenes had to be renamed.

Amino Acid Sequence↗

Prediction of functional sites in proteins using conserved functional group analysis.

A detailed knowledge of a protein's functional site is an absolute prerequisite for understanding its mode of action at the molecular level. However, the rapid pace at which sequence and structural information is being accumulated for proteins greatly exceeds our ability to determine their biochemical roles experimentally. As a result, computational methods are required which allow for the efficient processing of the evolutionary information contained in this wealth of data, in particular that related to the nature and location of functionally important sites and residues. The method presented here, referred to as conserved functional group (CFG) analysis, relies on a simplified representation of the chemical groups found in amino acid side-chains to identify functional sites from a single protein structure and a number of its sequence homologues. We show that CFG analysis can fully or partially predict the location of functional sites in approximately 96% of the 470 cases tested and that, unlike other methods available, it is able to tolerate wide variations in sequence identity. In addition, we discuss its potential in a structural genomics context, where automation, scalability and efficiency are critical, and an increasing number of protein structures are determined with no prior knowledge of function. This is exemplified by our analysis of the hypothetical protein Ydde_Ecoli, whose structure was recently solved by members of the North East Structural Genomics consortium. Although the proposed active site for this protein needs to be validated experimentally, this example illustrates the scope of CFG analysis as a general tool for the identification of residues likely to play an important role in a protein's biochemical function. Thus, our method offers a convenient solution to rapidly and automatically process the vast amounts of data that are beginning to emerge from structural genomics projects.

Algorithms↗

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article↗

Distribution, expression, and motif variability of ankyrin domain genes in Wolbachia pipientis.

The endosymbiotic bacterium Wolbachia pipientis infects a wide range of arthropods, in which it induces a variety of reproductive phenotypes, including cytoplasmic incompatibility (CI), parthenogenesis, male killing, and reversal of genetic sex determination. The recent sequencing and annotation of the first Wolbachia genome revealed an unusually high number of genes encoding ankyrin domain (ANK) repeats. These ANK genes are likely to be important in mediating the Wolbachia-host interaction. In this work we determined the distribution and expression of the different ANK genes found in the sequenced Wolbachia wMel genome in nine Wolbachia strains that induce different phenotypic effects in their hosts. A comparison of the ANK genes of wMel and the non-CI-inducing wAu Wolbachia strain revealed significant differences between the strains. This was reflected in sequence variability in shared genes that could result in alterations in the encoded proteins, such as motif deletions, amino acid insertions, and in some cases disruptions due to insertion of transposable elements and premature stops. In addition, one wMel ANK gene, which is part of an operon, was absent in the wAu genome. These variations are likely to affect the affinity, function, and cellular location of the predicted proteins encoded by these genes.

Ankyrins↗

Genetic and biochemical factors associated with variation in blood pressure in a genetic isolate.

We previously found an association between blood pressure and genetic variation of angiotensinogen in Canadian Hutterites. We hypothesized that variation in other candidate genes would also be associated with variation in blood pressure. We included genotypes of 12 candidate genes, along with clinical features and biochemical variables as covariates in an association analysis. We found that sex and body mass were significantly associated with variation in both systolic and diastolic blood pressures. We found that genotypes of APOB codon 4154 and AGT codon 174 were significantly associated with variation in systolic blood pressure. We found that genotypes of APOB codon 4154, AGT codon 174, and F7 codon 353 were significantly associated with variation in diastolic blood pressure. We found a significant association between age and variation in systolic but not diastolic blood pressure. We found a significant association between plasma apo B concentration and variation in diastolic but not systolic blood pressure. The association of genomic variation with resting blood pressure is consistent with the existence of important structural elements within or proximal to some genes in lipoprotein metabolism, the renin-angiotensin system, and the coagulation cascade. The association between plasma apo B concentration and diastolic blood pressure suggests that these traits may share some determinants.

Angiotensinogen↗

Pharmacogenomics and its potential impact on drug and formulation development.

Recent advances in genomic research have provided the basis for new insights into the importance of genetic and genomic markers during the different stages of drug development. A new field of research, pharmacogenomics, which studies the relationship between drug effects and the genome, has emerged. Structural pharmacogenomics maps the complete DNA sequences of whole genomes (genotypes) including individual variations, and functional pharmacogenomics assesses the expression levels of thousands of genes in one single experiment. Together, these two areas of pharmacogenomics have generated massive databases, which have become a challenge for the research field of informatics and have fostered a new branch of research, bioinformatics. If skillfully used, the databases generated by pharmacogenomics together with data mining on the Web promise to improve the drug development process in a variety of areas: identification of drug targets, evaluation of toxicity, classification of diseases, evaluation of formulations, assessment of drug response and treatment, post-marketing applications, and development of personalized medicines.

Animals↗

Characterization of phi 12, a bacteriophage related to phi 6: nucleotide sequence of the small and middle double-stranded RNA.

The isolation of additional bacteriophages containing segmented double-stranded RNA genomes has expanded the Cystoviridae family to nine members. Comparing the genomic sequences of these viruses has allowed evaluation of important genetic as well as structural motifs. These comparative studies are resulting in greater understanding of viral evolution and the role played by genetic and structural variation in the assembly mechanisms of the cystoviruses. In this regard, the small and middle double-stranded RNA genomic segments of bacteriophage phi 12 were copied as cDNA and their nucleotide sequences determined. This genome's organization is similar to that of the small and middle segments of bacteriophages phi 6, phi 8, and phi 13. Although there is little similarity in the nucleotide sequences, similarity exists in the amino acid sequence of the lysis cassette proteins to those of phi 6. The host cell attachment proteins are found to have marked similarity to the phi 13 attachment proteins.

Bacteriophage phi 6↗

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue C&#x3b1;-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins↗

Genomic structure of human anion exchanger 3 and its potential role in hereditary neurological disease.

Alterations in ion channel permeability or selectivity have been shown to cause neurological defects in humans. Anion exchanger isoform 3 (AE3) is prominently expressed in the brain and performs an electroneutral exchange of chloride and bicarbonate ions. In order to study the potential role of AE3 in human neurological disease, we characterized AE3 genomic structure and performed mutational analysis on patients with an episodic movement disorder that maps to the same genetic locus. AE3 genomic organization, including the nucleotide sequence of the 5'-untranslated region and intron/ exon boundaries, is highly conserved between humans and homologs from mouse and rat. Mutational analysis revealed no disease-causing defect in patients with familial paroxysmal dyskinesia, although several benign polymorphisms were identified. AE3 variation may prove useful for further genetic studies, such as finer resolution mapping. Characterization of genomic structure will facilitate mutational analysis of AE3 in studies of neurological diseases mapped to the same locus.

5' Untranslated Regions↗

[Characterization of 5S rRNA gene sequence and secondary structure in gymnosperms].

In higher plants the primary and the secondary structures of 5S ribosomal RNA gene are considered highly conservative. Little is known about the 5S rRNA gene structure, organization and variation in gyimnosperms. In this study we analyzed sequence and structure variation of 5S rRNA gene in Pinus through cloning and sequencing multiple copies of 5S rDNA repeats from individual trees of five pines, P. bungeana, P. tabulaeformis, P. yunnanensis, P. massoniana and P. densata. Pinus bungeana is from the subgenus Strobus while the other four are from the subgenus Pinus (diploxylon pines). Our results revealed variations in both primary and secondary structure among copies of 5S rDNA within individual genomes and between species. 5S rRNA gene in Pinus is 120 bp long in most of the 122 clones we sequenced except for one or two deletions in three clones. Among these clones 50 unique sequences were identified and they were shared by different pine species. Our sequences were compared to 13 sequences each representing a different gymnosperm species, and to six sequences representing both angiosperm monocots and dicots. Average sequence similarity was 97.1% among Pinus species and 94.3% between Pinus and other gymnosperms. Between gymnosperms and angiosperms the sequence similarity decreased to 88.1%. Similar to other molecular data, significant sequence divergence was found between the two Pinus subgenera. The 5S gene tree (neighbor-joining tree) grouped the four diploxylon pines together and separated them distinctly from P. bungeana. Comparison of sequence divergence within individuals and between species suggested that concerted evolution has been very weak especially after the divergence of the four diploxylon pines. The phylogenetic information contained in the 5S rRNA gene is limited due to its shorter length and the difficulties in identifying orthologous and paralogous copies of rDNA multigene family further complicate its phylogenetic application. Pinus densata is a diploid hybrid between P. tabulaeformis and P. yunnanensis. Its 5S rDNA composition is consistent with its hybrid origin. 5S rRNA of all gymnosperms published so far could be folded into a general secondary structure. Variation in this secondary structure was detected among species. About 55% of the 120 bp nucleotide positions was variable, in which 68% was on stem regions. Nevertheless, the positions at the end of the stems and those adjacent to loops are conserved. Their stability directly determines the size of the loops. Some mutations such as compensatory base-pair substitutions, and G-U pairing could be regarded as mechanisms for maintaining a stable secondary structure. The loops of the secondary structure are also relatively conserved. It seems that stable helices are necessary for the function of the gene. The conserved nucleotides in the loops are probably involved in the interaction with proteins and/or RNAs or with other nucleotide in the formation of the tertiary structure. However, unlike other reports, Loop E was found quite mutable among pines. These variations together with those on stems might be caused by the presence of pseudogenes among our clones. A preliminary evaluation indicates that only seven of 50 unique sequences are potentially functional genes.

Base Sequence↗

Optical genome mapping improves clinical interpretation of constitutional copy-number gains and reduces their VUS burden.

PURPOSE: Genomic structure of copy-number gains is critical for their clinical interpretation but cannot be determined by chromosomal microarray (CMA) analysis, which does not provide information about chromosomal location and orientation of multiplied regions. We thus hypothesized that in CMA testing gains have higher probability than losses to be classified as variants of uncertain significance (VUS) and that structural information from optical genome mapping (OGM) may improve their interpretation. METHODS: Using a &#x3c7;2 test, we assessed the association between classification of copy-number variants as VUS and their type (gains vs losses) in a cohort of 4073 CMA cases. Thirty-three VUS gains involving disease-associated genes were characterized by OGM to evaluate if OGM data enable their more conclusive clinical interpretation. RESULTS: The proportion of variants reported as VUS compared with likely pathogenic/pathogenic was significantly higher for gains than losses, confirming their increased VUS burden. OGM successfully determined genomic structure for all 33 copy-number gains, showing that 26 of 33 were tandem duplications and 7 of 33 were complex rearrangements. Structural information facilitated clinical interpretation in majority of the cases; it supported benign nature for 27 of 33 gains and was inconclusive or supported pathogenic role for 6 of 33. An estimated 20% of reported VUS gains would not have been reportable if we had OGM data. CONCLUSION: We illustrate a specific advantage of OGM compared with CMA: in addition to detecting both copy-number variants and balanced rearrangements, OGM improves clinical interpretation of copy-number gains by providing structural information and is thus expected to significantly decrease their VUS burden.

Humans↗

Sequence variation and evolution of nuclear DNA in man and the primates.

Recent advances in nucleic acid technology have facilitated the detection and detailed structural analysis of a wide variety of genes in higher organisms, including those in man. This in turn has opened the way to an examination of the evolution of structural genes and their surrounding and intervening sequences. In a study of the evolution of haemoglobin genes and neighbouring sequences in man and the primates, we have investigated gene arrangement and DNA sequence divergence both within and between species ranging from Old World monkeys to man. This analysis is beginning to reveal the evolutionary constraints that have acted on this region of the genome during primate evolution. Furthermore, DNA sequence variation, both within and between species, provides, in principle, a novel and powerful method for determining interspecific phylogenetic distances and also for analysing the structure of present-day human populations. Application of this new branch of molecular biology to other areas of the human genome should prove important in unravelling the history of genetic changes that have occurred during the evolution of man.

Animals↗

Perspectives on human genetic variation from the HapMap Project.

The completion of the International HapMap Project marks the start of a new phase in human genetics. The aim of the project was to provide a resource that facilitates the design of efficient genome-wide association studies, through characterising patterns of genetic variation and linkage disequilibrium in a sample of 270 individuals across four geographical populations. In total, over one million SNPs have been typed across these genomes, providing an unprecedented view of human genetic diversity. In this review we focus on what the HapMap Project has taught us about the structure of human genetic variation and the fundamental molecular and evolutionary processes that shape it.

Alleles↗