Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Characterization of phi 12, a bacteriophage related to phi 6: nucleotide sequence of the small and middle double-stranded RNA.

The isolation of additional bacteriophages containing segmented double-stranded RNA genomes has expanded the Cystoviridae family to nine members. Comparing the genomic sequences of these viruses has allowed evaluation of important genetic as well as structural motifs. These comparative studies are resulting in greater understanding of viral evolution and the role played by genetic and structural variation in the assembly mechanisms of the cystoviruses. In this regard, the small and middle double-stranded RNA genomic segments of bacteriophage phi 12 were copied as cDNA and their nucleotide sequences determined. This genome's organization is similar to that of the small and middle segments of bacteriophages phi 6, phi 8, and phi 13. Although there is little similarity in the nucleotide sequences, similarity exists in the amino acid sequence of the lysis cassette proteins to those of phi 6. The host cell attachment proteins are found to have marked similarity to the phi 13 attachment proteins.

Bacteriophage phi 6↗

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins↗

Genomic structure of human anion exchanger 3 and its potential role in hereditary neurological disease.

Alterations in ion channel permeability or selectivity have been shown to cause neurological defects in humans. Anion exchanger isoform 3 (AE3) is prominently expressed in the brain and performs an electroneutral exchange of chloride and bicarbonate ions. In order to study the potential role of AE3 in human neurological disease, we characterized AE3 genomic structure and performed mutational analysis on patients with an episodic movement disorder that maps to the same genetic locus. AE3 genomic organization, including the nucleotide sequence of the 5'-untranslated region and intron/ exon boundaries, is highly conserved between humans and homologs from mouse and rat. Mutational analysis revealed no disease-causing defect in patients with familial paroxysmal dyskinesia, although several benign polymorphisms were identified. AE3 variation may prove useful for further genetic studies, such as finer resolution mapping. Characterization of genomic structure will facilitate mutational analysis of AE3 in studies of neurological diseases mapped to the same locus.

5' Untranslated Regions↗

[Characterization of 5S rRNA gene sequence and secondary structure in gymnosperms].

In higher plants the primary and the secondary structures of 5S ribosomal RNA gene are considered highly conservative. Little is known about the 5S rRNA gene structure, organization and variation in gyimnosperms. In this study we analyzed sequence and structure variation of 5S rRNA gene in Pinus through cloning and sequencing multiple copies of 5S rDNA repeats from individual trees of five pines, P. bungeana, P. tabulaeformis, P. yunnanensis, P. massoniana and P. densata. Pinus bungeana is from the subgenus Strobus while the other four are from the subgenus Pinus (diploxylon pines). Our results revealed variations in both primary and secondary structure among copies of 5S rDNA within individual genomes and between species. 5S rRNA gene in Pinus is 120 bp long in most of the 122 clones we sequenced except for one or two deletions in three clones. Among these clones 50 unique sequences were identified and they were shared by different pine species. Our sequences were compared to 13 sequences each representing a different gymnosperm species, and to six sequences representing both angiosperm monocots and dicots. Average sequence similarity was 97.1% among Pinus species and 94.3% between Pinus and other gymnosperms. Between gymnosperms and angiosperms the sequence similarity decreased to 88.1%. Similar to other molecular data, significant sequence divergence was found between the two Pinus subgenera. The 5S gene tree (neighbor-joining tree) grouped the four diploxylon pines together and separated them distinctly from P. bungeana. Comparison of sequence divergence within individuals and between species suggested that concerted evolution has been very weak especially after the divergence of the four diploxylon pines. The phylogenetic information contained in the 5S rRNA gene is limited due to its shorter length and the difficulties in identifying orthologous and paralogous copies of rDNA multigene family further complicate its phylogenetic application. Pinus densata is a diploid hybrid between P. tabulaeformis and P. yunnanensis. Its 5S rDNA composition is consistent with its hybrid origin. 5S rRNA of all gymnosperms published so far could be folded into a general secondary structure. Variation in this secondary structure was detected among species. About 55% of the 120 bp nucleotide positions was variable, in which 68% was on stem regions. Nevertheless, the positions at the end of the stems and those adjacent to loops are conserved. Their stability directly determines the size of the loops. Some mutations such as compensatory base-pair substitutions, and G-U pairing could be regarded as mechanisms for maintaining a stable secondary structure. The loops of the secondary structure are also relatively conserved. It seems that stable helices are necessary for the function of the gene. The conserved nucleotides in the loops are probably involved in the interaction with proteins and/or RNAs or with other nucleotide in the formation of the tertiary structure. However, unlike other reports, Loop E was found quite mutable among pines. These variations together with those on stems might be caused by the presence of pseudogenes among our clones. A preliminary evaluation indicates that only seven of 50 unique sequences are potentially functional genes.

Base Sequence↗

Optical genome mapping improves clinical interpretation of constitutional copy-number gains and reduces their VUS burden.

PURPOSE: Genomic structure of copy-number gains is critical for their clinical interpretation but cannot be determined by chromosomal microarray (CMA) analysis, which does not provide information about chromosomal location and orientation of multiplied regions. We thus hypothesized that in CMA testing gains have higher probability than losses to be classified as variants of uncertain significance (VUS) and that structural information from optical genome mapping (OGM) may improve their interpretation. METHODS: Using a χ2 test, we assessed the association between classification of copy-number variants as VUS and their type (gains vs losses) in a cohort of 4073 CMA cases. Thirty-three VUS gains involving disease-associated genes were characterized by OGM to evaluate if OGM data enable their more conclusive clinical interpretation. RESULTS: The proportion of variants reported as VUS compared with likely pathogenic/pathogenic was significantly higher for gains than losses, confirming their increased VUS burden. OGM successfully determined genomic structure for all 33 copy-number gains, showing that 26 of 33 were tandem duplications and 7 of 33 were complex rearrangements. Structural information facilitated clinical interpretation in majority of the cases; it supported benign nature for 27 of 33 gains and was inconclusive or supported pathogenic role for 6 of 33. An estimated 20% of reported VUS gains would not have been reportable if we had OGM data. CONCLUSION: We illustrate a specific advantage of OGM compared with CMA: in addition to detecting both copy-number variants and balanced rearrangements, OGM improves clinical interpretation of copy-number gains by providing structural information and is thus expected to significantly decrease their VUS burden.

Humans↗

Sequence variation and evolution of nuclear DNA in man and the primates.

Recent advances in nucleic acid technology have facilitated the detection and detailed structural analysis of a wide variety of genes in higher organisms, including those in man. This in turn has opened the way to an examination of the evolution of structural genes and their surrounding and intervening sequences. In a study of the evolution of haemoglobin genes and neighbouring sequences in man and the primates, we have investigated gene arrangement and DNA sequence divergence both within and between species ranging from Old World monkeys to man. This analysis is beginning to reveal the evolutionary constraints that have acted on this region of the genome during primate evolution. Furthermore, DNA sequence variation, both within and between species, provides, in principle, a novel and powerful method for determining interspecific phylogenetic distances and also for analysing the structure of present-day human populations. Application of this new branch of molecular biology to other areas of the human genome should prove important in unravelling the history of genetic changes that have occurred during the evolution of man.

Animals↗

Perspectives on human genetic variation from the HapMap Project.

The completion of the International HapMap Project marks the start of a new phase in human genetics. The aim of the project was to provide a resource that facilitates the design of efficient genome-wide association studies, through characterising patterns of genetic variation and linkage disequilibrium in a sample of 270 individuals across four geographical populations. In total, over one million SNPs have been typed across these genomes, providing an unprecedented view of human genetic diversity. In this review we focus on what the HapMap Project has taught us about the structure of human genetic variation and the fundamental molecular and evolutionary processes that shape it.

Alleles↗

Involvement of cross-genus phages in bacterial resistance to chlorine disinfection.

Chlorine disinfection resistance in pathogenic microorganisms poses severe environmental concerns and public health risks. While phages play critical roles in host adaptation to environmental stress, how poly-host phages contribute to bacterial resistance to chlorine disinfectants remains poorly understood. Here, we investigated shifts in the population dynamics, transcriptional profiles, and function potentials of cross-genus phage-bacterial communities under exposure to chlorine disinfectants in a continuously operated anaerobic-anoxic-oxic system over a 92-day period, using integrated metagenomic and metatranscriptomic approaches. In the presence and absence of chlorine disinfectants, the genomic abundance and diversity of phage and bacterial communities showed similar variation trends, and the community structures of both exhibited clear differences. A strong significant positive correlation was observed between phage and bacterial diversity under chlorine exposure (R&#x202f;=&#x202f;0.975, p&#x202f;=&#x202f;0.00,057), whereas no significant correlation was detected in the absence of chlorine disinfection (R&#x202f;=&#x202f;-0.314, p&#x202f;=&#x202f;0.613), suggesting that chlorine disinfectants may enhance phage-bacteria interactions. Host-associated phages exhibited high consistency with their corresponding putative hosts in terms of genomic abundance (M2&#x202f;=&#x202f;0.0945, p&#x202f;=&#x202f;0.001) and transcript abundance (M2&#x202f;=&#x202f;0.3668, p&#x202f;=&#x202f;0.001), and they were also significantly correlated with cross-genus phages in both genomic abundance (R&#x202f;=&#x202f;0.97, p&#x202f;<&#x202f;2.2e-16) and transcript abundance (R&#x202f;=&#x202f;0.83, p&#x202f;<&#x202f;2.2e-16), which collectively suggests the critical role of cross-genus phages in the resistance of microbial communities to chlorine disinfectants. Bipartite association network analysis shows that cross-genus phages carry highly homologous genes to their putative hosts and may be involved in the horizontal transfer of these genes among bacteria. These homologous genes are involved in DNA repair, redox balance regulation, environmental stress adaptation and efflux pump functions, suggesting a synergistic role between cross-genus phages and their putative hosts in chlorine resistance. Our findings reveal that cross-genus phages can contribute to the resistance of bacterial communities to chlorine disinfectants, providing the theoretical foundation for evaluating the role of poly-host phages in microbial communities.

Chlorine resistance↗

Capturing genomic signatures of DNA sequence variation using a standard anonymous microarray platform.

Comparative genomics, using the model organism approach, has provided powerful insights into the structure and evolution of whole genomes. Unfortunately, only a small fraction of Earth's biodiversity will have its genome sequenced in the foreseeable future. Most wild organisms have radically different life histories and evolutionary genomics than current model systems. A novel technique is needed to expand comparative genomics to a wider range of organisms. Here, we describe a novel approach using an anonymous DNA microarray platform that gathers genomic samples of sequence variation from any organism. Oligonucleotide probe sequences placed on a custom 44 K array were 25 bp long and designed using a simple set of criteria to maximize their complexity and dispersion in sequence probability space. Using whole genomic samples from three known genomes (mouse, rat and human) and one unknown (Gonystylus bancanus), we demonstrate and validate its power, reliability, transitivity and sensitivity. Using two separate statistical analyses, a large numbers of genomic 'indicator' probes were discovered. The construction of a genomic signature database based upon this technique would allow virtual comparisons and simple queries could generate optimal subsets of markers to be used in large-scale assays, using simple downstream techniques. Biologists from a wide range of fields, studying almost any organism, could efficiently perform genomic comparisons, at potentially any phylogenetic level after performing a small number of standardized DNA microarray hybridizations. Possibilities for refining and expanding the approach are discussed.

Animals↗

Genetic variation of NSP1 and NSP4 genes among serotype G9 rotaviruses causing hospitalization of children in Melbourne, Australia, 1997-2002.

Serotype G9 rotaviruses have emerged as one of the leading causes of gastroenteritis in children worldwide. We examined 29 representative G9 rotavirus isolates from a 6-year collection (1997-2002) and determined the level of variation in genes encoding non-structural proteins, NSP1 and NSP4. Northern hybridization analysis with a whole genome probe derived from the prototype G9 strain, F45, revealed that the NSP1 gene (gene 5) of two isolates (R1 and R14) did not exhibit significant homology. Complementary DNA probes of R1 and R14 genes 5 were used in Northern blot hybridization and indicated the presence of at least two gene 5 alleles among Melbourne G9 rotaviruses. Nucleotide sequence analysis revealed that isolates carrying the R14 gene 5 shared 94-98% sequence identities with one another, while sequence identity to R1 was 78%. Surprisingly, R1 displayed 96% nucleotide identity with the prototype serotype G1 strain, Wa. The detection of different alleles of NSP1 genes prompted us to investigate the level of variation in another non-structural protein, NSP4, a multifunctional protein and the first viral-encoded enterotoxin. Phylogenetic analysis indicated that while all isolates clustered into one group containing the Wa NSP4 allele (genotype 1), isolate R1 was most closely related to Wa. This study reveals new information about the diversity of non-structural proteins of G9 rotaviruses.

Amino Acid Sequence↗

Haplotype structure and population genetic inferences from nucleotide-sequence variation in human lipoprotein lipase.

Allelic variation in 9.7 kb of genomic DNA sequence from the human lipoprotein lipase gene (LPL) was scored in 71 healthy individuals (142 chromosomes) from three populations: African Americans (24) from Jackson, MS; Finns (24) from North Karelia, Finland; and non-Hispanic Whites (23) from Rochester, MN. The sequences had a total of 88 variable sites, with a nucleotide diversity (site-specific heterozygosity) of .002+/-.001 across this 9.7-kb region. The frequency spectrum of nucleotide variation exhibited a slight excess of heterozygosity, but, in general, the data fit expectations of the infinite-sites model of mutation and genetic drift. Allele-specific PCR helped resolve linkage phases, and a total of 88 distinct haplotypes were identified. For 1,410 (64%) of the 2,211 site pairs, all four possible gametes were present in these haplotypes, reflecting a rich history of past recombination. Despite the strong evidence for recombination, extensive linkage disequilibrium was observed. The number of haplotypes generally is much greater than the number expected under the infinite-sites model, but there was sufficient multisite linkage disequilibrium to reveal two major clades, which appear to be very old. Variation in this region of LPL may depart from the variation expected under a simple, neutral model, owing to complex historical patterns of population founding, drift, selection, and recombination. These data suggest that the design and interpretation of disease-association studies may not be as straightforward as often is assumed.

Animals↗

A 9.1-kb gap in the genome reference map is shown to be a stable deletion/insertion polymorphism of ancestral origin.

We show a mute 9.1-kb gap in the human genome reference map, unraveled by RDA studies, to be a worldwide deletion/insertion polymorphism of stable type. The molecular and population data presented suggest its origin from a unique ancestral transposition event in chromosomal region 22q11.2, overlapping the IglambdaV genes at about 450 kb from the cluster of the IglambdaJ-C genes. These findings are not meant to be just another report of a polymorphic marker suitable for population studies. Rather, we wish to stress that a large number of inborn mute gaps may be spread all over the genome and that the many RDA-detected microdeletions already available are efficient tools for the discovery of this otherwise hidden category of genetic variation. Apart from their possible impact on expression of structural genes, mute gaps must be filled for the reference map of our genome to be truly completed.

Chromosome Deletion↗

Genomics and the Human Genome Project: implications for psychiatry.

In the past decade the Human Genome Project has made extraordinary strides in understanding of fundamental human genetics. The complete human genetic sequence has been determined, and the chromosomal location of almost all human genes identified. Presently, a large international consortium, the HapMap Project, is working to identify a large portion of genetic variation in different human populations and the structure and relationship of these variants to each other. The Human Genome Project has approached human genetics on a scale not previously seen in biology. This has been made possible by dramatic advances in high throughput technology and bio-informatics. Tools such as gene chips and micro-arrays have spawned an entirely new strategy to examine the function and expression of genes in a massively parallel fashion. Together these tools have dramatically advanced our knowledge about the human genome. They promise powerful new approaches to complex genetic traits such as psychiatric illness. The goals and progress of the Human Genome Project and the technology involved are reviewed. The implications of this science for psychiatric genetics are discussed.

Computational Biology↗

Heterogeneity in rates of recombination in the 6-Mb region telomeric to the human major histocompatibility complex.

Analysis of 784 informative meioses in the CEPH pedigrees revealed a total of 22 recombination events having occurred in the 6-Mb region between D6S265 (70 kb centromeric of HLA-A) and D6S276. These 22 breakpoints were localized with respect to anonymous polymorphic markers, leading to a detailed genetic map of the region telomeric to the human major histocompatibility complex. A nonrandom pattern of recombination was observed throughout this region: the low recombination rate of 0.19% within the 4-Mb interval centromeric to the HLA class I-like candidate gene for hemochromatosis indeed contrasts with the approximate 1% rate observed within the most telomeric two megabases. This reduced rate of recombination may be due to selective constraints depending on environmental factors related to immunity and iron status or to structural variations hampering proper meiotic pairing of homologous sequences. Population data from other human genome segments are now needed to determine whether linkage disequilibrium extending over 4 Mb is unique to this region.

Chromosome Mapping↗

Identification of two distinct subfamilies of alpha satellite DNA that are highly specific for human chromosome 15.

We report the isolation of two distinct subfamilies of alpha satellite DNA (pTRA-20 and -25) from human chromosome 15. In situ hybridization experiments indicated that both subfamilies are highly specific for this chromosome. Southern analysis of a somatic hybrid cell line carrying human chromosome 15 revealed a likely higher-order genomic band of 2.5 kb for pTRA-20. Similar analysis for pTRA-25 showed multiple higher-order bands of 3.5, 4.5, and 5 kb at moderately high hybridization stringency, but a predominance of the 4.5-kb species at very high stringency. Direct comparison with human genomic DNA confirmed the authenticity of these higher-order structures and demonstrated polymorphic variations using both probes. The origin of the different alphoid subfamilies on chromosome 15 is discussed. These sequences should be useful for the construction of centromere-based genetic linkage maps for human chromosome 15 and, in conjunction with the other alphoid sequences already reported for chromosomes 13, 14, 21, and 22, should allow a concerted analysis of the evolution and the possible etiological role of these DNAs in aberrations commonly seen in these chromosomes.

Blotting, Southern↗

Coalescent processes and relaxation of selective constraints leading to contrasting genetic diversity at paralogs AtHVA22d and AtHVA22e in Arabidopsis thaliana.

Duplicate loci offer a very powerful system for understanding the complicated genome structure and adaptive evolution of a gene family. In this study, the genetic variation at paralogs AtHVA22d and AtHVA22e, members of an ABA- and stress-inducible gene family, is examined in the selfing Arabidopsis thaliana. Population genetic analysis indicates contrasting levels of nucleotide diversity at overall exon sequence and nonsynonymous sites between AtHVA22d (pi = 0.00337, pi(rep) = 0.00158) and AtHVA22e (pi = 0.00054, pi(rep) = 0.00023). The fact of Ka/Ks ratios significantly less than 1 in all sequences indicates that both genes are functional and subjected to purifying selection. In addition, rooted at barley HVA22, accelerated evolution is detected at replacement changes in the AtHVA22d locus, indicating relaxation of purifying selection after gene duplication. However, relative rate tests reveal no deviation from the neutrality at synonymous sites between the two paralogs. Based on clock-like evolution, the rate of synonymous substitution is estimated at 1.83 x 10(-9) substitutions per site per year; and the divergence of the two paralogs is traced to 90 MYA, coinciding with a period of the diversification of angiosperms. Given no codon usage bias in both genes, natural selection alone cannot account for the 6.4-fold differences in the nucleotide variation at synonymous sites between the two paralogs. Random processes resulting in different coalescence times, 3.65 MYA at AtHVA22d vs. 1.20 MYA at AtHVA22e, may have predominantly contributed to the evident differences of the genetic diversity. Partially nonoverlapping modes of expression between the two functional paralogs suggest a subfunctionalization hypothesis for explaining the fates of duplicate loci.

Arabidopsis↗

Genotypic characteristics of bovine viral diarrhea virus 2 strains isolated in northern Italy.

Two strains of Bovine viral diarrhea virus 2 (BVDV-2) were isolated from calves in northern Italy. Variations in the 5'-untranslated region (UTR) of the genome were studied by primary structure alignment and neighbor-joining method based phylogenetic tree analyses and by palindromic nucleotide substitutions at the three variable loci in the 5'-UTR. Genetic analysis indicated their appurtenance to genovar BVDV-2a. Nucleotide sequence at the 5'-UTR of strain BS-95-II, one of the Italian isolates from healthy calves, showed 98% homology to that of the Japanese isolate OY89, a cytopathic strain derived from cattle with mucosal disease.

5' Untranslated Regions↗

A single-nucleotide natural variation (U4 to C4) in an influenza A virus promoter exhibits a large structural change: implications for differential viral RNA synthesis by RNA-dependent RNA polymerase.

The influenza A virus promoter is recognized by the influenza A virus RNA-dependent RNA polymerase, and directs both transcription and replication of the viral RNA genome. Within the sequence of this promoter, flu strains exhibit a natural, unique variation, either a U or a C, at the fourth position from the 3' end. Promoters that contain a C residue (C4 promoter), which are invariably found in genome segments that encode the three RNA polymerase subunits (PB1, PB2 and PA), down-regulate transcription but activate genome replication. Here, we have determined the structure of the C4 promoter by NMR spectroscopy and compared it with the structure of the U4 promoter, which was determined previously. The structure of the internal loop in the C4 promoter is similar to that of the U4 promoter. However, the terminal stem of the C4 promoter is strikingly different from that of the U4 promoter. These structural data suggest that the internal loop is important for polymerase binding to the promoter, and the terminal stem is crucial for differential regulation of transcription and replication.

Base Sequence↗