Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Structure and functional genomics of lipopolysaccharide expression in Haemophilus influenzae.

The involvement of genes in the lic loci in H. influenzae LPS expression has been known for some time. However, it was not until recently that it was shown that the lic1 locus contains genes required for phase variable expression of phosphocholine substituents, while genes in the lic2 locus and lgtC are required for expression of the globoside trisaccharide, alpha-D-Galp-(1 --> 4)-beta-D-Galp-(1 --> 4)-beta-D-Glcp (i.e., the pK blood group epitope). The availability of the complete sequence of the H. influenzae strain Rd genome has facilitated significant progress in understanding the role of these and other genes in the expression and biosynthesis of LPS. We have employed a comparative structural fingerprinting strategy to establish the structural relationships among LPS from H. influenzae mutant strains in which putative biosynthesis genes were inactivated. Using this functional genomics approach, we have gained considerable insight into the genetic basis for intra-strain and strain-to-strain variation in epitope expression.

Base Sequence↗

DNA structure constraint is probably a fundamental factor inducing CpG deficiency in bacteria.

MOTIVATION: It has been speculated that CpG dinucleotide deficiency in genomes is a consequence of DNA methylation. However, this hypothesis does not adequately explain CpG deficiency in bacteria. The hypothesis based on DNA structure constraint as an alternative explanation was therefore examined. RESULTS: By comparing real bacterial genomes and Markov artificial genomes in the second order, we found that the core structure of a restricted pattern, the TTCGAA pattern, was under represented in low GC content bacterial genomes regardless of CpG dinucleotide level. This is in contrast to the AACGTT pattern, indicating that the counterselection is context-dependent. Further study discovered nine underrepresented patterns that were supposed to be capable of inducing DNA structure constraint. In summary, most of them are in TTCGNA and TTCGAN patterns in both DNA strands. An explanation is also proposed for the strong correlation between GC content and CpG deficiency. The result of random sequence simulation showed that the occurrences of these patterns were correlated with GC content, as well as the percentage of CpG dinucleotides being trapped in these patterns. Finally, we suggest that the degree of counter-selection against these restricted patterns could be influenced by global GC content of a genome.

Base Sequence↗

Evolution of the cetacean mitochondrial D-loop region.

We sequenced the mitochondrial DNA D-loop regions from two cetacean species and compared these with the published D-loop sequences of several other mammalian species, including one other cetacean. Nucleotide substitution rates, DNA sequence simplicity, possible open reading frames (ORFs), and potential RNA secondary structure were investigated. The substitution rate is an order of magnitude lower than would be expected on the basis of reports on human sequence variation in this region but are consistent with interspecific primate and rodent D-loop sequence variation and with estimates of substitution rates from whole mitochondrial genomes. Deletions/insertions are less common in the cetacean D-loop than in other vertebrate species. Areas of high sequence simplicity (clusters of short repetitive motifs) across the region correspond to areas of high sequence divergence. Three regions predicted to form secondary structures are homologous to such putative structures in other species; however, the presumptive structures most conserved in cetaceans are different from those reported for other taxa. While all three species have possible long ORFs, only a short sequence of seven amino acids is shared with other mammalian species, and those changes that had occurred within it are all nonsynonymous. We conclude that DNA slippage, in addition to point mutation, contributes to the evolution of the D-loop and that regions of conserved secondary structure in cetaceans and an ORF are unlikely to contribute significantly to the conservation of the central region.

Amino Acid Sequence↗

Complete sequence of the mitochondrial DNA of the annelid worm Lumbricus terrestris.

We have determined the complete nucleotide (nt) sequence of the mitochondrial genome of an oligochaete annelid, the earthworm Lumbricus terrestris. This genome contains the 37 genes typical of metazoan mitochondrial DNA (mtDNA), including ATPase8, which is missing from some invertebrate mtDNAs. ATPase8 is not immediately upstream of ATPase6, a condition found previously only in the mtDNA of snails. All genes are transcribed from the same DNA strand. The largest noncoding region is 384 nt and is characterized by several homopolymer runs, a tract of alternating TA pairs, and potential secondary structures. All protein-encoding genes either overlap the adjacent downstream gene or end at an abbreviated stop codon. In Lumbricus mitochondria, the variation of the genetic code that is typical of most invertebrate mitochondrial genomes is used. Only the codon ATG is used for translation initiation. Lumbricus mtDNA is A + T rich, which appears to affect the codon usage pattern. The DHU arm appears to be unpaired not only in tRNAser(AGN), as is typical for metazoans, but perhaps also in tRNAser(UCN), a condition found previously only in a chiton and among nematodes. Relating the Lumbricus gene organization to those of other major protostome groups requires numerous rearrangements.

Amino Acid Sequence↗

Constrained genomic and conformational variability of the hypervariable region 1 of hepatitis C virus in chronically infected patients.

We analysed the genomic and conformational variability of the hypervariable region 1 (HVR1) of the hepatitis C virus (HCV) to evaluate the importance of its biological role. A total of 865 genotype 1b HVR1 subclones were collected from serially sampled sera in 11 patients with chronic hepatitis C, four of whom received interferon therapy. Consequently, 169 distinct sequences were examined for amino acid substitutions as well as hydrophilic or hydrophobic profile at each amino acid position within HVR1. Secondary structure of HVR1 was also predicted by the method of Robson in 90 distinct sequences from eight patients, including three interferon-treated patients. Some positions within the HVR1 were invariable or nearly so as to amino acid substitution. Hydrophilic or hydrophobic residues exclusively predominated at several positions. These constrained amino acid replacement and hydrophilic or hydrophobic profiles were conserved irrespective of interferon therapy, though the frequency of amino acid replacement was greater at almost all amino acid positions within the HVR1 in interferon-treated patients. The quasispecies of HCV showed various secondary structures of HVR1, but many sequences seemed to have common characteristics. beta sheet conformations around both the N-terminus and position 20 (numbered from the NH2 terminus of E2 envelope glycoprotein), and/or coil structures around the C-terminus of HVR1 could be identified. These results suggest that HVR1 amino acid replacements are strongly constrained by a well-ordered structure, in spite of being tolerant to amino acid substitutions, and imply an important biological role of the HVR1 protein in HCV replication.

Adult↗

The rat STSL locus: characterization, chromosomal assignment, and genetic variations in sitosterolemic hypertensive rats.

BACKGROUND: Elevated plant sterol accumulation has been reported in the spontaneously hypertensive rat (SHR), the stroke-prone spontaneously hypertensive rat (SHRSP) and the Wistar-Kyoto (WKY) rat. Additionally, a blood pressure quantitative trait locus (QTL) has been mapped to rat chromosome 6 in a New Zealand genetically hypertensive rat strain (GH rat). ABCG5 and ABCG8 (encoding sterolin-1 and sterolin-2 respectively) have been shown to be responsible for causing sitosterolemia in humans. These genes are organized in a head-to-head configuration at the STSL locus on human chromosome 2p21. METHODS: To investigate whether mutations in Abcg5 or Abcg8 exist in SHR, SHRSP, WKY and GH rats, we initiated a systematic search for the genetic variation in coding and non-coding region of Abcg5 and Abcg8 genes in these strains. We isolated the rat cDNAs for these genes and characterized the genomic structure and tissue expression patterns, using standard molecular biology techniques and FISH for chromosomal assignments. RESULTS: Both rat Abcg5 and Abcg8 genes map to chromosome band 6q12. These genes span ~40 kb and contain 13 exons and 12 introns each, in a pattern identical to that of the STSL loci in mouse and man. Both Abcg5 and Abcg8 were expressed only in liver and intestine. Analyses of DNA from SHR, SHRSP, GH, WKY, Wistar, Wistar King A (WKA) and Brown Norway (BN) rat strains revealed a homozygous G to T substitution at nucleotide 1754, resulting in the coding change Gly583Cys in sterolin-1 only in rats that are both sitosterolemic and hypertensive (SHR, SHRSP and WKY). CONCLUSIONS: The rat STSL locus maps to chromosome 6q12. A non-synonymous mutation in Abcg5, Gly583Cys, results in sitosterolemia in rat strains that are also hypertensive (WKY, SHR and SHRSP). Those rat strains that are hypertensive, but not sitosterolemic (e.g. GH rat) do not have mutations in Abcg5 or Abcg8. This mutation allows for expression and apparent apical targeting of Abcg5 protein in the intestine. These rat strains may therefore allow us to study the pathophysiological mechanisms involved in the human disease of sitosterolemia.

Animals↗

Sweeps in Space: Leveraging Geographic Data to Identify Beneficial Alleles in Anopheles gambiae.

As organisms adapt to environmental changes, natural selection modifies the frequency of nonneutral alleles. For beneficial mutations, the outcome of this process may be a selective sweep, in which an allele rapidly increases in frequency and perhaps reaches fixation within a population. Selective sweeps have well-studied effects on patterns of local genetic variation in panmictic populations, but much less is known about the dynamics of sweeps in continuous space. In particular, because limited movement across a landscape leads to unique patterns of population structure, spatial dynamics may influence the trajectory of selected mutations. Here, we use forward-in-time, individual-based simulations in continuous space to study the impact of space on beneficial mutations as they sweep through a population. In particular, we show that selection changes the joint distribution of allele frequency and geographic range occupied by a focal allele and demonstrate that this signal can be used to identify selective sweeps. We then leverage this signal to identify in-progress selective sweeps within the malaria vector Anopheles gambiae, a species under strong selection pressure from vector control measures. By considering space, we identify multiple previously undescribed variants with potential phenotypic consequences, including mutations impacting known IR-associated genes and altering protein structure and properties. Our results demonstrate a novel signal for detecting selection in spatial population genetic data that may have implications for genomic surveillance and understanding geographic patterns of genetic variation.

Animals↗

Sweeps in space: leveraging geographic data to identify beneficial alleles in Anopheles gambiae.

As organisms adapt to environmental changes, natural selection modifies the frequency of non-neutral alleles. For beneficial mutations, the outcome of this process may be a selective sweep, in which an allele rapidly increases in frequency and perhaps reaches fixation within a population. Selective sweeps have well-studied effects on patterns of local genetic variation in panmictic populations, but much less is known about the dynamics of sweeps in continuous space. In particular, because limited movement across a landscape leads to unique patterns of population structure, spatial dynamics may influence the trajectory of selected mutations. Here, we use forward-in-time, individual-based simulations in continuous space to study the impact of space on beneficial mutations as they sweep through a population. In particular, we show that selection changes the joint distribution of allele frequency and geographic range occupied by a focal allele and demonstrate that this signal can be used to identify selective sweeps. We then leverage this signal to identify in-progress selective sweeps within the malaria vector Anopheles gambiae , a species under strong selection pressure from vector control measures. By considering space, we identify multiple previously undescribed variants with potential phenotypic consequences, including mutations impacting known IR-associated genes and altering protein structure and properties. Our results demonstrate a novel signal for detecting selection in spatial population genetic data that may have implications for genomic surveillance and understanding geographic patterns of genetic variation.

Journal Article↗

Differences in the organization and methylation patterns of integrated avian sarcoma proviral DNA sequences in nonpermissive and permissive mammalian cells.

Four types of avian sarcoma virus (ASV)-transformed mammalian cells were analyzed for the presence of ASV-specific sequences in their genome DNA. A great variability in the number of proviral copies and their structure within the DNA of these lines was observed. In all cells tested gag and src sequences were present in a flexible arrangement. The greatest variation was detected within proviral sequences corresponding to pol and env regions. The number of integrated ASV proviral copies do not correlate with the capability of these cells to produce viral particles. Virus-producing (K2S and K12) and virogenic (XC) cells contain in their chromosomal DNA at least one complete proviral genome, whereas proviral sequences in helper-dependent nonvirogenic cells are substantially changed. Provirus expression level does not correlate with the number of integrated virus copies. In the non-virus-producing cells the proviral sequences are hypermethylated.

Animals↗

Pangenome of Streptomyces sampsonii and Relatives Highlights Horizontal Gene Transfer and Secondary Metabolism in Environmental Adaptation and Ecological Significance.

Streptomyces sampsonii is a promising biocontrol bacterium, but its genomic basis of adaptation and secondary metabolism remains unclear. Here, we present a chromosome-level genome assembly of S. sampsonii (7.20 Mb, 6015 protein-coding genes) and perform comparative analyses with 95 related Streptomyces species. Phylogenomic and synteny analyses revealed its closest relationship with S. albidoflavus, while extensive structural variations distinguished more distant lineages. Pangenome analysis uncovered 84,178 gene clusters, with pan_shell and pan_cloud genes predominantly enriched in xenobiotic biodegradation, metabolism, and antibiotic biosynthesis, highlighting their roles in ecological adaptation and biocontrol potential. Biosynthetic gene cluster (BGC) analysis identified numerous NRPS, PKS, and terpene pathways, many of which belong to pan_shell and pan_cloud regions, suggesting dynamic evolutionary origins. We further detected 66,260 horizontally transferred (HGT) genes, including 438 in BGCs, underscoring HGT as a major driver of metabolic innovation. Together, these findings provide novel insights into the genomic diversity, adaptive capacity, and secondary metabolic potential of S. sampsonii and its close relatives.

BGCs↗

Microgeographic variation in rDNA intergenic spacers of Anopheles gambiae in western Kenya.

The genetic population structure of Anopheles gambiae (Diptera: Culicidae) in western Kenya was investigated by hybridizing a rapidly evolving rDNA intergenic spacer sequence to restriction endonuclease digests of genomic DNA extracted from single mosquitoes from seven localities. Significantly different distributions of restriction fragment arrays were obtained from field sites less than 10 km apart, which suggests restricted gene flow and a subdivided population structure. Eight of twenty-one possible comparisons between pairs of populations yielded significant differences. An eastern Kenya coastal population did not share its restriction fragment arrays with any of the western populations, suggesting that isolation by distance can be complete on a relatively small geographic scale (700 km).

Animals↗

Low sequence variation among isolates of infectious hypodermal and hematopoietic necrosis virus (IHHNV) originating from Hawaii and the Americas.

A 2.9 kb fragment of the infectious hypodermal and hematopoietic necrosis virus (IHHNV) genome, which contains the coding sequence of putative non-structural and capsid proteins, was amplified and sequenced from each of 14 IHHNV isolates collected from cultured penaeid shrimp stocks in Hawaii and various sites in the Americas between 1982 and 1997. The sequence comparison indicates that the IHHNV genome is very stable, with 99.6 to 100% similarity among these 14 isolates. Only nucleotide substitutions were found. The percentage of substitution was higher in the putative capsid proteins region (1.3%) than in the putative non-structural proteins region (0.6%). Out of 25 substitutions found, 14 resulted in amino acid changes. There is no apparent association between clinical outcomes and particular amino acid substitutions. Based on genetic distances, the isolates were clustered into 3 groups that generally correspond with their geographic origins.

Animals↗

Numerous group I introns with variable distributions in the ribosomal DNA of a lichen fungus.

The length of the small subunit ribosomal DNA (SSU rDNA) differs significantly among individuals from natural populations of the ascomycetous lichen complex Cladonia chlorophaea. The sequence of the 3' region of the SSU rDNA from two individuals, chosen to represent the shortest and longest sequences, revealed multiple insertions within a region that otherwise aligned with a 520-nucleotide sequence of the SSU rDNA in Saccharomyces cerevisiae. The high degree of variability in SSU rDNA size can be accounted for by different numbers of insertions; one individual had two group I introns and the second had five introns, two of which were clearly related to introns at identical positions in the other individual. Yet, introns in different positions, whether within an individual or between individuals, were not similar in sequence. The distribution of introns at three of the positions is consistent with either intron loss or acquisition, and clearly indicates the dynamic variability in this region of the nuclear genome. All seven insertions, which ranged in size from 210 to 228 nucleotides, had the conserved sequence and secondary structural elements of group I introns. The variation in distribution and sequence of group I introns within a short highly conserved region of rDNA presents a unique opportunity for examining the molecular evolution and mobility of group I introns within a systematics framework.

Ascomycota↗

Evolutionary analysis of TATA-less proximal promoter function.

Many molecular studies describe how components of the proximal promoter affect transcriptional processes. However, these studies do not account for the likely effects of distant enhancers or chromatin structure, and thus it is difficult to conclude that the sequence variation in proximal promoters acts to modulate transcription in the natural context of the whole genome. This problem, the biological importance of proximal promoter sequence variation, can be addressed using a combination of molecular and evolutionary analyses. Provided here are molecular and evolutionary analyses of the variation in promoter function and sequence within and between populations of Fundulus heteroclitus for the lactate dehydrogenase-B (Ldh-B) proximal promoter. Approximately one third of the Ldh-B proximal promoter contains interspersed regions that are functionally important: (1) they bind transcription factors in vivo, (2) they effect a change in transcription as assayed by transient transfection into two different fish cell lines, and (3) they bind purified transcription factors in vitro. Evolutionary analyses that compare sequence variation in these functional regions versus the nonfunctional regions indicate that the changes in the Ldh-B proximal promoter sequences are due to directional selection. Thus, the Ldh-B proximal promoter sequence variations that affect transcriptional processes constitute a phenotypic change that is subject to natural selection, suggesting that proximal promoter sequence variation affects transcription in the natural context of the whole genome.

Animals↗

Antigenic and structural relatedness among non-capsid and capsid polypeptides of polioviruses belonging to different serotypes.

Antibodies were raised by immunization of Macaca fascicularis monkeys with extracts of M. fascicularis kidney cells separately infected with one of the three poliovirus serotypes. Preparations of antibodies were shown to be strictly type-specific in the neutralization test but formed immune complexes adsorbable on Staphylococcus aureus, Cowan I strain cells, with heterotypic as well as homotypic virus-specific polypeptides present in extracts from the virus-infected cells. The presence of intertypic antigenic determinants was demonstrated by this technique on both the non-capsid and capsid poliovirus polypeptides. Structural variations of poliovirus polypeptides were studied by analysing products of their partial proteolysis. Non-capsid polypeptides encoded in the central portion of the virus genome (polypeptides 5b and X) as well as in the 3'-terminal region (NCVP2 and NCVP4) were found to be highly conserved, whereas capsid polypeptides VP1, VP2, and VP3, which are encoded in the 5'-terminal region of the virus RNA, displayed a much greater variability.

Antigens, Viral↗

Involvement of gene products in bacterial evolution.

Three strategies of different quality contribute in parallel to the natural formation of genetic variants in bacteria: (1) small local alterations of DNA sequences; (2) recombinational reshuffling of segments of the genome; and (3) acquisition of DNA sequences by horizontal gene transfer. Key enzymes involved in these processes often act as variation generators by making use of structural flexibilities of biological macromolecules and of the effect of random encounter. In the theory of molecular evolution, genetic determinants of variation generators as well as of modulators of the frequency of genetic variation are defined as evolutionary genes. This postulate is consistent with the notion that spontaneous mutagenesis is in general not adaptive and that the direction of evolution depends on natural selection exerted on populations of genetic variants.

Bacteria↗

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing↗

Deciphering Campylobacter jejuni cell surface interactions from the genome sequence.

The completion of the Campylobacter jejuni genome sequence is a landmark in Campylobacter research. Discoveries directly arising from these data include the identification of a capsular polysaccharide, extensive capacity for phase variable gene expression and lipo-oligosaccharide structural phase variation. The recent identification of a unique system of general protein glycosylation in C. jejuni, a C. jejuni protein that is translocated into eukaryotic cells, and plasmid-encoded components of a putative type IV secretion system are likely to be significant in terms of the host-pathogen interaction.

Antigenic Variation↗