Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Structural genome variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Identification of two distinct subfamilies of alpha satellite DNA that are highly specific for human chromosome 15.

We report the isolation of two distinct subfamilies of alpha satellite DNA (pTRA-20 and -25) from human chromosome 15. In situ hybridization experiments indicated that both subfamilies are highly specific for this chromosome. Southern analysis of a somatic hybrid cell line carrying human chromosome 15 revealed a likely higher-order genomic band of 2.5 kb for pTRA-20. Similar analysis for pTRA-25 showed multiple higher-order bands of 3.5, 4.5, and 5 kb at moderately high hybridization stringency, but a predominance of the 4.5-kb species at very high stringency. Direct comparison with human genomic DNA confirmed the authenticity of these higher-order structures and demonstrated polymorphic variations using both probes. The origin of the different alphoid subfamilies on chromosome 15 is discussed. These sequences should be useful for the construction of centromere-based genetic linkage maps for human chromosome 15 and, in conjunction with the other alphoid sequences already reported for chromosomes 13, 14, 21, and 22, should allow a concerted analysis of the evolution and the possible etiological role of these DNAs in aberrations commonly seen in these chromosomes.

Blotting, Southern↗

Coalescent processes and relaxation of selective constraints leading to contrasting genetic diversity at paralogs AtHVA22d and AtHVA22e in Arabidopsis thaliana.

Duplicate loci offer a very powerful system for understanding the complicated genome structure and adaptive evolution of a gene family. In this study, the genetic variation at paralogs AtHVA22d and AtHVA22e, members of an ABA- and stress-inducible gene family, is examined in the selfing Arabidopsis thaliana. Population genetic analysis indicates contrasting levels of nucleotide diversity at overall exon sequence and nonsynonymous sites between AtHVA22d (pi = 0.00337, pi(rep) = 0.00158) and AtHVA22e (pi = 0.00054, pi(rep) = 0.00023). The fact of Ka/Ks ratios significantly less than 1 in all sequences indicates that both genes are functional and subjected to purifying selection. In addition, rooted at barley HVA22, accelerated evolution is detected at replacement changes in the AtHVA22d locus, indicating relaxation of purifying selection after gene duplication. However, relative rate tests reveal no deviation from the neutrality at synonymous sites between the two paralogs. Based on clock-like evolution, the rate of synonymous substitution is estimated at 1.83 x 10(-9) substitutions per site per year; and the divergence of the two paralogs is traced to 90 MYA, coinciding with a period of the diversification of angiosperms. Given no codon usage bias in both genes, natural selection alone cannot account for the 6.4-fold differences in the nucleotide variation at synonymous sites between the two paralogs. Random processes resulting in different coalescence times, 3.65 MYA at AtHVA22d vs. 1.20 MYA at AtHVA22e, may have predominantly contributed to the evident differences of the genetic diversity. Partially nonoverlapping modes of expression between the two functional paralogs suggest a subfunctionalization hypothesis for explaining the fates of duplicate loci.

Arabidopsis↗

Genotypic characteristics of bovine viral diarrhea virus 2 strains isolated in northern Italy.

Two strains of Bovine viral diarrhea virus 2 (BVDV-2) were isolated from calves in northern Italy. Variations in the 5'-untranslated region (UTR) of the genome were studied by primary structure alignment and neighbor-joining method based phylogenetic tree analyses and by palindromic nucleotide substitutions at the three variable loci in the 5'-UTR. Genetic analysis indicated their appurtenance to genovar BVDV-2a. Nucleotide sequence at the 5'-UTR of strain BS-95-II, one of the Italian isolates from healthy calves, showed 98% homology to that of the Japanese isolate OY89, a cytopathic strain derived from cattle with mucosal disease.

5' Untranslated Regions↗

Poliovirus type 3/Saukett: antigenic and structural correlates of sequence variation in the capsid proteins.

The Saukett/USA/50 strain is the type 3 component of the inactivated poliovirus vaccine. The capsid-coding region of genomic RNA of Saukett strains from five different sources was sequenced and the sequence differences were correlated with antigenic differences measurable with poliovirus type 3-specific neutralizing monoclonal antibodies. All strains appeared to have capsid protein genes identical in size to those of the entirely sequenced type 3 poliovirus strains. The nucleotide sequence identity between the strains was 91% on the average and the strains could be divided into three groups. Amino acid differences were seen in 30 positions located throughout the capsid region both within and outside the known antigenic sites. Substitutions at the known antigenic sites explained most of the observed antigenic differences. Use of the atomic coordinates of the crystal structure model of the Sabin 3 virus and prior data based on escape mutants and peptide scanning revealed that most of the exposed substitutions located outside the known antigenic sites are spatially associated with regions found to be antigenic by either or both of these methods.

Antigenic Variation↗

A single-nucleotide natural variation (U4 to C4) in an influenza A virus promoter exhibits a large structural change: implications for differential viral RNA synthesis by RNA-dependent RNA polymerase.

The influenza A virus promoter is recognized by the influenza A virus RNA-dependent RNA polymerase, and directs both transcription and replication of the viral RNA genome. Within the sequence of this promoter, flu strains exhibit a natural, unique variation, either a U or a C, at the fourth position from the 3' end. Promoters that contain a C residue (C4 promoter), which are invariably found in genome segments that encode the three RNA polymerase subunits (PB1, PB2 and PA), down-regulate transcription but activate genome replication. Here, we have determined the structure of the C4 promoter by NMR spectroscopy and compared it with the structure of the U4 promoter, which was determined previously. The structure of the internal loop in the C4 promoter is similar to that of the U4 promoter. However, the terminal stem of the C4 promoter is strikingly different from that of the U4 promoter. These structural data suggest that the internal loop is important for polymerase binding to the promoter, and the terminal stem is crucial for differential regulation of transcription and replication.

Base Sequence↗

Genomic structure of DNA encoding the lymphocyte homing receptor CD44 reveals at least 12 alternatively spliced exons.

The CD44 molecule is known to display extensive size heterogeneity, which has been attributed both to alternative splicing and to differential glycosylation within the extracellular domain. Although the presence of several alternative exons has been partly inferred from cDNA sequencing, the precise intron-exon organization of the CD44 gene has not been described to date to our knowledge. In the present study we describe the structure of the human CD44 gene, which contains at least 19 exons spanning some 50 kilobases of DNA. We have identified 10 alternatively spliced exons within the extracellular domain, including 1 exon that has not been previously reported. In addition to the inclusion or exclusion of whole exons, more diversity is generated through the utilization of internal splice donor and acceptor sites within 2 of the individual exons. The variation previously reported for the cytoplasmic domain is shown to result from the alternative splicing of 2 exons. The genomic structure of CD44 reveals a remarkable degree of complexity, and we confirm the role of alternative splicing as the basis of the structural and functional diversity seen in the CD44 molecule.

Alternative Splicing↗

Phylogenetic analysis of African horse sickness virus segment 10: sequence variation, virulence characteristics and cell exit.

African horse sickness virus (AHSV) genome segment 10 encodes the non-structural proteins NS3/NS3a, which is involved in release of virus from cells. Full length segment 10 cDNAs were amplified by reverse transcription-polymerase chain reaction, from isolates of AHSV serotypes 2, 3, 4, 5, 7, 8 and 9. These cDNAs were cloned, sequenced and their phylogenetic relationships analysed. High levels of sequence homology were detected in segment 10 from some isolates of different serotypes, confirming that they could be grouped on this basis (serotypes 4, 5, 6 and 9 (group alpha); serotypes 3 and 7 (group beta); serotypes 1, 2, and 8 (group gamma). However, data from bluetongue virus (the prototype orbivirus) indicate that the AHSV serotype is determined exclusively by the structural outer coat proteins VP2 and VP5, encoded by genome segments 2 and 5 respectively. Therefore, as a direct consequence of genome segment reassortment between AHSV strains from different serotypes, the differences observed in segment 10 do not give a reliable indication of virus serotype. Segment 10 of AHSV 3 (virulent) and AHSV 3att (attenuated) were also analysed. These strains, together with AHSV 8, have been used to study of the genetic basis of virulence using reassortment (O'Hara et al., this publication). Virus release studies, using Culicoides cell cultures, indicate that differences in segment 10 of AHSV 3att and 8 can influence the timing of virus release from the infected cell.

African Horse Sickness Virus↗

A novel method to calculate the G+C content of genomic DNA sequences.

The base composition of a DNA fragment or genome is usually measured by the proportion of A+T or G+C in the sequence. The G+C content along genomic sequences is usually calculated using an overlapping or non-overlapping sliding window method. The result and accuracy of such an approach depends on the size of the window and the moving distance adopted. In this paper, a novel windowless technique to calculate the G+C content of genomic sequences is proposed. By this method, the G+C content can be calculated at different "resolution". In an extreme case, the G+C content may be computed at a specific point, rather than in a window of finite size. This is particularly useful to analyze the fine variation of base composition along genomic sequences. As the first example, the variation of G+C content along each of 16 yeast chromosomes is analyzed. The G+C-rich regions with length larger than 5 kb sequences are detected and listed in details. It is found that each chromosome consists of several G+C-rich and G+C-poor regions alternatively, i.e., a mosaic structure. Another example is to analyze the G+C content for each of the two chromosomes of the Vibrio cholerae genome. Based on the variations of the G+C content in each chromosome, it is shown that some fragments in the Vibrio cholerae genome may have been transferred from other species. Especially, the position and size of the large integron island on the smaller chromosome was precisely predicted. This method would be a useful tool for analyzing genomic sequences.

Base Composition↗

Fine mapping of inherent flexibility variation along DNA molecules: validation by atomic force microscopy (AFM) in buffer.

Curvature and flexibility are structural properties of central importance to genome function. However, due to the difficulties in finding suitable experimental conditions, methods for studying one without the interference of the other have proven to be difficult. We propose a new approach that provides a measure of inherent flexibility of DNA by taking advantage of two powerful techniques, X-ray crystallography and nuclear magnetic resonance. Both techniques are able to detect local curvature on DNA fragments but, while the first analyzes DNA in the solid state, the second works on DNA in solution. Comparison of the two data sets allowed us to calculate the relative contribution to flexibility of the three rotations and three translations, which relate successive base pair planes for the ten different dinucleotide steps. These values were then used to compute the variation of flexibility along a given nucleotide sequence. This allowed us to validate the method experimentally through comparisons with maps of local fluctuations in DNA molecule trajectory constructed from atomic force microscopy imaging in solution. We conclude that the six dinucleotide-step parameters defined here provide a powerful tool for the exploration of DNA structure and, consequently will make an important contribution to our understanding of DNA-sequence-dependent biological processes.

Base Sequence↗

A locus on chromosome 13 influences levels of TAFI antigen in healthy Mexican Americans.

When activated, thrombin activatable fibrinolysis inhibitor (TAFI) inhibits fibrinolysis by modifying fibrin, depressing its plasminogen binding potential. Polymorphisms in the TAFI structural gene (CPB2) have been associated with variation in TAFI levels, but the potential occurrence of influential quantitative trait loci (QTLs) located elsewhere in the genome has been explored only in families ascertained in part through probands affected by thrombosis. We report the results of the first genome-wide linkage screen for QTLs that influence TAFI phenotypes. Data are from 635 subjects from 21 randomly ascertained Mexican American families participating in the San Antonio Family Heart Study. Potential QTLs were localized through a genome-wide multipoint linkage scan using 417 highly informative autosomal short tandem repeat markers spaced at approximately 10-cM intervals. We observed a maximum multipoint LOD score of 3.09 on chromosome 13q, the region of the TAFI structural gene. A suggestive linkage signal (LOD = 2.04) also was observed in this region, but may be an artifact. In addition, weak evidence for linkage occurred on chromosomes 17p and 9q. Our results suggest that polymorphisms in the TAFI structural gene or its nearby regulatory elements may contribute strongly to TAFI level variation in the general population, although several genes in other regions of the genome may also influence variation in this phenotype. Our findings support those of the Genetic Analysis of Idiopathic Thrombophilia (GAIT) project, which identified a potential TAFI QTL on chromosome 13q in a genome-wide linkage scan in Spanish thrombophilia families.

Adult↗

Ten years of bacterial genome sequencing: comparative-genomics-based discoveries.

It has been more than 10 years since the first bacterial genome sequence was published. Hundreds of bacterial genome sequences are now available for comparative genomics, and searching a given protein against more than a thousand genomes will soon be possible. The subject of this review will address a relatively straightforward question: "What have we learned from this vast amount of new genomic data?" Perhaps one of the most important lessons has been that genetic diversity, at the level of large-scale variation amongst even genomes of the same species, is far greater than was thought. The classical textbook view of evolution relying on the relatively slow accumulation of mutational events at the level of individual bases scattered throughout the genome has changed. One of the most obvious conclusions from examining the sequences from several hundred bacterial genomes is the enormous amount of diversity--even in different genomes from the same bacterial species. This diversity is generated by a variety of mechanisms, including mobile genetic elements and bacteriophages. An examination of the 20 Escherichia coli genomes sequenced so far dramatically illustrates this, with the genome size ranging from 4.6 to 5.5 Mbp; much of the variation appears to be of phage origin. This review also addresses mobile genetic elements, including pathogenicity islands and the structure of transposable elements. There are at least 20 different methods available to compare bacterial genomes. Metagenomics offers the chance to study genomic sequences found in ecosystems, including genomes of species that are difficult to culture. It has become clear that a genome sequence represents more than just a collection of gene sequences for an organism and that information concerning the environment and growth conditions for the organism are important for interpretation of the genomic data. The newly proposed Minimal Information about a Genome Sequence standard has been developed to obtain this information.

Bacterial Vaccines↗

Heterogeneity in regional GC content and differential usage of codons and amino acids in GC-poor and GC-rich regions of the genome of Apis mellifera.

The honeybee (Apis mellifera) has a genome with a wide variation in GC content showing 2 clear modal GC values, in some ways reminiscent of an isochore-like structure. To gain insight into causes and consequences of this pattern, we used a comparative approach to study the genome-wide alignment of primarily coding sequence of A. mellifera with Drosophila melanogaster and Anopheles gambiae. The latter 2 species show a higher average GC content than A. mellifera and no indications of bimodality, suggesting that the GC-poor mode is a derived condition in honeybee. In A. mellifera, synonymous sites of genes generally adopt the GC content of the region in which they reside. A large proportion of genes in GC-poor regions have not been assigned to the honeybee assembly because of the low sequence complexity of their genome neighborhood. The synonymous substitution rate between A. mellifera and the other species is very close to saturation, but analyses of nonsynonymous substitutions as well as amino acid substitutions indicate that the GC-poor regions are not evolving faster than the GC-rich regions. We describe the codon usage and amino acid usage and show that they are remarkably heterogeneous within the honeybee genome between the 2 different GC regions. Specifically, the genes located in GC-poor regions show a much larger deviation in both codon usage bias and amino acid usage from the Dipterans than the genes located in the GC-rich regions.

Amino Acids↗

Genome-Wide Profiling of Histone Modifications in Fission Yeast Using CUT&Tag.

Eukaryotic DNA is organized in the nucleus in the form of chromatin. Nucleosomes, the fundamental unit of chromatin, are subject to many posttranslational modifications (PTMs) as well as compositional variations through incorporation of histone variants. These alterations play important roles in regulation of genome structure and activity. Genome-wide profiling of these regulatory features is essential for understanding of genome function. Chromatin immunoprecipitation coupled with next-generation sequencing (ChIP-Seq) is a widely used method to assay genome-wide localization in fission yeast but suffers from the requirement for a large amount of input chromatin, antibodies, and a cumbersome experimental pipeline. New methods such as Cleavage Under Targets and Tagmentation (CUT&Tag), which combine the specificity of targeted cleavage and adapter insertion with the sensitivity of next-generation sequencing, enable identification and characterization of various epigenetic marks affording low input requirement as well as more streamlined protocols. However, these approaches have not been adapted for use in fission yeast, Schizosaccharomyces pombe. Here, we describe an adapted CUT&Tag protocol for epigenomic profiling in fission yeast using the heterochromatin-associated histone H3K9 methylation PTM for benchmarking.

Schizosaccharomyces↗

Structure and expression of metallothionein gene in ducks.

Metallothionein (MT) cDNAs were cloned and sequenced from two genera of ducks, Muscovy (Cairina muschata) and Tsai ya (Anas platyrhynchos). The two cDNAs show an extremely high sequence homology and contain an open reading frame encoding 63 amino acids. MT mRNA expressions were studied after metal induction using the cloned cDNA as a probe. Cadmium and copper induced MT gene efficiently, whereas zinc showed a markedly less effect. In addition, the MT mRNA accumulations in various developmental stages were also investigated. The result reveals a different pattern of expression from that of mammals. The discrepancy in MT gene between Tsai ya and Muscovy was further explored by examining genomic DNA structures. The duck MT showed three exons and two introns. The most significant variation of the genes occurs at intron II in which Tsai ya MT has 24 bases more than Muscovy MT. Moreover, MT expressions in the hybrids of Muscovy and Tsai ya were investigated using a reverse transcriptase-polymerase chain reaction. Those results demonstrated that parental MT genes are expressed in the hybrids after metal induction.

Amino Acid Sequence↗

Genetic characterization of ruminant pestiviruses: sequence analysis of viral genotypes isolated from sheep.

Historically, the genus pestivirus was believed to contain three species of viruses; bovine viral diarrhea virus (BVDV), border disease virus (BDV) and classical swine fever virus (CSFV). However, based on limited sequence analysis of a small number of pestiviral isolates from domestic livestock, evidence has recently emerged indicating that at least four distinct genotypes exist. In an attempt to gain a better understanding of the degree of viral variation among ruminant pestiviruses, the entire structural gene coding region of an ovine pestivirus. BD31, genome encompassing 3358 nucleotides was cloned and sequenced. Sequence analysis revealed that BD31 shares less than 71% nucleotide similarity with other pestiviruses, suggesting, that BD31 is distinct from BVDV, CSFV as well as other ovine and bovine pestiviruses currently referred to as BVDV type II. Based on this data, BD31 is the first North American pestivirus isolate that falls under the category true BDV. Results from the analysis of the nucleotide sequence of the E0-E1 coding region of six additional ruminant pestiviruses identified the existence of three distinct virus genotypes in North America. Thus, among ruminent pestiviruses, bovine isolates can be grouped into two genotypes, namely types 1 and 4, whereas ovine isolates fall into genotypes 1, 3 and 4.

Amino Acid Sequence↗

Modulation of virulence within a pathogenicity island in vancomycin-resistant Enterococcus faecalis.

Enterococci are members of the healthy human intestinal flora, but are also leading causes of highly antibiotic-resistant, hospital-acquired infection. We examined the genomes of a strain of Enterococcus faecalis that caused an infectious outbreak in a hospital ward in the mid-1980s (ref. 2), and a strain that was identified as the first vancomycin-resistant isolate in the United States, and found that virulence determinants were clustered on a large pathogenicity island, a genetic element previously unknown in this genus. The pathogenicity island, which varies only subtly between strains, is approximately 150 kilobases in size, has a lower G + C content than the rest of the genome, and is flanked by terminal repeats. Here we show that subtle variations within the structure of the pathogenicity island enable strains harbouring the element to modulate virulence, and that these variations occur at high frequency. Moreover, the enterococcal pathogenicity island, in addition to coding for most known auxiliary traits that enhance virulence of the organism, includes a number of additional, previously unstudied genes that are rare in non-infection-derived isolates, identifying a class of new targets associated with disease which are not essential for the commensal behaviour of the organism.

Base Composition↗

A transcriptomic analysis of the phylum Nematoda.

The phylum Nematoda occupies a huge range of ecological niches, from free-living microbivores to human parasites. We analyzed the genomic biology of the phylum using 265,494 expressed-sequence tag sequences, corresponding to 93,645 putative genes, from 30 species, including 28 parasites. From 35% to 70% of each species' genes had significant similarity to proteins from the model nematode Caenorhabditis elegans. More than half of the putative genes were unique to the phylum, and 23% were unique to the species from which they were derived. We have not yet come close to exhausting the genomic diversity of the phylum. We identified more than 2,600 different known protein domains, some of which had differential abundances between major taxonomic groups of nematodes. We also defined 4,228 nematode-specific protein families from nematode-restricted genes: this class of genes probably underpins species- and higher-level taxonomic disparity. Nematode-specific families are particularly interesting as drug and vaccine targets.

Animals↗

Identification of genomic organisation, sequence variants and analysis of the role of the human dishevelled 1 gene in late onset Alzheimer's disease.

Alzheimer's disease (AD) is a disorder characterised by a progressive deterioration in memory and other cognitive functions. Neurofibrillary tangles (NFT) are a major pathological hallmark of AD, these are aggregations of paired helical filaments (PHF) comprised of the hyperphosphorylated microtubule associated protein tau. Several kinases, such as glycogen synthase kinase 3 beta (GSK3beta) and c-Jun N-terminal kinase (JNK), phosphorylate tau at sites that are phosphorylated in PHF. Dishevelled 1 (DVL1) is thought to act as a positive regulator of the wnt signalling pathway, and inhibits GSK3beta activity preventing beta-catenin degradation and thus allowing wnt target gene expression. JNK activation is also regulated by DVL1, however it is unclear if this is via the wnt signalling pathway. These observations suggest a central role for DVL1 in tau phosphorylation and AD and led us to investigate DVL1 as a candidate gene for this disorder. We determined the genomic structure of the DVL1 gene by sequencing and data mining and searched for sequence variations in the coding sequences and flanking introns. The DVL1 gene spans a region of approximately 13.8 kb (not including the 5' untranslated region) and is encoded by 15 exons. Analysis of over 4.3 kb of sequence, including 98% of exonic sequences and introns 2, 3, 6, 7, 9, 10, 11 and 12, revealed there to be six rare (< or =6%) sequence variations. None of these had any association with late onset AD. This would suggest that polymorphic variations in the coding sequences of DVL1 are not important in AD. However further analysis of regulatory regions may lead to the identification of other sequence variations which may be implicated in AD.

Adaptor Proteins, Signal Transducing↗