Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genomic Structural Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Restricted variability of a 17 nucleotide stretch within the 5'-noncoding region of poliovirus genome.

The outbreak of poliomyelitis in Finland in 1984 was caused by a wild strain of poliovirus 3 with uncommon molecular and antigenic properties. We prepared a synthetic oligonucleotide probe complementary to nucleotides 494-510 in the 5'-noncoding part of the genome of a representative strain of the outbreak. This short nucleotide stretch was found to be relatively well conserved within the outbreak and uncommon among 82 independent poliovirus isolates. It may thus be a useful marker for screening isolates to identify those requiring more detailed genetic comparison. The sequences of the corresponding region of the genome are known for 32 separate poliovirus strains and 3 coxsackie B virus strains and show 6 fully conserved nucleotides that could assume a constant hairpin-loop position in a hypothetical secondary structure of the RNA. This could explain the persistence of a particular 17 nucleotide sequence for 40 years in nature in this highly variable region of the poliovirus genome.

Animals↗

Gene family evolution and homology: genomics meets phylogenetics.

With the advent of high-throughput DNA sequencing and whole-genome analysis, it has become clear that the coding portions of the genome are organized hierarchically in gene families and superfamilies. Because the hierarchy of genes, like that of living organisms, reflects an ancient and continuing process of gene duplication and divergence, many of the conceptual and analytical tools used in phylogenetic systematics can and should be used in comparative genomics. Phylogenetic principles and techniques for assessing homology, inferring relationships among genes, and reconstructing evolutionary events provide a powerful way to interpret the ever increasing body of sequence data. In this review, we outline the application of phylogenetic approaches to comparative genomics, beginning with the inference of phylogeny and the assessment of gene orthology and paralogy. We also show how the phylogenetic approach makes possible novel kinds of comparative analysis, including detection of domain shuffling and lateral gene transfer, reconstruction of the evolutionary diversification of gene families, tracing of evolutionary change in protein function at the amino acid level, and prediction of structure-function relationships. A marriage of the principles of phylogenetic systematics with the copious data generated by genomics promises unprecedented insights into the nature of biological organization and the historical processes that created it.

Animals↗

Existing variations on the gene structure of hepatitis E virus strains from some regions of China.

The isolation and identification of the 87A strain of hepatitis E virus (HEV) by means of cell culture have been described previously. This paper reports the nucleotide sequence of a portion of this HEV strain. The RNA extracted from the supernatants of the different passages of the 87A strain cultured in the A549 cell line was reverse-transcribed (RT) to cDNA, and then the polymerase chain reaction (PCR) amplification was carried out using the primers of HEV ET1.1 region. The PCR products from 1) the supernatant of the infected cells at the fourth passage, 2) the virus concentrated by polyethylene glycol (PEG) precipitation at the tenth passage, and 3) the virus purified by a sucrose gradient at the tenth passage were sequenced. In addition, three other PCR products obtained from sera of acute hepatitis E patients in Beijing (B-9) and Guangzhou (G-9 and G-20) were also sequenced. The nucleotide sequences of the above four strains of HEV (located in the genome from positions 4545-4754) were compared to those of some reported HEV strains. The nucleotide sequences of the B-9 strain and the 87A strain were similar to the Burmese strain and may belong to the same branch of HEV. The nucleotide sequences of the G-9 strain and the G-20 strain were a novel and unique branch. The Chinese HEV strains are multiplex and variable in gene structure.

Acute Disease↗

Molecular characterization of two isolates of human T cell leukaemia virus type II from Italian drug abusers and comparison of genome structure with other isolates.

The human T cell leukaemia virus type II (HTLV-II), whose pathogenicity is as yet unclear, was recently found to be associated with intravenous drug abuse in North America and Europe. HTLV-II was isolated from two Italian drug abusers belonging to the same cohort and coinfected with human immunodeficiency virus type 1. Two new isolates, HTLV-II Gu and Va, were established in a culture of BJAB cells, a continuous B cell line (Epstein-Barr virus-negative), and characterized by nucleotide sequence analysis of the long terminal repeat (LTR) and portions of the gag, env and X regions. These sequences were compared to those of the HTLV-II Mo isolate reported in the literature. No major variations were observed in important regulatory elements of LTR nor in the stem-bulge-loop configuration known to be essential for binding of rex protein. The results obtained from the sequence of the 1988 nucleotides examined indicated a 1.6% variability between the Gu and Va isolates and about 6% with respect to Mo. Notable differences were found in the structure of putative open reading frames of the X region when compared to those reported for the Mo isolate. Restriction analysis of proviral DNA of two isolates and comparison with the physical map of the Mo isolate confirmed the existence of genetic heterogeneity in the HTLV-II group and demonstrated that the new isolates Gu and Va belong to the HTLV-IIb subtype. The results of this study show that the new isolates have distinct features with respect to the Mo isolate though all important regulatory elements of the LTR appear to be well conserved.

Genes, Viral↗

The integration of genomic and structural information in the development of high affinity plasmepsin inhibitors.

The plasmepsins are key enzymes in the life cycle of the Plasmodium parasites responsible for malaria. Since plasmepsin inhibition leads to parasite death, these enzymes have been acknowledged to be important targets for the development of new antimalarial drugs. The development of effective plasmepsin inhibitors, however, is compounded by their genomic diversity which gives rise not to a unique target for drug development but to a family of closely related targets. Successful drugs will have to inhibit not one but several related enzymes with high affinity. Structure-based drug design against heterogeneous targets requires a departure from the classic 'lock-and-key' paradigm that leads to the development of conformationally constrained molecules aimed at a single target. Drug molecules designed along those principles are usually rigid and unable to adapt to target variations arising from naturally occurring genetic polymorphisms or drug-induced resistant mutations. Heterogeneous targets need adaptive drug molecules, characterised by the presence of flexible elements at specific locations that sustain a viable binding affinity against existing or expected polymorphisms. Adaptive ligands have characteristic thermodynamic signatures that distinguish them from their rigid counterparts. This realisation has led to the development of rigorous thermodynamic design guidelines that take advantage of correlations between the structure of lead compounds and the enthalpic and entropic components of the binding affinity. In this paper, we discuss the application of the thermodynamic approach to the development of high affinity (K(i) - pM) plasmepsin inhibitors. In particular, a family of allophenylnorstatine-based compounds is evaluated for their potential to inhibit a wide spectrum of plasmepsins.

Amino Acid Sequence↗

Phenotypic switching of variable surface lipoproteins in Mycoplasma bovis involves high-frequency chromosomal rearrangements.

Mycoplasma bovis, an important pathogen of cattle, was recently shown to possess a family of phase- and size-variable membrane surface lipoprotein antigens (Vsps). These proteins spontaneously undergo noncoordinate phase variation between ON and OFF expression states, generating surface antigenic variation. In the present study, we show that the spontaneously high rate of Vsp phenotypic switching involves DNA rearrangements that occur at high frequency in the M. bovis chromosome. A 1.5-kb HindIII genomic fragment carrying the vspA gene from M. bovis PG45 was cloned and sequenced. The deduced VspA amino acid sequence revealed that 80% of the VspA molecule is composed of reiterated intragenic coding sequences, creating a periodic polypeptide structure. Four distinct internal regions of repetitive sequences in the form of in-tandem blocks extending from the N-terminal to the C-terminal portion of the Vsp product were identified. Southern blot analysis of phenotypically switched isogenic lineages representing ON or OFF phase states of Vsp products suggested that changes in the Vsp expression profile were associated with detectable changes at the DNA level. By using a synthetic oligonucleotide representing a sequence complementary to the repetitive vspA gene region as a probe, we could identify the vspA-bearing restriction fragment undergoing high-frequency reversible rearrangements during oscillating phase transition of vspA. The 1.5-kb HindIII fragment carrying the vspA gene (on state) rearranged and produced a 2.3-kb HindIII fragment (OFF state) and vice versa. Two newly discovered vsp genes (vspE and vspF) were localized on two HindIII fragments flanking the vsp gene upstream and downstream. Southern blot hybridization with vspE- and vspF-specific oligonucleotides as probes against genomic DNA of VspA phase variants showed that the organization and size of the fragments adjacent to the vspA gene remained unchanged during VspA ON-OFF switching. The mechanisms regulating the vsp genes are yet unknown; our findings suggest that a recombinative mechanism possibly involving DNA inversions, DNA insertion, or mobile genetic elements may play a role in generating the observed high-frequency DNA rearrangements.

Amino Acid Sequence↗

The foldback-like transposon Galileo is involved in the generation of two different natural chromosomal inversions of Drosophila buzzatii.

Chromosomal inversions are the most common type of genome rearrangement in the genus Drosophila. Although the potential of transposable elements (TEs) for generating inversions has been repeatedly demonstrated in the laboratory, little is known on their role in the generation of natural inversions, which are those effectively contributing to the adaptation and/or evolution of species. We have cloned and sequenced the two breakpoints of the polymorphic inversion 2q7 of D. buzzatii. The sequence analysis of the breakpoint regions revealed the presence in the inverted chromosomes of large insertions, formed by complex assemblies of transposons, that are absent from the chromosomes without the inversion. Among the transposons inserted, the Foldback-like element Galileo, that was previously found responsible of the generation of the widespread inversion 2j of D. buzzatii, is present at both 2q7 breakpoints and is the most likely inducer of the inversion. A detailed study of the nucleotide and structural variation in the breakpoint regions of six chromosomal lines with the 2q7 inversion detected no nucleotide differences between them, which suggests a monophyletic and recent origin. In contrast, a remarkable degree of structural variation was observed in the same six chromosomal lines. It thus appears that the two breakpoints of the inverted chromosomes have become genetically unstable hotspots, as was previously found for the 2j inversion breakpoints. The possibility that this instability is caused by structural properties of Foldback elements is discussed.

Animals↗

SPA: a probabilistic algorithm for spliced alignment.

Recent large-scale cDNA sequencing efforts show that elaborate patterns of splice variation are responsible for much of the proteome diversity in higher eukaryotes. To obtain an accurate account of the repertoire of splice variants, and to gain insight into the mechanisms of alternative splicing, it is essential that cDNAs are very accurately mapped to their respective genomes. Currently available algorithms for cDNA-to-genome alignment do not reach the necessary level of accuracy because they use ad hoc scoring models that cannot correctly trade off the likelihoods of various sequencing errors against the probabilities of different gene structures. Here we develop a Bayesian probabilistic approach to cDNA-to-genome alignment. Gene structures are assigned prior probabilities based on the lengths of their introns and exons, and based on the sequences at their splice boundaries. A likelihood model for sequencing errors takes into account the rates at which misincorporation, as well as insertions and deletions of different lengths, occurs during sequencing. The parameters of both the prior and likelihood model can be automatically estimated from a set of cDNAs, thus enabling our method to adapt itself to different organisms and experimental procedures. We implemented our method in a fast cDNA-to-genome alignment program, SPA, and applied it to the FANTOM3 dataset of over 100,000 full-length mouse cDNAs and a dataset of over 20,000 full-length human cDNAs. Comparison with the results of four other mapping programs shows that SPA produces alignments of significantly higher quality. In particular, the quality of the SPA alignments near splice boundaries and SPA's mapping of the 5' and 3' ends of the cDNAs are highly improved, allowing for more accurate identification of transcript starts and ends, and accurate identification of subtle splice variations. Finally, our splice boundary analysis on the human dataset suggests the existence of a novel non-canonical splice site that we also find in the mouse dataset. The SPA software package is available at http://www.biozentrum.unibas.ch/personal/nimwegen/cgi-bin/spa.cgi.

Algorithms↗

Sequence divergence in a family of variant surface glycoprotein genes from trypanosomes: coding region hypervariability and downstream recombinogenic repeats.

The surface of the parasitic protozoan Trypanosoma brucei spp. is covered with a dense coat consisting of a single type of glycoprotein molecule, the variant surface glycoprotein (VSG). There may be as many as 1,000 genes for VSG within the genome of T. brucei, and the switch of expression from one to another is the phenomenon of antigenic variation. As an approach to understanding the evolution of VSG genes we have determined the genomic DNA sequences of the eight genes encoding the variant surface glycoprotein 117 (VSG) family. From these data we have observed a number of features concerning the relationships between these genes: (1) there is a region of high variability confined to the N-terminus of the coding sequence, and comparison of the sequences with the available X-ray diffraction crystal structures suggests that two of the most variable stretches within the N-terminal domain are present on surface-exposed loops, indicating a role for epitope selection in evolution of these genes; (2) the 29 nucleotides surrounding the splice acceptor site are absolutely conserved in all eight 117 VSG genes; (3) numerous insertion/deletion mutations are located within or immediately downstream of the C-terminal protein-coding sequences: (4) within 500 bp downstream of the insertion/deletion mutations are one or two copies of a repeat motif highly homologous to the recombinogenic 76-bp repeat sequences present upstream of many VSG basic copy genes and the expression-linked copy.

Amino Acid Sequence↗

No isochores in the human chromosomes 21 and 22?

The human genome is described in the literature as being composed of the isochores, i.e., long (hundreds of kilobases) segments with a homogeneous (G + C) content. We calculated the (G + C) content variations along the DNA molecules of the human chromosomes 21 and 22 and found the variations to be higher everywhere compared to the randomized sequences. Hence the (G + C) content is certainly not homogeneous on the isochore scale in the two human chromosomes. In addition, we found no significant difference between the two human molecules and the genome of E. coli regarding the (G + C) content variations. Hence no isochores are either present in the DNA molecules of the human chromosomes 21 and 22, or the isochores are also present in the genome of Escherichia coli. In any case, the present communication demonstrates that the isochores should be defined in unambiguous molecular terms if they are to be used for an up-to-date genome structure characterization.

Algorithms↗

Recurrent simple tandem repeat mutations during human Y-chromosome radiation in Caucasian subpopulations.

The haplotypes at four polymorphic loci of the Y chromosome were determined in 245 Caucasian males from 12 subpopulations. The data show that haplotype radiation occurred among Caucasians. Haplotype radiation was accompanied by recurrent mutations at STR loci that caused partial randomization of haplotype structure. The present distribution of alleles at short tandem repeats (STRs) can be explained by a mutation pattern similar to those described for autosomal STRs. The degree of variation among groups of subpopulations was assayed by using the Analysis of Molecular Variance. The results confirm a faster divergence of the Y chromosome as compared to the rest of the genome.

Haplotypes↗

Testing Pleistocene refugia theory: phylogeographical analysis of Desmognathus wrighti, a high-elevation salamander in the southern Appalachians.

During the colder climates of the Pleistocene, the ranges of high-elevation species in unglaciated areas may have expanded, leading to increased gene flow among previously isolated populations. The phylogeography of the pygmy salamander, Desmognathus wrighti, an endemic species restricted to the highest mountain peaks of the southern Appalachians, was examined to test the hypothesis that the range of D. wrighti expanded along with other codistributed taxa during the Pleistocene. Analyses of genetic variation at 14 allozymic loci and of the 12S rRNA gene in the mtDNA genome was conducted on individuals sampled from 14 population isolates throughout the range of D. wrighti. In contrast to the genetic patterns of many other high-elevation animals and plants, genetic distances derived from both molecular markers showed significant isolation by distance and genetic structuring of populations, suggesting long-term isolation of populations. Phylogeographical analyses revealed four genetically distinct population clusters that probably remained fragmented during the Pleistocene, although there was also evidence supporting recent gene flow among some population groups. Support for isolation by distance is rare among high-elevation species in unglaciated areas of North and Middle America, although not uncommon among Plethodontid Salamanders, and this pattern suggests that populations of D. wrighti did not expand entirely into suitable habitat during the Pleistocene. We propose that intrinsic barriers to dispersal, such as species interactions with other southern Appalachian plethodontid salamanders, persisted during the Pleistocene to maintain the fragmented distribution of D. wrighti and allow for significant genetic divergence of populations by restricting gene flow.

Animals↗

Evaluation of methods for detecting recombination from DNA sequences: computer simulations.

Recombination is a key evolutionary process that shapes the architecture of genomes and the genetic structure of populations. Although many statistical methods are available for the detection of recombination from DNA sequences, their absolute and relative performance is still unknown. Here we evaluated the performance of 14 different recombination detection algorithms. We used the coalescent with recombination to simulate DNA sequences with different levels of recombination, genetic diversity, and rate variation among sites. Recombination detection methods were applied to these data sets, and whether they detected or not recombination was recorded. Different recombination methods showed distinct performance depending on the amount of recombination, genetic diversity, and rate variation among sites. The model of nucleotide substitution under which the data were generated did not seem to have a significant effect. Most methods increase power with more sequence divergence. In general, recombination detection methods seem to capture the presence of recombination, but they are not very powerful. Methods that use substitution patterns or incompatibility among sites were more powerful than methods based on phylogenetic incongruence. Most methods do not seem to infer more false positives than expected by chance. Especially depending on the amount of diversity in the data, different methods could be used to attain maximum power while minimizing false positives. Results shown here will provide some guidance in the selection of the most appropriate method/s for the analysis of the particular data at hand.

Computer Simulation↗

The large-scale organization of the centromeric region in Beta species.

In higher eukaryotes, the DNA composition of centromeres displays a high degree of variation, even between chromosomes of a single species. However, the long-range organization of centromeric DNA apparently follows similar structural rules. In our study, a comparative analysis of the DNA at centromeric regions of Beta species, including cultivated and wild beets, was performed using a set of repetitive DNA sequences. Our results show that these regions in Beta genomes have a complex structure and consist of variable repetitive sequences, including satellite DNA, Ty3-gypsy-like retrotransposons, and microsatellites. Based on their molecular characterization and chromosomal distribution determined by fluorescent in situ hybridization (FISH), centromeric repeated DNA sequences were grouped into three classes. By high-resolution multicolor-FISH on pachytene chromosomes and extended DNA fibers we analyzed the long-range organization of centromeric DNA sequences, leading to a structural model of a centromeric region of the wild beet species Beta procumbens. The chromosomal mutants PRO1 and PAT2 contain a single wild beet minichromosome with centromere activity and provide, together with cloned centromeric DNA sequences, an experimental system toward the molecular isolation of individual plant centromeres. In particular, FISH to extended DNA fibers of the PRO1 minichromosome and pulsed-field gel electrophoresis of large restriction fragments enabled estimations of the array size, interspersion patterns, and higher order organization of these centromere-associated satellite families. Regarding the overall structure, Beta centromeric regions show similarities to their counterparts in the few animal and plant species in which centromeres have been analyzed in detail.

Amino Acid Sequence↗

Restriction fragment length polymorphisms in dairy and beef cattle at the growth hormone and prolactin loci.

Two bovine populations, a Holstein-Friesian dairy stock and a synthetic (Baladi X Hereford X Simmental X Charolais) beef stock, were screened for restriction fragment length polymorphisms (RFLPs) at the growth hormone and prolactin genes. Most RFLPs at the growth hormone gene are apparently the consequence of an insertion/deletion event which was localized to a region downstream of the structural gene. The restriction map for the genomic region including the growth hormone gene was extended. Two HindIII RFLPs at the growth hormone locus, as well as several RFLPs at the prolactin gene, seemed to be the consequence of a series of point mutations. The results are discussed in terms of the possibility that minor genomic variability underlies quantitative genetic variation.

Animals↗

Genetic variation and phylogeographic analyses of two species of Carpobrotus and their hybrids in California.

Despite the commonality and study of hybridization in plants, there are few studies between invasive and noninvasive species that examine the genetic variability and gene flow of cytoplasmic DNA. We describe the phylogeographical structure of chloroplast DNA (cpDNA) variation within and among several interspecific populations of the putative native, Carpobrotus chilensis and the introduced, Carpobrotus edulis (Aizoaceae). These species co-occur throughout much of coastal California and form several 'geographical hybrid populations'. Two hundred and thirty-seven individuals were analysed for variation in an approximate 7.0 kb region of the chloroplast genome using PCR-RFLP (polymerase chain reaction - restriction fragment length polymorphism) data. Phylogenetic analyses and cpDNA population differentiation were conducted for all morphotypes. Historic geographical dispersion and the coefficient of ancestry of the haplotypes were determined using nested clade analyses. Two haplotypic groupings (I and II) were represented in C. chilensis and C. edulis, respectively. The variation in cpDNA data is in agreement with the previously reported allozyme and morphological data; this supports relatively limited variation and high population differentiation among C. chilensis and hybrids and more wide-ranging variation in C. edulis and C. edulis populations backcrossed with C. chilensis. C. chilensis disproportionately contributes to the creation of hybrids with the direction of gene flow from C. chilensis into C. edulis. The cpDNA data support C. chilensis as the maternal contributor to the hybrid populations.

Aizoaceae↗

Phase and antigenic variation in bacteria.

Phase and antigenic variation result in a heterogenic phenotype of a clonal bacterial population, in which individual cells either express the phase-variable protein(s) or not, or express one of multiple antigenic forms of the protein, respectively. This form of regulation has been identified mainly, but by no means exclusively, for a wide variety of surface structures in animal pathogens and is implicated as a virulence strategy. This review provides an overview of the many bacterial proteins and structures that are under the control of phase or antigenic variation. The context is mainly within the role of the proteins and variation for pathogenesis, which reflects the main body of literature. The occurrence of phase variation in expression of genes not readily recognizable as virulence factors is highlighted as well, to illustrate that our current knowledge is incomplete. From recent genome sequence analysis, it has become clear that phase variation may be more widespread than is currently recognized, and a brief discussion is included to show how genome sequence analysis can provide novel information, as well as its limitations. The current state of knowledge of the molecular mechanisms leading to phase variation and antigenic variation are reviewed, and the way in which these mechanisms form part of the general regulatory network of the cell is addressed. Arguments both for and against a role of phase and antigenic variation in immune evasion are presented and put into new perspective by distinguishing between a role in bacterial persistence in a host and a role in facilitating evasion of cross-immunity. Finally, examples are presented to illustrate that phase-variable gene expression should be taken into account in the development of diagnostic assays and in the interpretation of experimental results and epidemiological studies.

Amino Acid Sequence↗

Anaplasma marginale major surface protein 3 is encoded by a polymorphic, multigene family.

The immunodominant surface protein, MSP3, is structurally and antigenically polymorphic among strains of Anaplasma marginale. In this study we show that a polymorphic multigene family is at least partially responsible for the variation seen in MSP3. The A. marginale msp3 gene msp3-12 was cloned and expressed in Escherichia coli. With msp3-12 as a probe, multiple, partially homologous gene copies were identified in the genomes of three A. marginale strains. These copies were widely distributed throughout the chromosome. Sequence analysis of three unique msp3 genes, msp3-12, msp3-11, and msp3-19, revealed both conserved and variant regions within the open reading frames. Importantly, msp3 contains amino acid blocks related to another polymorphic multigene family product, MSP2. These data, in conjunction with data presented in previous studies, suggest that multigene families are used to vary important antigenic surface proteins of A. marginale. These findings may provide a basis for studying antigenic variation of the organism in persistently infected carrier cattle.

Amino Acid Sequence↗