Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “sequence diversity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Sequence diversity of the Bacillus thuringiensis and B. cereus sensu lato flagellin (H antigen) protein: comparison with H serotype diversity.

We set out to analyze the sequence diversity of the Bacillus thuringiensis flagellin (H antigen [Hag]) protein and compare it with H serotype diversity. Some other Bacillus cereus sensu lato species and strains were added for comparison. The internal sequences of the flagellin (hag) alleles from 80 Bacillus thuringiensis strains and 16 strains from the B. cereus sensu lato group were amplified and cloned, and their nucleotide sequences were determined and translated into amino acids. The flagellin allele nucleotide sequences for 10 additional strains were retrieved from GenBank for a total of 106 Bacillus species and strains used in this study. These included 82 B. thuringiensis strains from 67 H serotypes, 5 B. cereus strains, 3 Bacillus anthracis strains, 3 Bacillus mycoides strains, 11 Bacillus weihenstephanensis strains, 1 Bacillus halodurans strain, and 1 Bacillus subtilis strain. The first 111 and the last 66 amino acids were conserved. They were referred to as the C1 and C2 regions, respectively. The central region, however, was highly variable and is referred to as the V region. Two bootstrapped neighbor-joining trees were generated: a first one from the alignment of the translated amino acid sequences of the amplified internal sequences of the hag alleles and a second one from the alignment of the V region amino acid sequences, respectively. Of the eight clusters revealed in the tree inferred from the entire C1-V-C2 region amino acid sequences, seven were present in corresponding clusters in the tree inferred from the V region amino acid sequences. With regard to B. thuringiensis, in most cases, different serovars had different flagellin amino acid sequences, as might have been expected. Surprisingly, however, some different B. thuringiensis serovars shared identical flagellin amino acid sequences. Likewise, serovars from the same H serotypes were most often found clustered together, with exceptions. Indeed, some serovars from the same H serotype carried flagellins with sufficiently different amino acid sequences as to be located on distant clusters. Species-wise, B. halodurans, B. subtilis, and B. anthracis formed specific branches, whereas the other four species, all in the B. cereus sensu lato group, B. mycoides, B. weihenstephanensis, B. cereus, and B. thuringiensis, did not form four specific clusters as might have been expected. Rather, strains from any of these four species were placed side by side with strains from the other species. In the B. cereus sensu lato group, B. anthracis excepted, the distribution of strains was not species specific.

Amino Acid Sequence↗

Unexpected sequence diversity in the amino-terminal ends of the coat proteins of strains of sugarcane mosaic virus.

The sequence of the 3'-terminal 1343 nucleotides of the SC strain of the sugarcane mosaic virus (SCMV-SC) genome was compared with the 1376 nucleotides at the 3' terminus of maize dwarf mosaic virus B (MDMV-B). The SCMV-SC sequence includes an open reading frame which codes for the viral coat protein of 313 amino acids (nucleotides 157 to 1116), followed by a 3' non-coding region of 235 nucleotides and a poly(A) tail. The MDMV-B sequence codes for the capsid protein (nucleotides 157 to 1139) of 328 amino acids and has a 3' non-coding region of 236 nucleotides. The coat protein of SCMV-SC has 92% identity with that of MDMV-B except for the region between amino acid residues 27 and 70 of SCMV-SC. This region of SCMV-SC is smaller (44 residues) than the equivalent region in MDMV-B (59 residues) and has only 22% identity with the MDMV-B sequence. Possible mechanisms for the generation of this sequence diversity are discussed. Despite this diversity, the sequence identities of both the major part of the coat proteins and the 3' non-coding regions confirm the proposal, based on previously described serological data, that SCMV-SC and MDMV-B are strains of SCMV.

Amino Acid Sequence↗

Sequence diversity of hepatitis C viral genomes.

The nucleotide sequences of cDNAs (275 base-pairs) in the non-structural protein 5 regions of Japanese isolates of hepatitis C virus (HCV-J) from the plasma of 11 patients with non-A, non-B hepatitis and the livers of five patients with hepatocellular carcinoma were analyzed. Approximately 14 to 17% of nucleotide sequences of the HCV-Js examined differed from that of the original isolate in the United States (HCV-US). Furthermore, 2.5 to 11% sequence diversity was found among the HCV-Js. The nucleotide sequences of the HCV-Js showed characteristic common differences from that of HCV-US, although they also showed some random substitutions. Plural HCV-J genomes were found in two of the cDNAs derived from liver specimens, and a deletion of 102 nucleotides was found in the cDNA derived from one plasma specimen. These results suggest that HCV-J is a strain different from the HCV-US and that mutation of the viral genome occurs at as high a frequency as in that of the human immunodeficiency virus.

Amino Acid Sequence↗

Analysis of the sequence diversity of the P1, HC, P3, NIb and CP genomic regions of several yam mosaic potyvirus isolates: implications for the intraspecies molecular diversity of potyviruses.

Partial sequences from serologically characterized yam mosaic potyvirus (YMV) isolates were determined in conserved (helper-component proteinase, HC; nuclear inclusion b, NIb) and variable (first protein, P1; third protein, P3; and coat protein, CP) regions of the potyviral genome in order to investigate the intraspecies molecular diversity of YMV. Multiple sequence alignments and pairwise comparisons were used to quantify the sequence polymorphism in these regions. Two levels of diversity were observed among YMV isolates: above 90% nucleotide (nt) sequence identities were found between YMV isolates of the same group (intragroup) regardless of the region considered, whereas identities between isolates from different groups (intergroup) were lower and depended upon the protein chosen. For instance, the average intergroup nt sequence identity between YMV isolates was about 65% in the P1 protein and the N terminus of the CP while there was more than 80% nt identity in the HC, P3 and NIb proteins. Thus P3 appeared to be conserved between YMV isolates even though this region was variable between potyvirus species. Similar analysis of the intraspecies molecular diversity of other potyviruses (potato virus Y, zucchini yellow mosaic virus, plum pox virus, pea seed-borne mosaic virus) led to the same results: (i) two levels of intraspecies molecular diversity were found (intragroup and intergroup); (ii) intraspecies molecular diversity differed from interspecies molecular diversity in the P3, P1 and N-terminal regions.

Amino Acid Sequence↗

Generation of sequence diversity in the kinetoplast DNA minicircles of Leishmania mexicana amazonensis.

In order to understand the mechanisms which generate minicircle sequence diversity, we sequenced three minicircles belonging to the same or closely related sequence classes from the kinetoplast DNA of Leishmania mexicana amazonensis strains, PH8, Raimundo, and Josefa. Closely related minicircles from PH8 and Raimundo were unexpectedly found to differ at 11% of positions within the evolutionarily conserved region, but at only 3.9% of positions in the variable region. It thus appears that accumulation of point mutations will not account for the wide intra-strain and intra-subspecies divergence of the variable region. Comparison of more distantly related minicircles from PH8 and Josefa revealed only two short stretches of 70% homology within the variable region. These stretches of homology are not located in the same positions relative to the conserved regions in their respective minicircles. They may represent vestiges of recombinational events responsible for the rapid divergence of minicircle variable regions.

Animals↗

TprK sequence diversity accumulates during infection of rabbits with Treponema pallidum subsp. pallidum Nichols strain.

The tprK gene in Treponema pallidum undergoes antigenic variation. In all T. pallidum isolates examined to date, except the Nichols type strain, heterogeneous tprK sequences have been identified. This heterogeneity is localized to seven variable (V) regions, and tprK sequence diversity accumulates with serial passage in naïve rabbits. The T. pallidum Nichols genome described a single tprK sequence, and after decades of independent passage, only minor tprK sequence diversity is seen among the Nichols strains from different laboratories. We hypothesized that T. pallidum Nichols is capable of only limited tprK diversification. To address this hypothesis, we passaged the T. pallidum Nichols strain in naïve rabbits at the peak of infection (rapid passage) or after the adaptive immune response had cleared most organisms in vivo (slow passage). After 22 rapid passages (9- to 10-day intervals), no tprK V region sequence changes were observed. In contrast, after two slow passages (30- to 35-day intervals), three V regions had sequences that were completely different from that of the original inoculum. New sequences were observed in all seven V regions by the fifth slow passage. In contrast to the rapid-passaged Nichols strain, rapid-passaged Chicago C, a clonal strain isolated from the highly diverse parent Chicago strain, developed significant tprK diversification. These findings suggest that tprK variation can occur, but at a lower rate, in Nichols and that immune pressure may be required for accumulation of bacteria with diverse tprK sequences. Adaptation to growth in rabbits may explain the limited repertoire of V region sequences seen in the Nichols strain.

Amino Acid Sequence↗

Quantitative assessment of peptide sequence diversity in M13 combinatorial peptide phage display libraries.

Novel statistical methods have been developed and used to quantitate and annotate the sequence diversity within combinatorial peptide libraries on the basis of small numbers (1-200) of sequences selected at random from commercially available M13 p3-based phage display libraries. These libraries behave statistically as though they correspond to populations containing roughly 4.0+/-1.6% of the random dodecapeptides and 7.9+/-2.6% of the random constrained heptapeptides that are theoretically possible within the phage populations. Analysis of amino acid residue occurrence patterns shows no demonstrable influence on sequence censorship by Escherichia coli tRNA isoacceptor profiles or either overall codon or Class II codon usage patterns, suggesting no metabolic constraints on recombinant p3 synthesis. There is an overall depression in the occurrence of cysteine, arginine and glycine residues and an overabundance of proline, threonine and histidine residues. The majority of position-dependent amino acid sequence bias is clustered at three positions within the inserted peptides of the dodecapeptide library, +1, +3 and +12 downstream from the signal peptidase cleavage site. Conformational tendency measures of the peptides indicate a significant preference for inserts favoring a beta-turn conformation. The observed protein sequence limitations can primarily be attributed to genetic codon degeneracy and signal peptidase cleavage preferences. These data suggest that for applications in which maximal sequence diversity is essential, such as epitope mapping or novel receptor identification, combinatorial peptide libraries should be constructed using codon-corrected trinucleotide cassettes within vector-host systems designed to minimize morphogenesis-related censorship.

Amino Acids↗

PCR-based RFLP analysis of DNA sequence diversity in the gastric pathogen Helicobacter pylori.

DNA sequence diversity among 60 independent isolates of the gastric pathogen Helicobacter pylori was assessed by testing for restriction fragment length polymorphisms (RFLPs) in several PCR-amplified gene segments. 18 Mbol and 27 HaeIII RFLPs were found in the 2.4 kb ureA-ureB (urease) segment from the 60 strains; this identified 44 separate groups, with each group containing one to four isolates. With one exception, each isolate not distinguished from the others by RFLPs in ureA-ureB was distinguished by Mbol digestion of the neighboring 1.7 kb ureC-ureD segment. The 1.5 kb flaA (flagellin) gene, which is not close to ure gene cluster, was also highly polymorphic. In contrast, isolates from initial and followup biopsies yielded identical restriction patterns in each of the three cases tested. The potential of this method for detecting population heterogeneity was tested by mixing DNAs from different strains before amplification: the arrays of restriction fragments obtained indicated co-amplification from both genomes in each of the five pairwise combinations tested. These results show that H. pylori is a very diverse species, that indicate PCR-based RFLP tests are almost as sensitive as arbitrary primer PCR (RAPD) tests, and suggest that such RFLP tests will be useful for direct analysis of H. pylori in biopsy and gastric juice specimens.

Base Sequence↗

Heterogeneous geographic patterns of nucleotide sequence diversity between two alcohol dehydrogenase genes in wild barley (Hordeum vulgare subspecies spontaneum).

Patterns of nucleotide sequence diversity in the predominantly self-fertilizing species Hordeum vulgare subspecies spontaneum (wild barley) are compared between the putative alcohol dehydrogenase 3 locus (denoted "adh3") and alcohol dehydrogenase 1 (adh1), two related but unlinked loci. The data consist of a sequence sample of 1,873 bp of "adh3" drawn from 25 accessions that span the species range. There were 104 polymorphic sites in the sequenced region of "adh3." The data reveal a strong geographic pattern of diversity at "adh3" despite geographic uniformity at adh1. Moreover, levels of nucleotide sequence diversity differ by nearly an order of magnitude between the two loci. Genealogical analysis resolved two distinct clusters of "adh3" alleles (dimorphic sequence types) that coalesce roughly 3 million years ago. One type consists of accessions from the Middle East, and the other consists of accessions predominantly from the Near East. The two "adh3" sequence types are characterized by a high level of differentiation between clusters ( approximately 2.2%), which induces an overall excess of intermediate frequency variants in the pooled sample. Finally, there is evidence of intralocus recombination in the "adh3" data, despite the high level of self-fertilization characteristic of wild barley.

Afghanistan↗

Multiple forms of alpha2-macroglobulin from a bony fish, the common carp (Cyprinus carpio): striking sequence diversity in functional sites.

Unlike mammals, bony fish possess multiple genes encoding the complement component C3, a member of the alpha2-macroglobulin (alpha2M) protein family, presumably expanding the diversity of immune recognition. To examine whether the alpha2M gene has also duplicated and diverged in the bony fish lineage, cDNA cloning of alpha2M from a pseudotetraploid teleost, the common carp (Cyprinus carpio), was conducted and resulted in the isolation of three distinct alpha2M sequences from a single individual, indicating the presence of multiple alpha2M genes in this species. The deduced amino acid sequences contained a post-translational cleavage signal, predicting a C3-like two-chain structure, as in lamprey alpha2M. Two distinct alpha2M proteins were purified from carp serum; both proved to be Mr 380,000 dimers, the subunits of which are composed of disulfide-linked alpha chains (Mr 93,000) and beta chains (Mr 85,000), as reported for the alpha2M from plaice, another teleost species. The presence of an internal thioester in the alpha chain was demonstrated by its autolytic fragmentation and direct incorporation of [14C]methylamine. Interestingly, the three forms of carp alpha2M exhibited outstanding sequence diversity in the bait region which displays target sequences for various proteases, and in the C-terminal region of the alpha chain assigned as the receptor-binding domain, while an Asn residue at the position corresponding to the catalytic His in C3 was completely conserved in the carp alpha2Ms, as in most alpha2Ms of other animals. The possible functional significance of the sequence diversity is discussed.

Amino Acid Sequence↗

Conservation of a potential metal binding motif despite extensive sequence diversity in the rotavirus nonstructural protein NS53.

The nucleotide sequence for the simian rotavirus SA11 gene segment 5 has been determined. The gene is 1611 nucleotides in length and contains a single open reading frame of 1485 nucleotides. The segment codes for the nonstructural protein NS53 which is predicted to be a polypeptide of 495 amino acids with a molecular weight of 58,484. When compared to the sequence of bovine RF gene segment 5 there are homologies of only 49 and 36% at the nucleotide and amino acid levels, respectively. This is in marked contrast to the situation with other rotavirus nonstructural proteins which are highly conserved between isolates. Nevertheless, there is a conserved region between amino acids 37-81 which contains a generalized motif for a metal binding domain. All eight cysteine and two histidine residues in this short sequence are conserved between the simian and bovine NS53 proteins. The conservation of this domain despite extensive sequence diversity in the remainder of the protein suggests that this region is functionally important.

Amino Acid Sequence↗

Selective escape from CD8+ T-cell responses represents a major driving force of human immunodeficiency virus type 1 (HIV-1) sequence diversity and reveals constraints on HIV-1 evolution.

The sequence diversity of human immunodeficiency virus type 1 (HIV-1) represents a major obstacle to the development of an effective vaccine, yet the forces impacting the evolution of this pathogen remain unclear. To address this issue we assessed the relationship between genome-wide viral evolution and adaptive CD8+ T-cell responses in four clade B virus-infected patients studied longitudinally for as long as 5 years after acute infection. Of the 98 amino acid mutations identified in nonenvelope antigens, 53% were associated with detectable CD8+ T-cell responses, indicative of positive selective immune pressures. An additional 18% of amino acid mutations represented substitutions toward common clade B consensus sequence residues, nine of which were strongly associated with HLA class I alleles not expressed by the subjects and thus indicative of reversions of transmitted CD8 escape mutations. Thus, nearly two-thirds of all mutations were attributable to CD8+ T-cell selective pressures. A closer examination of CD8 escape mutations in additional persons with chronic disease indicated that not only did immune pressures frequently result in selection of identical amino acid substitutions in mutating epitopes, but mutating residues also correlated with highly polymorphic sites in both clade B and C viruses. These data indicate a dominant role for cellular immune selective pressures in driving both individual and global HIV-1 evolution. The stereotypic nature of acquired mutations provides support for biochemical constraints limiting HIV-1 evolution and for the impact of CD8 escape mutations on viral fitness.

Acute Disease↗

Sequence diversity of human rotavirus strains investigated by northern blot hybridization analysis.

Rotavirus genomic RNAs, derived from a series of human isolates that exhibit variability in the pattern of migration of the double-stranded RNA on polyacrylamide gels, were transferred to diazobenzyloxymethyl paper, and their sequence diversity was investigated. Hybridization of cDNA probes prepared from the 11 segments of rotavirus RNA indicated that considerable sequence diversity exists among these viruses. Under conditions of both low and high stringency, hybridization analysis of virus collected between 1975 and 1980 suggested that the variation among rotavirus strains may have occurred by a process involving both "drift" and "shift" in the sequence of the rotavirus genomic segments.

Base Sequence↗

Sequence diversity in the glycoprotein B gene complicates real-time PCR assays for detection and quantification of cytomegalovirus.

Real-time quantitative PCR systems (Q-PCR) for the rapid detection and quantification of microorganisms in clinical specimens employ oligodeoxyribonucleotide primers and probes for specificity, which makes them vulnerable to false negatives caused by sequence diversity in the template. Schaade et al. (J. Clin. Microbiol. 39:3809, 2001) reported a sequence variant (C630T) in the cytomegalovirus (CMV) glycoprotein B (gB) gene that, although detectable in their Q-PCR assay, could not be accurately quantified. In an effort to evaluate the impact of CMV sequence variants in our patient population by use of a similar Q-PCR assay, we surveyed 54 isolates of CMV, each from a different patient. We detected evidence for the C630T variant in 4 of 54 (7.4%) patients. Furthermore, isolates from two additional patients were completely negative in the test. Sequencing of these false-negative isolates revealed multiple mutations within the probe hybridization sites. A Q-PCR that targeted the CMV polymerase gene instead of gB detected all 54 isolates. We suggest that Q-PCR assays for viral load be rigorously tested on large panels of viral isolates to assess the impact of sequence diversity on detection as well as quantification.

Base Sequence↗

Sequence diversity of T-superfamily conotoxins from Conus marmoreus.

Remarkable sequence diversity of T-superfamily conotoxins was found in a mollusk-hunting cone snail Conus marmoreus. The sequence of mr5a purified from the snail venom was determined, while six other sequences of Mr5.1a, Mr5.1b, Mr5.2, Mr5.3, Mr5.4a, and Mr5.4b were deduced from their corresponding cDNA cloned by RACE approach. mr5a of 10 amino acid residues is one of the shortest T-superfamily conotoxins ever found. They all share a typical (-CC-CC-) Cys pattern, a conserved signal peptide and a long 3'-untranslated region. A consensus Glu residue is preceded by the second two adjacent cysteines in all these toxins except in mr5a, whereas Mr5.1a, Mr5.1b, Mr5.4a and Mr5.4b are abundant in Trp residues. The identification of these highly divergent T-superfamily conotoxins will facilitate the understanding the relationship of their structure and function.

Amino Acid Sequence↗

Nucleotide sequence diversity of hypervariable region 1 of hepatitis C virus in Japanese hemophiliacs with chronic hepatitis C and patients with chronic posttransfusion hepatitis C.

Hemophiliac patients with chronic hepatitis C might be exposed to and become infected with multiple hepatitis C virus (HCV) strains by means of frequent use of blood products, even if they are infected with a single subtype of HCV. To test this hypothesis, we analyzed the genetic diversity of hypervariable region 1 (HVR1) of HCV in chronically infected hemophiliacs and in patients with chronic posttransfusion hepatitis with a single HCV inoculation. The diversity of nucleotide sequences in HVR1 of serum HCV RNA was compared between 21 hemophiliacs infected with a single HCV subtype and 16 patients with posttransfusion HCV infection. The number of HCV quasispecies was determined by fluorescence single-strand conformation polymorphism (SSCP) analysis. Direct sequencing was performed to determine the diversity in HVR1. The number of HCV quasispecies in the blood was 5.2 +/- 2.0 clones in hemophiliacs and 4.0 +/- 2.3 clones in posttransfusion patients, a nonsignificant difference (P = .0943). The number of sites at which the nucleotide was not homogenous in all quasispecies was significantly higher in hemophiliacs (13.0% +/- 7.4%) than in posttransfusion hepatitis patients (2.7% +/- 2.8%; P < .0001). In conclusion, there was a high degree of genetic variation in HVR1 of HCV specimens isolated from hemophiliacs compared with posttransfusion patients. These findings indicate the possibility that multiple infections of a single HCV subtype may occur among patients frequently exposed to blood products; single HCV subtypes may therefore derive from multiple origins.

Adult↗

An efficient test for comparing sequence diversity between two populations.

We address the problem of comparing interindividual genomic sequence diversity between two populations. Although the methods are general, for concreteness we focus on comparing two human immunodeficiency virus (HIV) infected populations. From a viral isolate(s) taken from each individual in a sample of persons from each population, suppose one or multiple measurements are made on the genetic sequence of a coding region of HIV. Given a definition of genetic distance between sequences, the goal is to test if the distribution of interindividual distances differs between populations. If distances between all pairs of sequences within each group are used, then data-dependencies arising from the use of multiple sequences from individuals invalidates the use of a standard two-sample test such as the t-test. Where this problem has been recognized, a typical solution has been to apply a standard test to a reduced dataset comprised of one sequence or a consensus sequence from each patient. Disadvantages of this procedure are that the conclusion of the test depends on the choice of utilized sequences, often an arbitrary decision, and exclusion of replicate sequences from the analysis may needlessly sacrifice statistical power. We present a new test free of these drawbacks, which is based on a statistic that linearly combines all possible standard test statistics calculated from independent sequence subsamples. We describe statistical power advantages of the test and illustrate its use by application to nucleotide sequence distances measured from HIV-1 infected populations in southern Africa (GenBank accession numbers AF110959--AF110981) and North America/Europe. The test makes minimal assumptions, is maximally efficient and objective, and is broadly applicable.

Africa, Southern↗

Sequence diversity of Pseudomonas aeruginosa: impact on population structure and genome evolution.

Comparative sequencing of Pseudomonas aeruginosa genes oriC, citS, ampC, oprI, fliC, and pilA in 19 environmental and clinical isolates revealed the sequence diversity to be about 1 order of magnitude lower than in comparable housekeeping genes of Salmonella. In contrast to the low nucleotide substitution rate, the frequency of recombination among different P. aeruginosa genotypes was high, leading to the random association of alleles. The P. aeruginosa population consists of equivalent genotypes that form a net-like population structure. However, each genotype represents a cluster of closely related strains which retain their sequence signature in the conserved gene pool and carry a set of genotype-specific DNA blocks. The codon adaptation index, a quantitative measure of synonymous codon bias of genes, was found to be consistently high in the P. aeruginosa genome irrespective of the metabolic category and the abundance of the encoded gene product. Such uniformly high codon adaptation indices of 0.55 to 0.85 fit the ubiquitous lifestyle of P. aeruginosa.

Biological Evolution↗