Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “downstream ORF”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Sequence analysis of the complete genome of rice black-streaked dwarf virus isolated from maize with rough dwarf disease.

The complete nucleotide sequences of 10 genomic segments (S1-S10) from an isolate of rice black-streaked dwarf virus causing rough dwarf disease on maize (RBSDV-Hbm) in China were determined, a total of 29,142 base pairs (bp). Each segment possessed the genus-specific termini with conserved nucleotide sequences of (+) 5'-AAGUUUUU......CAGCUNNNGUC-3' and a perfect or imperfect inverted repeat of seven to eleven nucleotides immediately adjacent to the terminal conserved sequence. While the coding strand of most RBSDV-Hbm segments contained one open reading frame (ORF), there were two non-overlapping ORFs in S7 and S9, and one small overlapping ORF downstream of the major ORF in S5. Homology comparisons suggest that S1 encodes a RNA-dependent RNA polymerase (RdRp), with 63.5% and 32.6% identity to the putative RdRp encoded by Fiji disease virus (FDV) and Nilaparvata lugens reovirus (NLRV), respectively. The proteins encoded by S2, S3, and S4 showed various degrees of similarity to those encoded by the corresponding segments of FDV or NLRV. In S5 and S6, low identities were found to those of FDV only, but not to NLRV. Sequence analyses showed that RBSDV-Hbm had the most similarities in the genome organizations and the coding assignments with a RBSDV isolated from rice in China, in which each pair of the corresponding segments shared sequence identities of 93.8-98.9% and 93.5-100% at nucleotide or amino acid levels, respectively. In addition, phylogenetic analyses suggested that RBSDV-Hbm had the closest evolutionary relationship to RBSDV in Fijivirus.

Conserved Sequence↗

Characterization of regulatory elements within the coat protein (CP) coding region of Tobacco mosaic virus affecting subgenomic transcription and green fluorescent protein expression from the CP subgenomic RNA promoter.

A replicon based on Tobacco mosaic virus that was engineered to express the open reading frame (ORF) of the green fluorescent protein (GFP) gene in place of the native coat protein (CP) gene from a minimal CP subgenomic (sg) RNA promoter was found to accumulate very low levels of GFP. Regulatory regions within the CP ORF were identified that, when presented as untranslated regions flanking the GFP ORF, enhanced or inhibited sg transcription and GFP expression. Full GFP expression from the CP sgRNA promoter required more than the first 20 nt of the CP ORF but not beyond the first 56 nt. Further analysis indicated the presence of an enhancer element between nt +25 and +55 with respect to the CP translation start site. The inclusion of this enhancer sequence upstream of the GFP ORF led to elevated sg transcription and to a 50-fold increase in GFP accumulation in comparison with a minimal CP promoter in which the entire CP ORF was displaced by the GFP ORF. Inclusion of the 3'-terminal 22 nt had a minor positive effect on GFP accumulation, but the addition of extended untranslated sequences from the 3' terminus of the CP ORF downstream of the GFP ORF was basically found to inhibit sg transcription. Secondary structure analysis programs predicted the CP sgRNA promoter to reside within two stable stem-loop structures, which are followed by an enhancer region.

Base Sequence↗

Cloning and sequencing of the genes encoding cyclic tetrasaccharide-synthesizing enzymes from Bacillus globisporus C11.

The genes for isomaltosyltransferase (CtsY) and 6-glucosyltransferase (CtsZ), involved in synthesis of a cyclic tetrasaccharide from alpha-glucan, have been cloned from the genome of Bacillus globisporus C11. The amino-acid sequence deduced from the ctsY gene is composed of 1093 residues having a signal sequence of 29 residues in its N-terminus. The ctsZ gene encodes a protein consisting of 1284 residues with a signal sequence of 35 residues. Both of the gene products show similarities to alpha-glucosidases belonging to glycoside hydrolase family 31 and conserve two aspartic acids corresponding to the putative catalytic residues of these enzymes. The two genes are linked together, forming ctsYZ. The DNA sequence of 16,515 bp analyzed in this study contains four open reading frames (ORFs) upstream of ctsYZ and one ORF downstream. The first six ORFs, including ctsYZ, form a gene cluster, ctsUVWXYZ. The amino-acid sequences deduced from ctsUV are similar in to a sequence permease and a sugar-binding protein for the sugar transport system from Thermococcus sp. B1001. The third ctsW encodes a protein similar to CtsY, suggested to be another isomaltosyltransferase preferring panose to high-molecular-mass substrates.

Amino Acid Sequence↗

Ribosome shunting in the cauliflower mosaic virus 35S RNA leader is a special case of reinitiation of translation functioning in plant and animal systems.

The shunt model predicts that small ORFs (sORFs) within the cauliflower mosaic virus (CaMV) 35S RNA leader and downstream ORF VII are translated by different mechanisms, that is, scanning-reinitiation and shunting, respectively. Wheat germ extract (WGE) and rabbit reticulocyte lysate (RRL) in vitro translation systems were used to discriminate between these two processes and to study the mechanism of ribosomal shunt. In both systems, expression downstream of the leader occurred via ribosomal shunt under the control of a stable stem and a small ORF preceding it. Shunting ribosomes were also able to initiate quite efficiently at non-AUG start codons just downstream of the shunt landing site in WGE but not in RRL. The short sORF MAGDIS from the mammalian AdoMetDC RNA, which conditionally suppresses reinitiation at a downstream ORF, prevented shunting if placed at the position of sORF A, the 5'-proximal ORF of the CaMV leader. We have demonstrated directly that sORF A is translated and that proper termination of translation at the 5'-proximal ORF is absolutely required for both shunting and linear ribosome migration. These findings strongly indicate that shunting is a special case of reinitiation.

Amino Acid Sequence↗

Molecular characterization of the vaccinia virus hemagglutinin gene.

The vaccinia virus hemagglutinin (HA) is a glycoprotein found on the plasma membrane of infected cells and the envelope of extracellular virus. Two forms of HA (85 and 68 kDa) are detected by immunoblot analysis. Although hemagglutination activity is only readily detectable late in infection, the 85-kDa HA appears early and accumulates throughout infection, whereas the 68-kDa form appears only late in the cycle. Production of the 68-kDa HA but not the 85-kDa HA was inhibited by either cytosine arabinoside or rifampin. Analysis of HA gene expression reveals a complex pattern of expression. The HA gene is transcribed early to yield a 1.65-kb dicistronic early transcript, consisting of the 945-bp HA open reading frame (ORF) fused to a 453-bp downstream ORF. Transcription from this site initiates 7 bases upstream of the AUG initiating codon of the HA ORF. Due to the discrepancy between the calculated size of the HA protein (33 kDa) and that reported for the unglycosylated HA protein derived from in vitro translation (58 kDa), we placed an early transcription termination signal (TTTTTAT) directly downstream of the 945-bp HA ORF. This led to a reduction in size of the early HA mRNA to 1.2 kb, as expected, but had no effect on the formation of either the 85- or 68-kDa protein. Transcripts originating from the early promoter are found throughout the infection cycle. However, after DNA replication, transcription from a second, late promoter ensues. The transcriptional start site of the late promoter is within a consensus TAAATG sequence located 135 bases upstream of the transcriptional start site of the first promoter. The late transcriptional start site is also found within an upstream ORF.

Amino Acid Sequence↗

Cloning and sequencing of kojibiose phosphorylase gene from Thermoanaerobacter brockii ATCC35047.

A gene encoding kojibiose phosphorylase was cloned from Thermoanaerobacter brockii ATCC35047. The kojP gene encodes a polypeptide of 775 amino acid residues. The deduced amino acid sequence was homologous to those of trehalose phosphorylase from T. brockii and maltose phosphorylases from Bacillus sp. and Lactobacillus brevis with 35%, 29% and 28% identities, respectively. Kojibiose phosphorylase was efficiently overexpressed in Escherichia coli JM109. The DNA sequence of 3956 bp analyzed in this study contains three open reading frames (ORFs) downstream of kojP. The four ORFs, kojP, kojE, kojF, and kojG, form a gene cluster. The amino acid sequences deduced from kojE and kojF are similar to those of the N-terminal and C-terminal regions of a sugar-binding periplasmic protein from Thermoanaerobacter tengcongensis MB4. Furthermore, the amino acid sequence deduced from kojG is similar to that of a permease of the ABC-type sugar transport systems from T. tengcongensis MB4. Each of three amino acid substitutions, D362N, K614Q and E642Q, caused a complete loss of kojibiose phosphorylase activity. These results suggest that D362, K614 and E642 play an important role in catalysis. Another mutation, D459N, increased K(m) values for kojibiose (7-fold that for the wild type), beta-G1P (11-fold) and glucose (7-fold), whereas K(m) for inorganic phosphate was minimally affected by this mutation, suggesting that D459 may be involved in the binding to saccharides.

Journal Article↗

Characterization of nra, a global negative regulator gene in group A streptococci.

During sequencing of an 11.5 kb genomic region of a serotype M49 group A streptococcal (GAS) strain, a series of genes were identified including nra(negative regulator of GAS). Transcriptional analysis of the region revealed that nra was primarily monocistronically transcribed. Polycistronic expression was found for the three open reading frames (ORFs) downstream and for the four ORFs upstream of nra. The deduced Nra protein sequence exhibited 62% homology to the GAS RofA positive regulator. In contrast to RofA, Nra was found to be a negative regulator of its own expression and that of the two adjacent operons by analysis of insertional inactivation mutants. By polymerase chain reaction and hybridization assays of 10 different GAS serotypes, the genomic presence of nra, rofA or both was demonstrated. Nra-regulated genes include the fibronectin-binding protein F2 gene (prtF2) and a novel collagen-binding protein (cpa). The Cpa polypeptide was purified as a recombinant maltose-binding protein fusion and shown to bind type I collagen but not fibronectin. In accordance with nra acting as a negative regulator of prtF2 and cpa, levels of attachment of the nra mutant strain to immobilized collagen and fibronectin was increased above wild-type levels. In addition, nra was also found to regulate negatively (four- to 16-fold) the global positive regulator gene, mga. Using a strain carrying a chromosomally integrated duplication of the nra 3' end and an nra-luciferase reporter gene transcriptional fusion, nra expression was observed to reach its maximum during late logarithmic growth phase, while no significant influence of atmospheric conditions could be distinguished clearly.

Adhesins, Bacterial↗

Cloning, sequencing, and expression of the genes encoding an isocyclomaltooligosaccharide glucanotransferase and an alpha-amylase from a Bacillus circulans strain.

The gene for a novel glucanotransferase, isocyclomaltooligosaccharide glucanotransferase (IgtY), involved in the synthesis of a cyclomaltopentaose cyclized by an alpha-1,6-linkage [ICG5; cyclo-{-->6)-alpha-D-Glcp-(1-->4)-alpha-D-Glcp-(1-->4)-alpha-D-Glcp-(1-->4)-alpha-D-Glcp-(1-->4)-alpha-D-Glcp-(1-->}] from starch, was cloned from the genome of B. circulans AM7. The IgtY gene, designated igtY, consisted of 2,985 bp encoding a signal peptide of 35 amino acids and a mature protein of 960 amino acids with a calculated molecular mass of 102,071 Da. The deduced amino-acid sequence showed similarities to 6-alpha-maltosyltransferase, alpha-amylase, and cyclomaltodextrin glucanotransferase. The four conserved regions common in the alpha-amylase family enzymes were also found in this enzyme, indicating that this enzyme should be assigned to this family. The DNA sequence of 8,325-bp analyzed in this study contained two open reading frames (ORFs) downstream of igtY. The first ORF, designated igtZ, formed a gene cluster, igtYZ. The amino-acid sequence deduced from igtZ exhibited no similarity to any proteins with known or unknown functions. IgtZ was expressed in Escherichia coli, and the enzyme was purified. The enzyme acted on maltooligosaccharides that have a degree of polymerization (DP) of 4 or more, amylose, and soluble starch to produce glucose and maltooligosaccharides up to DP5 by a hydrolysis reaction. The enzyme (IgtZ), which has a novel amino-acid sequence, should be assigned to alpha-amylase. It is notable that both IgtY and IgtZ have a tandem sequence similar to a carbohydrate-binding module belonging to a family 25. These two enzymes jointly acted on raw starch, and efficiently generated ICG5.

Amino Acid Sequence↗

Complete sequence of the citrus tristeza virus RNA genome.

The sequence of the entire genome of citrus tristeza virus (CTV), Florida isolate T36, was completed. The 19,296-nt CTV genome encodes 12 open reading frames (ORFs) potentially coding for at least 17 protein products. The 5'-proximal ORF 1a starts at nucleotide 108 and encodes a large polyprotein with calculated MW of 349 kDa containing domains characteristic of (from 5' to 3') two papain-like proteases (P-PRO), a methyltransferase (MT), and a helicase (HEL). Alignment of the putative P-PRO sequences of CTV with the related proteases of beet yellows closterovirus (BYV) and potyviruses allowed the prediction of catalytic cysteine and histidine residues as well as two cleavage sites, namely Val-Gly/Gly for the 5' proximal P-PRO domain and Met-Gly/Gly for the 5' distal P-PRO domain. The autoproteolytic cleavage of the polyprotein at these sites would release two N-terminal leader proteins of 54 and 55 kDa, respectively, and a 240-kDa C-terminal fragment containing MT and HEL domains. The apparent duplication of the leader domain distinguishes CTV from BYV and accounts for most of the size increase in the ORF 1a product of CTV. The downstream ORF 1b encodes a 57-kDa putative RNA-dependent RNA polymerase (RdRp), which is probably expressed via a +1 ribosomal frameshift. Sequence analysis of the frameshift region suggests that this +1 frameshift probably occurs at a rare arginine codon CGG and that elements of the RNA secondary structure are unlikely to be involved in this process. The complete polyprotein resulting from this frameshift event has a calculated MW of 401 kDa and after cleavage of the two N-terminal leaders would yield a 292-kDa protein containing the MT, HEL, and RdRp domains. Phylogenetic analysis of the three replication-associated domains, MT, HEL, and RdRp, indicates that CTV and BYV form a separate closterovirus lineage within the alpha-like supergroup of positive-strand RNA viruses. Two gene blocks or modules can be easily identified in the CTV genome. The first includes the replicative MT, HEL, and RdRp genes and is conserved throughout the entire alpha-like superfamily. The second block consists of five ORFs, 3 to 7, conserved among closteroviruses, including genes for the CTV homolog of HSP70 proteins and a duplicate of the coat protein gene. The 3'-terminal ORFs 8 to 11 encode a putative RNA-binding protein (ORF 11), and three proteins with unknown functions; this gene array is poorly conserved among closteroviruses.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

Molecular cloning and characterization of acrA and acrE genes of Escherichia coli.

The DNA fragment containing the acrA locus of the Escherichia coli chromosome has been cloned by using a complementation test. The nucleotide sequence indicates the presence of two open reading frames (ORFs). Sequence analysis suggests that the first ORF encodes a 397-residue lipoprotein with a 24-amino-acid signal peptide at its N terminus. One inactive allele of acrA from strain N43 was shown to contain an IS2 element inserted into this ORF. Therefore, this ORF was designated acrA. The second downstream ORF is predicted to encode a transmembrane protein of 1,049 amino acids and is named acrE. Genes acrA and acrE are probably located on the same operon, and both of their products are likely to affect drug susceptibilities observed in wild-type cells. The cellular localizations of these polypeptides have been analyzed by making acrA::TnphoA and acrE::TnphoA fusion proteins. Interestingly, AcrA and AcrE share 65 and 77% amino acid identity with two other E. coli polypeptides, EnvC and EnvD, respectively. Drug susceptibilities in one acrA mutant (N43) and one envCD mutant (PM61) have been determined and compared. Finally, the possible functions of these proteins are discussed.

Acriflavine↗

Isolation and preliminary characterization of the Streptococcus mutans rpsJ gene.

We previously reported that four putative open reading frames were identified in the regions flanking the Streptococcus mutans GS-5 fructosyltransferase gene. For one of these, ORF 4, only a small region had been isolated and the first 30 nucleotides had been sequenced. In order to determine whether this open reading frame is part of an expressed gene, we isolated a DNA fragment containing intact ORF 4 and a portion of the downstream ORF 5 by inverse polymerase chain reaction. A comparison of the deduced amino acid sequences of ORF 4 and ORF 5 with other proteins revealed that the ORF 4 and ORF 5 gene products were highly homologous to ribosomal proteins S10 and L3, respectively, of several bacteria. To identify the precise transcriptional start site for the ORF 4 gene, primer extension analysis was carried out. The results indicated initiation at a G residue with corresponding -10 and -35 regions homologous to the Escherichia coli consensus promoter sequences. These results indicate that the sequences of ORF 4 and ORF 5 are consistent with the structures of ribosomal proteins S10 and L3, respectively, and are present in a ribosomal protein operon.

Bacterial Proteins↗

Principles of 3' splice site selection and alternative splicing for an unusual group II intron from Bacillus anthracis.

We investigated the self-splicing properties of two introns from the bacterium Bacillus anthracis. One intron (B.a.I1) splices poorly in vitro despite having typical structural motifs, while the second (B.a.I2) splices well while having apparently degenerated features. The spliced exons of B.a.I2 were sequenced, and splicing was found to occur at a 3' site shifted one nucleotide from the expected position, thus restoring missing gamma-gamma' and IBS3-EBS3 pairings, but leaving the two conserved exonic ORFs out of frame. Because of the unexpected splice site, the principles for 3' intron definition were examined, which showed that the 3' splice site is flexible but contingent on gamma-gamma' and IBS3-EBS3 pairings, and can be as far away as four nucleotides from the wild-type site. Surprisingly, alternative splicing occurs at position +4 for wild-type B.a.I2 intron, both in vitro and in vivo, and the alternative event fuses the two conserved exon ORFs, presumably leading to translation of the downstream ORF. The finding suggests that the structural irregularities of B.a.I2 may be an adaptation to facilitate gene expression in vivo.

Alternative Splicing↗

Distance-dependent translational coupling and interference in Lactococcus lactis.

The possibility of raising the expression level of a heterologous gene in Lactococcus lactis by exploiting the principle of translational coupling was investigated. For this purpose, the Escherichia coli lacZ gene was transcriptionally fused to a short open reading frame (ORF) of lactococcal origin. A Shine-Dalgarno (SD) sequence was introduced at the boundary of the two ORFs. In a series of otherwise identical plasmids, the relative positions of the translational stop codon of the upstream ORF and the translational start codon of the downstream ORF (lacZ) were varied. The expression of lacZ gradually increased as the stop and start codons were placed in closer proximity. A concomitant switch from translational interference to translational coupling was observed. Best results were obtained with partially overlapping stop and start codons. It is concluded that the principle of translational coupling offers good possibilities to increase the level of heterologous gene expression in L. lactis.

Amino Acid Sequence↗

bilbo, a non-LTR retrotransposon of Drosophila subobscura: a clue to the evolution of LINE-like elements in Drosophila.

We used the repetitive character of transposable elements to isolate a non-LTR retrotransposon in Drosophila subobscura. bilbo, as we have called it, has homology to TRIM and LOA elements. Sequence analysis showed a 5' untranslated region (UTR), an open reading frame (ORF) with no RNA-binding domains, a downstream ORF that had structural homology to that of the I factor, and, finally, a 3' UTR which ended in several 5-nt repeats. The results of our phylogenetic and structural analyses shed light on the evolution of Drosophila non-LTR retrotransposons and support the hypothesis that an ancestor of these elements was structurally complex.

Amino Acid Sequence↗

DNA sequences of the tail fiber genes of bacteriophage P2: evidence for horizontal transfer of tail fiber genes among unrelated bacteriophages.

We have determined the DNA sequence of the bacteriophage P2 tail genes G and H, which code for polypeptides of 175 and 669 residues, respectively. Gene H probably codes for the distal part of the P2 tail fiber, since the deduced sequence of its product contains regions similar to tail fiber proteins from phages Mu, P1, lambda, K3, and T2. The similarities of the carboxy-terminal portions of the P2, Mu, ann P1 tail fiber proteins may explain the observation that these phages in general have the same host range. The P2 H gene product is similar to the products of both lambda open reading frame (ORF) 401 (stf, side tail fiber) and its downstream ORF, ORF 314. If 1 bp is inserted near the end of ORF 401, this reading frame becomes fused with ORF 314, creating an ORF that may represent the complete stf gene that encodes a 774-amino-acid-long side tail fiber protein. Thus, a frameshift mutation seems to be present in the common laboratory strain of lambda. Gene G of P2 probably codes for a protein required for assembly of the tail fibers of the virion. The entire G gene product is very similar to the products of genes U and U' of phage Mu; a region of these proteins is also found in the tail fiber assembly proteins of phages TuIa, TuIb, T4, and lambda. The similarities in the tail fiber genes of phages of different families provide evidence that illegitimate recombination occurs at previously unappreciated levels and that phages are taking advantage of the gene pool available to them to alter their host ranges under selective pressures.

Amino Acid Sequence↗

Analysis of the human cytomegalovirus genomic region from UL146 through UL147A reveals sequence hypervariability, genotypic stability, and overlapping transcripts.

BACKGROUND: Although the sequence of the human cytomegalovirus (HCMV) genome is generally conserved among unrelated clinical strains, some open reading frames (ORFs) are highly variable. UL146 and UL147, which encode CXC chemokine homologues are among these variable ORFs. RESULTS: The region of the HCMV genome from UL146 through UL147A was analyzed in clinical strains for sequence variability, genotypic stability, and transcriptional expression. The UL146 sequences in clinical strains from two geographically distant sites were assigned to 12 sequence groups that differ by over 60% at the amino acid level. The same groups were generated by sequences from the UL146-UL147 intergenic region and the UL147 ORF. In contrast to the high level of sequence variability among unrelated clinical strains, the sequences of UL146 through UL147A from isolates of the same strain were highly stable after repeated passage both in vitro and in vivo. Riboprobes homologous to these ORFs detected multiple overlapping transcripts differing in temporal expression. UL146 sequences are present only on the largest transcript, which also contains all of the downstream ORFs including UL148 and UL132. The sizes and hybridization patterns of the transcripts are consistent with a common 3'-terminus downstream of the UL132 ORF. Early-late expression of the transcripts associated with UL146 and UL147 is compatible with the potential role of CXC chemokines in pathogenesis associated with viral replication. CONCLUSION: Clinical isolates from two different geographic sites cluster in the same groups based on the hypervariability of the UL146, UL147, or the intergenic sequences, which provides strong evidence for linkage and no evidence for interstrain recombination within this region. The sequence of individual strains was absolutely stable in vitro and in vivo, which indicates that sequence drift is not a mechanism for the observed sequence hypervariability. There is also no evidence of transcriptional splicing, although multiple overlapping transcripts extending into the adjacent UL148 and UL132 open reading frames were detected using gene-specific probes.

Amino Acid Sequence↗

The complete nucleotide sequence and genome organization of red clover necrotic mosaic virus RNA-1.

The complete nucleotide sequence of red clover necrotic mosaic virus (RCNMV) RNA-1 has been determined. RNA-1 is 3889 nucleotides in length with a 5' terminal m7GpppA cap. The RNA contains three large open reading frames (ORFs): the 5' proximal ORF, encoding a 27-kDa polypeptide; the internal ORF, coding for a 57-kDa polypeptide; and the 3' terminal ORF, encoding the 37-kDa capsid protein. The sequence results confirm in vitro translation of 27-, 50-, and 37-kDa products but do not account for the observed 90-kDa product. A translational frameshift event from the 27- to the 57-kDa ORFs is proposed to explain the synthesis of the observed 90-kDa in vitro product. The putative translational frameshift region is structurally similar to several retrovirus frameshift regions and the putative barley yellow dwarf virus (BYDV) frameshift regions. Extensive amino acid homology was observed in the 57-kDa downstream ORF with the downstream domains of the carnation mottle virus (CarMV), turnip crinkle virus (TCV), maize chlorotic mottle virus (MCMV) readthrough, and BYDV fusion proteins. The 57-kDa ORF contained the conserved "GDD" motif. A significant alignment between the capsid proteins of RCNMV, CarMV, and TCV was also observed. Given the extensive amino acid sequence similarity of RCNMV, CarMV, and TCV polymerase and capsid proteins, we speculate that they are closely related, evolutionarily.

Amino Acid Sequence↗

Identification of a region of genetic variability among Bacillus anthracis strains and related species.

The identification of a region of sequence variability among individual isolates of Bacillus anthracis as well as the two closely related species, Bacillus cereus and Bacillus mycoides, has made a sequence-based approach for the rapid differentiation among members of this group possible. We have identified this region of sequence divergence by comparison of arbitrarily primed (AP)-PCR "fingerprints" generated by an M13 bacteriophage-derived primer and sequencing the respective forms of the only polymorphic fragment observed. The 1,480-bp fragment derived from genomic DNA of the Sterne strain of B. anthracis contained four consecutive repeats of CAATATCAACAA. The same fragment from the Vollum strain was identical except that two of these repeats were deleted. The Ames strain of B. anthracis differed from the Sterne strain by a single-nucleotide deletion. More than 150 nucleotide differences separated B. cereus and B. mycoides from B. anthracis in pairwise comparisons. The nucleotide sequence of the variable fragment from each species contained one complete open reading frame (ORF) (designated vrrA, for variable region with repetitive sequence), encoding a potential 30-kDa protein located between the carboxy terminus of an upstream ORF (designated orf1) and the amino terminus of a downstream ORF (designated lytB). The sequence variation was primarily in vrrA, which was glutamine- and proline-rich (30% of total) and contained repetitive regions. A large proportion of the nucleotide substitutions between species were synonymous. vrrA has 35% identity with the microfilarial sheath protein shp2 of the parasitic worm Litomosoides carinii.

Amino Acid Sequence↗