Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “downstream ORF”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Molecular cloning and sequence analysis of the genome of chicken anaemia agent.

The replicative form (RF) DNA of chicken anaemia agent (CAA) was isolated and cloned into bacterial plasmids. After religation of the cloned CAA DNA and transfection into MDCC-MSB1 cells, the DNA could induce c.p.e. characteristic of that caused by CAA, and an antigen was produced which gave positive immunofluorescence when detected with an anti-CAA serum. Sanger sequencing of the 2298 bp genome revealed several open reading frames (ORFs); the major ORF encoded a polypeptide of 51.8K. In SDS-PAGE of CAA viral particles a 50K protein has been reported as the only detectable viral protein. The genomic region downstream of the major ORF had several predicted GC-rich inverted repeats, a poly(A) signal and four copies of an 18 bp repeat element. Database searches did not reveal any sequence with homology to the viral genomic DNA, nor to the amino acid sequence of any of the ORFs, apart from the N-terminal 40 amino acids of the major ORF which showed a limited similarity to the structure of protamines.

Amino Acid Sequence↗

Cloning and characterization of genes responsible for metabolism of nitrile compounds from Pseudomonas chlororaphis B23.

The nitrile hydratase (NHase) of Pseudomonas chlororaphis B23, which is composed of two subunits, alpha and beta, catalyzes the hydration of nitrile compounds to the corresponding amides. The NHase gene of strain B23 was cloned into Escherichia coli by the DNA-probing method with the NHase gene of Rhodococcus sp. strain N-774 as the hybridization probe. Nucleotide sequencing revealed that an amidase showing significant similarity to the amidase of Rhodococcus sp. strain N-774 was also coded by the region just upstream of the subunit alpha-coding sequence. In addition to these three proteins, two open reading frames, P47K and OrfE, were found just downstream of the coding region of subunit beta. The direction and close locations to each other of these open reading frames encoding five proteins (amidase, subunits alpha and beta, P47K, and OrfE, in that order) suggested that these genes were cotranscribed by a single mRNA. Plasmid pPCN4, in which a 6.2-kb sequence covering the region coding for these proteins is placed under control of the lac promoter, directed overproduction of enzymatically active NHase and amidase in response to addition of isopropyl-beta-D-thiogalactopyranoside. Sodium dodecyl sulfate-polyacrylamide gel electrophoresis of the cell extract showed that the amount of subunits alpha and beta of NHase was about 10% of the total cellular proteins and that an additional 38-kDa protein probably encoded by the region upstream of the amidase gene was also produced in a large amount. The 38-kDa protein, as well as P47K and OrfE, appeared to be important for efficient expression of NHase activity in E. coli cells, because plasmids containing the NHase and amidase genes but lacking the region coding for the 38-kDa protein or the region coding for P47K and OrfE failed to express efficient NHase activity.

Amidohydrolases↗

Transcriptional analyses of the unique short segment of EHV-1 strain Kentucky A.

The unique short (Us) segment of the genome of equine herpesvirus type 1 (EHV-1) strain KyA is comprised of six open reading frames (ORFs) that encode: a) a homolog of the Us2 protein of herpes simplex virus type 1 (HSV-1); b) a serine threonine protein kinase that is a homolog of the HSV-1 Us3 protein; c) a homolog of pseudorabies virus glycoprotein gX and HSV-2 gG; d) a novel glycoprotein, EUS4, not encoded by other herpesviruses sequenced to date; e) a homolog of HSV-1 gD; and f) a homolog of HSV-1 Us9. The KyA strain is a deletion mutant that lacks Us sequences encoding gI, gE, and a potential 10 kD polypeptide, and thus may be useful as a parent virus for the generation of live virus vaccines. To complete the elucidation of the transcriptional program of the Us segment, Northern blot hybridization and S1 nuclease analyses were performed on poly(A)(+)-selected RNA isolated from infected cells maintained under early (phosphonoacetic acid-block) and late conditions. The findings revealed that the gene (EUS2 ORF) encoding the protein kinase is expressed as an early 2.9 kb transcript that overlaps and is 3' coterminal with a 1.6 kb early transcript that encodes the gG/gX homolog (EUS3 ORF). Two transcripts of 1.6 kb and 5.8 kb are 5' coterminal and may both encode the novel glycoprotein gene EUS4. The 1.6 kb transcript terminates at a poly(A) signal site downstream of the EUS4 ORF, and the 5.8 kb transcript terminates within the inverted repeat (IR) segment. Overall, the transcriptional program of the EHV-1 KyA Us segment is complex and exhibits similarities to that of HSV-1 Us segment: a) transcripts arise from both DNA strands; b) some transcripts, including those mapping at the termini of the Us segment, extend into the IR segments and are 3' coterminal with the 1.2 kb IR6 transcript; c) at least one transcript reads through a functional polyadenylation signal; d) some transcripts encoding genes that lie in different reading frames exist as a family of overlapping mRNAs, some in an anti-sense manner. Lastly, of the six Us genes of the EHV-1 KyA strain, only those encoding the EHV-1 protein kinase and the HSV-2 gG/gX homolog are members of the early kinetic class.

Base Sequence↗

ADP-ribosylating binary toxin genes of Clostridium difficile strain CCUG 20309.

The cdt genes that encode a binary ADP-ribosylating toxin in Clostridium difficile were first characterized from a toxigenic C. difficile strain CD196 in 1997. We report here C. difficile strain CCUG 20309 (ATCC 8864), a strain that produces toxin B but not toxin A, also carry a complete set of cdtA and cdtB genes. These genes were sequenced by cycle sequencing method. The 2 ORFs and the intergenic sequences of these 2 strains have a homology of 99.6%. Interestingly, 9 extra bases were found within the cdtA gene of strain CCUG 20309 which do not affect the downstream region of the ORF. Using the same homologous primers, the highly toxigenic reference strain VPI 10463 was found to carry only parts of the 2 ORFs while a nontoxigenic strain ATCC 8884 does not carry any of the cdt genes. Though it remains to be determined whether these genes are expressed, it is significant that strain CCUG 20309 contains the complete set of cdt genes. We speculate that the putative expressed proteins may contribute to pathogenesis, for example, enterotoxicity, of this unique strain of bacteria.

ADP Ribose Transferases↗

Analysis of large deletions in the Mauriceville and Varkud mitochondrial plasmids of Neurospora.

The Mauriceville and Varkud mitochondrial plasmids are closely related, closed-circular DNAs (3.6 and 3.7 kb, respectively) that have characteristics of mtDNA introns and retroid elements. Both plasmids contain a 710 amino acid open reading frame (ORF) that encodes an 81 kDa protein having reverse transcriptase activity. Here, we analyzed two mutant plasmids, V5-36 and M3-24, that have undergone relatively large deletions (approximately 0.35 and 0.5 kb, respectively). Both deletions occur downstream of the long ORF in a non-coding region of the plasmids that contains a direct repeat of 160 bp and a cluster of five PstI-palindromes, a repetitive sequence element in Neurospora mtDNA. In V5-36, the deletion end points are at the bases of two hairpin structures that are centered around PstI-palindromes and flank the deleted region. In M3-24, the deletion junction contains an extra T-residue that is not encoded in the plasmid. In both plasmids, the deletion end points do not correspond to homologous or directly repeated sequences of more than one nucleotide, whose pairing could account for the deletion junction. The characteristics of the deletion end points can be accounted for either by illegitimate recombination, possibly following double strand breaks at cruciform structures, or by interruption of reverse transcription followed by reinitiation downstream. The finding that the deletions encompass the 160 bp direct repeat and all five PstI-palindromes indicates that neither are required for propagation of the plasmids and supports the hypothesis that PstI-palindromes are selfish DNA elements that inserted into a nonessential region of the plasmid.

Amino Acid Sequence↗

Cloning, characterization, and expression of the murine cytomegalovirus homologue of the human cytomegalovirus 28-kDa matrix phosphoprotein (UL99).

We have identified, characterized, and expressed in bacteria and recombinant vaccinia viruses a protein which likely represents the murine cytomegalovirus (MCMV) homologue of the human cytomegalovirus (HCMV) 28-kDa matrix phosphoprotein, the product of the HCMV UL99 open reading frame (ORF). This protein, referred to as the MCMV UL99, is encoded by a 336-nucleotide ORF within the HindIII G fragment of MCMV strain Smith (K181). Using a DNA probe that corresponded to the amino terminus of the ORF, we detected a transcript of 4.8 kb at 8 hr and additional transcripts of 0.88, 2.4, and 5.7 kb at 24-48 hr postinfection (p.i.) of NIH 3T3 cells with MCMV. The smallest transcript is unspliced, initiates 235 nucleotides upstream from the start of the ORF, and utilizes a polyadenylation site located 62 nucleotides downstream from the end of the ORF. The ORF encodes a protein of 112 amino acids, with a predicted mass of 11.8 kDa. Comparison of the derived amino acid sequence with that of the HCMV UL99 gene product reveals 34.8% identity in an overlap of 66 amino acids. Within the amino acid sequence are at least two potential protein kinase C and one potential casein kinase II target motifs for phosphorylation. The ORF was cloned into the pGEX-KG prokaryotic expression vector and bacterially expressed protein was used to generate a specific rabbit antiserum against the protein. Western blotting of MCMV-infected NIH 3T3 cells showed that the ORF was expressed as a doublet of 16.3 and 15.2 kDa at 48 hr p.i. only in the absence of phosphonoacetic acid, thus demonstrating that this protein is a member of the true late gene class. The immunoreactive protein in MCMV-infected cells comigrated with that produced in cells infected with recombinant vaccinia virus containing the ORF. The protein appears to be part of the MCMV virion, is phosphorylated in vivo, and generates a strong humoral immune response following MCMV infection of BALB/c mice.

3T3 Cells↗

Baculovirus expression of proteins of porcine reproductive and respiratory syndrome virus strain Olot/91. Involvement of ORF3 and ORF5 proteins in protection.

Porcine reproductive and respiratory syndrome virus (PRRSV) is a new arterivirus that has spread rapidly all around the world in the last few years. The genomic region containing open reading frames (ORFs) 2 to 7 of PRRSV Spanish isolate Olot/91 was cloned and sequenced. The genomic sequence shared 95% identity with Lelystad and Tübingen isolates and between 61-64% with the ORF7 region of the American isolates. ORFs 2 to 7 were inserted into recombinant baculoviruses downstream of the polyhedrin promoter. Only ORFs 2, 3 5 and 7 were expressed in insect cells as detected by PRRS-specific pig antisera. To analyze the immunogenicity of these proteins and their ability to confer protection, Sf9 cells infected with recombinant baculoviruses expressing ORFs 3, 5 and 7 gene products were used to immunize pregnant sows, either individually or in combination. The results obtained indicate that ORFs 3 and 5 gene products could be major candidates for the development of a vaccine against PRRS since they conferred 68.4 and 50% protection, respectively, as evaluated by the number of piglets born alive and healthy at the time of weaning. In addition, piglets born to sows immunized with ORFs 3 and 5 proteins were seronegative to PRRSV after weaning, indicating absence of viral replication. ORF7 is the most immunogenic protein of PRRSV, but the antibodies induced in sows are non-protective and may even interfere with protection.

Animals↗

Members of the RTVL-H family of human endogenous retrovirus-like elements are expressed in placenta.

A cDNA clone homologous to the RTVL-H family of human retrovirus-like elements was isolated from a human placenta cDNA library. The nucleotide sequence of the 1084-bp cDNA revealed an open reading frame (ORF) that may encode a 146 amino acid protein with significant homology to retroviral proteases. Downstream from the putative protease ORF a 3' long terminal repeat (LTR) containing U3 and R regions was found. The cDNA sequence ends in a poly(A) tail appropriately positioned downstream from a polyadenylation signal in the LTR. Northern-blot analysis showed that several distinct RTVL-H homologous transcripts are present in human placenta. We also show that repetitive RTVL-H homologous sequences are present in the genomes of both gorilla and African green monkey.

Amino Acid Sequence↗

Unusual features of the retroid element PAT from the nematode Panagrellus redivivus.

The PAT retroid transposable elements differ from other retroids in that they have a 'split direct repeat' structure, i.e., and internal 300bp sequence is found repeated, about one half at each element extremity. A very abundant transcript of about 900 nt, the start of which maps to the preferentially deleted portion of PAT elements, is detected on total Panagrellus redivius RNA bearing Northern blots. A potentially corresponding ORF encodes a protein of 265 residues having a carboxy terminal Cystein motif, believed to be exclusively characteristic of the GAG protein in retoid elements. A much fainter, 1800nt long transcript, is also detected on Northern blots and maps slightly downstream of the first ORF. The predicted protein sequence of this region bears motifs typical of reverse transcriptase and RNaseH, as found in the Pol genes of retroid elements. Peptide motif similarities are greatest with the DIRS-1 element derived from Dictyostelium discoideum. The possibility of using PAT elements as transposon tagging system for Caenorhabditis elegans is discussed.

Amino Acid Sequence↗

Downstream ribosomal entry for translation of coronavirus TGEV gene 3b.

Gene 3b (ORF 3b) in porcine transmissible gastroenteritis coronavirus (TGEV) encodes a putative nonstructural polypeptide of 27.7 kDa with unknown function that during translation in vitro is capable of becoming a glycosylated integral membrane protein of 31 kDa. In the virulent Miller strain of TGEV, ORF 3b is 5'-terminal on mRNA 3-1 and is presumably translated following 5' cap-dependent ribosomal entry. For three other strains of TGEV, the virulent British FS772/70 and Taiwanese TFI and avirulent Purdue-116, mRNA species 3-1 is not made and ORF 3b is present as a non-overlapping second ORF on mRNA 3. ORF 3b begins at base 432 on mRNA 3 in Purdue strain. In vitro expression of ORF 3b from Purdue mRNA 3-like transcripts did not fully conform to a predicted leaky scanning pattern, suggesting ribosomes might also be entering internally. With mRNA 3-like transcripts modified to carry large ORFs upstream of ORF 3a, it was demonstrated that ribosomes can reach ORF 3b by entering at a distant downstream site in a manner resembling ribosomal shunting. Deletion analysis failed to identify a postulated internal ribosomal entry structure (IRES) within ORF 3a. The results indicate that an internal entry mechanism, possibly in conjunction with leaky scanning, is used for the expression of ORF 3b from TGEV mRNA 3. One possible consequence of this feature is that ORF 3b might also be expressed from mRNAs 1 and 2.

Animals↗

TGF-β1 regulation of human AT1 receptor mRNA splice variants harboring exon 2.

At least four alternatively spliced mRNAs can be synthesized from the human AT(1)R (hAT(1)R) gene that differ only in the inclusion or exclusion of exon 2 and/or 3. RT-PCR experiments demonstrate that splice variants harboring exon 2 accounts for at least 30% of all the hAT(1)R mRNA transcripts expressed in the human tissues investigated. Since exon 2 contains two upstream AUGs or open reading frames (uORFs), we hypothesized that these AUGs would inhibit the translation of the downstream hAT(1)R protein ORF harbored in exon 4. This study demonstrates that the inclusion of exon 2 in hAT(1)R mRNA transcripts dramatically reduces hAT(1)R protein levels (nine-fold) and significantly attenuates Ang II responsiveness ( approximately four-fold). Interestingly, only when both AUGs were mutated in combination were the hAT(1)R density and Ang II signaling levels comparable with those values obtained using mRNA splice variants that did not include exon 2. This observation is consistent with a model where the majority of the ribosomes likely translate uORF#1 and are then unable to reinitiate at the downstream hAT(1)R ORF, in part due to the presence of AUG#2 and to the short intercistronic spacing. Importantly, TGF-beta(1) treatment (4ng/ml for 4h) of fibroblasts up-regulated hAT(1)R mRNA splice variants, which harbored exon 2, six-fold. Since AT(1)R activation is closely associated with cardiovascular disease, the inclusion of exon 2 by alternative splicing represents a novel mechanism to reduce the overall production of the hAT(1)R protein and possibly limit the potential pathological effects of AT(1)R activation.

Alternative Splicing↗

Nucleotide sequence of the rubella virus capsid protein gene reveals an unusually high G/C content.

The nucleotide sequence of the rubella virus capsid protein (C) gene has been determined from a cDNA clone derived from the 40S genomic RNA. The sequence covers the coding region of the C protein (831 nucleotides), 70 nucleotides of the 5' untranslated region, and the 5' end of the downstream E2 membrane protein gene. The capsid gene is unusually rich in C (41.6%) and G (31.2%) residues (G + C 72.8%), and poor in A (15.4%) and U residues (11.8%). There are regions with long runs of up to 45% C or 35% G residues. The codon usage is non-random, with a strong preference for C and G residues in the third position. Starting from two in-frame AUG codons (seven amino acid residues apart) an open reading frame (ORF) was identified that extended in frame into the ORF coding for the downstream E2 membrane protein gene. Since the amino terminus of the capsid protein is blocked, we could not determine which of the AUGs serve as the initiating codon. To verify that the deduced ORF was correct, we have determined the amino acid sequence of 13 tryptic peptides corresponding to one-third of the C protein. Our data show that the C protein is about 277 residues in length (Mr about 30750). It is very hydrophilic and rich in prolines (14.1%) and arginines (14.4%). Clusters of these amino acids are concentrated in the amino-terminal third of the C protein. No sequence homology to the capsid protein of several alphaviruses was observed. Together with our previous sequence data we have now completed the sequence of the genes coding for the structural proteins C, E2 and E1 of rubella virus.

Amino Acid Sequence↗

A 250-nucleotide UA-rich element in the 3' untranslated region of Xenopus laevis Vg1 mRNA represses translation both in vivo and in vitro.

Xenopus laevis Vgl mRNA undergoes both localization and translational control during oogenesis. Vg1 protein does not appear until late stage IV, after localization is complete. To determine whether Vg1 translation is regulated by cytoplasmic polyadenylation, the RACE-PAT method was used. Vg1 mRNA has a constant poly(A) tail throughout oogenesis, precluding a role for cytoplasmic polyadenylation. To identify cis-acting elements involved in Vg1 translational control, the Vg1 3' UTR was inserted downstream of the luciferase ORF and in vitro transcribed, adenylated mRNA injected into stage III or stage VI oocytes. The Vg1 3' UTR repressed luciferase translation in both stages. Deletion analysis of the Vg1 3' UTR revealed that a 250-nt UA-rich fragment, the Vg1 translational element or VTE, which lies 118 nt downstream of the Vg1 localization element, could repress translation as well as the full-length Vg1 3' UTR. Poly(A)-dependent translation is not necessary for repression as nonadenylated mRNAs are also repressed, but cap-dependent translation is required as introduction of the classical swine fever virus IRES upstream of the luciferase coding region prevents repression by the VTE. Repression by the Vg1 3' UTR has been reproduced in Xenopus oocyte in vitro translation extracts, which show a 10-25-fold synergy between the cap and poly(A) tail. A number of proteins UV crosslink to the VTE including FRGY2 and proteins of 36, 42, 45, and 60 kDa. The abundance of p42, p45, and p60 is strikingly higher in stages I-III than in later stages, consistent with a possible role for these proteins in Vg1 translational control.

3' Untranslated Regions↗

Tissue specific expression of the retinoic acid receptor-beta 2: regulation by short open reading frames in the 5'-noncoding region.

The 40-S subunit of eukaryotic ribosomes binds to the capped 5'-end of mRNA and scans for the first AUG in a favorable sequence context to initiate translation. Most eukaryotic mRNAs therefore have a short 5'-untranslated region (5'-UTR) and no AUGs upstream of the translational start site; features that seem to assure efficient translation. However, approximately 5-10% of all eukaryotic mRNAs, particularly those encoding for regulatory proteins, have complex leader sequences that seem to compromise translational initiation. The retinoic-acid-receptor-beta 2 (RAR beta 2) mRNA is such a transcript with a long (461 nucleotides) 5'-UTR that contains five, partially overlapping, upstream open reading frames (uORFs) that precede the major ORF. We have begun to investigate the function of this complex 5'-UTR in transgenic mice, by introducing mutations in the start/stop codons of the uORFs in RAR beta 2-lacZ reporter constructs. When we compared the expression patterns of mutant and wild-type constructs we found that these mutations affected expression of the downstream RAR beta 2-ORF, resulting in an altered regulation of RAR beta 2-lacZ expression in heart and brain. Other tissues were unaffected. RNA analysis of adult tissues demonstrated that the uORFs act at the level of translation; adult brains and hearts of transgenic mice carrying a construct with either the wild-type or a mutant UTR, had the same levels of mRNA, but only the mutant produced protein. Our study outlines an unexpected role for uORFs: control of tissue-specific and developmentally regulated gene expression.

Amino Acid Sequence↗

Nucleotide sequence of the Escherichia coli rfe gene involved in the synthesis of enterobacterial common antigen. Molecular cloning of the rfe-rff gene cluster.

The genetic determinants of enterobacterial common antigen (ECA) include the rfe and rff genes located between ilv and cya near min 85 on the Escherichia coli chromosome. The rfe-rff gene cluster of E. coli K-12 was cloned in the cosmid pHC79. The cosmid clone complemented mutants defective in the synthesis of ECA due to lesions in the rfe, rffE, rffD, rffA, rffC, rffT, and rffM genes. Restriction endonuclease mapping combined with complementation studies of the original cosmid clone and six subclones revealed the order of genes in this region to be rfe-rffD/rffE-rffA/rffC-rffT-rffM . The rfe gene was localized to a 2.54-kilobase ClaI fragment of DNA, and the complete nucleotide sequence of this fragment was determined. The nucleotide sequencing data revealed two open reading frames, ORF-1 and ORF-2, located on the same strand of DNA. The putative initiation codon of ORF-1 was found to be 570 nucleotides downstream from the termination codon of rho. ORF-1 and ORF-2 specify putative proteins of 257 and 348 amino acids with calculated Mr values of 29,010 and 39,771, respectively. ORF-1 was identified as the rfe gene since ORF-1 alone was able to complement defects in the synthesis of ECA and 08-side chain synthesis in rfe mutants of E. coli. Data are also presented which suggest the possibility that the rfe gene is the structural gene for the tunicamycin sensitive UDP-GlcNAc:undecaprenylphosphate GlcNAc-1-phosphate transferase that catalyzes the synthesis of GlcNAc-pyrophosphorylundecaprenol (lipid I), the first lipid-linked intermediate involved in ECA synthesis.

Amino Acid Sequence↗

Topography of early HPV 16 transcription in high-grade genital precancers.

The extent to which human papillomavirus (HPV) type 16 is transcribed and the nature of the transcripts produced in genital precancers has not been clearly defined. The authors analyzed 28 cases of cervical (CIN) or vulvar (VIN) intraepithelial neoplasia by RNA-RNA in situ hybridization, using probes generated from HPV 16 open reading frames (ORFs) either upstream (E6-E7) or downstream (E2-E5-L2) to the E1 ORF, where HPV 16 genomic integration most commonly occurs. Hybridization signals corresponding to one or both probes were detected in a high proportion of cells throughout the lesional epithelium of low- and high-grade CIN, including basal layers. In serial sections analyzed with the two probes, hybridization signals were obtained from both, and in similar proportion, irrespective of CIN grade. The distribution and character of hybridization signals suggests that the morphologic progression of precancers is not associated with either cessation of HPV 16 early transcription or a change in the general character of the transcripts produced.

DNA, Neoplasm↗

Cloning and sequencing of a secY homolog from Streptomyces scabies.

The complete DNA sequence of the Streptomyces scabies (Ss) secY homolog and partial sequences of adjacent upstream and downstream open reading frames (ORFs) have been determined. The nucleotide sequence of a 2-kb region predicts a polypeptide of 437 amino acids in length with homology to the SecY protein family. The Ss secY homolog lies upstream from a sequence that has homology to the adenylate kinase gene (adk) family. The translational stop codon of the putative SecY ORF overlaps the predicted start codon for the Adk ORF. Another ORF that lies upstream from the secY homolog has sequence similarity to the genes that code for the L15 r-protein. Within the 243-bp intergenic region between the L15 and SecY coding sequences, the presence of a streptomycete-like promoter sequence and an 18-bp inverted repeat suggests that the secY homolog and the adjacent downstream sequences may be transcribed independently of the L15 coding sequence. Transcript analysis indicates that the secY homolog is expressed in both Ss and Streptomyces lividans. The proposed gene and transcript organization of the L15-SecY-Adk coding regions in the Ss clone resembles that of Micrococcus luteus which, like the streptomycetes, has a G+C-rich genome.

Adenylate Kinase↗

The p10 gene of Spodoptera littoralis nucleopolyhedrovirus: nucleotide sequence, transcriptional analysis and unique gene organization in the p10 locus.

The p10 gene of the Spodoptera littoralis (Spli) multicapsid nucleopolyhedrovirus (MNPV) was identified. With a coding sequence of 315 nucleotides (nt), corresponding to a protein of 104 amino acids, the SpliMNPV p10 gene is the longest p10 gene known. This gene codes for a putative protein with an Mr of 11130 and was found to be most closely related to the Spodoptera exigua (Se) MNPV p10 (49.4% amino acid identity) and most distant from the Autographa californica (Ac) MNPV p10 (20.0% amino acid identity). Characterization of the protein's secondary structure and a comparison with other p10 protein species suggested that this p10 has an extended alpha-helical domain with high probability of forming a large coiled-coil structure. The p10 mRNA was about 1500 nt long, as determined by Northern blot analysis. Primer extension assay mapped three transcription start sites to a conserved baculovirus late promoter motif, TAAG. In the SpliMNPV genome, the p10 gene is not flanked by genes similar to p26 and p74, as found in SeMNPV, AcMNPV, Choristoneura fumiferana MNPV and Orgyia pseudotsugata MNPV. Instead, an open reading frame (ORF) of 945 bp is located downstream from the p10 gene and is followed by another ORF in opposite orientation, encoding the p74 protein. Upstream of the p10 sequences, an ORF of 552 bp was identified that potentially encodes a 184 amino acid protein of Mr 20925, which showed 52.2% identity with the encoded product of the SeMNPV xb187 gene.

Amino Acid Sequence↗