Search PubMedSearch

SEARCH · Search PubMed

Results for “upstream ORF”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Translation of the downstream ORF from bicistronic mRNAs by human cells: Impact of codon usage and splicing in the upstream ORF.

Biochemistry textbooks describe eukaryotic mRNAs as monocistronic. However, increasing evidence reveals the widespread presence and translation of upstream open reading frames preceding the "main" ORF. DNA and RNA viruses infecting eukaryotes often produce polycistronic mRNAs and viruses have evolved multiple ways of manipulating the host's translation machinery. Here, we introduce an experimental model to study gene expression regulation from virus-like bicistronic mRNAs in human cells. The model consists of a short upstream ORF and a reporter downstream ORF encoding a fluorescent protein. We have engineered synonymous variants of the upstream ORF to explore large parameter space, including codon usage preferences, mRNA folding features, and splicing propensity. We show that human translation machinery can translate the downstream ORF from bicistronic mRNAs, albeit reporter protein levels are thousand times lower than those from the upstream ORF. Furthermore, synonymous recoding of the upstream ORF exclusively during elongation significantly influences its own translation efficiency, reveals cryptic splice signals, and modulates the probability of downstream ORF translation. Our results are consistent with a leaky scanning mechanism facilitating downstream ORF translation from bicistronic mRNAs in human cells, offering new insights into the role of upstream ORFs in translation regulation.

Humans

Escherichia coli orfE (upstream of pyrE) encodes RNase PH.

RNase PH from extracts of Escherichia coli was purified to homogeneity and subjected to NH2-terminal sequencing. Comparison of this sequence with all open reading frames in the GenBank data base revealed at least 95% identity to an unidentified open reading frame (orfE) upstream of pyrE at 81.7 min on the E. coli chromosome. Clones of orfE overexpress RNase PH activity, verifying that orfE encodes this ribonuclease. We suggest that orfE be renamed rph.

Amino Acid Sequence

Eukaryotic initiation factor (eIF)-4F. Implications for a role in internal initiation of translation.

In order to study the eukaryotic translation initiation mechanisms of "internal initiation," "re-initiation," and/or "coupled internal initiation," a series of model mRNAs have been constructed which contain two non-overlapping open reading frames (ORFs) that encode different lengths of rabbit alpha globin. These mRNAs, along with the bicistronic constructs TK/CAT and TK/P2CAT developed by Pelletier and Sonenberg (Pelletier, J., and Sonenberg, N. (1988) Nature 334, 320-325, 1988), were used to program an in vitro rabbit reticulocyte lysate translation system. Cap-dependent and cap-independent translation were distinguished by monitoring translation in the presence or absence of exogenously added cap analog (m7GTP). Messenger RNAs which translate both ORF1 and ORF2 by a cap-dependent mechanism, as well as mRNAs that translate ORF2 by a cap-independent mechanism while still translating ORF1 in a cap-dependent fashion have been obtained. These same alpha globin mRNAs differ by no more than 45 nucleotides in intercistronic length. Initiation factor addition studies were performed in this same in vitro translation system. Both eukaryotic initiation factor (eIF)-4F and, to a lesser extent, eIF-4B can stimulate translation of an internally located ORF independent of upstream ORF translation and in a manner not dependent on mRNA cap recognition. This indicates that the cap-recognition initiation factor, eIF-4F, and eIF-4B facilitate cap-independent and internal initiation of an open reading frame.

Animals

Sequence encoding ribosomal protein L33 of Lactococcus lactis.

A cloned fragment from Lactococcus lactis chromosome encoding the L33 ribosomal protein was sequenced. Two incomplete open reading frames (ORFs) were also found: the upstream ORF shows similarity to the tetracycline-resistance protein (Tet) of Bacillus stearothermophilus, and the downstream ORF shows homology to a protein of Bacillus subtilis participating in sporulation (SpoVE), and to proteins of Escherichia coli involved in cell division (FtsW) and the maintenance of cell shape (RodA).

Amino Acid Sequence

Cloning and characterization of the groESL operon from Bacillus subtilis.

The sequence of the 10 N-terminal amino acids of a Bacillus subtilis protein that cross-reacts with antibody to Escherichia coli GroEL was used to design a set of degenerate oligonucleotide probes. These probes identified a clone which carries almost the entire groESL operon from a B. subtilis subgenomic library. By chromosomal walking, an additional fragment carrying the 3' end of groESL and its flanking sequence was isolated. Sequence analysis revealed two open reading frames (ORFs) in the cloned DNA. The upstream ORF encodes a 10-kDa protein which has 47% amino acid identity with E. coli GroES. The downstream ORF encodes a 58-kDa protein which is 62% identical to E. coli GroEL. A 2.1-kb groESL mRNA from B. subtilis was detected independently by Northern (RNA) blot analyses with a groES- and a groEL-specific probe. This demonstrated that groES and groEL are in an operon. The groESL promoter was located by using a promoter-probing plasmid, and the apparent transcription start site was mapped by primer extension analysis. The same promoter is utilized under normal and heat shock conditions. This promoter has the same features as a typical sigma A promoter. A strain in which the groESL operon was under the control of the sucrose-inducible sacB promoter was created. With this strain, it was possible to show that both groES and groEL are essential genes under both normal and heat shock conditions.

Amino Acid Sequence

The transcriptional and translational landscape of HCoV-OC43 infection.

The coronavirus HCoV-OC43 circulates continuously in the human population and is a frequent cause of the common cold. Here, we generated a high-resolution atlas of the transcriptional and translational landscape of OC43 during a time course following infection of human lung fibroblasts. Using ribosome profiling, we quantified the relative expression of the canonical open reading frames (ORFs) and identified previously unannotated ORFs. These included several potential short upstream ORFs and a putative ORF nested inside the M gene. In parallel, we analyzed the cellular response to infection. Endoplasmic reticulum (ER) stress response genes were transcriptionally and translationally induced beginning 12 and 18 hours post infection, respectively. By contrast, conventional antiviral genes mostly remained quiescent. At the same time points, we observed accumulation and increased translation of noncoding transcripts normally targeted by nonsense mediated decay (NMD), suggesting NMD is suppressed during the course of infection. This work provides resources for deeper understanding of OC43 gene expression and the cellular responses during infection.

Humans

Clostridium difficile toxin A carries a C-terminal repetitive structure homologous to the carbohydrate binding region of streptococcal glycosyltransferases.

A detailed analysis of the 8130-bp open reading frame (ORF) of gene toxA and of an upstream ORF designated utxA, indicates the presence of a transcription terminator stem-loop for toxA, promoter sequences, and Shine-Dalgarno boxes for toxA and utxA. No transcription terminator between toxA and utxA is suggested by the sequence. ToxA contains two domains, one-third (C-terminal) with a repetitive structure and the residual two-thirds with no repetitions. The 2499-bp sequence encoding the repetitive structure is composed of nine groups of different short repetitive oligodeoxyribonucleotides (SRONs). A combination of these SRONs codes for five groups of combined repetitive oligopeptides (CROPs). Seven 50-amino acid (aa) CROPs and 23 CROPs of 21 aa in length are noticed. The CROPs are generally highly conserved, but four exhibit variability and possibly represent 'hot spots' of the repetitive structure. The reactivity of the C-terminal repeat with monoclonal antibody 1337C8 indicates that this part contains the carbohydrate-binding domain of ToxA. In this region homology exists between the ToxA repeats and the glucosyltransferases of Streptococci. We propose that binding of ToxA to cells occurs via the C-terminal repeat domain, with the N-terminal domain being responsible for toxic function.

Amino Acid Sequence

Molecular cloning and characterization of comC, a late competence gene of Bacillus subtilis.

comC is a Bacillus subtilis gene required for the development of genetic competence. We have cloned a fragment from the B. subtilis chromosome that carries comC and contains all the information required to complement a Tn917lac insertion in comC. Genetic tests further localized comC to a 2.0-kilobase HindIII fragment. Northern (RNA) blotting experiments revealed that an 800-base-pair comC-specific transcript appeared at the time of transition from exponential to stationary phase during growth through the competence regimen. The DNA sequence of the comC region revealed two open reading frames (ORFs), transcribed in the same direction. The upstream ORF encoded a protein with apparent sequence similarity to the folC gene of Escherichia coli. Insertion of a chloramphenicol resistance determinant into this ORF and integration of the disrupted construct into the bacterial chromosome by replacement did not result in competence deficiency. The downstream ORF, which contained the Tn917lac insertion that resulted in a lack of competence, is therefore the comC gene. The predicted protein product of comC consisted of 248 amino acid residues and was quite hydrophobic. The comC gene product was not required for the expression of any other com genes tested, and this fact, together with the marked hydrophobicity of ComC, suggests that it may be a component of the DNA-processing apparatus of competent cells.

Amino Acid Sequence

Role of the open reading frames of Rous sarcoma virus leader RNA in translation and genome packaging.

The Rous sarcoma virus (RSV) RNA leader sequence carries three open reading frames (uORFs) upstream of the AUG initiator of the gag gene. We studied, in vivo, the role of these uORFs by changing two or three nucleotides of the three AUGs or by deleting the first uORF. Our results show that (i) unlike most previously characterized uORFs, which decrease translation, the first uORF (AUG1) of RSV acts as an enhancer of translation, since absence of the first AUG decreased translation; AUG3 also modulates translation, probably by interfering with scanning ribosomes as described for other upstream ORFs, and mutation of AUG2 had no effect on translation. (ii) Mutation of each of the upstream AUGs lowered the infectivity of progeny virions. (iii) Unexpectedly, mutation of AUG1 and/or AUG3 dramatically reduced RNA packaging by 50-to 100-fold, unlike mutation of AUG2 which did not alter RNA packaging efficiency. Additional mutants in the vicinity of uORF1 and uORF3 were constructed in order to elucidate the mechanism by which uORFs affect RNA packaging: a translation model requiring uORFs 1 and 3, and involving ribosome pausing at AUG 3 is discussed.

Amino Acid Sequence

Distance-dependent translational coupling and interference in Lactococcus lactis.

The possibility of raising the expression level of a heterologous gene in Lactococcus lactis by exploiting the principle of translational coupling was investigated. For this purpose, the Escherichia coli lacZ gene was transcriptionally fused to a short open reading frame (ORF) of lactococcal origin. A Shine-Dalgarno (SD) sequence was introduced at the boundary of the two ORFs. In a series of otherwise identical plasmids, the relative positions of the translational stop codon of the upstream ORF and the translational start codon of the downstream ORF (lacZ) were varied. The expression of lacZ gradually increased as the stop and start codons were placed in closer proximity. A concomitant switch from translational interference to translational coupling was observed. Best results were obtained with partially overlapping stop and start codons. It is concluded that the principle of translational coupling offers good possibilities to increase the level of heterologous gene expression in L. lactis.

Amino Acid Sequence

Nucleotide sequence and transcript organization of a region of the vaccinia virus genome which encodes a constitutively expressed gene required for DNA replication.

A vaccinia virus (VV) gene required for DNA replication has been mapped to the left side of the 16-kilobase (kb) VV HindIII D DNA fragment by marker rescue of a DNA- temperature-sensitive mutant, ts17, using cloned fragments of the viral genome. The region of VV DNA containing the ts17 locus (3.6 kb) was sequenced. This nucleotide sequence contains one complete open reading frame (ORF) and two incomplete ORFs reading from left to right. Analysis of this region at early times revealed that transcription from the incomplete upstream ORF terminates coincidentally with the complete ORF encoding the ts17 gene product, which is directly downstream. The predicted proteins encoded by this region correlate well with polypeptides mapped by in vitro translation of hybrid-selected early mRNA. The nucleotide sequences of a 1.3-kb BglII fragment derived from ts17 and from two ts17 revertants were also determined, and the nature of the ts17 mutation was identified. S1 nuclease protection studies were carried out to determine the 5' and 3' ends of the transcripts and to examine the kinetics of expression of the ts17 gene during viral infection. The ts17 transcript is present at both early and late times postinfection, indicating that this gene is constitutively expressed. Surprisingly, the transcriptional start throughout infection occurs at the proposed late regulatory element TAA, which immediately precedes the putative initiation codon ATG. Although the biological activity of the ts17-encoded polypeptide was not identified, it was noted that in ts17-infected cells, expression of a nonlinked VV immediate-early gene (thymidine kinase) was deregulated at the nonpermissive temperature. This result may indicate that the ts17 gene product is functionally required at an early step of the VV replicative cycle.

Amino Acid Sequence

Sequence of the Bacillus subtilis glutamine synthetase gene region.

The nucleotide sequence of the glutamine synthetase (GS) region of Bacillus subtilis has been determined and found to contain several unique features. An open reading frame (ORF) upstream of the GS structural gene is part of the same operon as GS and is involved in regulation. Two downstream ORFs are separated from glnA by an apparent Rho-independent termination site. One of the downstream ORFs encodes a very hydrophobic polypeptide and contains its own potential RNA polymerase and ribosome-binding sites. The derived amino acid (aa) sequence of B. subtilis GS is similar to that of several other prokaryotes, especially to the GS of Clostridium acetobutylicum. The B. subtilis and C. acetobutylicum enzymes differ from the others in the lack of a stretch of about 25 aa as well as the presence of extra cysteine residues in a region known to contain regulatory as well as catalytic mutations. The region around the tyrosine residue that is adenylylated in GS from many species is fairly similar in the B. subtilis GS despite its lack of adenylylation.

Amino Acid Sequence

Partial conservation of the 5' ndhE-psaC-ndhD 3' gene arrangement of chloroplasts in the cyanobacterium Synechocystis sp. PCC 6803: implications for NDH-D function in cyanobacteria and chloroplasts.

The psaC gene, which encodes the 8.9 kDa iron-sulfur containing subunit of Photosystem I, has been sequenced from Synechocystis sp. PCC 6803 and shows greater similarity to reported plant sequences than other cyanobacterial psaC sequences. The deduced amino acid sequence of the protein encoded by the Synechocystis psaC gene is identical to the tobacco PSA-C sequence. In plants psaC is located in the small single-copy region of the chloroplast genome between two genes (designated ndhE and ndhD) with similarity to genes encoding subunits of the mitochondrial NADH Dehydrogenase Complex I. The 5' ndhE-psaC-ndhD3' gene arrangement of higher plants is only partially conserved in Synechocystis. An open reading frame (ORF) upstream of the Synechocystis psaC gene has 85% identity to the tobacco ndhE gene. Downstream of psaC there is a 273 bp ORF with 48% identity to the 5' portion of the tobacco ndhD gene (1527 bp). psaC, ndhE and the region of similarity to ndhD are present in a single copy in the Synechocystis genome. Part of the wheat ndhD gene was sequenced and used as a probe for the presence of the 3' portion of the ndhD gene. The wheat ndhD probe did not hybridize to Synechocystis or Anabaena sp. PCC 7120 genomic DNA, but did hybridize to Oenothera chloroplast DNA. These results indicate the complete ndhD gene is absent in two cyanobacteria, and raises the question of what role, if any, the ndhD gene product plays in the facultative heterotroph Synechocystis sp. PCC 6803.

Amino Acid Sequence

The precise structure and coding capacity of mRNAs from early region 2B of human adenovirus serotype 2.

Replication of human adenovirus (Ad) DNA requires three virus-encoded proteins that are coordinately transcribed from a single promoter at early times after infection. The mRNAs for two of these proteins, the precursor to the terminal protein (pTP) and the Ad DNA polymerase (Ad Pol), share several exons, including one encoded near Ad genome coordinate 39. The positions of the splice points of these mRNAs have been mapped by S1 nuclease mapping, by RNA sequencing, and by cDNA cloning. As a result of RNA splicing events, a short open reading frame (ORF) encoded at genome coordinate 39 is connected to the beginning of both the pTP and Ad Pol coding sequences; inclusion of this upstream ORF is essential for expression of functional pTP and Ad Pol proteins.

Adenoviruses, Human

Posttranscriptional trans-activation in cauliflower mosaic virus.

The ability of plant cells to translate dicistronic mRNAs that mimic a segment of the polycistronic 35S RNA from cauliflower mosaic virus has been tested. The chloramphenicol acetyltransferase and beta-glucuronidase open reading frames (ORFs) were fused in-frame to the second viral cistron (ORF I). Efficient reporter expression from the corresponding plasmids in plant protoplasts was observed only upon cotransfection with viral DNA. The trans-activating gene maps at ORF VI, which is expressed from a separate, monocistronic messenger (19S RNA). Deletion analysis shows that trans-activation selectively enhances downstream gene expression; the high expression of the upstream ORF is not further increased. The major reporter transcript remained bicistronic upon trans-activation, and its abundance varied only to a limited extent. Results indicate that trans-activation enhances the translation of downstream ORFs on polycistronic mRNAs derived from cauliflower mosaic virus.

Base Sequence

Cloning, sequence and characterization of m5C-methyltransferase-encoding gene, hgiDIIM (GTCGAC), from Herpetosiphon giganteus strain Hpa2.

We have cloned the gene (hgiDIIM) encoding the methyltransferase (MTase) of the SalI isoschizomeric restriction-modification (R-M) system, HgiDII (GTCGAC), into Escherichia coli. The hgiDIIM gene has been isolated from the same plasmid library of Herpetosiphon giganteus strain Hpa2, as was the previously cloned R-M system, HgiDI [AcyI/GRCGYC; Düsterhöft et al., Nucleic Acids Res. 19 (1991) 1049-1056]. Sequencing and functional localization of hgiDIIM revealed an open reading frame (ORF) of 354 codons (39786 Da) with significant homologies to the group of m5C-, rather than the m4C-/m6A-, MTases. Subsequent cloning and analysis of adjacent chromosomal segments led to the identification of two additional ORFs upstream (ORF15, 139 codons) and downstream (ORF68, 611 codons) from hgiDIIM with the same transcriptional orientation as the hgiDIIM gene. However, the expected restriction enzyme function was not found in either of these ORFs.

Amino Acid Sequence

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames