Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “upstream ORF”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The relationship between eukaryotic translation and mRNA stability. A short upstream open reading frame strongly inhibits translational initiation and greatly accelerates mRNA degradation in the yeast Saccharomyces cerevisiae.

A new strategy was developed to study the relationship between the translation and degradation of a specific mRNA in the yeast Saccharomyces cerevisiae. A series of 5'-untranslated regions (UTR) was combined with the cat gene from the bacterial transposon Tn9, allowing us to test the influence of upstream open reading frames (uORFs) on translation and mRNA stability. The 5'-UTR sequences were designed so that the minimum possible sequence alteration, a single nucleotide substitution, could be used to create a 7-codon ORF upstream of the cat gene. The uORF was translated efficiently, but at the same time inhibited translation of the cat ORF and destabilized the cat mRNA. Investigations of various derivatives of the 5'-UTR indicated that cat translation was primarily attributable to leaky scanning of ribosomes past the uORF rather than to reinitiation. Therefore, these data directly demonstrate destabilization of a specific mRNA linked to changes in translational initiation on the same transcript. In contrast to the previously proposed nonsense-mediated mRNA decay pathway, destabilization was not triggered by premature translational termination in the main ORF and was not discernibly dependent upon a reinitiation-driven mechanism. This suggests the existence of an as yet not described pathway of translation-linked mRNA degradation.

Base Sequence↗

Three genes encoding for putative methyl- and acetyltransferases map adjacent to the wzm and wzt genes and are essential for O-antigen biosynthesis in Rhizobium etli CE3.

The elucidation of the structure of the O-antigen of Rhizobium etli CE3 predicts that the R. etli CE3 genome must contain genes encoding acetyl- and methyltransferases to confer the corresponding modifications to the O-antigen. We identified three open reading frames (ORFs) upstream of wzm, encoding the membrane component of the O-antigen transporter and located in the lps alpha-region of R. etli CE3. The ORFs encode two putative acetyltransferases with similarity to the CysE-LacA-LpxA-NodL family of acetyltransferases and one putative methyltransferase with sequence motifs common to a wide range of S-adenosyl-L-methionine-dependent methyltransferases. Mutational analysis of the ORFs encoding the putative acetyltransferases and methyltransferase revealed that the acetyl and methyl decorations mediated by these specific enzymes are essential for O-antigen synthesis. Composition analysis and high performance anion exchange chromatography analysis of the lipopolysaccharides (LPSs) of the mutants show that all of these LPSs contain an intact core region and lack the O-antigen polysaccharide. The possible role of these transferases in the decoration of the O-antigen of R. etli is discussed.

ATP-Binding Cassette Transporters↗

The major open reading frame of the beta2.7 transcript of human cytomegalovirus: in vitro expression of a protein posttranscriptionally regulated by the 5' region.

beta2.7 is the major early transcript produced during human cytomegalovirus infection. This abundantly expressed RNA is polysome associated, but no protein product has ever been detected. In this study, a stable peptide of 24 kDa was produced in vitro from the major open reading frame (ORF), TRL4. Following transient transfection, the intracellular localization was nucleolar and the expression was posttranscriptionally inhibited by the 5' sequence of the transcript, which harbors two short upstream ORFs.

Animals↗

Posttranscriptional trans-activation in cauliflower mosaic virus.

The ability of plant cells to translate dicistronic mRNAs that mimic a segment of the polycistronic 35S RNA from cauliflower mosaic virus has been tested. The chloramphenicol acetyltransferase and beta-glucuronidase open reading frames (ORFs) were fused in-frame to the second viral cistron (ORF I). Efficient reporter expression from the corresponding plasmids in plant protoplasts was observed only upon cotransfection with viral DNA. The trans-activating gene maps at ORF VI, which is expressed from a separate, monocistronic messenger (19S RNA). Deletion analysis shows that trans-activation selectively enhances downstream gene expression; the high expression of the upstream ORF is not further increased. The major reporter transcript remained bicistronic upon trans-activation, and its abundance varied only to a limited extent. Results indicate that trans-activation enhances the translation of downstream ORFs on polycistronic mRNAs derived from cauliflower mosaic virus.

Base Sequence↗

Cloning, sequence and characterization of m5C-methyltransferase-encoding gene, hgiDIIM (GTCGAC), from Herpetosiphon giganteus strain Hpa2.

We have cloned the gene (hgiDIIM) encoding the methyltransferase (MTase) of the SalI isoschizomeric restriction-modification (R-M) system, HgiDII (GTCGAC), into Escherichia coli. The hgiDIIM gene has been isolated from the same plasmid library of Herpetosiphon giganteus strain Hpa2, as was the previously cloned R-M system, HgiDI [AcyI/GRCGYC; Düsterhöft et al., Nucleic Acids Res. 19 (1991) 1049-1056]. Sequencing and functional localization of hgiDIIM revealed an open reading frame (ORF) of 354 codons (39786 Da) with significant homologies to the group of m5C-, rather than the m4C-/m6A-, MTases. Subsequent cloning and analysis of adjacent chromosomal segments led to the identification of two additional ORFs upstream (ORF15, 139 codons) and downstream (ORF68, 611 codons) from hgiDIIM with the same transcriptional orientation as the hgiDIIM gene. However, the expected restriction enzyme function was not found in either of these ORFs.

Amino Acid Sequence↗

ORFannotate: reproducible coding sequence annotation of transcriptome assemblies.

SUMMARY: Accurate annotation of coding sequences and translational features within transcript models is essential for interpreting assembled transcriptomes and their functional potential. Existing open reading frame (ORF) prediction tools typically operate on transcript FASTA files and do not reintegrate coding sequence (CDS) information back into transcript models, limiting their utility in long-read sequencing workflows where GTF/GFF annotations are the primary output. We present ORFannotate, a lightweight, GTF-native Python command-line tool that predicts ORFs from transcript annotations and reinserts precise, exon-aware CDS and UTR features into the original GTF/GFF file. In addition, ORFannotate provides biologically informative translational context by annotating Kozak sequence strength, detecting non-overlapping upstream ORFs (uORFs) with coding probabilities, characterising 5' and 3' untranslated regions (UTRs), and predicting nonsense-mediated decay (NMD) susceptibility. All annotations are consolidated in a transcript-level summary to support downstream analysis. By generating GTF files with accurate CDS annotations, ORFannotate facilitates reproducible analysis of both long- and short-read transcriptomes and integrates seamlessly with visualization tools, genome browsers, and comparative transcript analysis workflows. ORFannotate is fast, scalable and provides a practical solution for transcriptome annotation beyond coding potential prediction alone. AVAILABILITY AND IMPLEMENTATION: ORFannotate is implemented in Python and freely available under the GNU General Public License v3 (GPL-3.0) at: https://github.com/egustavsson/ORFannotate (DOI: https://doi.org/10.5281/zenodo.16812866).

Open Reading Frames↗

Transcriptional analysis of the mtA idiomorph of Neurospora crassa identifies two genes in addition to mtA-1.

In Neurospora crassa, mating and heterokaryon formation between opposite mating-types is controlled by a single locus with two alternate forms termed mt A and mt a. Previously, an open reading frame (mt A-1) that confers mating identity and heterokaryon incompatibility was characterized in the 5.3 kb mt A idiomorph. In this study, we describe the structural and transcriptional characterization of two additional genes in the mt A idiomorph, Mt A-2 and mt A-3. The 373 amino acid mt A-2 ORF has 23% identity to the SMR1 ORF of Podospora anserina. DNA sequence analysis of a mutation affecting ascospore to 129 amino acids. The 324 amino acids mt A-3 ORF has an HMG domain and shows 22% amino acid identity to SMR2 of P. anserina. Transcripts from mt A-2 and mt A-3 are constitutively expressed during both vegetative and sexual reproduction. The presence of upstream ORFs in the mt A-2 and mt A-3 transcripts suggests the possibility of post-transcriptional regulation of the expression mt A-2 and mt A-3 polypeptides.

Amino Acid Sequence↗

An operon encoding aspartokinase and purine phosphoribosyltransferase in Thermus flavus.

The nucleotide sequence of a 1.1 kb XhoI-HindIII fragment downstream of the malate dehydrogenase (mdh) gene of Thermus flavus revealed the presence of an ORF and an incomplete ORF lacking its NH2-terminal portion, in the opposite orientation to that of the mdh gene. These two genes overlapped with each other, sharing two base pairs, suggesting that these genes are co-transcribed in a single mRNA. One ORF (termed gpt) encoded a protein of 154 amino acids showing significant amino acid sequence similarity to purine phosphoribosyltransferases, such as xanthine-guanine phosphoribosyltransferase of Escherichia coli and human hypoxanthine phosphoribosyltransferase. Cloning and sequencing of the upstream region of the gpt gene, together with sequence comparison of the gene product encoded by the region upstream of gpt, suggested that the upstream ORF encoded two in-frame overlapping aspartokinase genes, askA, encoding the alpha-subunit of 405 amino acids, and askB, encoding the beta-subunit of 161 amino acids, which was part of the 3' portion of askA. Consistent with the sequence data, the askAB and the gpt genes conferred the heat-stable enzyme activities of aspartokinase and phosphoribosyltransferase, respectively, on E. coli. Preliminary characterization of these enzymes produced in E. coli is described.

Amino Acid Sequence↗

Sequence and expression of a halobacterial beta-galactosidase gene.

Studies of gene expression in haloarchaea have been greatly hindered by the lack of a convenient reporter gene. In a previous study, a beta-galactosidase from Haloferax alicantei was purified and several peptide sequences determined. The peptide sequences have now been used to clone the entire beta-galactosidase gene (designated bgaH) along with some flanking chromosomal DNA. The deduced amino acid sequence of BgaH was 665 amino acids (74 kDa) and showed greatest amino acid similarity to members of glycosyl hydrolase family 42 [classification of Henrissat, B., and Bairoch, A. (1993) New families in the classification of glycosyl hydrolases based on amino acid sequence similarities. Biochem J 293: 781-788]. Within this family, BgaH was most similar (42-43% aa identity) to enzymes from extremely thermophilic bacteria such as Thermotoga and Thermus. Family 42 enzymes are only distantly related to the Sulfolobus LacS and Escherichia coli LacZ enzymes (families one and two respectively). Three open reading frames (ORFs) upstream of bgaH were readily identified by database searches as glucose-fructose oxidoreductase, 2-dehydro-3-deoxyphosphogluconate aldolase and 2-keto-3-deoxygluconate kinase, enzymes that are also involved in carbohydrate metabolism. Downstream of bgaH there was an ORF which contained a putative fibronectin III motif. The bgaH gene was engineered into a halobacterial plasmid vector and introduced into Haloferax volcanii, a widely used strain that lacks detectable beta-galactosidase activity. Transformants were shown to express the enzyme; colonies turned blue when sprayed with Xgal and enzyme activity could be easily quantitated using a standard ONPG assay. In an accompanying publication, Patenge et al. (2000) have demonstrated the utility of bgaH as a promoter reporter in Halobacterium salinarum.

Amino Acid Sequence↗

Characterization of the cis-acting elements controlling subgenomic mRNAs of citrus tristeza virus: production of positive- and negative-stranded 3'-terminal and positive-stranded 5'-terminal RNAs.

Citrus tristeza virus (CTV), a member of the Closteroviridae, has an approximately 20-kb positive-sense RNA genome with two 5' ORFs translated from the genomic RNA and 10 3' genes expressed via nine or ten 3'-terminal subgenomic (sg) RNAs. The expression of the 3' genes appears to have properties intermediate between the smaller viruses of the "alphavirus supergroup" and the larger viruses of the Coronaviridae. The sgRNAs are contiguous with the genome, without a common 5' leader, and are associated with large amounts of complementary sgRNAs. Production of the different sgRNAs is regulated temporally and quantitatively, with the highly expressed genes having noncoding regions (NCR) 5' of the ORFs. The cis-acting elements that control the highly expressed major coat protein (CP) gene and the intermediately expressed minor coat protein (CPm) gene were mapped and compared. Mutational analysis showed that the CP sgRNA controller element mapped within nts -47 to -5 upstream of the transcription start site, entirely within the NCR, while the CPm control region mapped within a 57 nt sequence within the upstream ORF. Although both regions were predicted to fold into two stem-loop structures, mutagenesis suggested that primary structure might be more important than the secondary structure. Because each controller element produced large amounts of 3'-terminal positive- and negative-stranded sgRNAs, we could not differentiate whether the cis-acting element functioned as a promoter or terminator, or both. Reversal of the control element unexpectedly produced large amounts of a negative-stranded sgRNA apparently by termination of negative-stranded genomic RNA synthesis. Further examination of controller elements in their native orientation showed normal production of abundant amounts of positive-stranded sgRNAs extending to near the 5'-terminus, corresponding to termination at each controller element. Thus, each controller element produced three sgRNAs, a 5'-terminal positive strand and both positive- and negative-stranded 3'-terminal RNAs. Therefore, theoretically CTV could produce 30-33 species of RNAs in infected cells.

Base Sequence↗

Analysis of astrovirus serotype 1 RNA, identification of the viral RNA-dependent RNA polymerase motif, and expression of a viral structural protein.

We report the results from sequence analysis and expression studies of the gastroenteritis agent astrovirus serotype 1. We have cloned and sequenced 5,944 nucleotides (nt) of the estimated 7.2-kb RNA genome and have identified three open reading frames (ORFs). ORF-3, at the 3' end, is 2,361 nt in length and is fully encoded in both the genomic and subgenomic viral RNAs. Expression of ORF-3 in vitro yields an 87-kDa protein that is immunoprecipitated with a monoclonal antibody specific for viral capsids. This protein comigrates with an authentic 87-kDa astrovirus protein immunoprecipitated from infected cells, indicating that this region encodes a viral structural protein. The adjacent upstream ORF (ORF-2) is 1,557 nt in length and contains a viral RNA-dependent RNA polymerase motif. The viral RNA-dependent RNA polymerase motifs from four astrovirus serotypes are compared. Partial sequence (2,018 nt) of the most 5' ORF (ORF-1) reveals a 3C-like serine protease motif. The ORF-1 sequence is incomplete. These results indicate that the astrovirus genome is organized with nonstructural proteins encoded at the 5' end and structural proteins at the 3' end. ORF-2 has no start methionine and is in the -1 frame compared with ORF-1. We present sequence evidence for a ribosomal frameshift mechanism for expression of the viral polymerase.

Amino Acid Sequence↗

Rapid acquisition of entire DNA polymerase gene of a novel herpesvirus from green turtle fibropapilloma by a genomic walking technique.

A 4837-bp sequence of a newfound green turtle herpesvirus (GTHV), implicated in the etiology of green turtle fibropapilloma, was obtained from tumor tissues of a green turtle with fibropapilloma using a genomic walking method based on restriction enzyme digestion, self-ligation and inverse polymerase chain reaction (IPCR). The 4837-bp sequence was 56.23% G/C rich and contained three nonoverlapping open reading frames (ORF). The largest ORF (3507-bp) encoded the DNA polymerase gene (pol gene), which exhibited a high degree of homology at both amino acid and nucleotide levels with the DNA pol genes of human and animal herpesviruses, with a predicted protein of 1169 amino acids and molecular weight of 132.6 kilodaltons. The ATG at 518 to 520 was the first initiation codon in the ORF and was presumed to be the first methionine codon of the pol gene. Phylogenetic analysis, based on the amino acid sequence of the GTHV DNA pol gene and the corresponding regions of other known human and animal herpesviruses, indicated that GTHV belonged to the Alphaherpesvirinae subfamily. The upstream ORF of the pol gene encoded the N-terminal region of the GTHV homologue of the DNA-binding protein gene, whereas the downstream ORF was the C-terminal region of a gene which was homologous to ORFs conserved in human and animal herpesviruses, i.e., herpes simplex virus 1 (HSV1) gene UL31, Epstein-Barr virus (EBV) gene BFLF2, equine herpesvirus 1 (EHV1) gene 29, and alcelaphine herpesvirus 1 (AHV1) hypothetical protein 69 gene. The arrangement of these three genes in GTHV genome was identical to that seen in other alphaherpesviruses. The sequence and location of the DNA pol gene in the GTHV genome should greatly facilitate future studies of the viral life cycle.

Alphaherpesvirinae↗

Upstream AUGs in embryonic proinsulin mRNA control its low translation level.

Proinsulin is expressed prior to development of the pancreas and promotes cell survival. Here we study the mechanism affecting the translation efficiency of a specific embryonic proinsulin mRNA. This transcript shares the coding region with the pancreatic form, but presents a 32 nt extended leader region. Translation of proinsulin is markedly reduced by the presence of two upstream AUGs within the 5' extension of the embryonic mRNA. This attenuation is lost when the two upstream AUGs are mutated to AAG, leading to translational efficiency similar to that of the pancreatic mRNA. The upstream AUGs are recognized as initiator codons, because expression of upstream ORF is detectable from the embryonic transcript, but not from the mutated or the pancreatic mRNAs. Strict regulation of proinsulin biosynthesis appears to be necessary, since exogenous proinsulin added to embryos in ovo decreased apoptosis and generated abnormal developmental traits. A novel mechanism for low level proinsulin expression thus relies on upstream AUGs within a specific form of embryonic proinsulin mRNA, emphasizing its importance as a tightly regulated developmental signal.

3T3 Cells↗

Translational efficiency of polycistronic mRNAs and their utilization to express heterologous genes in mammalian cells.

The translation of polycistronic mRNAs in mammalian cells was studied. Transcription units, constructed to contain one, two or three open reading frames (ORFs), were introduced stably into Chinese hamster ovary cells and transiently into COS monkey cells. The analysis of mRNA levels and protein synthesis in these cells demonstrated that the mRNAs transcribed were translated to generate multiple proteins. The efficiency of translation was reduced approximately 40- to 300-fold by the insertion of an upstream ORF. The results support a modified 'scanning' model for translation initiation which allows for translation initiation at internal AUG codons. High-level expression of human granulocyte-macrophage colony stimulating factor was achieved utilizing a vector that contains a polycistronic transcription unit encoding an amplifiable dihydrofolate reductase marker gene in its 3' end. Thus, polycistronic expression vectors can be exploited to obtain high-level expression of foreign genes in mammalian cells.

Animals↗

The HrpZ proteins of Pseudomonas syringae pvs. syringae, glycinea, and tomato are encoded by an operon containing Yersinia ysc homologs and elicit the hypersensitive response in tomato but not soybean.

The Pseudomonas syringae pathovars are composed of host-specific plant pathogens that characteristically elicit the defense-associated hypersensitive response (HR) in nonhost plants. P. s. pv. syringae 61 secretes an HR elicitor, harpinPss (HrpZPss), in a hrp-dependent manner. An internal fragment of the P. s. pv. syringae 61 hrpZ gene was used to clone the hrpZ locus from P. s. pv. glycinea race 4 (bacterial blight of soybean) and P. s. pv. tomato DC3000 (bacterial speck of tomato). DNA sequence analysis revealed that hrpZ is the second ORF in a polycistronic operon. The amino acid sequence identities of HrpZPss/HrpZPsg and HrpZPss/HrpZPst were 79 and 63%, respectively. Although none of the HrpZ proteins showed significant overall sequence similarity with other known proteins, HrpZPst contained a 24-amino acid sequence that is homologous with a region of the PopA1 elicitor protein of the tomato pathogen, Pseudomonas solanacearum GMI1000. hrpA, the upstream ORF, was highly divergent: The amino acid sequence identities of HrpAPss/HrpAPsg and HrpAPss/HrpAPst were 91 and 28%, respectively, and no HrpA sequence showed similarity to known proteins. In contrast, the predicted products of the downstream ORFs in P. s. pv. syringae and P. s. pv. tomato, hrpB, hrpC, hrpD, and hrpE showed varying levels of similarity to those of yscI, yscJ, yscK, and yscL. These are colinearly arranged genes in the virC locus of Yersinia spp., which are involved in the secretion of the Yop virulence proteins via the type III pathway. The similarity of the Ysc proteins was generally stronger in comparisons with the P. s. pv. tomato Hrp proteins.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Distribution patterns of over-represented k-mers in non-coding yeast DNA.

MOTIVATION: Over-represented k-mers in genomic DNA regions are often of particular biological interest. For example, over-represented k-mers in co-regulated families of genes are associated with the DNA binding sites of transcription factors. To measure over-representation, we introduce a statistical background model based on single-mismatches, and apply it to the pooled 500 bp ORF Upstream Regions (USRs) of yeast. More importantly, we investigate the context and spatial distribution of over-represented k-mers in yeast USRs. RESULTS: Single and double-stranded spatial distributions of most over-represented k-mers are highly non-random, and predominantly cluster into a small number of classes that are robust with respect to over-representation measures. Specifically, we show that the three most common distribution patterns can be related to DNA structure, function, and evolution and correspond to: (a) homologous ORF clusters associated with sharply localized distributions; (b) regulatory elements associated with a symmetric broad hill-shaped distribution in the 50-200 bp USR; and (c) runs of As, Ts, and ATs associated with a broad hill-shaped distribution also in the 50-200 bp USR, with extreme structural properties. Analysis of over-representation, homology, localization, and DNA structure are essential components of a general data-mining approach to finding biologically important k-mers in raw genomic DNA and understanding the 'lexicon' of regulatory regions.

Amino Acid Motifs↗

Sequencing, chromosomal inactivation, and functional expression in Escherichia coli of ppsR, a gene which represses carotenoid and bacteriochlorophyll synthesis in Rhodobacter sphaeroides.

Sequencing of a DNA fragment that causes trans suppression of bacteriochlorophyll and carotenoid levels in Rhodobacter sphaeroides revealed two genes: orf-192 and ppsR. The ppsR gene alone is sufficient for photopigment suppression. Inactivation of the R. sphaeroides chromosomal copy of ppsR results in overproduction of both bacteriochlorophyll and carotenoid pigments. The deduced 464-amino-acid protein product of ppsR is homologous to the CrtJ protein of Rhodobacter capsulatus and contains a helix-turn-helix domain that is found in various DNA-binding proteins. Removal of the helix-turn-helix domain renders PpsR nonfunctional. The promoter of ppsR is located within the coding region of the upstream orf-192 gene. When this promoter is replaced by a lacZ promoter, ppsR is expressed in Escherichia coli. An R. sphaeroides DNA fragment carrying crtD', -E, and -F and bchC, -X, -Y, and -Z' exhibited putative promoter activity in E. coli. This putative promoter activity could be suppressed by PpsR in both E. coli and R. sphaeroides. These results suggest that PpsR is a transcriptional repressor. It could potentially act by binding to a putative regulatory palindrome found in the 5' flanking regions of a number of R. sphaeroides and R. capsulatus photosynthesis genes.

Aerobiosis↗

Translational regulation of ornithine decarboxylase and other enzymes of the polyamine pathway.

It has long been known that polyamines play an essential role in the proliferation of mammalian cells, and the polyamine biosynthetic pathway may provide an important target for the development of agents that inhibit carcinogenesis and tumor growth. The rate-limiting enzymes of the polyamine pathway, ornithine decarboxylase (ODC) and S-adenosylmethionine decarboxylase (AdoMetDC), are highly regulated in the cell, and much of this regulation occurs at the level of translation. Although the 5' leader sequences of ODC and AdoMetDC are both highly structured and contain small internal open reading frames (ORFs), the regulation of their translation appears to be quite different. The translational regulation of ODC is more dependent on secondary structure, and therefore responds to the intracellular availability of active eIF-4E, the cap-binding subunit of the eIF-4F complex, which mediates translation initiations. Cell-specific translation of AdoMetDC appears to be regulated exclusively through the internal ORF, which causes ribosome stalling that is independent of eIF-4E levels and decreases the efficiency with which the downstream ORF encoding AdoMetDC protein is translated. The translation of both ODC and AdoMetDC is negatively regulated by intracellular changes in the polyamines spermidine and spermine. Thus, when polyamine levels are low, the synthesis of both ODC and AdoMetDC is increased, and an increase in polyamine content causes a corresponding decrease in protein synthesis. However, an increase in active eIF-4E may allow for the synthesis of ODC even in the presence of polyamine levels that repress ODC translation in cells with lower levels of the initiation factor. In contrast, the amino acid sequence that is encoded by the upstream ORF is critical for polyamine regulation of AdoMetDC synthesis and polyamines may affect synthesis by interaction with the putative peptide, MAGDIS.

Adenosylmethionine Decarboxylase↗