Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “downstream ORF”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Identification and molecular characterization of two novel Trypanosoma cruzi genes encoding polypeptides sharing sequence motifs found in proteins involved in RNA editing reactions.

We have previously identified a Trypanosoma cruzi cDNA encoding a protein named Tc52 sharing structural and functional properties with the thioredoxin and glutaredoxin protein family involved in thiol-disulphide redox reactions. Furthermore, we reported that Tc52 also plays a role in T. cruzi-associated immunosuppression observed during Chagas' disease. Moreover, Tc52 gene targeting deletion strategy allowed us to demonstrate that monoallelic disruption of Tc52 resulted in the alteration of the metacyclogenesis process and the production of less virulent parasites. Sequence analysis of a 7358 bp genomic fragment containing the Tc52 encoding gene revealed two additional open reading frames (ORF-A and C). The ORFs are likely to have protein coding function by a number of criteria, including reverse transcriptase polymerase chain reaction (RT-PCR), Western blot and immunofluorescence analyses. The deduced amino-acid (aa) sequence of the ORF-A localized upstream of the Tc52 gene revealed that it contains within its N-terminus (aa 1 to 170) four RGG boxes known to act as RNA binding motifs in some proteins that interact with RNA, interspersed with a high density of glycine with regular spacing of tryptophan (WX(9-10)) in which X is often a glycine. Moreover, the C-terminal part of the ORF-C (aa 253-289) contains a motif that is strikingly similar (7-35% identity, 14-46% similarity over 28aa) to a short sequence (RNP1) comprising the consensus sequence RNA binding domain (CS-RBD) found in a number of proteins that interact with RNA. The aa sequence from the ORF-C localized downstream of the Tc52 gene showed significant homology to human adenosine deaminase acting on RNA (hADAT1) that specifically deaminates adenosine 37 to inosine in eukaryotic tRNA(Ala) and to its homologue yeast protein (Tad1p) (22-25% identity and an additional 38-40% similarity over 177aa). Moreover, highly similar motifs of the deaminase domain are present in the T. cruzi ORF-C. Furthermore, the 5' flanking regions of the genes contained repeat TATA and CAAT nucleotide sequences which resemble the motifs found upstream of the transcription initiation sites in eukaryotic promoters. Therefore, the characterization of novel T. cruzi genes encoding proteins which show similarity to components of RNA processing reactions provides new tools to investigate the gene expression regulation in these parasitic organisms. Moreover, our recent findings on the Tc52 encoding gene underline the interest of genetic manipulation of T. cruzi, not only making it possible to use more closely an in vitro approach to find out how genes function, but also to obtain 'attenuated' strains that could be used in the development of vaccinal strategies.

Adenosine Deaminase↗

Replication regions from plant-pathogenic Pseudomonas syringae plasmids are similar to ColE2-related replicons.

Many strains of the phytopathogen Pseudomonas syringae contain mutually compatible plasmids that share extensive regions of sequence homology and essential replication determinants. The replication regions of two compatible large plasmids involved in virulence or pathogenicity, pPT23A from P. syringae pv. tomato strain PT23 and pAV505 from P. syringae pv. phaseolicola strain HRI1302A, were isolated. DNA sequencing of the origins of replication revealed homologous ORFs, designated ORF-Pto and ORF-Pph, respectively. Both ORFs are 1311 bp long and encode peptides of 437 amino acids with predicted molecular masses of 48259 (Pto) and 48334 (Pph) Da. Expression of the two ORFs in Escherichia coli produced peptides of 50 kDa (Pto) and 56 kDa (Pph). The predicted peptides showed an overall identity of 897 %, being highly conserved from residues 1 to 373, but showing considerable variation in their C-terminal regions (50% identity over the last 64 aa). The two ORFs had significant similarity with the putative replication protein from plasmid pTiK12 of Thiobacillus intermedius and other CoIE2-related plasmids. However, both peptides were 100 residues longer than any of the known CoIE2-related rep sequences. Subcloning of fragments from the replication region of pPT23A revealed the presence of at least three incompatibility determinants, designated IncA, IncB and IncC. Partial sequencing of the region downstream of ORF-Pto revealed homology to the ru/AB genes, involved in UV resistance, from plasmid pPSR1. It is proposed that the replication origin of pPT23A serves as the prototype of a family of related plasmids.

Amino Acid Sequence↗

Metal-binding, nucleic acid-binding finger sequences in the CDC16 gene of Saccharomyces cerevisiae.

The CDC16 gene is involved in the process of chromosome segregation in mitosis and a cdc16ts mutant accumulates the predominant microtubule-associated protein at the nonpermissive temperature. We find that the CDC16 gene open reading frame (ORF) is capable of encoding a protein whose calculated molecular weight and pI are 94,967 and 6.60, respectively. This hypothetical protein contains 16 cysteine residues; five are clustered at the N-terminal, 4 are placed about 3 residues apart in the middle of the peptide, and 3 are located close to the C-terminal. Each of these could form a metal-binding, nucleic acid-binding domain, suggesting this protein acts either as a repressor of the microtubule-associated protein gene or as a component necessary for spindle elongation, possibly interacting with the DNA. The start of the CDC16 ORF is only 95 bp downstream from the end of the MAK11 ORF. In this region there are two TATA boxes in tandem, but there is no room for a UAS or other regulatory sequences. An ATG is present 5 bp upstream of the start of the large ORF. Its frame terminates after only two amino acids.

Amino Acid Sequence↗

Cloning, sequencing and genetic mapping of a Bacillus subtilis cell wall hydrolase gene.

We have cloned DNA fragments from Bacillus subtilis 168S into Escherichia coli, which produced a lytic zone on an agar medium containing B. subtilis cell wall. Sequencing of the fragments showed the presence of an open reading frame (ORF) which encodes a polypeptide of 272 amino acids with a molecular mass of 29919 Da. The deduced amino acid sequence showed considerable homology with that of the cell wall hydrolase gene of Bacillus sp. (Potvin, C., Leclerc, D., Tremblay, G., Asselin, A. & Bellemare, G. (1988). Molecular and General Genetics 214, 241-248). Accordingly, the gene was designated cwlA, for cell wall lysis. The N-terminal amino acid sequence of cwlA gene product prepared from a E. coli clone was AIKVVKNLVSKSKYGLKCPN, which is consistent with that of the deduced sequence starting from Ala at second position from the initiation codon of the cwlA gene. A presumed sigma A promoter and a rho-independent terminator were found upstream and downstream of the ORF, respectively. A chloramphenicol-resistance determinant integrated into the ORF was mapped by PBS1 transduction, which indicated the gene sequence dnaE-aroD-cwlA.

Amino Acid Sequence↗

The p6.5 gene region of a nuclear polyhedrosis virus of Orgyia pseudotsugata: DNA sequence and transcriptional analysis of four late genes.

The gene encoding the basic DNA-binding protein (p6.5) of the multicapsid nuclear polyhedrosis virus of Orgyia pseudotsugata (OpMNPV) was localized by Southern blot analysis using a cDNA probe containing the Autographa californica virus (AcMNPV) p6.9 gene. The OpMNPV p6.5 gene was mapped to the HindIII G fragment at map unit 67. Nucleotide sequence and transcriptional analysis of a 3.26 kb region encompassing this area revealed four open reading frames (ORFs 1 to 4) oriented in the same direction. ORF 1 demonstrated a seven codon overlap with ORF 2. Messenger RNAs initiated upstream of each of the four ORFs late in infection and were coterminal at a single site downstream of the fourth ORF. The conserved late gene promoter/mRNA start site sequence (ATAAG) was present upstream of all the ORFs, but did not appear to be the major site of mRNA initiation for the third ORF, as determined by primer extension analysis. The fourth ORF in this series encoded a predicted peptide of 51 amino acids (6.5K), which was 80% similar to the p6.9 basic DNA-binding protein of AcMNPV.

Amino Acid Sequence↗

Vaccinia virus encodes a family of genes with homology to serine proteinase inhibitors.

Nucleotide sequencing of a region of the vaccinia virus genome proximal to the right inverted terminal repeat (ITR) identified two open reading frames (ORFs) encoding proteins of 39K and 40K with amino acid homology to each other, to another vaccinia virus gene near the opposite end of the virus genome and to the superfamily of serine proteinase inhibitors (serpins). Serpins have now been found in poxviruses from the genera orthopox (cowpox and vaccinia viruses), leporipox (myxoma virus) and avipox (fowlpox virus). One of the vaccinia virus serpins identified here (B13R) shares 92% amino acid identity with the serpin from cowpox virus and 46% and 19% identity with vaccinia serpins B24R and K2L, respectively. The amino acid sequence of B13R reported here differs at 11 positions from a recently reported sequence and contains an additional three internal residues. The serpin genes near the right ITR are separated by 8 kb of DNA. Both genes contain early transcriptional termination signals just downstream of the ORFs and are transcribed in a rightward direction towards the end of the genome. Analysis of mRNAs from virus-infected cells demonstrated that all three vaccinia virus serpin genes are transcribed early during infection. The amino acid sequences at the active sites of these serpins suggest that they may inhibit serine proteinases of differing biochemical specificities. The possible functions of these genes are discussed.

Amino Acid Sequence↗

Three surface layer homology domains at the N terminus of the Clostridium cellulovorans major cellulosomal subunit EngE.

The gene engE, coding for endoglucanase E, one of the three major subunits of the Clostridium cellulovorans cellulosome, has been isolated and sequenced. engE is comprised of an open reading frame (ORF) of 3,090 bp and encodes a protein of 1,030 amino acids with a molecular weight of 111,796. The amino acid sequence derived from engE revealed a structure consisting of catalytic and noncatalytic domains. The N-terminal-half region of EngE consisted of a signal peptide of 31 amino acid residues and three repeated surface layer homology (SLH) domains, which were highly conserved and homologous to an S-layer protein from the gram-negative bacterium Caulobacter crescentus. The C-terminal-half region, which is necessary for the enzymatic function of EngE and for binding of EngE to the scaffolding protein CbpA, consisted of a catalytic domain homologous to that of family 5 of the glycosyl hydrolases, a domain of unknown function, and a duplicated sequence (DS or dockerin) at its C terminus. engE is located downstream of an ORF, ORF1, that is homologous to the Bacillus subtilis phosphomethylpyrimidine kinase (pmk) gene. The unique presence of three SLH domains and a DS suggests that EngE is capable of binding both to CbpA to form a CbpA-EngE cellulosome complex and to the surface layer of C. cellulovorans.

Amino Acid Sequence↗

Genes encoding proteins homologous to halobacterial Gvps N, J, K, F & L are located downstream of gvpC in the cyanobacterium Anabaena flos-aquae.

Only two gas vesicle genes have been previously identified in the cyanobacteria, gvpA and gvpC, both of which encode structural gas vesicle proteins. Analysis of the nucleotide sequence immediately downstream of gvpC in the cyanobacterium Anabaena flos-aquae has revealed the presence of 4 ORFs (open reading frames) the products of which share significant homology with a number of the gene products derived from halobacterial gvp gene clusters. In halobacteria the gas vesicle gene clusters consist of 14 genes involved in gas vesicle synthesis and assembly. The product of Anabaena ORF 1, located immediately downstream of gvpC is homologous to halobacterial GvpNs. For the remaining ORFs the predicted gene products show homology to both GvpJ and GvpA for ORF 2, to GvpK and GvpA for ORF 3, and to both GvpF and GvpL for ORF 4.

Amino Acid Sequence↗

Lambda clone B22 contains a 7676 bp genomic fragment of Saccharomyces cerevisiae chromosome VII spanning the VAM7-SPM2 intergenic region and containing three novel transcribed open reading frames.

A genomic clone of 7676 bp designated B22 from Saccharomyces cerevisiae has been sequenced. The 5' end matches the previously described gene, VAM7, and the 3' end matches the previously described gene, SPM2, both of which have been assigned to the left arm of chromosome VII. The intergenic region contains three transcribed open reading frames (ORFs). The first is related to an uncharacterized ORF of Bacillus subtilis and more weakly to MesJ in Escherichia coli; this is found as a single transcript of 1.1 kb by Northern blotting. The second ORF encodes a small ras-like GTPase of 222 residues with strong homology to yeast Ypt8p and to mammalian Rab11; this is found as a single transcript of 1.1 kb by Northern blotting. The third ORF generates a transcript of 1.6 kb and encodes a protein of 382 residues including a perfect match to the consensus sequence of a C2H2 zinc finger domain; it shares a strong homology with yeast Mig1p and Cre-A from Aspergillus, Emericella and E. coli. This ORF also has a striking similarity to a putative 43 kDa zinc finger protein encoded by an ORF (YEL8) immediately downstream of YPT8, raising the possibility that a region between VAM7 and SPM2 on chromosome VII arose as a duplication of the YPT8-YEL8 region of chromosome V, followed by a translocation.

Amino Acid Sequence↗

Cloning and preliminary characterization of lh3 gene encoding a putative acetyltransferase from a rifamycin SV-producing strain Amycolatopsis mediterranei.

An ORF located immediately downstream of glnR gene was cloned from Amycolatopsis mediterranei U32 and was named lh3. Sequence analysis revealed that lh3 encodes a putative acetyltransferase, which shows high amino acid sequence similarities to the mycothiol synthase (MshD) from other actinomycetes. For functional analysis, mutation in lh3 gene was generated by gene replacement with an apramycin resistance gene through homologous recombination. Compared with the wild type strain, the resulting mutant was more sensitive to H2O2, apramycin and erythromycin by two- to three-fold. These results suggest that the lh3 gene plays an important role in the course of detoxification in A. mediterranei U32.

Acetyltransferases↗

Structural and phylogenetic analysis of the MotA and MotB families of bacterial flagellar motor proteins.

MotA and MotB are two well-characterized proteins in Escherichia coli which are believed to function as the proton channel and the anchor, respectively, of the motor component of the bacterial flagellum. We have identified and analysed all currently sequenced members of the MotA and MotB families. Members of these families include (1) these E. coli proteins, (2) their pmf-interacting motor homologues in other bacteria, (3) two ORFs which map downstream of the gene encoding the catabolite repression-mediating CepA protein in Bacillus species and (4) unidentified open reading frames. With one exception (the MotB protein of Rhodobactec sphaeroides), members of the MotB family exhibit a C-terminal domain that is homologous to peptidoglycan-interaction domains of numerous sequenced lipoproteins and outer membrane proteins. Multiple alignments, average hydropathy and similarity plots, and phylogenetic trees have allowed (1) identification of regions of relative conservation, (2) definition of signature sequences for these protein families and (3) determination of relative phylogenetic distances relating all members of each family. The phylogenies of these proteins do not follow those of the organisms from which they were isolated, suggesting the presence of divergent isoforms in many bacteria. Phylogenetic analyses of the peptidoglycan-interaction domains of MotB proteins indicated that, except for MotB of R. sphaeroides, these domains became associated with the MotB proteins early during evolutionary history, before members of the MotB family or members of the outer membrane protein family diverged from each other.

Amino Acid Sequence↗

Dual functions of ribosome recycling factor in protein biosynthesis: disassembling the termination complex and preventing translational errors.

We summarize in this communication the data supporting the two functions of ribosome recycling factor (RRF, originally called ribosome releasing factor). The first described role involves the disassembly of the termination complex which consists of mRNA, tRNA and the ribosome bound to the mRNA at the termination codon. This process is catalyzed by two factors, elongation factor G (EF-G) and RRF. RRF stimulated protein synthesis as much as eight-fold in the in vitro lysozyme synthesis system, when ribosomes were limiting. In the absence of RRF, ribosomes remain mRNA-bound at the termination codon and translate downstream codons. In the in vitro system, the site of reinitiation is the triplet codon 3' to the termination codon. RRF is an essential protein for bacterial life. Temperature sensitive (ts) RRF mutants were isolated and in vivo translational reinitiation due to inactivation of ts RRF was demonstrated using the beta-galactosidase reporter gene placed downstream from the termination codon. A second function of RRF involves preventing errors in translation. In polyphenylalanine synthesis programmed by polyuridylic acid, misincorporation of isoleucine, leucine or a mixture of amino acids was stimulated upto 17-fold when RRF was omitted from the in vitro system. RRF did not influence the large error (10-fold increase) induced by streptomycin. This means that RRF participates not only in the disassembly of the termination complex but also in peptide elongation. Extending this concept and its conventional role for releasing ribosomes from mRNA, involvement of RRF in the reinitiation in the 3A' system (a construct using S aureus protein A, a collaborative work with Dr Isaksson), in programmed frame shifting, in trans-translation with 10Sa RNA (collaborative work with Dr Muto), and in the reinitiation downstream from the ORF A of the IS 3 (insertion sequence of a transposon, collaborative work with Dr Sekine) are discussed on the basis of preliminary data to be published elsewhere. Finally, we review the known RRF sequences from various organisms including eukaryotes and discuss the possible mechanism for disassembly of the eukaryotic termination complex.

Bacteria↗

A ColE1-type plasmid from Salmonella enteritidis encodes a DNA cytosine methyltransferase.

The multicopy plasmid pFM366 was isolated from a virulent Salmonella enteritidis strain and was found to code for DNA methylase activity (Ibáñez and Rotger, 1993). The present work was aimed at characterizing the genetic organization and functional features of this 5.6 kb plasmid. We found pFM366 almost identical to the plasmid P4 isolated from Shigella sonnei, that encodes the SsoII restriction-modification system (Karyagina et al., 1993), and related to other ColE1-type plasmids. Examination of these plasmids revealed a common organization which suggests they were the result of similar recombinational events. The cytosine methylase of pFM366 is nearly identical to M. SsoII, whereas the gene encoding the restrictase homologous to R. SsoII is truncated and its product is inactive. The expression of the cytosine methylase encoded by pFM366 is strongly affected by deletion of regions located upstream and downstream of its ORF, and is negatively controlled by the rpoS gene in Escherichia coli. The methylase activity encoded by pFM366 induces the SOS response, which could be responsible for the observed delay in the growth of E. coli.

Amino Acid Sequence↗

The opgGIH and opgC genes of Rhodobacter sphaeroides form an operon that controls backbone synthesis and succinylation of osmoregulated periplasmic glucans.

Osmoregulated periplasmic glucans (OPGs) of Rhodobacter sphaeroides are anionic cyclic molecules that accumulate in large amounts in the periplasmic space in response to low osmolarity of the medium. Their anionic character is provided by the substitution of the glucosidic backbone by succinyl residues. A wild-type strain was subject to transposon mutagenesis, and putative mutant clones were screened for changes in OPGs by thin layer chromatography. One mutant deficient in succinyl substitution of the OPGs was obtained and the gene inactivated in this mutant was characterized and named opgC. opgC is located downstream of three ORFs, opgGIH, two of which are similar to the Escherichia coli operon, mdoGH, governing OPG backbone synthesis. Inactivation of opgG, opgI or opgH abolished OPG production and complementation analysis indicated that the three genes are necessary for backbone synthesis. In contrast, inactivation of a gene similar to ndvB, encoding the OPG-glucosyl transferase in Sinorhizobium meliloti, had no consequence on OPG synthesis in Rhodobacter sphaeroides. Cassette insertions in opgH had a polar effect on glucan substitution, indicating that opgC is in the same transcription unit. Expression of opgIHC in E. coli mdoB/mdoC and mdoH mutants allowed the production of slightly anionic and abnormally long linear glucans.

Bacterial Proteins↗

Recruitment of mRNA cleavage/polyadenylation machinery by the yeast chromatin protein Sin1p/Spt2p.

The yeast chromatin protein Sin1p/Spt2p has long been studied, but the understanding of its function has remained elusive. The protein has sequence similarity to HMG1, specifically binds crossing DNA structures, and serves as a negative transcriptional regulator of a small family of genes that are activated by the SWI/SNF chromatin-remodeling complex. Recently, it has been implicated in maintaining the integrity of chromatin during transcription elongation. Here we present experiments whose results indicate that Sin1p/Spt2 is required for, and is directly involved in, the efficient recruitment of the mRNA cleavage/polyadenylation complex. This conclusion is based on the following findings: Sin1p/Spt2 frequently binds specifically downstream of many ORFs but almost always upstream of the first polyadenylation site. It directly interacts with Fir1p, a component of the cleavage/polyadenylation complex. Disruption of Sin1p/Spt2p results in foreshortened poly(A) tracts on mRNA. It is synthetically lethal with Cdc73p, which is involved in the recruitment of the complex. This report shows that a chromatin component is involved in 3' end processing of RNA.

3' Untranslated Regions↗

A method for construction of long randomized open reading frames and polypeptides.

A method is presented for construction of randomized open reading frame sequences (ORFs) and gene libraries containing them. The building blocks for the ORFs were 75 bp long DNA fragments generated by cloning sequences from a single synthetic oligonucleotide preparation by bridge mutagenesis. The fragments had the property that, regardless of their orientation in the ligated product, the ORF of the construct was maintained. The heterogeneity of the ORFs resulted from the random ligation of 2000 different DNA fragments. The randomized ORFs were cloned downstream from the lac promoter in a multicopy plasmid in Escherichia coli. To test the method, a library of 10(6) clones was constructed.

Amino Acid Sequence↗

The flagellin N-methylase gene fliB and an adjacent serovar-specific IS200 element in Salmonella typhimurium.

The cloning and molecular genetic analysis of a locus mapping within the flagellar gene (fli) complex of Salmonella typhimurium is reported. A copy of the insertion element IS200 was located in a noncoding stretch of DNA upstream of the fliA gene. Comparative nucleotide sequence analysis showed that this copy of IS200 was 711 bp long and that its flanking regions contained no features common to other characterized insertion sites of this element. The element was located 37 bp downstream of an ORF whose product was shown by interspecific transfer and amino acid analysis to carry out N-methylation of selected lysine residues in Salmonella flagellin. The sequence and phenotype of this ORF identified it as fliB, encoding the only prokaryotic N-methylase acting on amino groups to have been characterized to date. It was found to be conserved among all clinically significant serovars of Salmonella. The IS200 insertion site is of particular interest since it was conserved in all but two rare evolutionary lines of S. typhimurium, and was absent from 85 Salmonella strains belonging to 37 other serovars. It is thus a phylogenetically significant marker at the serovar level.

Amino Acid Sequence↗

A Clostridium difficile gene encoding flagellin.

Six strains of Clostridium difficile examined by electron microscopy were found to carry flagella. The flagella of these strains were extracted and the N-terminal sequences of the flagellin proteins were determined. Four of the strains carried the N-terminal sequence MRVNTNVSAL exhibiting up to 90% identity to numerous flagellins. Using degenerate primers based on the N-terminal sequence and the conserved C-terminal sequence of several flagellins, the gene encoding the flagellum subunit (fliC) was isolated and sequenced from two virulent strains. The two gene sequences exhibited 91% inter-strain identity. The gene consists of 870 nt encoding a protein of 290 amino acids with an estimated molecular mass of 31 kDa, while the extracted flagellin has an apparent molecular mass of 39 kDa on SDS-PAGE. The FliC protein displays a high degree of identity in the N- and C-terminal amino acids whereas the central region is variable. A second ORF is present downstream of fliC displaying homology to glycosyltransferases. The fliC gene was expressed in fusion with glutathione S-transferase, purified and a polyclonal monospecific antiserum was obtained. Flagella of C. difficile do not play a role in adherence, since the antiserum raised against the purified protein did not inhibit adherence to cultured cells. PCR-RFLP analysis of amplified flagellin gene products and Southern analysis revealed inter-strain heterogeneity; this could be useful for epidemiological and phylogenetic studies of this organism.

Amino Acid Sequence↗