Search PubMed⌕ Search

Biomedical subjects

Pavel V Baranov

Publications and source records attributed to Pavel V Baranov.

16 recordsLinked to original sources

Expanding the human proteome with microproteins and peptideins.

A major scientific drive is to characterize the protein-coding genome, which is a primary basis for studying human health. But the fundamental question remains of what has been missed in previous analyses. Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states1-3, with major implications for biomedical science. However, a key gap in knowledge has been which ncORFs produce small microproteins or alternative protein molecules that contribute to the human proteome. Here we report the collaborative efforts of the TransCODE Consortium4 to produce a consensus landscape of protein-level evidence for ncORFs. We show that about 25% of a set of 7,264 ncORFs gives rise to detectable peptides in a large-scale analysis of 95,520 proteomics experiments. We develop an annotation framework for ncORF-encoded microproteins as human proteins and codify the new conceptual model of 'peptideins' as microproteins that have indeterminate potential as functional proteins. To probe the biological implications of peptideins, we create an evolutionary analysis approach, termed ORF relative branch length (ORBL), and determine that evolutionary constraint is common and associates with observation of ncORF-derived peptides. We then characterize a pan-essential cellular phenotype for one peptidein from the OLMALINC long non-coding RNA. Overall, we generate public research tools supported by GENCODE and PeptideAtlas and advance biomedical discovery for understudied components of the human proteome.

Humans↗

Programmed ribosomal frameshifting during PLEKHM2 mRNA decoding generates a constitutively active proteoform that supports myocardial function.

Programmed ribosomal frameshifting is a process where a proportion of ribosomes change their reading frame on an mRNA. While frameshifting is commonly used by viruses, very few phylogenetically conserved examples are known in nuclear encoded genes. Here, we report a +1 frameshifting event during decoding of the human gene PLEKHM2 that provides access to a second internally overlapping ORF. The new carboxyl-terminal domain of this frameshift protein forms an α helix, which relieves PLEKHM2 from autoinhibition and allows it to move to the tips of cells without activation by ARL8. Reintroducing both the canonically translated and frameshifted protein are necessary to restore normal contractile function of PLEKHM2 knockout cardiomyocytes, demonstrating the necessity of frameshifting for normal cardiac activity.

Frameshifting, Ribosomal↗

Diverse bacterial genomes encode an operon of two genes, one of which is an unusual class-I release factor that potentially recognizes atypical mRNA signals other than normal stop codons.

BACKGROUND: While all codons that specify amino acids are universally recognized by tRNA molecules, codons signaling termination of translation are recognized by proteins known as class-I release factors (RF). In most eukaryotes and archaea a single RF accomplishes termination at all three stop codons. In most bacteria, there are two RFs with overlapping specificity, RF1 recognizes UA(A/G) and RF2 recognizes U(A/G)A. THE HYPOTHESIS: First, we hypothesize that orthologues of the E. coli K12 pseudogene prfH encode a third class-I RF that we designate RFH. Second, it is likely that RFH responds to signals other than conventional stop codons. Supporting evidence comes from the following facts: (i) A number of bacterial genomes contain prfH orthologues with no discernable interruptions in their ORFs. (ii) RFH shares strong sequence similarity with other class-I bacterial RFs. (iii) RFH contains a highly conserved GGQ motif associated with peptidyl hydrolysis activity (iv) residues located in the areas supposedly interacting with mRNA and the ribosomal decoding center are highly conserved in RFH, but different from other RFs. RFH lacks the functional, but non-essential domain 1. Yet, RFH-encoding genes are invariably accompanied by a highly conserved gene of unknown function, which is absent in genomes that lack a gene for RFH. The accompanying gene is always located upstream of the RFH gene and with the same orientation. The proximity of the 3' end of the former with the 5' end of the RFH gene makes it likely that their expression is co-regulated via translational coupling. In summary, RFH has the characteristics expected for a class-I RF, but likely with different specificity than RF1 and RF2. TESTING THE HYPOTHESIS: The most puzzling question is which signals RFH recognizes to trigger its release function. Genetic swapping of RFH mRNA recognition components with its RF1 or RF2 counterparts may reveal the nature of RFH signals. IMPLICATIONS OF THE HYPOTHESIS: The hypothesis implies a greater versatility of release-factor like activity in the ribosomal A-site than previously appreciated. A closer study of RFH may provide insight into the evolution of the genetic code and of the translational machinery responsible for termination of translation. REVIEWERS: This article was reviewed by Daniel Wilson (nominated by Eugene Koonin), Warren Tate (nominated by Eugene Koonin), Yoshikazu Nakamura (nominated by Eugene Koonin) and Eugene Koonin.

Journal Article↗

ARFA: a program for annotating bacterial release factor genes, including prediction of programmed ribosomal frameshifting.

UNLABELLED: Correct annotation of genes encoding release factors in bacterial genomes is often complicated by utilization of +1 programmed ribosomal frameshifting during synthesis of release factor 2, RF2. In the absence of robust computational approaches for predicting ribosomal frameshifting, the success of proper annotation depends on annotators' familiarity with this phenomenon. Here we describe a novel computer tool that allows automatic discrimination of genes encoding class-I bacterial release factors, RF1, RF2 and RFH. Most usefully, this program identifies and automatically annotates +1 frameshifting in RF2 encoding genes. Comparison of ARFA performance with existing annotations of bacterial genomes revealed that only 20% of RF2 genes utilizing ribosomal frameshifting during their expression are annotated correctly. AVAILABILITY: The PHP based web interface of ARFA and the source code are located at http://recode.genetics.utah.edu/arfa

Base Sequence↗

Recoding in bacteriophages and bacterial IS elements.

Dynamic shifts between open reading frames and the redefinition of codon meaning at specific sites, programmed by signals in mRNA, permits versatility of gene expression. Such alterations are characteristic of organisms in all domains of life and serve a variety of functional purposes. In this article, we concentrate on programmed ribosomal frameshifting, stop codon read-through and transcriptional slippage in the decoding of phage genes and bacterial mobile elements. Together with their eukaryotic counterparts, the genes encoding these elements are the richest known source of nonstandard decoding. Recent analyses revealed several novel sequences encoding programmed alterations in gene decoding and provide a glimpse of the emerging picture.

Bacteriophages↗

Pyrrolysine and selenocysteine use dissimilar decoding strategies.

Selenocysteine (Sec) and pyrrolysine (Pyl) are known as the 21st and 22nd amino acids in protein. Both are encoded by codons that normally function as stop signals. Sec specification by UGA codons requires the presence of a cis-acting selenocysteine insertion sequence (SECIS) element. Similarly, it is thought that Pyl is inserted by UAG codons with the help of a putative pyrrolysine insertion sequence (PYLIS) element. Herein, we analyzed the occurrence of Pyl-utilizing organisms, Pyl-associated genes, and Pyl-containing proteins. The Pyl trait is restricted to several microbes, and only one organism has both Pyl and Sec. We found that methanogenic archaea that utilize Pyl have few genes that contain in-frame UAG codons, and many of these are followed with nearby UAA or UGA codons. In addition, unambiguous UAG stop signals could not be identified. This bias was not observed in Sec-utilizing organisms and non-Pyl-utilizing archaea, as well as with other stop codons. These observations as well as analyses of the coding potential of UAG codons, overlapping genes, and release factor sequences suggest that UAG is not a typical stop signal in Pyl-utilizing archaea. On the other hand, searches for conserved Pyl-containing proteins revealed only four protein families, including methylamine methyltransferases and transposases. Only methylamine methyltransferases matched the Pyl trait and had conserved Pyl, suggesting that this amino acid is used primarily by these enzymes. These findings are best explained by a model wherein UAG codons may have ambiguous meaning and Pyl insertion can effectively compete with translation termination for UAG codons obviating the need for a specific PYLIS structure. Thus, Sec and Pyl follow dissimilar decoding and evolutionary strategies.

Amino Acid Sequence↗

Programmed ribosomal frameshifting in decoding the SARS-CoV genome.

Programmed ribosomal frameshifting is an essential mechanism used for the expression of orf1b in coronaviruses. Comparative analysis of the frameshift region reveals a universal shift site U_UUA_AAC, followed by a predicted downstream RNA structure in the form of either a pseudoknot or kissing stem loops. Frameshifting in SARS-CoV has been characterized in cultured mammalian cells using a dual luciferase reporter system and mass spectrometry. Mutagenic analysis of the SARS-CoV shift site and mass spectrometry of an affinity tagged frameshift product confirmed tandem tRNA slippage on the sequence U_UUA_AAC. Analysis of the downstream pseudoknot stimulator of frameshifting in SARS-CoV shows that a proposed RNA secondary structure in loop II and two unpaired nucleotides at the stem I-stem II junction in SARS-CoV are important for frameshift stimulation. These results demonstrate key sequences required for efficient frameshifting, and the utility of mass spectrometry to study ribosomal frameshifting.

Base Sequence↗

Transcriptional slippage in bacteria: distribution in sequenced genomes and utilization in IS element gene expression.

BACKGROUND: Transcription slippage occurs on certain patterns of repeat mononucleotides, resulting in synthesis of a heterogeneous population of mRNAs. Individual mRNA molecules within this population differ in the number of nucleotides they contain that are not specified by the template. When transcriptional slippage occurs in a coding sequence, translation of the resulting mRNAs yields more than one protein product. Except where the products of the resulting mRNAs have distinct functions, transcription slippage occurring in a coding region is expected to be disadvantageous. This probably leads to selection against most slippage-prone sequences in coding regions. RESULTS: To find a length at which such selection is evident, we analyzed the distribution of repetitive runs of A and T of different lengths in 108 bacterial genomes. This length varies significantly among different bacteria, but in a large proportion of available genomes corresponds to nine nucleotides. Comparative sequence analysis of these genomes was used to identify occurrences of 9A and 9T transcriptional slippage-prone sequences used for gene expression. CONCLUSIONS: IS element genes are the largest group found to exploit this phenomenon. A number of genes with disrupted open reading frames (ORFs) have slippage-prone sequences at which transcriptional slippage would result in uninterrupted ORF restoration at the mRNA level. The ability of such genes to encode functional full-length protein products brings into question their annotation as pseudogenes and in these cases is pertinent to the significance of the term 'authentic frameshift' frequently assigned to such genes.

Adenosine↗

Expression levels influence ribosomal frameshifting at the tandem rare arginine codons AGG_AGG and AGA_AGA in Escherichia coli.

The rare codons AGG and AGA comprise 2% and 4%, respectively, of the arginine codons of Escherichia coli K-12, and their cognate tRNAs are sparse. At tandem occurrences of either rare codon, the paucity of cognate aminoacyl tRNAs for the second codon of the pair facilitates peptidyl-tRNA shifting to the +1 frame. However, AGG_AGG and AGA_AGA are not underrepresented and occur 4 and 42 times, respectively, in E. coli genes. Searches for corresponding occurrences in other bacteria provide no strong support for the functional utilization of frameshifting at these sequences. All sequences tested in their native context showed 1.5 to 11% frameshifting when expressed from multicopy plasmids. A cassette with one of these sequences singly integrated into the chromosome in stringent cells gave 0.9% frameshifting in contrast to two- to four-times-higher values obtained from multicopy plasmids in stringent cells and eight-times-higher values in relaxed cells. Thus, +1 frameshifting efficiency at AGG_AGG and AGA_AGA is influenced by the mRNA expression level. These tandem rare codons do not occur in highly expressed mRNAs.

Arginine↗

P-site tRNA is a crucial initiator of ribosomal frameshifting.

The expression of some genes requires a high proportion of ribosomes to shift at a specific site into one of the two alternative frames. This utilized frameshifting provides a unique tool for studying reading frame control. Peptidyl-tRNA slippage has been invoked to explain many cases of programmed frameshifting. The present work extends this to other cases. When the A-site is unoccupied, the P-site tRNA can be repositioned forward with respect to mRNA (although repositioning in the minus direction is also possible). A kinetic model is presented for the influence of both, the cognate tRNAs competing for overlapping codons in A-site, and the stabilities of P-site tRNA:mRNA complexes in the initial and new frames. When the A-site is occupied, the P-site tRNA can be repositioned backward. Whether frameshifting will happen depends on the ability of the A-site tRNA to subsequently be repositioned to maintain physical proximity of the tRNAs. This model offers an alternative explanation to previously published mechanisms of programmed frameshifting, such as out-of-frame tRNA binding, and a different perspective on simultaneous tandem tRNA slippage.

Animals↗

Sequences that direct significant levels of frameshifting are frequent in coding regions of Escherichia coli.

It is generally believed that significant ribosomal frameshifting during translation does not occur without a functional purpose. The distribution of two frameshift-prone sequences, A_AAA_AAG and CCC_TGA, in coding regions of Escherichia coli has been analyzed. Although a moderate level of selection against the first sequence is evident, 68 genes contain A_AAA_AAG and 19 contain CCC_TGA. The majority of those tested in their genomic context showed >1% frameshifting. Comparative sequence analysis was employed to assess a potential biological role for frameshifting in decoding these genes. Two new candidates, in pheL and ydaY, for utilized frameshifting have been identified in addition to those previously known in dnaX and nine insertion sequence elements. For the majority of the shift-prone sequences no functional role can be attributed to them, and the frameshifting is likely erroneous. However, none of frameshift sequences is in the 306 most highly expressed genes. The unexpected conclusion is that moderate frameshifting during expression of at least some other genes is not sufficiently harmful for cells to trigger strong negative evolutionary pressure.

Base Sequence↗

RECODE 2003.

The RECODE database is a compilation of translational recoding events (programmed ribosomal frameshifting, codon redefinition and translational bypass). The database provides information about the genes utilizing these events for their expression, recoding sites, stimulatory sequences and other relevant information. The Database is freely available at http://recode.genetics.utah.edu/.

Animals↗

Maintenance of the correct open reading frame by the ribosome.

During translation, a string of non-overlapping triplet codons in messenger RNA is decoded into protein. The ability of a ribosome to decode mRNA without shifting between reading frames is a strict requirement for accurate protein biosynthesis. Despite enormous progress in understanding the mechanism of transfer RNA selection, the mechanism by which the correct reading frame is maintained remains unclear. In this report, evidence is presented that supports the idea that the translational frame is controlled mainly by the stability of codon-anticodon interactions at the P site. The relative instability of such interactions may lead to dissociation of the P-site tRNA from its codon, and formation of a complex with an overlapping codon, the process known as P-site tRNA slippage. We propose that this process is central to all known cases of +1 ribosomal frameshifting, including that required for the decoding of the yeast transposable element Ty3. An earlier model for the decoding of this element proposed 'out-of-frame' binding of A-site tRNA without preceding P-site tRNA slippage.

Anticodon↗

Translational recoding signals between gag and pol in diverse LTR retrotransposons.

Because of their compact genomes, retroelements (including retrotransposons and retroviruses) employ a variety of translational recoding mechanisms to express Gag and Pol. To assess the diversity of recoding strategies, we surveyed gag/pol gene organization among retroelements from diverse host species, including elements exhaustively recovered from the genome sequences of Caenorhabditis elegans, Drosophila melanogaster, Schizosaccharomyces pombe, Candida albicans, and Arabidopsis thaliana. In contrast to the retroviruses, which typically encode pol in the -1 frame relative to gag, nearly half of the retroelements surveyed encode a single gag-pol open reading frame. This was particularly true for the Ty1/copia group retroelements. Most animal Ty3/gypsy retroelements, on the other hand, encode gag and pol in separate reading frames, and likely express Pol through +1 or -1 frameshifting. Conserved sequences conforming to slippery sites that specify viral ribosomal frameshifting were identified among retroelements with pol in the -1 frame. None of the plant retroelements encoded pol in the -1 frame relative to gag; however, two closely related plant Ty3/gypsy elements encode pol in the +1 frame. Interestingly, a group of plant Ty1/copia retroelements encode pol either in a +1 frame relative to gag or in two nonoverlapping reading frames. These retroelements have a conserved stem-loop at the end of gag, and likely express pol either by a novel means of internal ribosomal entry or by a bypass mechanism.

Base Sequence↗

Recoding: translational bifurcations in gene expression.

During the expression of a certain genes standard decoding is over-ridden in a site or mRNA specific manner. This recoding occurs in response to special signals in mRNA and probably occurs in all organisms. This review deals with the function and distribution of recoding with a focus on the ribosomal frameshifting used for gene expression in bacteria.

Animals↗

Release factor 2 frameshifting sites in different bacteria.

The mRNA encoding Escherichia coli polypeptide chain release factor 2 (RF2) has two partially overlapping reading frames. Synthesis of RF2 involves ribosomes shifting to the +1 reading frame at the end of the first open reading frame (ORF). Frameshifting serves an autoregulatory function. The RF2 gene sequences from the 86 additional bacterial species now available have been analyzed. Thirty percent of them have a single ORF and their expression does not require frameshifting. In the approximately 70% that utilize frameshifting, the sequence cassette responsible for frameshifting is highly conserved. In the E. coli RF2 gene, an internal Shine-Dalgarno (SD) sequence just before the shift site was shown earlier to be important for frameshifting. Mutagenic data presented here show that the spacer region between the SD sequence and the shift site influences frameshifting, and possible mechanisms are discussed. Internal translation initiation occurs at the shift site, but any functional role is obscure.

DNA Mutational Analysis↗