Search PubMed⌕ Search

PubMed · 14616028

Protein sequences yield a proteomic code.

Abstract

Analysis of crystallized protein structures suggests that globular proteins are organized as consecutively connected units of 25-35 residues. These units are closed loops, that is returns of the polypeptide chain trajectory to a close contact with itself. This universal feature of apparently polymer-statistical nature is a basis for a principally novel view on the globular proteins as loop fold structures. The same unit size has been detected in protein sequences translated from complete prokaryotic genomes by positional autocorrelation analysis, which strongly indicates the evolutionary connection of the units. The units are further characterized by prototype sequences matching to their numerous derivatives in the translated genomes. The matches to five strongest prokaryotic prototypes and three prototypes of C. elegans are identified in the sequences of crystallized proteins, and their structures analyzed. Corresponding segments of the polypeptide chains in majority of cases form closed loops, though evolutionary fate of every prototype element is shown to be rather diverse. Then loop ends can be separated by a sequence-wise distant segments and stabilized by the spatial interactions in the context of the overall globular structure. The units belong to a presumably limited spectrum of the sequence prototypes, full repertoire of which would constitute a proteomic code.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Igor N Berezovsky, Alla Kirzhner, Valery M Kirzhner, Vladimir R Rosenfeld, Edward N Trifonov. 2003. Protein sequences yield a proteomic code.. https://doi.org/10.1080/07391102.2003.10506928

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs↗

Signalling thresholds and negative B-cell selection in acute lymphoblastic leukaemia.

B cells are selected for an intermediate level of B-cell antigen receptor (BCR) signalling strength: attenuation below minimum (for example, non-functional BCR) or hyperactivation above maximum (for example, self-reactive BCR) thresholds of signalling strength causes negative selection. In ∼25% of cases, acute lymphoblastic leukaemia (ALL) cells carry the oncogenic BCR-ABL1 tyrosine kinase (Philadelphia chromosome positive), which mimics constitutively active pre-BCR signalling. Current therapeutic approaches are largely focused on the development of more potent tyrosine kinase inhibitors to suppress oncogenic signalling below a minimum threshold for survival. We tested the hypothesis that targeted hyperactivation--above a maximum threshold--will engage a deletional checkpoint for removal of self-reactive B cells and selectively kill ALL cells. Here we find, by testing various components of proximal pre-BCR signalling in mouse BCR-ABL1 cells, that an incremental increase of Syk tyrosine kinase activity was required and sufficient to induce cell death. Hyperactive Syk was functionally equivalent to acute activation of a self-reactive BCR on ALL cells. Despite oncogenic transformation, this basic mechanism of negative selection was still functional in ALL cells. Unlike normal pre-B cells, patient-derived ALL cells express the inhibitory receptors PECAM1, CD300A and LAIR1 at high levels. Genetic studies revealed that Pecam1, Cd300a and Lair1 are critical to calibrate oncogenic signalling strength through recruitment of the inhibitory phosphatases Ptpn6 (ref. 7) and Inpp5d (ref. 8). Using a novel small-molecule inhibitor of INPP5D (also known as SHIP1), we demonstrated that pharmacological hyperactivation of SYK and engagement of negative B-cell selection represents a promising new strategy to overcome drug resistance in human ALL.

Amino Acid Motifs↗

Ribosomal protein L7a binds RNA through two distinct RNA-binding domains.

The human ribosomal protein L7a is a component of the major ribosomal subunit. We previously identified three nuclear-localization-competent domains within L7a, and demonstrated that the domain defined by aa (amino acids) 52-100 is necessary, although not sufficient, to target the L7a protein to the nucleoli. We now demonstrate that L7a interacts in vitro with a presumably G-rich RNA structure, which has yet to be defined. We also demonstrate that the L7a protein contains two RNA-binding domains: one encompassing aa 52-100 (RNAB1) and the other encompassing aa 101-161 (RNAB2). RNAB1 does not contain any known nucleic-acid-binding motif, and may thus represent a new class of such motifs. On the other hand, a specific region of RNAB2 is highly conserved in several other protein components of the ribonucleoprotein complex. We have investigated the topology of the L7a-RNA complex using a recombinant form of the protein domain that encompasses residues 101-161 and a 30mer poly(G) oligonucleotide. Limited proteolysis and cross-linking experiments, and mass spectral analyses of the recombinant protein domain and its complex with poly(G) revealed the RNA-binding region.

Amino Acid Motifs↗