Search PubMedSearch

PubMed · 41848554

The proteomic origin of the genetic code.

Abstract

INTRODUCTION: The origin and evolution of the genetic code is a central problem in molecular biology. Classical models have emphasized stereochemistry, frozen accidents, or adaptive optimization, often treating proteins as passive products of preexisting codes. More recent views instead portray the code as a dynamic, coevolving system shaped by reciprocal interactions among amino acids, RNA, and early catalysts. AREAS COVERED: Here, I review efforts of phylogeny reconstruction of the history of tRNA, protein structural domains, and dipeptide sequences in proteomes. These complementary approaches allow exploration of the entry of amino acids and codons into the code, and the transition from an operational RNA code in the tRNA acceptor arm to the canonical code in the anticodon loop. Evidence for ancestral synthetase enzymes with dual functions in aminoacylation and peptide-bond formation, as well as early bidirectional (sense-antisense) coding reflected in dipeptide-antidipeptide emergence is also discussed. EXPERT OPINION: The genetic code is best viewed as a proteome-driven, evolvable system in which early peptides actively shaped coding rules by stabilizing structure, expanding chemical diversity, and enhancing catalysis. This perspective connects origin-of-life studies with modern efforts of code expansion, translational engineering, and peptide-based therapeutics, highlighting the impact of the code's proteomic origin.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gustavo Caetano-Anollés. 2026-03-23. The proteomic origin of the genetic code.. https://doi.org/10.1080/14789450.2026.2646677

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Alternative genetic codes in bacteria and archaea identified with a fast k-mer-based algorithm.

The genetic code is conserved across all domains of life and is often described as universal. Nevertheless, many exceptions to the "universal" code have now been documented, most of these through manual or semiautomated inspection of highly conserved genes. Modern bioinformatics tools improved our ability to find alternative genetic codes but remain computationally expensive, preventing widespread use on thousands of new species identified by sequencing environmental samples. Here, I report a >100-fold accelerated method for inferring the genetic code directly from assembled genomes and apply it to thousands of previously uncharacterized assemblies from archaea and bacteria. I describe three candidate genetic code variations, one of which, an alternative genetic code used by a family of Asgard archaea, is a unique example of sense codon reassignments for this domain. Identifying genetic code variations is important for understanding evolution of the standard code and improving accuracy of protein databases and open reading frame identification.

Genetic Code

The genetic code at the balance point of error and demand.

The origin and organizing principles of the genetic code remain central problems in molecular evolution. The low probability of the natural codon-to-amino acid mapping arising by chance has spurred the hypothesis that its structure is optimized for robustness to mutations and translational errors. For the construction of effective molecular machines, the repertoire of encoded amino acids must also be diverse enough in physicochemical features. Here, we examine whether the standard genetic code can be understood as a near-optimal solution balancing these two objectives: minimizing error load and aligning codon assignments with the naturally occurring amino acid composition. Using simulated annealing, we explore this trade-off across a broad range of parameters. We find that the standard genetic code resides near an optimum in the fitness landscape of possible genetic codes. The degeneracy of the code plays a dual role, minimizing mistranslation errors while matching codon multiplicity to amino acid usage frequencies. As a result, uniform codon usage alone is sufficient to recover the empirical amino acid composition, without any additional bias. It is a highly effective solution that balances fidelity against resource availability constraints. A comparative analysis of natural variants also reveals a functional decoupling: error robustness acts as a rigid global constraint determined by code topology, whereas compositional alignment serves as a more flexible variable that adapts to lineage-specific demands. These results support a multi-objective optimization framework in which the genetic code reflects a balance between translational fidelity and proteomic demand.

Genetic Code

Reading another hidden message in the genetic code.

The genetic code determines not only the amino acid sequences of proteins but also mRNA stability. How is this hidden message read? Hia and colleagues have now identified human DHX29 as a reader of the mRNA stability code carried by codons, providing new mechanistic insights into translation-coupled gene regulation.

Genetic Code