Search PubMedSearch

PubMed · 42709807

Alternative genetic codes in bacteria and archaea identified with a fast k-mer-based algorithm.

Abstract

The genetic code is conserved across all domains of life and is often described as universal. Nevertheless, many exceptions to the "universal" code have now been documented, most of these through manual or semiautomated inspection of highly conserved genes. Modern bioinformatics tools improved our ability to find alternative genetic codes but remain computationally expensive, preventing widespread use on thousands of new species identified by sequencing environmental samples. Here, I report a >100-fold accelerated method for inferring the genetic code directly from assembled genomes and apply it to thousands of previously uncharacterized assemblies from archaea and bacteria. I describe three candidate genetic code variations, one of which, an alternative genetic code used by a family of Asgard archaea, is a unique example of sense codon reassignments for this domain. Identifying genetic code variations is important for understanding evolution of the standard code and improving accuracy of protein databases and open reading frame identification.

Explore related subjects

Keep this discovery

BibTeXRIS

Artem V Melnykov. 2026-09-08. Alternative genetic codes in bacteria and archaea identified with a fast k-mer-based algorithm.. https://doi.org/10.1073/pnas.2610659123

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related citations

Depth-dependent microbial succession and interspecies hydrogen transfer drive pit mud maturation in Chinese strong-flavor baijiu fermentation.

Microbial communities in fermentation pit mud play a key role in determining the quality of Chinese strong-flavor baijiu (CSFB). However, the ecological processes underlying pit mud maturation across spatial and temporal scales remain unclear. In this study, amplicon sequencing and metagenomic analyses were employed to investigate the taxonomic succession, community assembly, and metabolic functions of bacterial and archaeal communities during the transition from fresh pit mud (FPM) to new pit mud (NPM) and old pit mud (OPM). A pronounced depth-dependent succession pattern was observed, with 4 cm representing a critical ecological boundary separating distinct community structures and maturation trajectories. During surface-layer maturation, community assembly shifted from stochastic to deterministic processes, accompanied by homogeneous selection and increasing network complexity. In contrast, stochastic processes remained dominant throughout deep-layer maturation. Metagenomic analyses revealed a functional transition from lactate and acetate production, primarily associated with Lactobacillus in FPM and NPM, to butyrate and caproate production associated with Clostridium and Caproiciproducens in OPM. This functional transition was accompanied by enhanced amino acid metabolism, which was associated with the enrichment of Proteiniphilum and Aminobacterium. Notably, methanogen-mediated interspecies hydrogen transfer (IHT) emerged as a key ecological feature during pit mud maturation. In OPM, IHT networks primarily involving Methanobacterium and Methanosarcina linked methanogenesis with reverse β-oxidation through diverse hydrogen-transfer pathways, reinforcing metabolic interactions underlying caproate production. These findings provide new insights into the ecological mechanisms underlying pit mud maturation and offer a theoretical basis for the directed cultivation of high-quality pit mud in CSFB production.

Hydrogen

Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)6.

The standard genetic code reduces the impact of point mutations, but the robustness of this property across physicochemical metrics, naturally occurring variant codes, and codon-reassignment mechanisms remains incompletely quantified. Embedding the 64 codons in GF(2)6 represents the hypercube Q6 as a coordinate-dependent subgraph of the encoding-independent single-nucleotide mutation graph H(3,4), and enables continuous &#x3c1;-interpolation between the two. Under a quartet-pattern shuffle null (n=10,000), the standard code is significantly low-cost across four established, code-independent physicochemical distance metrics with partially overlapping content (Grant ham p=0.0062; Miyata p<0.001; Woese polar requirement p=0.003; Kyte-Doolittle hydropathy p=0.001), and the signal strengthens monotonically as &#x3c1; moves Q6&#x2192;H(3,4). A structure-aware sensitivity analysis under the alignment-derived ProtSub matrix (Jia & Jernigan 2021) yields the most extreme percentile of any measure tested (p=0.0004; all five p-values pass Bonferroni at &#x3b1;=0.05). Across the 27 NCBI translation tables, near-optimality is preserved: 11 of 12 informative-distance variants retain top-5% placement after BH-FDR correction. Natural codon reassignments avoid disrupting codon-family connectivity: under the encoding-independent H(3,4) adjacency, observed events are topology-breaking at relative risk 0.32 versus the candidate landscape (permutation p&#x2264;10-4). The H(3,4) result is stable by construction; the Q6 decomposition is representation-specific and fails to show depletion under 8 of 24 base-to-bit encodings, so we report H(3,4) as the primary test and Q6 as a sensitivity. Event-level conditional-logit modelling shows that topology avoidance and local physicochemical cost provide complementary, only weakly correlated signal (rs=0.15), and that topology adds explanatory value beyond physicochemistry under both Q6 and encoding-independent H(3,4) adjacency. Retrospective reanalysis of nine genome-recoding datasets is consistent with codon-family topology operating as an evolutionary-trajectory constraint distinct from acute engineering fitness. The contribution is the second axis: code evolution is jointly constrained by physicochemical smoothness and codon-family topological integrity, and these two constraints are partly independent.

Codon reassignment

Fructophilic lactic acid bacteria as a window into multi-scale convergent evolution.

Fructophilic lactic acid bacteria (FLAB) are a group of lactic acid bacteria with unique growth characteristics, that is, poor growth on glucose. Their growth is enhanced in the presence of fructose or external electron acceptors. These organisms inhabit fructose-rich environments such as flowers, fruits, and pollinating insects, particularly honey bees. Apilactobacillus spp. and Fructobacillus spp. are representatives of FLAB, although they belong to phylogenetically distant clades. These organisms commonly possess markedly small genomes with a low number of coding DNA sequences. Furthermore, their genomes are characterized by a markedly reduced number of genes involved in carbohydrate transport and metabolism. Genome reduction in FLAB reflects convergent adaptation to fructose-rich environments rather than general genome streamlining. The two distinct FLAB genera, Fructobacillus and Apilactobacillus, independently lost more than 100 genes in statistically similar orders. In contrast, genes involved in carbohydrate and amino acid metabolism exhibited reversed orders of loss between the two genera. Furthermore, FLAB genomes lack an intact bifunctional alcohol/aldehyde dehydrogenase gene (adhE), which causes their poor growth on glucose. A comparative genomic study suggested the evolutionary process underlying adhE gene decay during adaptation to the fructose-rich environments, including pollinating insects. In conclusion, FLAB represent a unique example of habitat-driven convergent reductive evolution that can be investigated across multiple biological scales - from individual genes to whole genomes - in the diverse LAB group with a wide range of habitats, and partially share the fructophilic evolution with eukaryotic yeasts found in fructose-rich habitats.

Fructose