Search PubMedSearch

SEARCH · Search PubMed

Results for “genetic code expansion”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Accessing isotopically labeled proteins containing genetically encoded phosphoserine for NMR with optimized expression conditions.

Phosphoserine (pSer) sites are primarily located within disordered protein regions, making it difficult to experimentally ascertain their effects on protein structure and function. Therefore, the production of 15N- (and 13C)-labeled proteins with site-specifically encoded pSer for NMR studies is essential to uncover molecular mechanisms of protein regulation by phosphorylation. While genetic code expansion technologies for the translational installation of pSer in Escherichia coli are well established and offer a powerful strategy to produce site-specifically phosphorylated proteins, methodologies to adapt them to minimal or isotope-enriched media have not been described. This shortcoming exists because pSer genetic code expansion expression hosts require the genomic ΔserB mutation, which increases pSer bioavailability but also imposes serine auxotrophy, preventing growth in minimal media used for isotopic labeling of recombinant proteins. Here, by testing different media supplements, we restored normal BL21(DE3) ΔserB growth in labeling media but subsequently observed an increase of phosphatase activity and mis-incorporation not typically seen in standard rich media. After rounds of optimization and adaption of a high-density culture protocol, we were able to obtain ≥10 mg/L homogenously labeled, phosphorylated superfolder GFP. To demonstrate the utility of this method, we also produced the intrinsically disordered serine/arginine-rich region of the SARS-CoV-2 Nucleocapsid protein labeled with 15N and pSer at the key site S188 and observed the resulting peak shift due to phosphorylation by 2D and 3D heteronuclear single quantum correlation analyses. We propose this cost-effective methodology will pave the way for more routine access to pSer-enriched proteins for 2D and 3D NMR analyses.

Humans

High throughput screening of eukaryotic release factor 1 variants to enhance noncanonical amino acid incorporation.

Noncanonical amino acids (ncAAs) enable diversification of protein functions, but the efficiency of genetic code expansion (GCE) in eukaryotes is hindered by competition between suppressor tRNAs and release factors. Prior work has identified eukaryotic release factor 1 (eRF1) mutants that improve ncAA incorporation, suggesting that screens for improved variants may lead to further enhancements. Here, we developed a high-throughput system to screen eRF1 mutants in Saccharomyces cerevisiae where eRF1 mutants are coexpressed on a plasmid alongside genomically encoded, wild-type eRF1. This strategy enabled recovery of live cells expressing eRF1 variants that enhance ncAA incorporation, even with mutants known to severely affect cell viability in the absence of WT eRF1 expression. We prepared and screened a million-member library of randomly mutated eRF1 variants for clones exhibiting improved ncAA integration phenotypes. Deep sequencing revealed a diverse set of enriched mutations across all three major domains of eRF1. Interestingly, several enriched mutations identified here are also found in naturally occurring eRF1 homologs from species that recode canonical stop codons. When eRF1 variants were combined with yeast knockout strains also known to enhance ncAA incorporation, this resulted in further improvements to efficiency, highlighting the complementarity of release factor engineering to other GCE enhancement strategies. This work demonstrates that high-throughput engineering of the eukaryotic translational apparatus is a powerful approach to identify previously unknown solutions for enhancing ncAA incorporation, with implications for elucidating and precisely manipulating the molecular functions of essential translational machinery.

Noncanonical amino acids

Selenoprotein synthesis: an expansion of the genetic code.

A number of enzymes employ the unusual amino acid selenocysteine as part of their active site because of its high chemical reactivity. Selenocysteine is incorporated into these proteins co-translationally: biosynthesis occurs on a specific tRNA and insertion into a growing polypeptide is directed by a UGA codon in the mRNA. In E. coli, this requires a specific translation factor. Selenocysteine thus represents a unique expansion of the genetic code.

Animals

Ribosome-mediated incorporation of a non-standard amino acid into a peptide through expansion of the genetic code.

One serious limitation facing protein engineers is the availability of only 20 'proteinogenic' amino acids encoded by natural messenger RNA. The lack of structural diversity among these amino acids restricts the mechanistic and structural issues that can be addressed by site-directed mutagenesis. Here we describe a new technology for incorporating non-standard amino acids into polypeptides by ribosome-based translation. In this technology, the genetic code is expanded through the creation of a 65th codon-anticodon pair from unnatural nucleoside bases having non-standard hydrogen-bonding patterns. This new codon-anticodon pair efficiently supports translation in vitro to yield peptides containing a non-standard amino acid. The versatility of the ribosome as a synthetic tool offers new possibilities for protein engineering, and compares favourably with another recently described approach in which the genetic code is simply rearranged to recruit stop codons to play a coding role.

Amino Acid Sequence

Genetic Incorporation of a Thioxanthone-Containing Amino Acid for the Design of Artificial Photoenzymes.

Genetically encodable photosensitizers allow the design of artificial photoenzymes to expand the scope of abiological reactions. Herein, we report the genetic incorporation of a thioxanthone-containing amino acid into a protein scaffold via an engineered pyrrolysyl-tRNA/pyrrolysyl-tRNA synthetase pair. The designer enzyme was engineered to catalyze a dearomative [2+2] cycloaddition reaction in high yields (up to>99 % yield) with excellent enantioselectivity (up to 98 : 2 e.r.). This work provides a robust and facile method for photoenzyme design and lays the foundation for the development of further photoenzymatic reactions.

Xanthones

Selenocysteine: the 21st amino acid.

Great excitement was elicited in the field of selenium biochemistry in 1986 by the parallel discoveries that the genes encoding the selenoproteins glutathione peroxidase and bacterial formate dehydrogenase each contain an in-frame TGA codon within their coding sequence. We now know that this codon directs the incorporation of selenium, in the form of selenocysteine, into these proteins. Working with the bacterial system has led to a rapid increase in our knowledge of selenocysteine biosynthesis and to the exciting discovery that this system can now be regarded as an expansion of the genetic code. The prerequisites for such a definition are co-translational insertion into the polypeptide chain and the occurrence of a tRNA molecule which carries selenocysteine. Both of these criteria are fulfilled and, moreover, tRNASec even has its own special translation factor which delivers it to the translating ribosome. It is the aim of this article to review the events leading to the elucidation of selenocysteine as being the 21st amino acid.

Bacterial Proteins

The advent of DNA databanks: implications for information privacy.

Genetic identification tests -- better known as DNA profiling -- currently allow criminal investigators to connect suspects to physical samples retrieved from a victim or the scene of a crime. A controversial yet acclaimed expansion of DNA analysis is the creation of a massive databank of genetic codes. This Note explores the privacy concerns arising out of the collection and retention of extremely personal information in a central database. The potential for unauthorized access by those not investigating a particular crime compels the implementation of national standards and stringent security measures.

Confidentiality

Somatic mutation, affinity maturation and the antibody repertoire: a computer model.

Somatic mutation has been implicated as a significant and possibly primary factor in the maturation of antibody affinity in the humoral immune response. B cells stimulated by antigen experience a hyper-mutation in the gene segments that code for the antigen-binding site of the antibody, creating antibody specificities that did not exist at the time of immunization. Although most of the mutations are likely to be disadvantageous, new specificities with a higher affinity for the antigen are sometimes created. These higher-affinity cells are preferentially selected for proliferation and eventual antibody secretion, resulting in a progressively higher average affinity over time. In this paper we present the results of an investigation of somatic mutation through the use of a computer model. At the basis of the model is a large repertoire of discrete antibodies and antigens, having three-dimensional structures, that exhibit properties similar to those of the real populations. The key factor is that the binding strength between any antibody/antigen pair can be calculated as a function of the complementarity of the (a) size, (b) shape and (c) functional groups that comprise the two structures. The created repertoires are imbedded in a dynamical system model of the immune response to directly evaluate the affect of somatic mutation on affinity maturation. We also present an expanded hypothesis of clonal selection and development to explain how the mutational restrictions imposed by the genetic code and the structure of the antibody repertoire, along with antigen concentration, affinity, and probabilistic factors may interact and contribute to the expansion of specific clones as the response develops over time.

Animals

A Unified Mechanism of +1 Ribosomal Frameshifting.

Ribosomes decode 3-nucleotide codons and move in 1-codon increments to maintain the messenger RNA (mRNA) frame thereby accurately producing the encoded protein. In special cases, including viral genomes and regulatory cellular proteins, frameshifting occurs to expand the coding repertoire of an mRNA to make more than one protein. How these frameshifting events are induced and regulated is an active area of research. Here, we discuss recent progress in the understanding of +1 frameshifting (+1FS), during which the ribosome shifts by 1 mRNA nucleotide in the 3' direction. Structural and biochemical studies yielded insights into +1FS induced by mRNA slippery sequences and transfer RNA (tRNA) stem-loop expansion or modifications. tRNAs with an additional anticodon nucleotide are explored as a biotechnology tool for expanding the genetic code in an approach termed quadruplet decoding. We revisit the challenges of the quadruplet decoding model, discuss +1FS scenarios in bacteria and eukaryotes, and propose a unifying structural mechanism for +1FS.

Frameshifting, Ribosomal

Regulatory Evolution and the Genetic Basis of Human Brain Expansion.

The evolution of the human brain is characterized by profound changes in structure and function, despite relatively limited divergence in protein-coding genes compared to other primates. This paradox has led to increasing recognition of gene regulatory elements (GREs) as primary drivers of evolutionary innovation. In this review, we synthesize current knowledge on the role of conserved noncoding elements (CNEs), human accelerated regions (HARs), and transposable element (TE)-derived sequences in shaping gene regulatory networks (GRNs) underlying brain development. Comparative analyses across humans and closely related primates, including the chimpanzee, gorilla, and orangutan, reveal that while core regulatory architectures are highly conserved, subtle changes in regulatory elements drive species-specific gene expression patterns. We highlight how CNEs provide a stable regulatory framework, whereas HARs and TE-derived elements introduce lineage-specific modifications that fine-tune neurodevelopmental processes. Advances in functional genomics, including CRISPR-based perturbations, massively parallel reporter assays, and single-cell multi-omics, have enabled direct interrogation of regulatory function, linking sequence variation to cellular phenotypes. Furthermore, we discuss how regulatory evolution contributes to both cognitive innovation and susceptibility to neurological disorders. Despite significant progress, challenges remain in establishing causal relationships between regulatory variation and phenotypic outcomes. Future integration of multi-omics data and comparative models will be essential for resolving these complexities. Together, this review provides a comprehensive framework for understanding the molecular basis of primate brain evolution through the lens of gene regulation.

Brain evolution

The proteomic origin of the genetic code.

INTRODUCTION: The origin and evolution of the genetic code is a central problem in molecular biology. Classical models have emphasized stereochemistry, frozen accidents, or adaptive optimization, often treating proteins as passive products of preexisting codes. More recent views instead portray the code as a dynamic, coevolving system shaped by reciprocal interactions among amino acids, RNA, and early catalysts. AREAS COVERED: Here, I review efforts of phylogeny reconstruction of the history of tRNA, protein structural domains, and dipeptide sequences in proteomes. These complementary approaches allow exploration of the entry of amino acids and codons into the code, and the transition from an operational RNA code in the tRNA acceptor arm to the canonical code in the anticodon loop. Evidence for ancestral synthetase enzymes with dual functions in aminoacylation and peptide-bond formation, as well as early bidirectional (sense-antisense) coding reflected in dipeptide-antidipeptide emergence is also discussed. EXPERT OPINION: The genetic code is best viewed as a proteome-driven, evolvable system in which early peptides actively shaped coding rules by stabilizing structure, expanding chemical diversity, and enhancing catalysis. This perspective connects origin-of-life studies with modern efforts of code expansion, translational engineering, and peptide-based therapeutics, highlighting the impact of the code's proteomic origin.

Genetic Code

A DNA segment encoding two genes very tightly linked to Huntington's disease.

The discovery of D4S10, an anonymous DNA marker genetically linked to Huntington's disease (HD), introduced the capacity for limited presymptomatic diagnosis in this late-onset neurodegenerative disorder and raised the hope of cloning and characterizing the defect based on its chromosomal location. Progress on both fronts has been limited by the absence of additional DNA markers closer to the HD gene. An anonymous DNA locus, D4S43, has now been found that shows extremely tight linkage to HD. Like the disease gene, D4S43 is located in the most distal region of the chromosome 4 short arm, flanked by D4S10 and the telomere. In three extended HD kindreds, D4S43 displays no recombination with HD, placing it within 0 to 1.5 centimorgans of the genetic defect. Expansion of the D4S43 region to include 108 kilobases of cloned DNA has allowed identification of eight restriction fragment length polymorphisms and at least two independent coding segments. In the absence of crossovers, these genes must be considered candidates for the site of the HD defect, although the D4S43 restriction fragment length polymorphisms do not display linkage disequilibrium with the disease gene.

Alleles

Evolution of the genetic code.

The structure of the genetic code suggests that amino acid biosynthesis and hydrophobicity were important factors in shaping the genetic code, as the primitive code coevolved with new varieties of amino acids generated by the expanding pathways of biosynthesis. The current code is exceptionally stable. Deviant codes nonetheless have been observed in a number of mitochondrial and cellular genomes. Even the membership of encoded amino acids is undergoing expansion to include phosphoserine and selenocysteine. Experimental mutation of the code also has proven feasible, in a replacement of tryptophan by 4-fluorotryptophan as a component constituent of proteins. Such mutations, introducing novel varieties of encoded amino acids, will open up a new dimension in protein engineering and design.

Amino Acid Sequence

The complete sequence of the silkworm W chromosome uncovers its rapid evolution by large-scale duplications/deletions and translocation of W-linked genes.

The complete sequence of the W chromosome, which carries feminization activity in the silkworm, is crucial for understanding the sex-determination system in Lepidoptera. However, extensive accumulation of transposons due to lack of recombination, the very rare protein-coding genes and almost no information about molecular markers has hindered full W sequencing. We report the first complete silkworm W sequence (T2T_W, 11683305 bp) obtained by combining sequencing-assembly technologies and newly developed error detection methods, evaluated with genetically mapped W-RAPD markers, W-mutants, and W-derived BAC clones. The T2T_W sequence showed that the W is composed of a massive 92% accumulation of transposons and repeat sequences, among which the main constituents are intact LTR/LINE retrotransposons indicating recent expansions. In addition to Fem clusters producing Fem piRNA (Feminizer-derived PIWI-interacting RNA), we found 26 protein-coding genes in the W sequence. These include four gene pairs encoding zinc-finger motifs designated z1:z20 and a gene encoding serine/arginine repetitive matrix protein 1-like (SRRM1-like). To identify candidate genes for female sex-determination and differentiation we also sequenced the shortest W (3.8 Mb) from a translocation mutant with feminizing activity, which harbored four conventional genes: a Fem cluster, a pair of z1:z20 isoforms, z20-S, and a SRRM1-like gene. Phylogenetic analysis revealed that z1:z20 originated from a copy of an autosomal zinc-finger gene pair, z2:z21, translocated onto the W around 2.43 Mya and subsequently amplified to yield 4 W-linked zinc-finger gene pairs. The complete W sequence revealed that large-scale deletions and amplifications played a significant role in W chromosome evolution.

Animals

3D epigenome of glial cell types in developing human cortex.

The human cortex is complex and heterogeneous, undergoing extensive expansion during development1,2. Our prior study of neurogenesis, including radial glia (RG), intermediate progenitor cells, excitatory neurons and interneurons demonstrated that chromatin looping underlies transcriptional regulation for lineage-specific genes, shedding light on how non-coding genetic variants contribute to neuropsychiatric disorders by means of cell-type-specific gene regulation3. RG have a crucial role in generating cellular diversity through both neurogenesis and gliogenesis and can be further classified into ventricular RG (vRG) and outer RG (oRG)4,5. Given their significance in cortical development, we conducted a comprehensive three-dimensional (3D) epigenomic analysis of four main glial populations, including vRG, oRG, oligodendrocyte precursor cells and microglia, from the mid-gestational human neocortex. By integrating gene expression, chromatin accessibility, DNA methylation and 3D chromatin interactions, we identified cell-type-specific candidate cis-regulatory elements (cCREs) and validated their regulatory function using transgenic mouse embryos. Using machine learning, we prioritized 112 schizophrenia risk variants within glia cCREs and further confirmed the predicted vRG enhancer disruption by the rs4449074 risk allele in vivo. Finally, oRG cCREs are enriched for human accelerated regions compared with other cCREs and a subset of human accelerated regions show activity differences from their chimpanzee orthologues that interact with genes involved in neuronal development. Our findings advance the understanding of human-specific gene regulation during corticogenesis.

Journal Article

Normal and altered phenotypic expression of immunoglobulin genes.

Genetically controlled intraspecific differences between immunoglobulins (allotypes) provide valuable markers for the study of the quantitative expression of allelic and nonallelic alternative forms of immunoglobulins (Igs) during the normal development of rabbits. Heterozygous rabbits are mosaics of cells expressing different Ig-genes since fully differentiated productive cells generally secrete only one of alternative forms of Ig. The proportions of cells that differentiate to produce allelic forms of immunoglobulins during normal development depend on the particular heterozygous genotype. The normal proportions of some markers can be drastically altered if the differentiation of lymphoid cells in the young rabbit occurs in the milieu of antibody specific for one form (allotype suppression). An initiating step in the establishment of persistent allotype suppression is probably the interaction of antiallotype antibody with allotype-bearing receptors on lymphoid cell surfaces, but the mechanism for the maintenance of a state of chronic suppression may well be more complex. Allotype suppression can be viewed as one example of numerous immunological phenomena that reflect specific and finely tuned regulatory mechanisms governing the differentiation and clonal expansion of lymphoid cells destined to secrete immunoglobulins.

Alleles

Unraveling the genomic blueprint of the Indian black soldier fly: From genome assembly to evolutionary insights.

The black soldier fly (BSF) (Hermetia illucens) has been renowned for its sustainable bioconversion capabilities, resulting in smart protein production with wide applications in animal feed, bioenergy, and biofertilizer. However, the genetic mechanisms underlying efficient bioconversion and productivity remain poorly understood. To advance strain-specific applications and strengthen genetic resource availability, we present the whole genome sequencing (WGS) data for an Indian isolate of black soldier fly. The assembled genome was 1.46 Gb with a scaffold N50 of 172.7 Mb, and a GC content of 42.6%. Furthermore, 64.17% of genomic sequences were masked as repeated, and 14,317 protein-coding sequences were identified. Variant analysis against the reference genome identified 34.44 million variants (∼33.25 million SNPs and ∼ 1.18 million INDELs), with the majority (99.3%) classified as MODIFIER, 0.54% as LOW impact, 0.14% as MODERATE, and only 0.003% as HIGH impact. Comparative genomic analysis with other related species revealed expansions of gene families in BSF associated with Immune effector (Antimicrobial peptides (AMPs), Lysozymes, and Peptidoglycan Recognition Protein (PGRP) and Detoxification (cytochrome P450 enzymes). Notably, AMPs in the Indian isolate showed enhanced copy number variation in defensin (27) and PGRP (40) compared to reference BSF, suggesting potential regional adaptations to pathogen exposure. Collectively, this genomic data provides an improved resource for evolutionary studies, functional genomics, and targeted genetic improvement of BSF for sustainable bioconversion applications.

Comparative genomics