Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Multiple effects of codon usage optimization on expression and immunogenicity of DNA candidate vaccines encoding the human immunodeficiency virus type 1 Gag protein.

We have analyzed the influence of codon usage modifications on the expression levels and immunogenicity of DNA vaccines, encoding the human immunodeficiency virus type 1 (HIV-1) group-specific antigen (Gag). In the presence of Rev, an expression vector containing the wild-type (wt) gag gene flanked by essential cis-acting sites such as the 5'-untranslated region and 3'-Rev response element supported substantial Gag protein expression and secretion in human H1299 and monkey COS-7 cells. However, only weak Gag production was observed from the murine muscle cell line C2C12. In contrast, optimization of the Gag coding sequence to that of highly expressed mammalian genes (syngag) resulted in an obvious increase in the G+C content and a Rev-independent expression and secretion of Gag in all tested mammalian cell lines, including murine C2C12 muscle cells. Mice immunized intramuscularly with the syngag plasmid showed Th1-driven humoral and cellular responses that were substantially higher than those obtained after injection of the Rev-dependent wild-type (wt) gag vector system. In contrast, intradermal immunization of both wt gag and syngag vector systems with the particle gun induced a Th2-biased antibody response and no cytotoxic T lymphocytes. Deletion analysis demonstrated that the CpG motifs generated within syngag by codon optimization do not contribute significantly to the high immunogenicity of the syngag plasmid. Moreover, low doses of coadministered stimulatory phosphorothioate oligodeoxynucleotides (ODNs) had only a weak effect on antibody production, whereas at higher doses immunostimulatory and nonstimulatory ODNs showed a dose-dependent suppression of humoral responses. These results suggest that increased Gag expression, rather than modulation of CpG-driven vector immunity, is responsible for the enhanced immunogenicity of the syngag DNA vaccine.

AIDS Vaccines↗

Preferential codons enhancing the expression level of human beta-defensin-2 in recombinant Escherichia coli.

Human beta-defensin-2 (hBD2) is a small antimicrobial peptide with potential as a therapeutic agent. The effect of codon usage on the expression of hBD2 in Escherichia coli was studied. Two coding sequences encoding the same hBD2 precursor were both expressed as fusion protein with thioredoxin in E. coli BL21 (DE3). One is the wild-type human cDNA and the other is a gene synthesized by a PCR-based method in which rare codons were altered to those frequently used in E. coli. The expression level of recombinant hBD2 was over 50% of the total cellular protein when the synthetic gene with preferential codons was employed which was a 9-fold enhancement over the wild-type cDNA. The result shows the codon bias of the host was a major barrier in high-level expression of recombinant hBD2 and suggests a similar approach may be used in the expression of other defensins in E. coli.

Amino Acid Sequence↗

The effective number of codons for individual amino acids: some codons are more optimal than others.

The aim of this study was to evaluate the codon bias using the effective number of codons for individual amino acids (N(c)(AA)) and to assess the codon bias in relation to the definition of optimal codons, using Escherichia coli as a model organism. We show that a general correlation exists between the effective number of codons (Ncirc(c)) or codon adaptation index (CAI) and N(c)(AA), but that this correlation is not equally strong for all amino acids within a degeneracy group. For example, leucine codons contribute more to Ncirc(c) and the codon adaptation index than serine codons. A possible explanation is that some optimal codons are more optimal than others, in terms of the selectional advantage they offer. This hypothesis is confirmed by further analysis on the correlations that exist between values for relative synonymous codon usage (RSCU), N(c)(AA), and the codon adaptation index.

Amino Acids↗

Neighboring-nucleotide effects on the rates of germ-line single-base-pair substitution in human genes.

The spectrum of single-base-pair substitutions logged in The Human Gene Mutation Database (HGMD), comprising 7,271 different lesions in the coding regions of 547 different human genes, was analyzed for nearest-neighbor effects on relative mutation rates. Owing to its retrospective nature, HGMD allows mutation rates to be estimated only in relative terms. Therefore, a novel methodology was devised in order to obtain these estimates in iterative fashion, correcting, at the same time, for the confounding effects of differential codon usage and for the fact that different types of amino acid replacement come to clinical attention with different probabilities. Over and above the hypermutability of CpG dinucleotides, reflected in transition rates five times the base mutation rate, only a subtle and locally confined influence of the surrounding DNA sequence on relative single-base-pair substitution rates was observed, which extended no farther than 2 bp from the substitution site. A disparity between the two DNA strands was evidenced by the fact that, when substitution rates were estimated conditional on the 5' and 3' flanking nucleotides, a significant rate difference emerged for 10 of 96 possible pairs of complementary substitutional events. Mutational bias, favoring substitutions toward flanking bases, a phenomenon reminiscent of misalignment mutagenesis, was apparent and exhibited both directionality and reading-frame sensitivity. No specific preponderance of repeat-sequence motifs was observed in the vicinity of nucleotide substitutions, but a moderate correlation between the relative mutability and thermodynamic stability of DNA triplets emerged, suggesting either inefficient DNA replication in regions of high stability or the transient stabilization of misaligned intermediates.

Base Pairing↗

Proteome composition in Plasmodium falciparum: higher usage of GC-rich nonsynonymous codons in highly expressed genes.

The parasite Plasmodium falciparum, responsible for the most deadly form of human malaria, is one of the extremely AT-rich genomes sequenced so far and known to possess many atypical characteristics. Using multivariate statistical approaches, the present study analyzes the amino acid usage pattern in 5038 annotated protein-coding sequences in P. falciparum clone 3D7. The amino acid composition of individual proteins, though dominated by the directional mutational pressure, exhibits wide variation across the proteome. The Asn content, expression level, mean molecular weight, hydropathy, and aromaticity are found to be the major sources of variation in amino acid usage. At all stages of development, frequencies of residues encoded by GC-rich codons such as Gly, Ala, Arg, and Pro increase significantly in the products of the highly expressed genes. Investigation of nucleotide substitution patterns in P. falciparum and other Plasmodium species reveals that the nonsynonymous sites of highly expressed genes are more conserved than those of the lowly expressed ones, though for synonymous sites, the reverse is true. The highly expressed genes are, therefore, expected to be closer to their putative ancestral state in amino acid composition, and a plausible reason for their sequences being GC-rich at nonsynonymous codon positions could be that their ancestral state was less AT-biased. Negative correlation of the expression level of proteins with respective molecular weights supports the notion that P. falciparum, in spite of its intracellular parasitic lifestyle, follows the principle of cost minimization.

Amino Acids↗

Frequencies of codons in histones, tubulins and fibrinogen: bias due to interference between transcription signals and protein function.

The distribution of codons was studied in 65 proteins: 48 histones, 14 tubulins, and three fibrinogens, With the methodology used, (1) we confirmed that the preterminator state of a codon has no detectable effect on codon bias. (2) The well-known effect of CG suppression was visible. We also found that (3) some codons which are very rare, are equal to parts of known transcription signals. Thus, we advanced that to avoid signal interference, the use of these codons is suppressed when a synonymous codon is available. In addition we found that in the whole series of codons, transcription signals are less frequent than in a random sequence of equal composition. Finally we observed (4) that tryptophan is absent in histones. This absence was related not to the TGG codon itself, but to characteristics of the amino acid. We conclude that the functional constraints of a protein can influence, at least for synonymous codon usage, the evolution of its own coding sequence.

Animals↗

[Analysis of apolipoprotein gene family in codon space--non-random selection of nucleotide changes in evolution].

The choice of nucleotide changes in DNA evolution can be either selectively neutral or biased. To study how apolipoprotein gene selects the nucleotide substitutions in the course of evolution, a codon space is constructed in which its DNA sequence can be mapped as a matrix of nucleotide frequencies in three codon positions. Accordingly, a number of methods that measure the nonrandomness of nucleotide distribution in codon space are developed based on maximum entropy techniques to define the nature of nucleotide change selection in evolution. By these methods, we demonstrated that the nucleotide composition in 1st and 3rd codon position of apolipoprotein genes is highly nonrandom, which appears to be a result of non-neutral selection of codon positions by adenosine and thymidine. In addition, this paper is also concerned in the divergence of synonymos codon usage and its correlation to taxonomic distances among species. As a result, a codon usage clock was reported in apolipoprotein A-I. Our studies suggest that non-random selection of nucleotide changes in codon space may represent an evolutionary characteristics of apolipoprotein genes.

Animals↗

The B cell response to autologous type II collagen: biased V gene repertoire with V gene sharing and epitope shift.

Collagen-induced arthritis is an autoimmune model disease induced in the DBA/1 mouse immunized with type II collagen (CII). Both T and B cells play a critical role for the induction of arthritis. Draining lymph nodes from CII-immunized mice contain high numbers of CII-specific B cells, which are isotype switched and V gene selected. In the present study we analyze the V region gene usage and epitope specificity of CII-reactive B cell hybridomas, randomly isolated from the primary and the secondary response in mice immunized with rat CII we make the following conclusions. 1) There are major epitopes in the native CII molecule to which the B cells preferentially respond. 2) B cells specific for the same epitope show a preferential pairing of certain VH/VK genes or a biased usage of individual VH (VHJ558 and VHX24) or VK genes (VK21). 3) The V genes are germ line encoded in the primary response and somatically mutated in the secondary response. Somatic mutations give the Abs cross-reactivity between CII epitopes, and epitope shift, i.e., another epitope within the CII molecule is recognized. 4) There is a sharing of certain V genes in B cell clones specific for different epitopes, indicating structural similarities of the different CII epitopes.

Amino Acid Sequence↗

Gene expression levels influence amino acid usage and evolutionary rates in endosymbiotic bacteria.

Most endosymbiotic bacteria have extremely reduced genomes, accelerated evolutionary rates, and strong AT base compositional bias thought to reflect reduced efficacy of selection and increased mutational pressure. Here, we present a comparative study of evolutionary forces shaping five fully sequenced bacterial endosymbionts of insects. The results of this study were three-fold: (i) Stronger conservation of high expression genes at not just nonsynonymous, but also synonymous, sites. (ii) Variation in amino acid usage strongly correlates with GC content and expression level of genes. This pattern is largely explained by greater conservation of high expression genes, leading to their higher GC content. However, we also found indication of selection favoring GC-rich amino acids that contrasts with former studies. (iii) Although the specific nutritional requirements of the insect host are known to affect gene content of endosymbionts, we found no detectable influence on substitution rates, amino acid usage, or codon usage of bacterial genes involved in host nutrition.

AT Rich Sequence↗

Biased usages of arginines and lysines in proteins are correlated with local-scale fluctuations of the G + C content of DNA sequences.

Amino acid residues arginine (R) and lysine (K) have similar physicochemical characteristics and are often mutually substituted during evolution without affecting protein function. Statistical examinations on human proteins show that more R than K residues are used in the proximity of R residues, whereas more K than R are used near K residues. This biased use occurs on both a global and a local scale (shorter than approximately 100 residues). Even within a given exon, G + C-rich and A + T-rich short DNA segments preferentially encode R and K, respectively. The biased use of R and K on a local scale is also seen in Saccharomyces cerevisiae and Caenorhabdidtis elegans, which lack global-scale mosaic structures with varying GC%, or isochores. Besides R and K, several amino acids are also used with a positive or negative correlation with the local GC% of third codon bases. The local-, or "within-gene"-, scale heterogeneity of the DNA sequence may influence the sequence of the encoded protein segment.

Amino Acid Substitution↗

Synonymous substitution rates in Drosophila: mitochondrial versus nuclear genes.

Synonymous substitution rates in mitochondrial and nuclear genes of Drosophila were compared. To make accurate comparisons, we considered the following: (1) relative synonymous rates, which do not require divergence time estimates, should be used; (2) methods estimating divergence should take into account base composition; (3) only very closely related species should be used to avoid effects of saturation; (4) the heterogeneity of rates should be examined. We modified the methods estimating synonymous substitution numbers to account for base composition bias. By using these methods, we found that mitochondrial genes have 1.7-3.4 times higher synonymous substitution rates than the fastest nuclear genes or 4.5-9.0 times higher rates than the average nuclear genes. The average rate of synonymous transversions was 2.7 (estimated from the melanogaster species subgroup) or 2.9 (estimated from the obscura group) times higher in mitochondrial genes than in nuclear genes. Synonymous transversions in mitochondrial genes occurred at an approximately equivalent rate to those in the fastest nuclear genes. This last result is not consistent with the hypothesis that the difference in turnover rates between mitochondrial and nuclear genomes is the major factor determining higher synonymous substitution rates in mtDNA. We conclude that the difference in synonymous substitution rates is due to a combination of two factors: a higher transitional mutation rate in mtDNA and constraints on nuclear genes due to selection for codon usage.

Animals↗

OspA, a lipoprotein antigen of the obligate intracellular bacterial pathogen Piscirickettsia salmonis.

No effective recombinant vaccines are currently available for any rickettsial diseases. In this regard the first non-ribosomal DNA sequences from the obligate intracellular pathogen Piscirickettsia salmonis are presented. Genomic DNA isolated from Percoll density gradient purified P. salmonis, was used to construct an expression library in lambda ZAP II. In the absence of preexisting DNA sequence, rabbit polyclonal antiserum raised against P. salmonis, with a bias toward P. salmonis surface antigens, was used to identify immunoreactive clones. Catabolite repression of the lac promoter was required to obtain a stable clone of a 4,983 bp insert in Escherichia coli due to insert toxicity exerted by the accompanying radA open reading frame (ORF). DNA sequence analysis of the insert revealed 1 partial and 4 intact predicted ORF's. A 486 bp ORF, ospA, encoded a 17 kDa antigenic outer surface protein (OspA) with 62% amino acid sequence homology to the genus common 17 kDa outer membrane lipoprotein of Rickettsia prowazekii, previously thought confined to members of the genus Rickettsia. Palmitate incorporation demonstrated that OspA is posttranslationally lipidated in E. coli, albeit poorly expressed as a lipoprotein even after replacement of the signal sequence with the signal sequence from lpp (Braun lipoprotein) or the rickettsial 17 kDa homologue. To enhance expression, ospA was optimized for codon usage in E. coli by PCR synthesis. Expression of ospA was ultimately improved (approximately 13% of total protein) with a truncated variant lacking a signal sequence. High level expression (approximately 42% tot. prot.) was attained as an N-terminal fusion protein with the fusion product recovered as inclusion bodies in E. coli BL21. Expression of OspA in P. salmonis was confirmed by immunoblot analysis using polyclonal antibodies generated against a synthetic peptide of OspA (110-129) and a strong antibody response against OspA was detected in convalescent sera from coho salmon (Oncorhynchus kisutch).

Amino Acid Sequence↗

Evidence of biased immunoglobulin variable gene usage in highly stable B-cell chronic lymphocytic leukemia.

Recognition of biased immunoglobulin variable (IgV) gene usage in B-cell chronic lymphocytic leukemia (B-CLL) may yield insight into leukemogenesis and may help to refine prognostic categories. We explored Ig variable heavy (VH) and light (VL) chain gene usage in highly stable and indolent B-CLL (n=25) who never required treatment over 10 or more years. We observed an unexpectedly high usage of mutated VH3-72 (6/25; 24.0%), a gene that was otherwise rare in B-CLL (7/805; 0.87%; P<0.01), including mutated cases (6/432; 1.39%; P<0.01) and was exceptional among indolent (1/230, 0.435%; P<0.01), and aggressive B-cell lymphomas (0/105; P<0.01). Three of six VH3-72 B-CLL cases utilized the same VL Vkappa4-1 gene. Two V(H)3-72 B-CLL cases had highly homologous VH complementarity determining regions 3 (CDR3s), encoding Cys-XXXX-Cys domains, and utilized Vkappa4-1 genes with homologous IgVL CDR3s. An identical threonine to isoleucine change at codon 84 of V(H)3-72 framework region 3 (FR3) recurred in four cases of highly stable VH3-72 B-CLL. This mutation is expected to cause a conformational change of FR3 proximal to CDR3 that might critically affect high-affinity antigen binding. B-cell receptors encoded by VH3-72 may identify a specific B-CLL group and be implicated in leukemogenesis through an antigen-driven expansion of B cells.

Amino Acid Sequence↗

VH gene analysis of primary cutaneous B-cell lymphomas: evidence for ongoing somatic hypermutation and isotype switching.

Primary cutaneous B-cell lymphomas are B-cell non-Hodgkin's lymphomas that arise in the skin. The major subtypes discerned are follicle center cell lymphomas, immunocytomas (marginal zone B-cell lymphomas), and large B-cell lymphomas of the leg. In this study, we analyzed the variable heavy chain (VH) genes of 7 of these lymphomas, ie, 4 follicle center cell lymphomas (diffuse large-cell lymphomas) and 3 immunocytomas. We show that all these lymphomas carry heavily mutated VH genes, with no obvious bias in VH gene usage. The low ratios of replacement versus silent mutations observed in the framework regions of 5 of the 7 lymphomas suggest that the structure of the B-cell antigen receptor was preserved, as in normal B cells that are selected for antibody expression. Moreover, evidence for ongoing mutation was obtained in 3 immunocytomas and in one lymphoma of large-cell type. In addition, in 1 immunocytoma, both IgG- and IgA-expressing clones were found, indicative of isotype switching. Our data provide insight into the biology of primary cutaneous B-cell lymphomas and may be of significance for their classification.

Codon↗

CRITICA: coding region identification tool invoking comparative analysis.

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Algorithms↗

Immunoglobulin light chain variable region gene sequences for human antibodies to Haemophilus influenzae type b capsular polysaccharide are dominated by a limited number of V kappa and V lambda segments and VJ combinations.

The immune repertoire to Haemophilus influenzae type b capsular polysaccharide (Hib PS) appears to be dominated by certain light chain variable region genes (IgVL). In order to examine the molecular basis underlying light chain bias, IgVL genes have been cloned from a panel of heterohybridomas secreting human anti-Hib PS (antibody) (anti-Hib PS Ab). One hybridoma, representative of the predominant serum clonotype of anti-Hib PS Ab in older children and adults following immunization or Hib infection, uses a V kappa II segment identical to the germline gene A2, and a JK3 segment. A second kappa hybridoma uses a member of the V kappa I family and a JK4 segment. Four lambda antibodies, all cross-reactive with the structurally related antigen Escherichia coli K100 PS, use V lambda VII segments which are 96-98% homologous to one another, and may originate from a single germline gene. Two additional lambda antibodies, not K100-cross-reactive, are encoded by members of the V lambda II family. All lambda antibodies use highly homologous J lambda 2 or J lambda 3 segments. The VJ joints of all lambda antibodies and the V kappa II-encoded antibody are notable for the presence of an arginine codon, suggesting an important role in antigen binding. Although more complex than heavy chain variable region gene usage, a significant portion of serum anti-Hib PS Ab is likely to be encoded by a limited number of V kappa and V lambda segments and VJ combinations, which may be selectively expressed during development, or following antigen exposure.

Adult↗

Reduction of wobble-position GC bases in Corynebacteria genes and enhancement of PCR and heterologous expression.

Corynebacteria codon usage exhibits an overall GC content of 67%, and a wobble-position GC content of 88%. Escherichia coli, on the other hand has an overall GC content of 51%, and a wobble-position GC content of 55%. The high GC content of Corynebacteria genes results in an unfavorable codon preference for heterologous expression, and can present difficulties for polymerase-based manipulations due to secondary-structure effects. Since these characteristics are due primarily to base composition at the wobble-position, synthetic genes can, in principle, be designed to eliminate these problems and retain the wild-type amino acid sequence. Such genes would obviate the need for special additives or bases during in vitro polymerase-based manipulation and mutant host strains containing uncommon tRNA's for heterologous expression. We have evaluated synthetic genes with reduced wobble-position G/C content using two variants of the enzyme 2,5-diketo-D-gluconic acid reductase (2,5-DKGR A and B) from Corynebacterium. The wild-type genes are refractory to polymerase-based manipulations and exhibit poor heterologous expression in enteric bacteria. The results indicate that a subset of codons for five amino acids (alanine, arginine, glutamate, glycine and valine) contribute the greatest contribution to reduction in G/C content at the wobble-position. Furthermore, changes in codons for two amino acids (leucine and proline) enhance bias for expression in enteric bacteria without affecting the overall G/C content. The synthetic genes are readily amplified using polymerase-based methodologies, and exhibit high levels of heterologous expression in E. coli.

Base Composition↗

A strong propensity toward loop formation characterizes the expressed reading frames of the D segments at the Ig H and T cell receptor loci.

A compilation of murine and human Ig H and TcR beta D segment sequences was used to estimate the relative usage of the various reading frames and to look for associated sequence patterns. We confirm a strong bias in the expression of the Ig H D segments, with more than 90% (murine) and 85% (human) expressed peptides resulting from a preferred reading frame. Remarkably, 86% (mouse) and 90% (human) of those peptides contain at least one glycine residue. All but one of the atypical preferred D peptides contain serine or proline residues and are found in the immediate vicinity of glycine residues provided by specific JH segments. The presence of tyrosine residues is also a characteristic feature of expressed reading frames in both mouse (75%) and human (90%). These results suggest that the constraints of forming a flexible loop within the third complementarity-determining region, is a factor in the preference for a particular reading frame in Ig H D. For the TcR beta D segments, glycine is specified in most reading frames, and no significant preference is observed.

Amino Acid Sequence↗