Search PubMedSearch

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569 bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A + T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU > 1) with a pronounced preference for A/C-ending codons. The ΔRSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean = 1.209 ± 0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae

Design, synthesis and expression of a human interleukin-2 gene incorporating the codon usage bias found in highly expressed Escherichia coli genes.

A synthetic gene encoding human interleukin-2 (IL-2) was designed such that the codon usage bias resembled that found in highly expressed Escherichia coli genes. The percentage of preferred codons was increased from 43% in the native cDNA sequence to 85% in the synthetic sequence. The cDNA and synthetic IL-2 genes were placed under the control of the trc promoter and expressed in E. coli JM101. While Northern blot analysis of IL-2 mRNA from each genetic construct demonstrated equivalent message half-lives, immunoblot and bioactivity analyses showed the synthetic gene to direct the synthesis of up to 16 times more IL-2 than the native cDNA sequence.

Amino Acid Sequence

Codon usage determines translation rate in Escherichia coli.

We wish to determine whether differences in translation rate are correlated with differences in codon usage or with differences in mRNA secondary structure. We therefore inserted a small DNA fragment in the lacZ gene either directly or flanked by a few frame-shifting bases, leaving the reading frame of the lacZ gene unchanged. The fragment was chosen to have "infrequent" codons in one reading frame and "common" codons in the other. The insert in these constructs does not seem to give mRNAs that are able to form extensive secondary structures. The translation time for these modified lacZ mRNAs was measured with a reproducibility better than plus or minus one second. We found that the mRNA with infrequent codons inserted has an approximately three-seconds longer translation time than the one with common codons. In another set of experiments we constructed two almost identical lacZ genes in which the lacZ mRNAs have the potential to generate stem structures with stabilities of about -75 kcal/mol. In this way we could investigate the influence of mRNA structure on translation rate. This type of modified gene was generated in two reading frames with either common or infrequent codons similar to our first experiments. We find that the yield of protein from these mRNAs is reduced, probably due to the action in vivo of an RNase. Nevertheless, the data do not indicate that there is any effect of mRNA secondary structure on translation rate. In contrast, our data persuade us that there is a difference in translation rate between infrequent codons and common codons that is of the order of sixfold.

Bacterial Proteins

Codon usage can affect efficiency of translation of genes in Escherichia coli.

By inserting synthetic oligonucleotides into a highly expressed gene in E. coli it has been shown that unfavourable codon usage can reduce the maximum translation rate of a protein. However, in the case of the codon used (AGG), a significant effect on translation was only seen at very high transcription rates from a gene containing multiple copies of the unfavourable codon.

Bacterial Proteins

The relationship between base composition and codon usage in bacterial genes and its use for the simple and reliable identification of protein-coding sequences.

Bacterial genes that code for proteins appear to possess a codon usage characteristic of their overall base composition. This results in different but predictable non-random distributions of nucleotides within codons, permitting the recognition of protein-coding sequences in a wide range of bacterial species. The nature of this distribution depends on the base composition of the coding sequence. The position-specific differences are especially conspicuous in genes of extreme G + C content, allowing the particularly reliable prediction of the reading frame and coding strand of experimentally determined DNA sequences. This finding has been exploited to identify the coding sequence of the viomycin phosphotransferase (vph) gene of Streptomyces vinaceus. An easily applied computer program ("Frame") has been written to carry out and display such analyses.

Bacterial Proteins

Nucleotide sequence and codon usage of the elongation factor Tu(EF-Tu) gene from Mycoplasma pneumoniae.

The Mycoplasma pneumoniae tuf gene, encoding the elongation factor protein Tu, was cloned and sequenced. The nucleotide sequence of the mycoplasmal gene showed about 60% homology to the sequences of tuf genes of other prokaryotes, yeast mitochondria and Euglena gracilis chloroplasts, and about 75% similarity was found when comparing the deduced amino acid sequences of the various Tu proteins. The relatively low G + C content (40%) of the M. pneumoniae DNA was reflected in a low G + C content (44.6%) of the tuf gene, and in a preferential use of adenine and uracil at the third position of codons, yet codon usage analysis revealed the presence of almost all of the codons of the genetic code in the mycoplasmal gene. Southern blot hybridization of digested DNAs of 11 Mollicutes species with the entire M. pneumoniae tuf gene and with its 5' part suggested the presence of one copy only of this gene in the representative species of the Mollicutes. In this respect, the Mollicutes resemble Gram-positive bacteria and differ from the Gram-negative bacteria, which carry two copies of the tuf gene.

Amino Acid Sequence

Codon usage in Plasmodium falciparum.

The codon frequencies used in 7874 codons from 17 sequences of Plasmodium falciparum have been examined. The frequency distribution is markedly biased. A and C occur with similar frequency in all positions but G is predominantly in the first base and T is predominantly in the last position. This information can be used to predict the coding strand and reading frame of P. falciparum genes.

Animals

Comparison of dinucleotide frequency and codon usage in Toxoplasma and Plasmodium: evolutionary implications.

The weight-averaged observed/expected dinucleotide frequencies for the sum total of the coding regions of five Toxoplasma genes were compared with the same parameters previously determined for the coding regions of 21 Plasmodium genes. In addition, codon usage in the five Toxoplasma genes was compared with that in the 21 Plasmodium genes, and the percent distribution of amino acids in the Toxoplasma protein pool and the Plasmodium protein pool were compared with that in a general protein pool of 314 proteins. The results are consistent with the hypothesis that, contrary to currently held opinion, the genera Toxoplasma and Plasmodium are not especially closely related.

Animals

Evident diversity of codon usage patterns of human genes with respect to chromosome banding patterns and chromosome numbers; relation between nucleotide sequence data and cytogenetic data.

The sequences of the human genome compiled in DNA databases are now about 10 megabase pairs (Mb), and thus the size of the sequences is several times the average size of chromosome bands at high resolution. By surveying this large quantity of data, it may be possible to clarify the global characteristics of the human genome, that is, correlation of gene sequence data (kb-level) to cytogenetic data (Mb-level). By extensively searching the GenBank database, we calculated codon usages in about 2000 human sequences. The highest G + C percentage at the third codon position was 97%, and that of about 250 sequences was 80% or more. The lowest G + C% was 27%, and that in about 150 sequences was 40% or less. A major portion of the GC-rich genes was found to be on special subsets of R-bands (T-bands and/or terminal R-bands). AT-rich genes, however, were mainly on G-bands or non-T-type internal R-bands. Average G + C% at the third position for individual chromosomes differed among chromosomes, and were related to T-band density, quinacrine dullness, and mitotic chiasmata density in the respective chromosomes.

Base Composition

Correlation between molecular clock ticking, codon usage fidelity of DNA repair, chromosome banding and chromatin compactness in germline cells.

The vertebrate genome is built of long DNA regions, relatively homogeneous in GC content, which likely correspond to bands on stained chromosomes. Large differences in composition have been found among DNA regions belonging to the same genome. They are paralleled by differences in codon usage in genes differently localized. The hypothesis presented here asserts that these differences in composition are caused by different mutational bias of alpha and beta DNA polymerases, these polymerases being involved to different extents in the repair of DNA lesions in compact and relaxed chromatin, respectively, in germline cells.

Animals

Spectinomycin operon of Micrococcus luteus: evolutionary implications of organization and novel codon usage.

The complete DNA sequence of the Micrococcus luteus spectinomycin (spc) operon and its adjacent regions has been determined. The sequence has revealed the presence of genes that are homologous to those of the Escherichia coli ribosomal and related proteins, L14, L24, L5, S8, L6, L18, S5, L30, L15, and secretion protein Y (sec Y), and the gene for adenylate kinase (adk). The gene arrangement in the spc operon is essentially the same as that of E. coli except for the absence in the M. luteus spc operon of the genes for S14 and X protein that exist in the E. coli spc operon. SecY and adk seem to be composed of another operon (adk operon) with at least an open reading frame. The deduced amino acid sequences for these ribosomal proteins are well conserved among the two species (40-65% identity). Reflecting the high genomic guanine and cytosine (GC) content of M. luteus (74%), the codon usage of the genes is extremely biased toward use of G and C, about 94% of the codon third positions being G or C. Seven codons, AUA, AAA, AGA, UUA, GUA, CUA, and CAA, all of which have A at the codon third positions, are completely absent in the M. luteus genes examined. Out of 11 genes in the M. luteus spc and adk operons, 5 (10) use GUG (UGA) and 6 (1) use AUG (UAA) as an initiation (termination) codon.

Amino Acid Sequence

Structural features of multiple nifH-like sequences and very biased codon usage in nitrogenase genes of Clostridium pasteurianum.

The structural gene (nifH1) encoding the nitrogenase iron protein of Clostridium pasteurianum has been cloned and sequenced. It is located on a 4-kilobase EcoRI fragment (cloned into pBR325) that also contains a portion of nifD and another nifH-like sequence (nifH2). C. pasteurianum nifH1 encodes a polypeptide (273 amino acids) identical to that of the isolated iron protein, indicating that the smaller size of the C. pasteurianum iron protein does not result from posttranslational processing. The 5' flanking region of nifH1 or nifH2 does not contain the nif promoter sequences found in several gram-negative bacteria. Instead, a sequence resembling the Escherichia coli consensus promoter (TTGACA-N17-TATAAT) is present before C. pasteurianum nifH2, and a TATAAT sequence is present before C pasteurianum nifH1. Codon usage in nifH1, nifH2, and nifD (partial) is very biased. A preference for A or U in the third position of the codons is seen. nifH2 could encode a protein of 272 amino acid residues, which differs from the iron protein (nifH1 product) in 23 amino acid residues (8%). Another nifH-like sequence (nifH3) is located on a nonadjacent EcoRI fragment and has been partially sequenced. C. pasteurianum nifH2 and nifH3 may encode proteins having several amino acids that are conserved in other proteins but not in C. pasteurianum iron protein, suggesting a possible role for the multiple nifH-like sequences of C. pasteurianum in the evolution of nifH. Among the nine sequenced iron proteins, only the C. pasteurianum protein lacks a conserved lysine residue which is near the extended C terminus of the other iron proteins. The absence of this positive charge in the C. pasteurianum iron protein might affect the cross-reactivity of the protein in heterologous systems.

Amino Acid Sequence

Human hemoglobin expression in Escherichia coli: importance of optimal codon usage.

The overexpression of a nonfusion product of human beta-globin in Escherichia coli from its cDNA sequence has been accomplished for the first time. Expression of beta-globin from its native cDNA required the use of the strong bacteriophage T7 promoter. In this system, beta-globin accumulated to approximately 10% of total E. coli proteins. alpha-Globin was not expressed in the T7 system using the native cDNA. For the expression of alpha-globin, synthetic genes containing optimal E. coli codons were constructed. Neither synthetic alpha- nor beta-globin gene alone was expressed from the lac or tac promoter. Globin expression was achieved when the two synthetic alpha- and beta-globin genes were combined as an operon downstream of the lac promoter. The two proteins combined intracellularly with endogenous heme, which was concomitantly overproduced to yield tetrameric hemoglobin as roughly 5-10% of total E. coli protein. Cloning the alpha- and beta-globin cDNAs in a construct identical with the lac promoter did not yield globin production, establishing the requirement for optimal codon usage. The recombinant beta-globin from the T7 expression system was purified and reconstituted in vitro with heme and native alpha chains. N-terminal analyses showed that the beta-globin produced in the T7 system and the tetrameric hemoglobin produced from the synthetic genes contained an additional beta 1 methionine residue. Two additional mutants, beta 1 Val----Met and beta 1 Val----Ala were produced using the T7 system. Functional and structural properties of the purified hemoglobins will be discussed in the following papers.

Amino Acid Sequence

Nucleotide sequence of simian virus 40 DNA: structure of the middle segment of the HindII + III restriction fragment B (sixth part of the T antigen gene) and codon usage.

We report here the nucleotide sequence of the simian virus 40 DNA region that lies between the EcoRII restriction endonuclease cleavage sites at map positions 0.214 and 0.281. The sequence was determined by partial chemical degradation of terminally labeled DNA fragments according to the procedure of Maxam and Gilbert. This region represents 6.7% of the SV40 genome and is located in the middle of HindII + III restriction fragment B. It is expressed as part of the early 19-S messenger RNA, which codes for the large-T antigen protein. Only one open reading frame for translation can be deduced from the message strand of the DNA and this reading frame connects in phase with the one of both neighboring fragments. This publication is the last in a series of papers about the T-antigen gene, and several properties of this gene and its product are discussed. The non-randomness of codon usage is similar to that previously discussed for the late part of the genome. Moreover, it appears that the choice of a third letter can be determined by the nature of the following codon; some codons which start with a pyrimidine are almost never preceded by an adenosine and some ANN-type codons are almost never preceded by a guanosine.

Amino Acid Sequence

Codon usage and genome composition.

The GC levels of codon third positions from 49 genomes covering a wide phylogenetic range are linearly correlated with the GC levels of the corresponding genomes. Three different relationships have been found: one for prokaryotes and viruses, one for lower eukaryotes, and one for vertebrates. All points not fitting the first relationship can be brought into quasi coincidence with it when plotted against GC levels of coding sequences.

Animals

On codon usage.

Explore the source record for details and available documents.

Codon