Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon usage bias”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Reassortment and concerted evolution in banana bunchy top virus genomes.

The nanovirus Banana bunchy top virus (BBTV) has six standard components in its genome and occasionally contains components encoding additional Rep (replication initiation protein) genes. Phylogenetic network analysis of coding sequences of DNA 1 and 3 confirmed the two major groups of BBTV, a Pacific and an Asian group, but show evidence of web-like phylogenies for some genes. Phylogenetic analysis of 102 major common regions (CR-Ms) from all six components showed a possible concerted evolution within the Pacific group, which is likely due to recombination in this region. The CR-M of additional Rep genes is close to that of DNA 1 and 2. Comparison of tree topologies constructed with DNA 1 and DNA 3 coding sequences of 14 BBTV isolates showed distinct phylogenetic histories based on Kishino-Hasegawa and Shimodaira-Hasegawa tests. The results of principal component analysis of amino acid and codon usages indicate that DNA 1 and 3 have a codon bias different from that of all other genes of nanoviruses, including all currently known additional Rep genes of BBTV, which suggests a possible ancient genome reassortment event between distinctive nanoviruses.

Asia↗

The complex evolution and genomic dynamics of mating-type loci in Cryptococcus and Kwoniella.

Sexual reproduction in basidiomycete fungi is governed by MAT loci (P/R and HD), which exhibit remarkable evolutionary plasticity, characterized by expansions, rearrangements, and gene losses often associated with mating system transitions. The sister genera Cryptococcus and Kwoniella provide a powerful framework for studying MAT loci evolution owing to their diverse reproductive strategies and distinct architectures, spanning bipolar and tetrapolar systems with either linked or unlinked MAT loci. Building on recent comparative genomic analyses, we generated additional chromosome-level assemblies, uncovering distinct trajectories shaping MAT loci organization. Contrasting with the small-scale expansions and gene acquisitions observed in Kwoniella, our analyses revealed independent expansions of the P/R locus in tetrapolar Cryptococcus, possibly driven by pheromone gene duplications. Notably, these expansions coincided with a pronounced GC-content reduction best explained by reduced GC-biased gene conversion following recombination suppression, rather than relaxed codon usage selection. Diverse modes of MAT locus linkage were also identified, including three previously unrecognized transitions: one resulting in a pseudobipolar arrangement and two leading to bipolarity. All three transitions involved translocations. In the pseudobipolar configuration, the P/R and HD loci remained on the same chromosome but genetically unlinked, whereas the bipolar transitions additionally featured rearrangements that fused the two loci into a nonrecombining region. Mating assays confirmed a sexual cycle in Cryptococcus decagattii, demonstrating its ability to undergo mating and sporulation. Progeny analysis in Kwoniella mangrovensis revealed substantial ploidy variation and aneuploidy, likely stemming from haploid-diploid mating, yet evidence of recombination and loss of heterozygosity indicates that meiotic exchange occurs despite irregular chromosome segregation. Our findings underscore the importance of continued diversity sampling and provide further evidence for convergent evolution of fused MAT loci in basidiomycetes, offering new insights into the genetic and chromosomal changes driving reproductive transitions.

Genes, Mating Type, Fungal↗

The preferential codon usages in variable and constant regions of immunoglobulin genes are quite distinct from each other.

The pattern of codon utilization in the variable and constant regions of immunoglobulin genes are compared. It is shown that, in these regions, codon utilizations are quite distinct from one another: For most degenerate codons, there is a selective bias that prefers C and/or G ending codons to U and/or A ending codons in the constant region compared with the bias in the variable region. This would strongly suggest that, in immunoglobulin genes, the bias in code word usage is determined by other factors than those concerning with the translational mechanism such as tRNA availability and codon-anticodon interaction. A possibility is also suggested that this differance of code word usage between them is due to the existence of secondary structure in the constant region but not in the variable region.

Anticodon↗

Nucleotide sequence of the genomic region encompassing Adh and Adh-dup genes of D. lebanonensis (Scaptodrosophila): gene expression and evolutionary relationships.

The region of the genome of D. lebanonensis that contains the Adh gene and the downstream Adh-dup gene was sequenced. The structure of the two genes is the same as has been described for D. melanogaster. Adh has two promoters and Adh-dup has only one putative promoter. The levels of expression of the two genes in this species are dramatically different. Hybridizing the same Northern blots with a specific probe for Adh-dup, we did not find transcripts for this gene in D. lebanonensis. The level of Adh distal transcript in adults of D. lebanonensis is five times greater than that of D. melanogaster adults. The maximum levels of proximal transcript are attained at different larval stages in the two species, being three times higher in D. melanogaster late-second-instar larvae than in D. lebanonensis first-instar larvae. The level of Adh transcripts allowed us to determine distal and proximal initiation transcription sites, the position of the first intron, the use of two polyadenylation signals, and the heterogeneity of polyadenylation sites. Temporal and spatial expression profiles of the Adh gene of D. lebanonensis show qualitative differences compared with D. melanogaster. Adh and Adh-dup evolve differently as shown by the synonymous and nonsynonymous substitution rates for the coding region of both genes when compared across two species of the melanogaster group, two of the obscura group of the subgenus Sophophora and D. lebanonensis of the victoria group of the subgenus Scaptodrsophila. Synonymous rates for Adh are approximately half those for Adh-dup, while nonsynonymous rates for Adh are generally higher than those for Adh-dup. Adh shows 76.8% identities at the protein level and 70.2% identities at the nucleotide level while Adh-dup shows 83.7% identities at the protein level and 67.5% identities at the nucleotide level. Codon usage for Adh-dup is shown to be less biased than for Adh, which could explain the higher synonymous rates and the generally lower nonsynonymous substitution rates in Adh-dup compared with Adh. Phylogenetic trees reconstructed by distance matrix and parsimony methods show that Sophophora and Scaptodrosophila subgenera diverged shortly after the separation from the Drosophila subgenus.

Alcohol Dehydrogenase↗

A cluster of vitellogenin genes in the Mediterranean fruit fly Ceratitis capitata: sequence and structural conservation in dipteran yolk proteins and their genes.

Four genes encoding the major egg yolk polypeptides of the Mediterranean fruit fly Ceratitis capitata, vitellogenins 1 and 2 (VG1 and VG2), were cloned, characterized and partially sequenced. The genes are located on the same region of chromosome 5 and are organized in pairs, each encoding the two polypeptides on opposite DNA strands. Restriction and nucleotide sequence analysis indicate that the gene pairs have arisen from an ancestral pair by a relatively recent duplication event. The transcribed part is very similar to that of the Drosophila melanogaster yolk protein genes Yp1, Yp2 and Yp3. The Vg1 genes have two introns at the same positions as those in D. melanogaster Yp3; the Vg2 genes have only one of the introns, as do D. melanogaster Yp1 and Yp2. Comparison of the five polypeptide sequences shows extensive homology, with 27% of the residues being invariable. The sequence similarity of the processed proteins extends in two regions separated by a nonconserved region of varying size. Secondary structure predictions suggest a highly conserved secondary structure pattern in the two regions, which probably correspond to structural and functional domains. The carboxy-end domain of the C. capitata proteins shows the same sequence similarities with triacyglycerol lipases that have been reported previously for the D. melanogaster yolk proteins. Analysis of codon usage shows significant differences between D. melanogaster and C. capitata vitellogenins with the latter exhibiting a less biased representation of synonymous codons.

Alleles↗

Sequence comparison of ACE-1, the gene encoding acetylcholinesterase of class A, in the two nematodes Caenorhabditis elegans and Caenorhabditis briggsae.

The ace-1 gene, which encodes acetylcholinesterase of class A, has been cloned and sequenced in C. briggsae and compared to its homologue in C. elegans. Both genes present an open reading frame of 1860 nucleotides. The percentages of identity are 80% and 95% at the nucleotide and aminoacid levels respectively. All residues characteristic of an acetylcholinesterase are found in conserved positions in C. briggsae ACE-1. The deduced C-terminus is hydrophilic, thus resembling the catalytic peptide T of vertebrate cholinesterases. Codon usage in both ace-1 genes appears to be lowly biased. This may indicate that these genes are lowly expressed. The splicing sites of the eight introns of ace-1 in C. elegans are conserved in C. briggsae, but introns are shorter in C. briggsae. No homology was found between intronic sequences in both species, except for the consensus border sequences.

Acetylcholinesterase↗

Evidence for codon bias selection at the pre-mRNA level in eukaryotes.

We investigated codon usage patterns across eukaryotic exons. We have shown that in humans codon preference varies with distance from the splice sites. This is consistent with the distribution of RNA elements involved in splicing regulation. Our results provide the first evidence that selection at the pre-mRNA level influences codon usage in humans. We also show that systematic trends in codon usage are found in other eukaryotes, suggesting that pre-mRNA level selection for codon usage could be a widespread phenomenon in organisms that undergo RNA splicing.

Animals↗

Low diversity and divergence in the fil1 gene family of Antirrhinum (Scrophulariaceae).

Detailed nucleotide diversity studies revealed that the fil1 gene of Antirrhinum, which has been reported to be single copy, is a member of a gene family composed of at least five genes. In four Antirrhinum majus populations with different mating systems and one A. graniticum population, diversity within populations is very low. Divergence among Antirrhinum species and between Antirrhinum and Digitalis is also low. For three of these genes we also obtained sequences from a more divergent member of the Scrophulariaceae, Verbascum nigrum. Compared with Antirrhinum, little divergence is again observed. These results, together with similar data obtained previously for five cycloidea genes, suggest either that these gene families (or the Antirrhinum genome) are unusually constrained or that there is a low rate of substitution in these lineages. Using a sample of 52 genes, based on two measures of codon usage (ENC and GC3 content), we show that cyc and fil1 are among the least biased Antirrhinum genes, so that their low diversity is not due to extreme codon bias.

Arabidopsis Proteins↗

Equal G and C contents in histone genes indicate selection pressures on mRNA secondary structure.

Protein-specific versus taxon-specific patterns of nucleotide frequencies were studied in histone genes. The third positions of codons have a (well-known) taxon-specific G+C level and a histone type-specific G/C ratio. This ratio counterbalances the G/C ratio in the first and second positions so that the overall G and C levels in the coding region become approximately equal. The compensation of the G/C ratio indicates a selection pressure at the mRNA level rather than a selection pressure or mutation bias at the DNA level or a selection pressure on codon usage. The structure of histone mRNAs is compatible with the hypothesis that the G/C compensation is due to selection pressures on mRNA secondary structure. Nevertheless, no specific motifs seem to have been selected, and the free energy of the secondary structures is only slightly lower than that expected on the basis of nucleotide frequencies.

Animals↗

Codon usage patterns suggest independent evolution of two catabolic operons on toluene-degradative plasmid TOL pWW0 of Pseudomonas putida.

TOL plasmid pWW0 of Pseudomonas putida encodes a set of enzymes responsible for the degradation of toluene. The structural genes for these catobolic enzymes are clustered into two operons--namely, the xy/CMAB and xy/XYZLTEGFJQKIH operons. We examined the codon usage patterns of these catabolic genes by measuring the codon-usage distances between pairs of these catabolic genes. The codon-usage distance, d, between gene 1 and gene 2 was defined as d = [sigma(pj-qj)2]1/2, are the frequencies of the j-th codon in gene 1 and 2, respectively, j being any one of the 64 possible codons. We found that the genes in the same operon exhibit similar codon-usage patterns while genes in the different operons exhibit different codon bias. This observation suggests that genes in the same operon have coevolved, and that the ancestors of the xy/CMAB and xy/XYZLTEGFJQKIH operons evolved in different organisms.

Biodegradation, Environmental↗

Sequence analysis of rice dwarf phytoreovirus genome segments S4, S5, and S6: comparison with the equivalent wound tumor virus segments.

The complete nucleotide sequences of genome segments S4, S5, and S6 of rice dwarf phytoreovirus (RDV) were determined. S4 and S5 consist of 2468 and 2570 base pairs, respectively, S5 thus being larger in size than S4, contrary to the situation suggested by their relative migration in a polyacrylamide gel. S6 is 1699 nucleotides long. The individual segments have segment-specific inverted repeats adjacent to the conserved terminal sequences (5'GGUAAA---UGAU3' for S4, 5'GGCAAA---UGAU3' for S5 and S6). S4, S5, and S6 each have single long open reading frames encoding 727, 801, and 509 amino acids, respectively. A low level of amino acid sequence homology was observed between RDV S4 and wound tumor virus (WTV) S4 (22.4%), and between RDV S6 and WTV S6 (20.2%). On the other hand, RDV S5 and WTV S5 show 52.0% amino acid sequence similarity, indicating that S5 is much more conserved than any other segments of RDV and WTV reported so far. Further comparative analyses indicate that the RDV segment shows a greater frequency of usage of codons XYG and XYC, and much less frequent usage of codon XYA than the equivalent WTV segment, this codon preference bias being more conspicuous than expected from the base contents.

Amino Acid Sequence↗

Organization and codon usage of the streptomycin operon in Micrococcus luteus, a bacterium with a high genomic G + C content.

The DNA sequence of the Micrococcus luteus str operon, which includes genes for ribosomal proteins S12 (str or rpsL) and S7 (rpsG) and elongation factors (EF) G (fus) and Tu (tuf), has been determined and compared with the corresponding sequence of Escherichia coli to estimate the effect of high genomic G + C content (74%) of M. luteus on the codon usage pattern. The gene organization in this operon and the deduced amino acid sequence of each corresponding protein are well conserved between the two species. The mean G + C content of the M. luteus str operon is 67%, which is much higher than that of E. coli (51%). The codon usage pattern of M. luteus is very different from that of E. coli and extremely biased to the use of G and C in silent positions. About 95% (1,309 of 1,382) of codons have G or C at the third position. Codon GUG is used for initiation of S12, EF-G, and EF-Tu, and AUG is used only in S7, whereas GUG initiates only one of the EF-Tu's in E. coli. UGA is the predominant termination codon in M. luteus, in contrast to UAA in E. coli.

Amino Acid Sequence↗

Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages.

In the plant chloroplast genome the codon usage of the highly expressed psbA gene is unique and is adapted to the tRNA population, probably due to selection for translation efficiency. In this study the role of selection on codon usage in each of the fully sequenced chloroplast genomes, in addition to Chlamydomonas reinhardtii, is investigated by measuring adaptation to this pattern of codon usage. A method is developed which tests selection on each gene individually by constructing sequences with the same amino acid composition as the gene and randomly assigning codons based on the nucleotide composition of noncoding regions of that genome. The codon bias of the actual gene is then compared to a distribution of random sequences. The data indicate that within the algae selection is strong in Cyanophora paradoxa, affecting a majority of genes, of intermediate intensity in Odontella sinensis, and weaker in Porphyra purpurea and Euglena gracilis. In the plants, selection is found to be quite weak in Pinus thunbergii and the angiosperms but there is evidence that an intermediate level of selection exists in the liverwort Marchantia polymorpha. The role of selection is then further investigated in two comparative studies. It is shown that average relative codon bias is correlated with expression level and that, despite saturation levels of substitution, there is a strong correlation among the algae genomes in the degree of codon bias of homologous genes. All of these data indicate that selection for translation efficiency plays a significant role in determining the codon bias of chloroplast genes but that it acts with different intensities in different lineages. In general it is stronger in the algae than the higher plants, but within the algae Euglena is found to have several unusual features which are noted. The factors that might be responsible for this variation in intensity among the various genomes are discussed.

Chloroplasts↗

Structure of the gene encoding the exoglucanase of Cellulomonas fimi.

In Cellulomonas fimi the cex gene encodes an exoglucanase (Exg) involved in the degradation of cellulose. The gene now has been sequenced as part of a 2.58-kb fragment of C. fimi DNA. The cex coding region of 1452 bp (484 codons) was identified by comparison of the DNA sequence to the N-terminal amino acid (aa) sequence of the Exg purified from C. fimi. The Exg sequence is preceded by a putative signal peptide of 41 aa, a translational initiation codon, and a sequence resembling a ribosome-binding site five nucleotides (nt) before the initiation codon. The nt sequence immediately following the translational stop codon contains four inverted repeats, two of which overlap, and which can be arranged in stable secondary structures. The codon usage in C. fimi appears to be quite different from that of Escherichia coli. A dramatic (98.5%) bias occurs for G or C in the third position for the 35 codons utilized in the cex gene.

Amino Acid Sequence↗

Codon usage in yeast: cluster analysis clearly differentiates highly and lowly expressed genes.

Codon usage data has been compiled for 110 yeast genes. Cluster analysis on relative synonymous codon usage revealed two distinct groups of genes. One group corresponds to highly expressed genes, and has much more extreme synonymous codon preference. The pattern of codon usage observed is consistent with that expected if a need to match abundant tRNAs, and intermediacy of tRNA-mRNA interaction energies are important selective constraints. Thus codon usage in the highly expressed group shows a higher correlation with tRNA abundance, a greater degree of third base pyrimidine bias, and a lesser tendency to the A+T richness which is characteristic of the yeast genome. The cluster analysis can be used to predict the likely level of gene expression of any gene, and identifies the pattern of codon usage likely to yield optimal gene expression in yeast.

Base Composition↗

Putidaredoxin reductase and putidaredoxin. Cloning, sequence determination, and heterologous expression of the proteins.

The oxidation of camphor by cytochrome P-450cam requires the participation of a flavoprotein, putidaredoxin reductase, and an iron-sulfur protein, putidaredoxin, to mediate the transfer of electrons from NADH to P-450 for oxygen activation. A 2.2-kilobase pair BamHI-StuI fragment from whole cell DNA of camphor-grown Pseudomonas putida has been cloned and sequenced. Translation of the sequence revealed two open reading frames that could code for putidaredoxin reductase and putidaredoxin. In the case of putidaredoxin, the translated sequence matched the published sequence (Tanaka, M., Haniu, M., Yasunobu, K. T., Dus, K., and Gunsalus, I. C. (1974) J. Biol. Chem. 249, 3689-3701) with the exception of one amino acid. Codon usage in these proteins, like the proteins of other Pseudomonads, is strongly biased to G + C in the third nucleotide. A potential transcription termination site was found 3' to the putidaredoxin coding region. The "FAD-binding" amino acid consensus sequence, present in other flavoproteins, was found in putidaredoxin reductase beginning at residue 11 and a second occurrence of this sequence was found beginning with amino acid 156. The second sequence could represent the NAD-binding site. The regions encoding putidaredoxin reductase and putidaredoxin were subcloned and independently expressed in Escherichia coli at the level of 0.4 and 4.8 mg of enzymatically active protein/g wet weight of cells, respectively. Site-directed mutagenesis was used to change the rare start codon, GTG, of putidaredoxin reductase to ATG which resulted in an 18-fold increase in the level of expression of this protein to 7.4 mg/g wet weight of cells. The construction of these two clones, which express these important proteins, will facilitate studies of their interaction with each other and with P-450cam.

Amino Acid Sequence↗

Nucleotide sequence of Candida pelliculosa beta-glucosidase gene.

The nucleotide sequence of the DNA fragment containing the beta-glucosidase gene of Candida pelliculosa was determined. Analysis of the sequence revealed three open reading frames which could encode 65,825, and 412 amino acid residues. The presence of the second frame was found to be sufficient for the expression of the beta-glucosidase gene in a heterologous host Saccharomyces cerevisiae. Putative protein encoded by this gene had hydrophobic amino acids, resembling a signal peptide, at its N-terminal region and 19 potential glycosylation sites. Codon usage of Candida genes had the similar pattern shown in S.cerevisiae. Codon bias of the beta-glucosidase gene of Candida was relatively low, compared with that of the highly expressed genes of S. cerevisiae.

Amino Acid Sequence↗

Analysis of pFQ31, a 8551-bp cryptic plasmid from the symbiotic nitrogen-fixing actinomycete Frankia.

The actinomycete Frankia has never been transformed genetically. To favour the development of Frankia cloning vectors, we have fully sequenced the Frankia alni pFQ31 cryptic plasmid and performed analyses to characterise its coding and non-coding regions. This plasmid is 8551 bp-long and contains 72% G+C. Computer-assisted analyses identified 18 open reading frames (ORFs). These ORFs show a synonymous codon usage different from the one of Frankia chromosomal genes, suggesting an evolutionary bias linked to the nature of the replicon or a horizontal transfer. Three ORFs were found to encode genes likely to be involved in plasmid replication and stability: parFA (partition protein), ptrFA (transcriptional repressor of the GntR family) and repFA (initiation of replication). DNA signatures of a replication origin were identified in the ptrFA-repFA intergenic region. These structural motifs are similar to those observed among origins of iteron-containing plasmids replicating via a θ mode.

Actinomycetales↗