Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Nucleotide sequence of the coding portion of human alpha globin messenger RNA.

The nucleotide sequence of the coding portion of human alpha globin mRNA has been determined by sequence analysis using human alpha globin cDNA cloned in bacterial plasmids. The sequence was obtained by a combination of direct sequence analysis of the cloned cDNA and analysis of cDNA obtained by primer extension, using short restriction endonuclease fragments of cloned alpha cDNA that were hybridized to human globin mRNA and elongated on the mRNA template by viral reverse transcriptase. The human alpha globin mRNA has an unexpectedly high G + C base composition (64.7%), similar to that observed for rabbit globin alpha mRNA, and displays a striking bias in the use of synonym codons for various amino acids. The bias in codon usage of human alpha globin mRNA is similar, with some exceptions, to that previously observed for rabbit alpha globin mRNA as well as for human and rabbit beta globin mRNAs. A detailed restriction endonuclease map of the human alpha globin cDNA is presented.

Amino Acid Sequence↗

[Studies on the molecular evolution of apolipoprotein multigene family].

The apolipoprotein genes represent a large family of genes encoding various binding proteins for plasma lipid transport. Because of their long divergence history, it is not known whether and how these genes have evolved through gene duplication from a common ancestor. To test this possibility and reconstruct a reliable phylogenetic tree, a simple method to evaluate the branch length and its divergence time in unrooted parsimony tree under the condition of non-even evolutionary rate was developed. The tree built from the 26 apolipoprotein sequences by above method clearly shows: (1) The common ancestor of ApoA-I ApoA-II, ApoA-IV, ApoE may appear 460 million years ago in an ordovician vertebrate which may be related with the major apolipoprotein LAL1 and LAL2 in Lamprey from the evidence of sequence alignment; (2) The central role of different selection pressure upon the ancestor gene of apolipoprotein made them evolved into different subgroups; (3) The high evolution rate in rodent ApoE molecules may be related with the existence of a large amount of hidden substitutions and the disruption of synonym codon usage clock in their genome; (4) The evolutionary rate of various branches in parsimony tree is significantly different in which the average UEP of ApoA-I, ApoA-IV is 2.0 MY, ApoA-II 1.7MY, ApoE 2.4MY; (5) The receptor domain in ApoE seems to be more conservative than other fragments. These data suggest a long, complex evolutionary history for apolipoprotein genes in which the gene duplication events of different origins took place.

Amino Acid Sequence↗

[Molecular evolution of MHC DQA genes. II. Phylogenetic analysis based on nucleotide substitution and SCU bias].

Phylogenetics of 23 alleles at MHC DQA loci in 7 mammalian species was studied based on their nucleotide (NT) substitution and synonymous codon usage (SCU) bias. (1) It was demonstrated that the NT substitution rates are 1.0 x 10(-9) NT/site/yr for exon2 and 1.3 x 10(-9) NT/site/yr for exon2-4 in a large time scale, which is similar to other nuclear genes, while for mouse and rat the rates are nearly twice as high as above mentioned. (2) The DQA locus diversity and their interallelic diversity developed long after the radiation of mammalian 80Mya (million years ago). The bovine counterpart, of, and with the same recent ancestor of ovine DQA2, remains to be discovered. HLA-DQA2 locus split from HLA-DQA1 ancestor at the time between 12 approximately 20 Mya while allele diversity of HLA-DQA1 emerged and developed from 24 Mya to less than 1 Mya. (3) The phylogenetic trees based on SCU divergence reflect the phylogenetics of MHC DQA genes quite well generally in a new respect and reveal that HLA-DQA2 has a distinctive SCU bias different from all other MHC DQA locianalyzed. It indicates that SCU statistics plays an important and unique role in phylogenetic analysis of orthologous genes. The method to estimate the SCU divergence and SCU similarity was improved in this research.

Animals↗

Divergence of the yellow gene between Drosophila melanogaster and D. subobscura: recombination rate, codon bias and synonymous substitutions.

The yellow (y) gene maps near the telomere of the X chromosome in Drosophila melanogaster but not in D. subobscura. Thus the strong reduction in the recombination rate associated with telomeric regions is not expected in D. subobscura. To study the divergence of a gene whose recombination rate differs between two species, the y gene of D. subobscura was sequenced. Sequence comparison between D. melanogaster and D. subobscura revealed several elements conserved in noncoding regions that may correspond to putative cis-acting regulatory sequences. Divergence in the y gene coding region between D. subobscura and D. melanogaster was compared with that found in other genes sequenced in both species. Both, yellow and scute exhibit an unusually high number of synonymous substitutions per site (ps). Also for these genes, the extent of codon bias differs between both species, being much higher in D. subobscura than in D. melanogaster. This pattern of divergence is consistent with the hitchhiking and background selection models that predict an increase in the fixation rate of slightly deleterious mutations and a decrease in the rate of fixation of slightly advantageous mutations in regions with low recombination rates such as in the y-sc gene region of D. melanogaster.

Amino Acid Sequence↗

[Variations of apolipoprotein A IV gene in Chinese endogenous hypertriglyceridemics].

OBJECTIVE: The aim of this study was to investigate variations of apolipoprotein A IV (apo A IV) gene and its relation to endogenous hypertriglyceridemia(HTG) in Chinese population. METHODS: One hundred and six endogenous hypertriglyceridemics and 171 healthy subjects from a population of Chinese Han nationality in Chengdu area were studied using restriction fragment length polymorphisms (RFLPs) and sequencing of apoA IV gene amplified by polymerase chain reaction (PCR). The polymorphic sites of apo A IV gene studied included codon 9 (A to G, synonymous mutation), codon 347 (A to T, non-synonymous mutation), codon 360 (G to T, non-synonymous mutation), and Msp I polymorphism (CC/TGG) within intron 2. RESULTS: The frequency of G allele at codon 9 in HTG group was higher than that in healthy controls(0.453 vs 0.366, P<0.05). The other polymorphic sites showed no significant differences of the allele frequencies between the two groups. The frequencies of rare alleles, such as G allele at codon 9, T allele at codon 347 and T allele at codon 360 polymorphic site were significantly different from those reported in European Caucasians (0.366 vs 0.032, P<0.001, 0.000 vs 0.160, P<0.001; 0.000 vs 0.070, P<0.001), but no differences were found when compared with those in Japanese, including Msp I site (P>0.05). In the healthy male control group, subjects with genotype G/G of codon 9 had a higher serum mean concentration of apoA I when compared with that of genotype A/A(P<0.01). In the HTG group, subjects with genotype C/T of Msp I site had a higher serum mean concentration of TG with compared with those with genotype C/C and T/T (P<0.05). This difference was only observed in male HTG group when male and female subgroups were further separated. CONCLUSION: These results suggest that Msp I and codon 9 polymorphism in apoA IV gene are associated with endogenous hypertriglyceridemia to some extent in Chinese population.

Adult↗

Patterns of nucleotide substitution among simultaneously duplicated gene pairs in Arabidopsis thaliana.

We characterized rates and patterns of synonymous and nonsynonymous substitution in 242 duplicated gene pairs on chromosomes 2 and 4 of Arabidopsis thaliana. Based on their collinear order along the two chromosomes, the gene pairs were likely duplicated contemporaneously, and therefore comparison of genetic distances among gene pairs provides insights into the distribution of nucleotide substitution rates among plant nuclear genes. Rates of synonymous substitution varied 13.8-fold among the duplicated gene pairs, but 90% of gene pairs differed by less than 2.6-fold. Average nonsynonymous rates were approximately fivefold lower than average synonymous rates; this rate difference is lower than that of previously studied nonplant lineages. The coefficient of variation of rates among genes was 0.65 for nonsynonymous rates and 0.44 for synonymous rates, indicating that synonymous and nonsynonymous rates vary among genes to roughly the same extent. The causes underlying rate variation were explored. Our analyses tentatively suggest an effect of physical location on synonymous substitution rates but no similar effect on nonsynonymous rates. Nonsynonymous substitution rates were negatively correlated with GC content at synonymous third codon positions, and synonymous substitution rates were negatively correlated with codon bias, as observed in other systems. Finally, the 242 gene pairs permitted investigation of the processes underlying divergence between paralogs. We found no evidence of positive selection, little evidence that paralogs evolve at different rates, and no evidence of differential codon usage or third position GC content.

Arabidopsis↗

Nucleotide substitution pattern in rice paralogues: implication for negative correlation between the synonymous substitution rate and codon usage bias.

Understanding the correlation between synonymous substitution rate and GC content is essential to decipher the gene evolution. However, it has been controversial on their relationship. We analyzed the GC content and synonymous substitution rate in 1092 paralogues produced by two large-scale duplication events in the rice genome. According to the GC content at the third codon sites (GC3), the paralogues were classified into GC3-rich and GC3-poor genes. By referring to their outgroup sequences, we inferred the last common ancestor of sister paralogues and, consequently, calculated the average synonymous substitution rate for two gene classes. The results suggest that average synonymous substitution rate is lower in GC3-rich genes than that in GC3-poor genes, indicating that the synonymous substitution rate is negatively correlated with GC content in the rice genome. Through characterizing the synonymous nucleotide substitution pattern, we found a strong synonymous nucleotide substitution frequency bias from AT to GC in GC3-rich genes. This indicates possible limitations of commonly used methods developed to estimate the synonymous substitution rate. Their estimates might produce misleading results on correlation between the synonymous substitution rate and GC content.

Base Composition↗

Directional mutation pressure and transfer RNA in choice of the third nucleotide of synonymous two-codon sets.

Bacterial species have diverged into a series of families, some with high G + C content in their DNA, and other with high A + T content, resulting, respectively, from G.C- and A.T-directional mutation pressures. Such mutation pressure (G.C/A.T pressure) may be an important determinant for codon usage. It has also been suggested that tRNA acts as a selective constraint for determining codon usage. We have studied the relation between G.C/A.T pressure and tRNA constraints in determining choice of the third nucleotide of eight two-codon sets, using codon usage data obtained from protein genes in four bacterial species, Mycoplasma capricolum, Bacillus subtilis, Escherichia coli, and Micrococcus luteus, and in liverwort (Marchantia polymorpha) chloroplasts. The genomic G + C contents of these range from 25% to 74%. The results demonstrate that tRNA levels act additively to A.T and G.C pressure in affecting contents of A (pairing with *UNN anticodons, in which *U indicates a 2-thiouridine derivative) and C (pairing with GNN anticodons) or G (pairing with CNN anticodons), respectively, in third nucleotide positions of codons.

Chloroplasts↗

Small regions of preferential codon usage and their effect on overall codon bias--the case of the plp gene.

Preferential codon usage within synonymous groups (codon bias), is hypothesised to be a consequence of either mutational pressure or translational selection. This paper examines PLP and DM-20, two transcripts differentially spliced from the plp gene, using a variety of previously described indices measuring codon bias as well as a novel index for quantifying the response to directional mutation pressure. The results demonstrate that small regions of extreme codon bias may have appreciable effects on the quantification of total bias, and that this effect may be either an increase or a decrease depending which index is used.

Animals↗

Comparative analysis of the base composition and codon usages in fourteen mycobacteriophage genomes.

To study the possible codon usage and base composition variation in the bacteriophages, fourteen mycobacteriophages were used as a model system here and both the parameters in all these phages and their plating bacteria, M. smegmatis had been determined and compared. As all the organisms are GC-rich, the GC contents at third codon positions were found in fact higher than the second codon positions as well as the first + second codon positions in all the organisms indicating that directional mutational pressure is strongly operative at the synonymous third codon positions. Nc plot indicates that codon usage variation in all these organisms are governed by the forces other than compositional constraints. Correspondence analysis suggests that: (i) there are codon usage variation among the genes and genomes of the fourteen mycobacteriophages and M. smegmatis, i.e., codon usage patterns in the mycobacteriophages is phage-specific but not the M. smegmatis-specific; (ii) synonymous codon usage patterns of Barnyard, Che8, Che9d, and Omega are more similar than the rest mycobacteriophages and M. smegmatis; (iii) codon usage bias in the mycobacteriophages are mainly determined by mutational pressure; and (iv) the genes of comparatively GC rich genomes are more biased than the GC poor genomes. Translational selection in determining the codon usage variation in highly expressed genes can be invoked from the predominant occurrences of C ending codons in the highly expressed genes. Cluster analysis based on codon usage data also shows that there are two distinct branches for the fourteen mycobacteriophages and there is codon usage variation even among the phages of each branch.

Bacteriophages↗

Large-scale analyses of synonymous substitution rates can be sensitive to assumptions about the process of mutation.

A popular approach to examine the roles of mutation and selection in the evolution of genomes has been to consider the relationship between codon bias and synonymous rates of molecular evolution. A significant relationship between these two quantities is taken to indicate the action of weak selection on substitutions among synonymous codons. The neutral theory predicts that the rate of evolution is inversely related to the level of functional constraint. Therefore, selection against the use of non-preferred codons among those coding for the same amino acid should result in lower rates of synonymous substitution as compared with sites not subject to such selection pressures. However, reliably measuring the extent of such a relationship is problematic, as estimates of synonymous rates are sensitive to our assumptions about the process of molecular evolution. Previous studies showed the importance of accounting for unequal codon frequencies, in particular when synonymous codon usage is highly biased. Yet, unequal codon frequencies can be modeled in different ways, making different assumptions about the mutation process. Here we conduct a simulation study to evaluate two different ways of modeling uneven codon frequencies and show that both model parameterizations can have a dramatic impact on rate estimates and affect biological conclusions about genome evolution. We reanalyze three large data sets to demonstrate the relevance of our results to empirical data analysis.

Amino Acid Substitution↗

Nonsynonymous polymorphic sites in the apolipoprotein (apo) A-IV gene are associated with changes in the concentration of apo B- and apo A-I-containing lipoproteins in a normal population.

The aims of this study were to detect polymorphic sites in the apolipoprotein (apo) A-IV gene, to establish their frequencies, to determine potential haplotypes, and to investigate the role of these polymorphisms in lipid metabolism. A sequencing study of four individuals led to the identification of two synonymous mutations (codons 9 and 54) and three nonsynonymous mutations (Val-8----Met, Gln360----His, and Thr347----Ser) and of a VNTR polymorphism within a series of three or four CTGT repeats in the noncoding region of exon 3. Frequencies of these polymorphisms were determined in 291 students by using naturally occurring (BstEII for the synonymous mutation in codon 54, HinfI for Thr347----Ser, and Fnu4HI for Gln360----His) or artificially introduced restriction-enzyme cutting sites (BstEII for the synonymous mutation in codon 9 and MamI for Val-8----Met), subsequent to PCR amplification. The four-base deletion/insertion polymorphism and its localization cis or trans to the mutations in codons 347 and 360 were studied by direct sequencing of PCR-amplified DNA from 87 students. Frequencies of the rarer alleles were .007 for apo A-IV-8:Met, .04 for the synonymous mutation in codon 9, .14 for the synonymous mutation in codon 54, .16 for apo A-IV347:Ser, .07 for apo A-IV360:His, and .39 for the four-base of insertion. Apo A-IV360:His in all cases was cis-localized to the (CTGT)3 repeat and apo A-IV347:Thr; and apo A-IV347:Ser was cis-localized to the (CTGT)4 repeat and apo A-IV360:Gln. Four haplotypes formed from these three polymorphic sites were thus found. The apo A-IV347:Ser allele was associated both with significantly lower plasma apo B concentrations in both sexes and with significantly lower LDL-cholesterol concentrations in men. Heterozygous carriers of apo A-IV360:His exhibited significantly higher concentrations of LDL-cholesterol and lower Lp(a) concentrations, compared with apo A-IV360:Gln homozygotes. We could not confirm the previously reported association of apo A-IV360:His with elevated HDL-cholesterol concentrations. In the population, the Val-8----Met polymorphism was not associated with significantly different lipid concentrations, but in a family study the Met-8 allele was associated with lower HDL-cholesterol and higher LDL-cholesterol concentrations. In conclusion, our results indicate an important role of the apo A-IV gene locus in the metabolism of apo B and, to a lesser extent, apo A-I containing lipoproteins.

Adult↗

Codon-usage based regulation of colicin K synthesis by the stress alarmone ppGpp.

The molecular mechanism of the upregulation of Escherichia coli colicin K (Cka) synthesis during stress conditions was studied. Nutrient starvation experiments and the use of relA spoT mutant strains, IPTG-regulated overproduction of ppGpp and lacZ fusions revealed that the stringent response alarmone guanosine 3',5'-bispyrophosphate (ppGpp) is the main positive effector of Cka synthesis. Comparison of the amounts of protein produced (Western blotting) and specific mRNA (Northern blotting) before and after nutrient starvation demonstrated increases in Cka protein with unaltered specific mRNA levels, suggesting a post-transcriptional regulatory mechanism. Reporter (beta-galactosidase) assays using truncated cka of variable length fused to lacZ located the key regulatory region close to the 5' end of the cka mRNA. Closer analysis of this region indicated the presence of several rare codons, including the leucine-encoding codon CUA. Synonymous exchange of the rare codons with more frequently used ones abolished the regulatory effect of ppGpp. Supplementation of the strain with the plasmid CodonPlus carrying several rare tRNA genes yielded similar results, indicating that codon usage (in particular, the fifth codon for the amino acid leucine) and tRNA availability (i.e. tRNAleu) are the key elements of the regulatory function of ppGpp. We conclude that ppGpp regulates Cka synthesis via a novel post-transcriptional mechanism that is based on rare codon usage and variable cognate tRNA availability.

Codon↗

Novel alleles HLA-B*7802 and B*51022: evidence for convergency in the HLA-B5 family.

We have characterized two novel HLA-B alleles, B*7802 and B*51022. The Caucasian-derived variant B*7802 most resembles the African-derived variant B*7801, from which B*7802 differs by two nucleotides. Only one of these modifications, however, is translated: a tyrosine for aspartate substitution occurs at residue 74 in B*7802, while the second nucleotide difference reflects a proximal synonymous substitution in codon 23. A second variant, B*51022, differs synonymously only at codon 23 from B*51021. Comparative analysis of the B5 CREG demonstrates that other pairs of B5 alleles differ synonymously only at codon 23 or synonymously at codon 23 and non-synonymously at a second more distal location. Contrary to the genesis of like pairs of B5 alleles via introduction of coordinate yet distant mutagenic events onto a single B5 progenitor, we postulate that synonymously different B5 progenitor molecules, B5ATT and B5ATC, are evolving in convergence to generate homologous B5 allele pairs differing silently at codon 23. Our finding that B*7802 is a single amino acid away from complete convergence with B*7801 and that B*51022 and B*51021 are in complete convergence is exemplary of such evolution.

Alleles↗

Effect of tandem rare codon substitution and vector-host combinations on the expression of the EBV gp110 C-terminal domain in Escherichia coli.

Gp110 of Epstein-Barr virus (EBV) is a glycoprotein that functions exclusively during the assembly of EBV nucleocapsid and the release of infectious EBV. Its C-terminal tail domain (gp110 CTD) is essential for gp110's function and may provide signals that are responsible for the assembly and release of EBV. In the present study, to get large amounts of gp110 CTD for structural analysis, the effects of vector system, codon usage, and host strain on expression levels of gp110 CTD in Escherichia coli have been investigated. The coding region of gp110 CTD (11 kDa) was subcloned into the expression vectors pSE 280, pET-15b, pET-29a, pMAL-c2x, and pGEX-4T-1. Except the pMAL-c2x construct, all the others failed to express detectable amounts of recombinant gp110 CTD. Substituting a tandem rare AGA (Arg) codon with a synonymous CGC (Arg) codon facilitated expression of the recombinant protein, while a protease-deficient host E. coli strain helped in the accumulation of a soluble form of gp110 CTD fusion. The secondary structures of the obtained recombinant gp110 CTD purified from soluble extracts and inclusion bodies were compared using circular dichroism analysis. In aqueous solutions, both samples equally adopt a mixed alpha-helix and beta-sheet conformation as well as a partly unordered structure. Notably, in the membrane-mimicking environments the helical propensity of gp110 CTD increased up to the previously predicted level based on its sequence, suggesting that gp110 CTD may fold into a more stable conformation through interactions with the cell membrane.

Circular Dichroism↗

The rate of synonymous substitution in enterobacterial genes is inversely related to codon usage bias.

Genes sequences from Escherichia coli, Salmonella typhimurium, and other members of the Enterobacteriaceae show a negative correlation between the degree of synonymous-codon usage bias and the rate of nucleotide substitution at synonymous sites. In particular, very highly expressed genes have very biased codon usage and accumulate synonymous substitutions very slowly. In contrast, there is little correlation between the degree of codon bias and the rate of protein evolution. It is concluded that both the rate of synonymous substitution and the degree of codon usage bias largely reflect the intensity of selection at the translational level. Because of the high variability among genes in rates of synonymous substitution, separate molecular clocks of synonymous substitution might be required for different genes.

Biological Evolution↗

Nucleic acid composition, codon usage, and the rate of synonymous substitution in protein-coding genes.

Based on the rates of synonymous substitution in 42 protein-coding gene pairs from rat and human, a correlation is shown to exist between the frequency of the nucleotides in all positions of the codon and the synonymous substitution rate. The correlation coefficients were positive for A and T and negative for C and G. This means that AT-rich genes accumulate more synonymous substitutions than GC-rich genes. Biased patterns of mutation could not account for this phenomenon. Thus, the variation in synonymous substitution rates and the resulting unequal codon usage must be the consequence of selection against A and T in synonymous positions. Most of the variation in rates of synonymous substitution can be explained by the nucleotide composition in synonymous positions. Codon-anticodon interactions, dinucleotide frequencies, and contextual factors influence neither the rates of synonymous substitution nor codon usage. Interestingly, the nucleotide in the second position of codons (always a nonsynonymous position) was found to affect the rate of synonymous substitution. This finding links the rate of nonsynonymous substitution with the synonymous rate. Consequently, highly conservative proteins are expected to be encoded by genes that evolve slowly in terms of synonymous substitutions, and are consequently highly biased in their codon usage.

Animals↗