Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “synonymous codon”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Structuring of the genetic code took place at acidic pH.

I have observed that in multiple regression the number of codons specifying amino acids in the genetic code is positively correlated with the isoelectric point of amino acids and their molecular weight. Therefore basic amino acids are, on average, codified in the genetic code by a larger number of codons, which seems to imply that the genetic code originated in an acidic 'intracellular' environment. Moreover, I compare the proteins from Picrophilus torridus and Thermoplasma volcanium, which have different intracellular pH and I define the ranks of acidophily for the amino acids. A simple index of acidophily (AI), which can be easily obtained from acidophily ranks, can be associated to any protein and, therefore, can also be associated to the genetic code if the number of synonymous codons attributed to the amino acids in the code is assumed to be the frequency with which the amino acids appeared in ancestral proteins. Finally, the sampling of the variable AI among organisms having an intracellular pH less than or equal to 6.6 and those having a non-acidic intracellular pH leads to the conclusion that the value of the genetic code's AI is not typical of proteins of the latter organisms. As the genetic code's AI value is also statistically not different from that of proteins of the organisms having an acidic intracellular pH, this supports the hypothesis that the structuring of the genetic code took place in acidic pH conditions.

Amino Acid Sequence↗

The genomic rate of adaptive amino acid substitution in Drosophila.

The proportion of amino acid substitutions driven by adaptive evolution can potentially be estimated from polymorphism and divergence data by an extension of the McDonald-Kreitman test. We have developed a maximum-likelihood method to do this and have applied our method to several data sets from three Drosophila species: D. melanogaster, D. simulans, and D. yakuba. The estimated number of adaptive substitutions per codon is not uniformly distributed among genes, but follows a leptokurtic distribution. However, the proportion of amino acid substitutions fixed by adaptive evolution seems to be remarkably constant across the genome (i.e., the proportion of amino acid substitutions that are adaptive appears to be the same in fast-evolving and slow-evolving genes; fast-evolving genes have higher numbers of both adaptive and neutral substitutions). Our estimates do not seem to be significantly biased by selection on synonymous codon use or by the assumption of independence among sites. Nevertheless, an accurate estimate is hampered by the existence of slightly deleterious mutations and variations in effective population size. The analysis of several Drosophila data sets suggests that approximately 25% +/- 20% of amino acid substitutions were driven by positive selection in the divergence between D. simulans and D. yakuba.

Adaptation, Biological↗

Influence of parasitic life style on the patterns of codon usage and base frequencies of Ancylostoma and Necator species.

Parametric analyses were used to investigate the nucleotide, codon, and amino acid composition of coding sequences corresponding to hook-worms. Ancylostoma caninum and Necator americanus. Although genomic research has become prevalent within the scientific community, few studies have dealt directly with parasitic species. Parasites have existed throughout the history of mankind due to their wide range of distribution in nature and their ability to evade immune detection. An AT nucleotide bias was identified in both A. caninum and N. americanus sequences. A similar AT bias was also identified in both datasets when considering relative synonymous codon usage. However, the codon bias was much more pronounced in N. americanus as compared to A. caninum. Bias was also present at the amino acid level, and appeared to be partially independent of the nucleotide-based biases. Analysis of parasite genomes will facilitate the development of vaccines against larval forms of parasites. Moreover, the examination of the parasite genes in general, will allow for a more in-depth understanding of the evolution of the parasites and parasitism.

Ancylostoma↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

Effective structure of a leader open reading frame for enhancing the expression of GC-rich genes.

To overexpress broad kinds of GC-rich genes in Escherichia coli, we examined how the structures of leader open reading frames (leader ORFs) affect the expression of GC-rich genes, such as polA, trpA, and trpB, from Thermus thermophilus. When a leader ORF overlapped with the polA-initiation codon by 1 bp in the TGATG motif, gene expression increased by more than 3-fold compared to when a leader ORF was several-bp distant from the initiation codon. A 4-bp overlap with the ATGA motif was more effective than a 1-bp overlap with the TGATG motif. When a 4-bp overlapping leader ORF was placed in front of the successive trpB and trpA genes, the trpA gene was poorly expressed whereas the trpB gene was overexpressed. Mutation analysis revealed that the expression of the trpA gene was strongly enhanced by replacing G and C in the translation termination region of the leader ORF with A and T. In contrast, other mutations, such as alterations between synonymous codons in the trpA-coding region, produced diminished gene expression. Using the most effective leader ORF obtained from these results, new expression vectors were constructed.

Amino Acids↗

Some aspects of the organization and evolution of the genetic code.

In this paper, I define a measure of the relative position of each amino acid in the genetic code by means of a 21-dimensional vector describing its potential for mutation, in a single step, to each of the other amino acids, or to a chain termination codon. This measure allows us to make a systematic investigation of the type and number of the physicochemical properties of the amino acids that were involved in evolution. The polar character and size of amino acids are identified in this analysis as properties that played a leading role in the evolutionary history of the genetic code. The application of cluster analysis and discriminant analysis reveals the characteristics of the structural organization of the genetic code. Finally, I suggest the existence of a relationship between the molecular weight of the amino acids and the number of synonymous codons.

Amino Acids↗

Codon preference in corynebacteria.

The codon usage (CU) of 34 genes from the closely related species, Brevibacterium lactofermentum and Corynebacterium glutamicum (BLCG), was analysed and compared with that of 23 genes from other Brevibacterium and Corynebacterium species. The G+C content of the BLCG genes ranged from 50 to 62%. A wider range was found in other corynebacterial genes (25-71%). The G+C contents of non-coding regions in glutamic acid bacteria are lower than those of the coding regions and both values are lower than the G+C content of ribosomal RNA (rRNA) sequences, suggesting an unusual biased mutation pressure. The CU and synonymous codon usage (SCU) analysis showed several common characteristics among the sequenced corynebacterial genes, consistent with the close relatedness of B. lactofermentum and C. glutamicum. A subset of 25 preferred codons were deduced from the presumably highly expressed genes and they encode most of the amino acid (aa) residues of the BLCG group. An analysis of the effective number of codons (Nc) was carried out in order to check the GC3s (G+C content at the silent third position of sense codons) dependence of the CU in corynebacteria. Nc values showed differences between the BLCG group and other corynebacterial sequences. A comparison of the most used codons for each aa showed a stronger similarity to Streptomyces than to Escherichia coli. The CU/SCU tables of corynebacteria are useful for identification of protein-coding regions, including start codons when they are uncertain, and for designing oligodeoxyribonucleotide probes from an aa sequence.

Base Sequence↗

Characteristics and clustering of human ribosomal protein genes.

BACKGROUND: The ribosome is a central player in the translation system, which in mammals consists of four RNA species and 79 ribosomal proteins (RPs). The control mechanisms of gene expression and the functions of RPs are believed to be identical. Most RP genes have common promoters and were therefore assumed to have a unified gene expression control mechanism. RESULTS: We systematically analyzed the homogeneity and heterogeneity of RP genes on the basis of their expression profiles, promoter structures, encoded amino acid compositions, and codon compositions. The results revealed that (1) most RP genes are coordinately expressed at the mRNA level, with higher signals in the spleen, lymph node dissection (LND), and fetal brain. However, 17 genes, including the P protein genes (RPLP0, RPLP1, RPLP2), are expressed in a tissue-specific manner. (2) Most promoters have GC boxes and possible binding sites for nuclear respiratory factor 2, Yin and Yang 1, and/or activator protein 1. However, they do not have canonical TATA boxes. (3) Analysis of the amino acid composition of the encoded proteins indicated a high lysine and arginine content. (4) The major RP genes exhibit a characteristic synonymous codon composition with high rates of G or C in the third-codon position and a high content of AAG, CAG, ATC, GAG, CAC, and CTG. CONCLUSION: Eleven of the RP genes are still identified as being unique and did not exhibit at least some of the above characteristics, indicating that they may have unknown functions not present in other RP genes. Furthermore, we found sequences conserved between human and mouse genes around the transcription start sites and in the intronic regions. This study suggests certain overall trends and characteristic features of human RP genes.

Amino Acids↗

Transcriptional regulation of gene expression by the coding sequence: An attempt to enhance expression of human AChE.

In a previous report, Morel and Massoulié showed that Bungarus AChE (bBAChE) is produced more efficiently than rat AChE in various expression systems, mainly because the Bungarus coding sequence exerts a stimulatory effect on transcription (Morel and Massoulié, 2000). They reported that a 5' Bungarus fragment could partially transfer this property to a CAT expression vector. This appeared to offer the possibility of increasing the production of recombinant proteins. In the present paper, we show that insertion of this fragment in the transcribed region, before the polyadenylation site, may have either stimulatory or inhibitory effects, depending on the vector and on the reporter gene. Since the stimulatory effect of Bungarus coding region could not be attached to a small number of discrete motifs, we reasoned that it might result from a general feature of the sequence. Therefore it might be possible to partially transfer this property to the very homologous human AChE (hHAChE) coding sequence by modifications based on synonymous codons, which increased nucleotide identity between the 5' fragment (721 nucleotides) of bBAChE and hHAChE from 71% to 85%. The production of human AChE in transfected COS cells was increased nearly 2-fold with this modified construct, but still remained about 4-fold smaller than that of Bungarus AChE. There was no change in expression level in transformed Pichia pastoris. We thus confirm that coding sequences can strongly influence gene expression, but in a manner that depends on the context and cannot yet be predicted.

Acetylcholinesterase↗

Sequence diversity of KIAA0027/MLC1: are megalencephalic leukoencephalopathy and schizophrenia allelic disorders?

The aim of the study is to validate the etiological role of KIAA0027/MLC1 in childhood-onset megalencephalic leukoencephalopathy with subcortical cysts (MLC) and in schizophrenia, particularly the catatonic subtype, which were reported to be allelic diseases. Among a series of five patients with MLC, four mutant alleles were detected: one case of compound heterozygosity for a splice site mutation and a six-base-pair in-frame deletion, one patient with a homozygous frameshifting insertion-deletion, and a further case heterozygous for a A157E substitution. A systematic mutation screening in 140 index cases with schizophrenia revealed 13 different single nucleotide polymorphisms (SNPs): one SNP in the 5'-UTR, seven SNPs in intronic regions, two synonymous codon variants (T52, Y199), and three coding variants. Two of them, C171F and N218K, were observed in controls at a significant frequency. The L309M variant that was previously supposed to be the causative factor for chromosome 22q(tel) linked-periodic catatonia was found nonsegregating in a further multiplex pedigree. Furthermore, a complicated 33-bp insertion/deletion polymorphism at the 5'-end of exon 11 of MLC1 was found at equal frequency among schizophrenic patients and controls. In summary, our study provides further evidence for allelic heterogeneity in megalencephalic leukoencephalopathy, excludes MLC1 as a susceptibility locus for schizophrenia, and thereby rules out that MLC and schizophrenia are allelic disorders.

Adolescent↗

Mutations of the Nogo-66 receptor (RTN4R) gene in schizophrenia.

Schizophrenia (SCZD) or schizoaffective disorders are quite common features in patients with DiGeorge/velo-cardio-facial syndrome (DGS/VCFS) as a result of chromosome 22q11.2 aploinsufficiency. We evaluated the Nogo-66 receptor gene (RTN4R), which maps within the DGS/VCFS critical region, as a potential candidate for schizophrenia susceptibility. RTN4R encodes for a functional cell surface receptor, a glycosylphosphatidylinositol (GPI)-linked protein, with multiple leucine-rich repeats (LRR), which is implicated in axonal growth inhibition. One hundred and twenty unrelated Italian schizophrenic patients were screened for mutations in the RTN4R gene using denaturing high performance liquid chromatography (DHPLC). Three mutant alleles were detected, including two missense changes (c.355C>T; R119W and c.587G>A; R196H), and one synonymous codon variant (c.54G>A; L18L). The two schizophrenic patients with the missense changes were strongly resistant to the neuroleptic treatment at any dosage. Both missense changes were absent in 300 control subjects. Molecular modeling revealed that both changes lead to putative structural alterations of the native protein.

Adult↗

The tendency of lentiviral open reading frames to become A-rich: constraints imposed by viral genome organization and cellular tRNA availability.

Human immunodeficiency virus type 1 (HIV-1) and other lentiviridae demonstrate a strong preference for the A-nucleotide, which can account for up to 40% of the viral RNA genome. The biological mechanism responsible for this nucleotide bias is currently unknown. The increased A-content of these viral genomes corresponds to the typical use of synonymous codons by all members of the lentiviral family (HIV, SIV, BIV, FIV, CAEV, EIAV, visna) and the human spuma retrovirus, but not by other retroviruses like the human T-cell leukemia viruses HTLV-1 and HTLV-II. In this article, we analyzed A-bias for all codon groups in all open reading frames of several lentiviruses. The extent of lentiviral codon bias could be related to host cellular translation. By calculating codon bias indices (CBIs), we were able to demonstrate an inverse correlation between the extent of codon bias and the rate of translation of individual reading frames in these viruses. Specifically, the shift toward A-rich codons is more pronounced in pol than in gag lentiviral genes. Since it is known that Gag synthesis exceeds Pol synthesis by a factor of 20 due to infrequent ribosomal frame-shifting during translation of the gap-pol mRNA molecule, we propose that the aminoacyl-tRNA availability in the host cell restricts the lentiviral preference for A-rich codons. In addition, less A-nucleotides were found in regions of the viral genome encoding multiple functions; e.g., overlapping reading frames (tat-rev-env) or in genes that overlap regulatory sequences (nef-LTR region).(ABSTRACT TRUNCATED AT 250 WORDS)

Adenine↗

Conservation of the mammalian RNA polymerase II largest-subunit C-terminal domain.

We have isolated and sequenced a portion of the gene encoding the carboxy-terminal domain (CTD) of the largest subunit of RNA polymerase II from three mammals. These mammalian sequences include one rodent and two primate CTDs. Comparisons of the new sequences to mouse and Chinese hamster show a high degree of conservation among the mammalian CTDs. Due to synonymous codon usage, the nucleotide differences between hamster, rat, ape, and human result in no amino acid changes. The amino acid sequence for the mouse CTD appears to have one different amino acid when compared to the other four sequences. Therefore, except for the one variation in mouse, all of the known mammalian CTDs have identical amino acid sequences. This is in marked contrast to the situation among more divergent species. The present study suggests that there is a strong evolutionary pressure to maintain the primary structure of the mammalian CTD.

Animals↗

Mammalian gene evolution: nucleotide sequence divergence between mouse and rat.

As a paradigm of mammalian gene evolution, the nature and extent of DNA sequence divergence between homologous protein-coding genes from mouse and rat have been investigated. The data set examined includes 363 genes totalling 411 kilobases, making this by far the largest comparison conducted between a single pair of species. Mouse and rat genes are on average 93.4% identical in nucleotide sequence and 93.9% identical in amino acid sequence. Individual genes vary substantially in the extent of nonsynonymous nucleotide substitution, as expected from protein evolution studies; here the variation is characterized. The extent of synonymous (or silent) substitution also varies considerably among genes, though the coefficient of variation is about four times smaller than for nonsynonymous substitutions. A small number of genes mapped to the X-chromosome have a slower rate of molecular evolution than average, as predicted if molecular evolution is "male-driven." Base composition at silent sites varies from 33% to 95% G+C in different genes; mouse and rat homologues differ on average by only 1.7% in silent-site G+C, but it is shown that this is not necessarily due to any selective constraint on their base composition. Synonymous substitution rates and silent site base composition appear to be related (genes at intermediate G+C have on average higher rates), but the relationship is not as strong as in our earlier analyses. Rates of synonymous and nonsynonymous substitution are correlated, apparently because of an excess of substitutions involving adjacent pairs of nucleotides. Several factors suggest that synonymous codon usage in rodent genes is not subject to selection.

Amino Acid Sequence↗

Nucleotide sequences from the colicin E8 operon: homology with plasmid ColE2-P9.

The primary structures of the immunity (Imm) and lysis (Lys) proteins, and the C-terminal 205 amino acid residues of colicin E8 were deduced from nucleotide sequencing of the 1,265 bp ClaI-PvuI DNA fragment of plasmid ColE8-J. The gene order is col-imm-lys confirming previous genetic data. A comparison of the colicin E8 peptide sequence with the available colicin E2-P9 sequence shows an identical receptor-binding domain but 20 amino acid replacements and a clustering of synonymous codon usage in the nuclease-active region. Sequence homology of the two colicins indicates that they are descended from a common ancestral gene and that colicin E8, like colicin E2, may also function as a DNA endonuclease. The native ColE8 imm (resident copy) is 258 bp long and is predicted to encode an acidic protein of 9,604 mol. wt. The six amino acid replacements between the resident imm and the previously reported non-resident copy of the ColE8 imm ([E8 imm]) found in the ribonuclease-producing ColE3-CA38 plasmid offer an explanation for the incomplete protection conferred by [E8 Imm] to exogenously added colicin E8. Except for one nucleotide and amino acid change in the putative signal peptide sequence, the ColE8 lys structure is identical to that present in ColE2-P9 and ColE3-CA38.

Amino Acid Sequence↗

Two genes encoding gas vacuole proteins in Halobacterium halobium.

The archaebacterium Halobacterium halobium contains two related gas vacuole protein-encoding genes (vac). One of these genes encodes a protein of 76 amino acids and resides on the major plasmid. The second gene is located on the chromosome in a (G + C)-rich DNA fraction and encodes a slightly larger but highly homologous protein consisting of 79 amino acids. The plasmid encoded vac gene is transcribed constitutively throughout the growth cycle while the chromosomal vac gene is expressed during the stationary phase of growth. Comparison of the nucleotide sequences of the two genes indicates differences in the putative promoter regions as well as 35 single base-pair exchanges within the coding regions of the two genes. The majority of the nucleotide exchanges in the coding region occur in the third position of a codon triplet generating the codon synonym. The only differences between the two encoded proteins are the exchange of 2 amino acids (positions 8 and 29) and a deletion of 3 amino acids near the carboxy-terminus of the plasmid encoded vac protein. The genomic DNAs from other halobacterial isolates (Halobacterium sp. SB3, GN101 and YC819-9) were found to contain only a chromosomal vac gene copy. There is a high conservation of the chromosomal vac gene and the genomic region surrounding it among the halobacterial strains investigated.

Amino Acid Sequence↗

On concerted origin of transfer RNAs with complementary anticodons.

Pairs of antiparallely oriented consensus tRNAs with complementary anticodons show surprisingly small numbers of mispairings within the 17-bp- long anticodon stem and loop region. Even smaller such complementary distances are shown by illegitimately complementary anticodons, i.e. those with allowed pairing between G and U bases. Accordingly, we suppose that transfer RNAs have emerged concertedly as complementary strands of primordial double helix-like RNA molecules. Replication of such molecules with illegitimately complementary anticodons might generate new synonymous codons for the same pair of amino acids. Logically, the idea of tRNA concerted origin dictates very ancient establishment of direct links between anticodons and the type of amino acids with which pre-tRNAs were to be charged. More specifically, anticodons (first of all, the 2nd base) could selectively target 'their' amino acids, reaction of acylating itself being performed by another non-specific site of pre-tRNA or even by another ribozyme. In all, the above findings and speculations are consistent to the hypercyclic concept (Eigen and Schuster, 1979), and throw new light on the genetic code origin and associated problems. Also favoring this idea are data on complementary codon usage patterns in different genomes.

Amino Acids↗

Prokaryotic genetic code.

The prokaryotic genetic code has been influenced by directional mutation pressure (GC/AT pressure) that has been exerted on the entire genome. This pressure affects the synonymous codon choice, the amino acid composition of proteins and tRNA anticodons. Unassigned codons would have been produced in bacteria with extremely high GC or AT genomes by deleting certain codons and the corresponding tRNAs. A high AT pressure together with genomic economization led to a change in assignment of the UGA codon, from stop to tryptophan, in Mycoplasma.

Anticodon↗