Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

A Markov analysis of DNA sequences.

We present a model by which we look at the DNA sequence as a Markov process. It has been suggested by several workers that some basic biological or chemical features of nucleic acids stand behind the frequencies of dinucleotides (doublets) in these chains. Comparing patterns of doublet frequencies in DNA of different organisms was shown to be a fruitful approach to some phylogenetic questions (Russel & Subak-Sharpe, 1977). Grantham (1978) formulated mRNA sequence indices, some of which involve certain doublet frequencies. He suggested that using these indices may provide indications of the molecular constraints existing during gene evolution. Nussinov (1981) has shown that a set of dinucleotide preference rules holds consistently for eukaryotes, and suggested a strong correlation between these rules and degenerate codon usage. Gruenbaum, Cedar & Razin (1982) found that methylation in eukaryotic DNA occurs exclusively at C-G sites. Important biological information thus seems to be contained in the doublet frequencies. One of the basic questions to be asked (the "correlation question") is to what extent are the 64 trinucleotide (triplet) frequencies measured in a sequence determined by the 16 doublet frequencies in the same sequence. The DNA is described here as a Markov process, with the nucleotides being outcomes of a sequence generator. Answering the correlation question mentioned above means finding the order of the Markov process. The difficulty is that natural sequences are of finite length, and statistical noise is quite strong. We show that even for a 16000 nucleotide long sequence (like that of the human mitochondrial genome) the finite length effect cannot be neglected. Using the Markov chain model, the correlation between doublet and triplet frequencies can, however, be determined even for finite sequences, taking proper account of the finite length. Two natural DNA sequences, the human mitochondrial genome and the SV40 DNA, are analysed as examples of the method.

Base Sequence↗

Mutagenic damage to mammalian cells by therapeutic alkylating agents.

Cytotoxic alkylating agents used as therapeutics include nitrogen mustards, ethyleneimines, alkyl sulfonates, nitrosoureas and triazenes. Their reactivity with DNA, RNA and proteins can cause cell death. Side-effects of treatment include tissue toxicity and secondary malignancies, likely due to the genetic damage induced. The full mutagenic potential of alkylating agents may only be realised after they undergo metabolic activation, principally by cytochromes P450. Mutagenicity is related to the ability of alkylating agents to form crosslinks and/or transfer an alkyl group to form monoadducts in DNA. The most frequent location of adducts in the DNA is at guanines. Expressed mutations involve different base substitutions, including all types of transitions and transversions. The mutational spectra of alkylating agents on mammalian cells is distinct from that induced in bacterial cells, reflecting the different codon usage by bacteria and differences in DNA repair and replication enzymes. Mutations are induced by busulfan, chlorambucil (CAB), cyclophosphamide (CP, or its metabolite), dacarbazine, mechlorethamine, melphalan, mitomycin-C (MMC), nitrosoureas and thiotepa. Although dose-dependent, the relationship is not always linear. The molarities at which alkylating agents induce cell killing and mutations vary over three orders of magnitude. The mutagenic efficiency, of alkylating agents also varies, with some agents inducing three times more mutations for equivalent cell killing. The induction of micronuclei, sister chromatid exchanges, or chromosome aberrations is variable, but has been observed for CP, CAB, MMC, melphalan and triethylenemelamine. There is insufficient information to determine whether any synergistic effects of alkylating agents used in combination will influence the cytotoxic and mutagenic damage equally. Understanding the potential synergy of alkylating agents at the cellular and molecular level should allow improvement of the therapeutic efficacy of alkylating agents without increasing the unwanted mutation induction.

Alkylating Agents↗

Tobacco mosaic virus coat protein and the large subunit of the host protein ribulose-1,5-biphosphate carboxylase share a common antigenic determinant.

An immunological relationship was detected between the coat protein of the common (U1) strain of tobacco mosaic virus (TMV) and the large subunit of the ubiquitous CO2-fixing host enzyme, ribulose-1,5-biphosphate carboxylase (RuBisCo). When assayed by Western immunoblotting or indirect ELISA, polyclonal antisera to TMV coat protein and to RuBisCo reacted with both antigens. In addition, a monoclonal antibody specific for the C-terminal antigenic determinant of TMV coat protein reacted with RuBisCo. Conversely, several monoclonal antibodies generated to the large subunit of RuBisCo reacted with TMV coat protein. This cross-reactivity was verified by an examination of the amino acid sequences of both proteins. A region of homology was found between the carboxy proximal portion of coat protein and the sequence 60-73 residues from the amino terminus of RuBisCo large subunit. This homology was not mirrored at the nucleic acid level because of different codon usages for the two proteins.

Amino Acid Sequence↗

Rift Valley fever virus M segment: phlebovirus expression strategy and protein glycosylation.

The M segment RNA of Rift Valley fever virus (RVFV) encodes four gene products: the two viral envelop glycoproteins G2 and G1, a glycosylated 78-kDa protein, and a nonglycosylated 14-kDa protein. These proteins are generated from a single open reading frame (ORF) by a strategy involving independent translational initiations at both the first and second in-phase ATG codons and co-translational processing of primary polyprotein products. The ORF encodes six sites for N-linked glycosylation: one present in the "preglycoprotein region" preceding the coding sequences of the mature envelop glycoproteins, and within the coding sequences of both the 78- and 14-kDa proteins; one site in the glycoprotein G2 coding region, also present in the 78-kDa protein; and four sites within glycoprotein G1. From analyses of RVFV proteins produced in cells infected with recombinant vaccinia viruses expressing various M segment regions, we show glycoprotein G2 was glycosylated at its single site and glycoprotein G1 at at least three sites. Both sites for N-linked glycosylation in the 78-kDa protein were occupied with glycan. This latter result indicated the preglycoprotein region glycosylation site was utilized in the 78-kDa protein, but this same site within the 14-kDa protein was not. Further analysis showed utilization of this glycosylation site, as well as proteolytic processing at the amino terminus of the mature glycoprotein G2, appeared to be determined by initiation codon usage. The two-site translational initiation expression strategy of this phlebovirus M segment and its role in the control of post-translational protein modification and processing are discussed.

Bunyaviridae↗

Dynamic structures and functions of transfer ribonucleic acids from extreme thermophiles.

tRNA species from an extreme thermophile T. thermophilus that grows up to 85 degrees C have been found to be more thermostable than those from moderate thermophiles and mesophiles. Such thermostability of T. thermophilus tRNA species is partly due to the high contents of G.C base pairs in the stem regions. In addition, a novel modified nucleoside s2T has been found that substitutes T in position 54. The extent of 2-thiolation of T(54) has been found to depend on environmental temperatures from 50 to 80 degrees C. Two tRNA(Ile) species have been isolated from T. thermophilus HB8, tRNA(1aIle) with s2T(54) and tRNA(1bIle) with T(54), which have the identical nucleotide sequence except for position 54. However, the melting temperature of tRNA(1aIle) is higher by 3 degrees C than that of tRNA(1bIle). This clearly indicates that the 2-thiolation of T(54) contributes directly to the thermostability of T. thermophilus tRNA species. Proton NMR analyses have shown that the nucleoside s2T is "rigid" and predominantly takes the C3'-endo-gg-anti form of A-RNA, because of the steric effect of the bulky 2-thiocarbonyl groups and the 2'-hydroxyl group. Thus, the inherent rigidity of s2T in position 54 significantly enhances the stability of the tertiary structure of tRNA. In protein synthesis of T. thermophilus, s2T(54)-bearing tRNA and T(54)-bearing tRNA species are selectively utilized depending on environmental temperature. In the anticodons of major tRNA species from T. thermophilus, G or C exclusively appears in the first position, and GGN and CCN are favored over synonymous GCN or CGN. These characteristic anticodon sequences correspond to the characteristic codon usage in thermophile genes.

Anticodon↗

Nucleotide sequence of yeast LEU2 shows 5'-noncoding region has sequences cognate to leucine.

The LEU2 structural gene and its regulatory sequences were isolated on a 2200 bp Xho I-Sal I fragment. Sequencing of the 5'-noncoding region showed that at -151 there is an open reading frame of 23 codons of which six are for leucine. The leucine codon usage in this reading frame follows exactly that of other yeast genes. At the carboxy-terminal end and immediately after the peptide reading frame, a 14 bp hairpin (followed by a T-rich segment) can form in the putative mRNA; this arrangement closely resembles an RNA polymerase terminator. These and other features suggest a model for regulation. Preceding this is a gene (which starts at -463) for tRNALeu3, the major tRNALeu isoacceptor. RNA polymerase III transcription start and termination signals flank 5' and 3' ends, respectively, of the structural gene. The features noted above are in the same DNA strand that codes for the LEU2 gene product.

Base Sequence↗

Sea urchin (lytechinus pictus) late-stage histone H3 and H4 genes: characterization and mapping of a clustered but nontandemly linked multigene family.

We have cloned and characterized members of a small multigene family that encodes late-stage histone H3 and H4 mRNAs from the sea urchin Lytechinus pictus. Unlike their highly repetitive histone gene counterparts, which are expressed at an earlier developmental stage, late H3 and H4 histone genes are not present in tandem repeats. In addition, the late stage H3 and H4 genes are not always tightly clustered together with the H1, H2A and H2B genes as they are in early histone genes. The spacer DNA that separates adjoining H3 and H4 coding regions is not conserved between nonallelic members of the late histone gene family. We have determined the nucleotide sequence of a continuous 2100 bp segment of DNA including both H3 and H4 coding sequences, the entire spacer DNA separating the genes and surrounding nonhistone DNA. The late histone H3 and H4 genes encode proteins identical to their early gene counterparts; however, the 5' leader sequence is shorter in late genes and the codon usage is different.

Animals↗

Yeast mitochondrial RNA polymerase is homologous to those encoded by bacteriophages T3 and T7.

Analysis of the nucleotide sequence of the genetic locus for yeast mitochondrial RNA polymerase (RPO41) reveals a continuous open reading frame with the coding potential for a polypeptide of 1351 amino acids, a size consistent with the electrophoretic mobility of this enzymatic activity. The transcription product from this gene spans the singular reading frame. In vivo transcript abundance reflects codon usage and growth under stringent conditions for mitochondrial biogenesis and function results in a several fold higher level of gene expression than growth under glucose repression. A comparison of the yeast mitochondrial RNA polymerase amino acid sequence to those of E. coli RNA polymerase subunits failed to demonstrate any regions of homology. Interestingly, the mitochondrial enzyme is highly homologous to the DNA-directed RNA polymerases of bacteriophages T3 and T7, especially in regions most highly conserved between the T3 and T7 enzymes themselves.

Amino Acid Sequence↗

A general approach to isolating Plasmodium falciparum genes using non-redundant oligonucleotides inferred from protein sequences of other organisms.

We have constructed a number of oligonucleotide probes and tested their utility in identifying various genes in Plasmodium falciparum. The probe sequences were based on known conserved regions of proteins from other organisms, coupled with an analysis of the codon usage of the parasite. By using long single oligonucleotides, we have successfully isolated the DHFR-TS gene, two actin genes and two tubulin genes from the K1 (Thailand) isolate of P. falciparum. We compare these single probes to multiply-redundant short oligonucleotide probes and to heterologous probes. We also present a detailed quantitative analysis of optimal probe design, and of how this approach can best be implemented as a general method of isolating plasmodial genes.

Animals↗

A species-specific antigen of Trypanosoma (Duttonella) vivax detectable in the course of infection is encoded by a differentially expressed tandemly reiterated gene.

A monoclonal antibody that is used as a Trypanosoma vivax species-specific diagnostic reagent on antigen-trapping enzyme-linked immunosorbent assay recognized an 8-kDa peptide on western blots. The 8-kDa species-specific antigen was isolated and employed in raising rabbit polyclonal antibodies, which were used in the immunoscreening of a T. vivax cDNA library in lambda gt11.2. A clone containing a 0.8-kb insert was isolated. The cloned gene is tandemly repeated, with a monomeric unit length of 900 bp, in the genomes of all T. vivax isolates from diverse geographic locations in Africa and South America. The gene is differentially expressed, since both the transcript and antigen are present in bloodstream-stage parasites, but not in the epimastigotes of T. vivax. Although the gene is found in all T. vivax isolates so far tested, it either exists in low copy number or in a divergent form in one isolate from Kilifi at the Kenya Coast. Sequence translation revealed a remarkable degree of bias in codon usage with preference for G and C (82%) in the wobble position. Using the deduced amino acid sequence to search the databases for any structurally related peptides, revealed no significant identity with any known proteins. The function of the species-specific antigen of T. vivax is thus unknown. Nevertheless the identification and characterization of proteins released into the circulation of protozoan parasite-infected animals is important and should allow the determination of what role such molecules may play in the modulation of disease pathology.

Amino Acid Sequence↗

Purification and characterization of dihydrofolate reductase of Plasmodium falciparum expressed by a synthetic gene in Escherichia coli.

We have expressed the dihydrofolate reductase (DHFR) part of the DHFR-thymidylate synthetase complex of P. falciparum in Escherichia coli, by constructing a gene with synthetic oligonucleotides that changed the gene's codon usages. The induced expression in an E. coli cell of the synthetic gene yielded a product that constituted about 30% of the total bacterial protein. The product was precipitated in an inclusion body in a cell. Its enzymatic activity was restored after denaturation and renaturation procedures with guanidine-HCl. Recombinant DHFRs with Ser or Thr at position 108 were prepared. Kinetic characterization showed that the DHFRSer108 has less of an affinity for NADPH and dihydrofolate than the DHFRThr108.

Amino Acid Sequence↗

Cloning of two isozymes of Trichoderma koningii glyceraldehyde-3-phosphate dehydrogenase with different sensitivity to koningic acid.

Koningic acid inhibits glyceraldehyde-3-phosphate dehydrogenase (GAPDH) by binding to the SH group in the active center. The fungus Trichoderma koningii, the producer of koningic acid, contains two GAPDH isozymes (GAPDHs I and II). GAPDH I is inhibited 50% by 1.1.10(-3) M koningic acid, while GAPDH II is inhibited 50% at 6.8 x 10(-6) M. cDNAs of the two isozymes were cloned from T. koningii and their nucleotide sequences were determined. The sequence of coding region and codon usage in both clones were compared with each other and with those of the gene for Aspergillus nidulans GAPDH (enzyme activity is inhibited 50% by 2.7 x 10(-7) M koningic acid). Results indicated that GAPDH II is more closely related to A. nidulans GAPDH than GAPDH I. All essential amino acid residues, except 174 and 181, which are implicated in catalysis and binding of NAD and substrates, were conserved among A. nidulans GAPDH and GAPDHs I and II. Residues 174 and 181 are threonine in both A. nidulans GAPDH and GAPDH II, but alanine and serine, respectively, in GAPDH I. The side-chain of alanine-174 in GAPDH I can not replace threonine-174 functionally as threonine-174 side-chain forms a hydrogen bond with the catalytically essential histidine-176.

Amino Acid Sequence↗

Baculovirus expression of human basic fibroblast growth factor from a synthetic gene: role of the Kozak consensus and comparison with bacterial expression.

Synthetic genes encoding the 146 and 155 amino acid forms of human basic fibroblast growth factor (bFGF) were constructed with codon usage biased towards the polyhedrin-encoding gene of Autographa californica nuclear polyhedrosis virus (AcNPV). Expression of both bFGF genes in Spodoptera frugiperda (SF-21) suspension cell culture using a recombinant baculovirus yielded approximately 2.5 mg of mitogenically fully active protein per 10(9) cells following heparin-affinity chromatography. To improve translational efficiency, the Kozak consensus sequence was introduced and it was found that neither the replacement of a pyrimidine by a purine at position -3, nor the nature of the base at position +4 had any noticeable effect on the final levels of bFGF expression in SF-21 cells. The bases at these critical points in the consensus do not therefore play a major role in expression levels of the bFGF synthetic genes. The two synthetic genes were also expressed in Escherichia coli as native proteins using the T7 expression system. 5 mg of mitogenically fully active bFGF were obtained from 1 l of bacterial culture. Both insect cell- and E. coli-derived bFGF were equally mitogenic for Swiss 3T3 fibroblasts.

3T3 Cells↗

The nucleotide sequence of the gene coding for the elongation factor 1 alpha in Sulfolobus solfataricus. Homology of the product with related proteins.

The cloning and sequencing of the gene coding for the archaebacterial elongation factor 1 alpha (aEF-1 alpha) was performed by screening a Sulfolobus solfataricus genomic library using a probe constructed from the eptapeptide KNMITGA that is conserved in all the EF-1 alpha/EF-Tu known so far. The isolated recombinant phage contained the part of the aEF-1 alpha gene from amino acids 1 to 171. The other part (amino acids 162-435) was obtained through the amplification of the S. solfataricus DNA by PCR. The codon usage by the aEF-1 alpha gene showed a preference for triplets ending in A and/or T. This behavior was almost identical to that of the S. acidocaldarius EF-1 alpha gene but differed greatly from that of EF-1 alpha/EF-Tu genes in other archaebacteria eukaryotes and eubacteria. The translated protein is made of 435 amino acid residues and contains sequence motifs for the binding of GTP, tRNA and ribosome. Alignments of aEF-1 alpha with several EF-1 alpha/EF-Tu revealed that aEF-1 alpha is more similar to its eukaryotic than to its eubacterial counterparts.

Amino Acid Sequence↗

Production of recombinant SERA proteins of Plasmodium falciparum in Escherichia coli by using synthetic genes.

We expressed two regions of the serine repeat antigen (SERA) protein of Plasmodium falciparum in Escherichia coli by synthesizing the genes with a changed codon usage. One of the synthetic gene sequences encodes amino acid residues 17-382 (SE47') and the other encodes amino acid residues 586-802 (SE50A). The products produced by the synthetic gene sequences in E. coli accounted for 15-30% of the total bacterial protein. Antisera against both the purified gene products prepared in rats inhibited malaria parasite growth in vitro. The anti-SE47' serum was significantly more inhibitory than the anti-SE50A serum. The described methods provide a large scale preparation of recombinant antigens for improving and producing malaria vaccine.

Amino Acid Sequence↗

Utilization of acetate in Escherichia coli: structural organization and differential expression of the ace operon.

Growth of Escherichia coli on acetate as the sole source of carbon and energy requires operation of the glyoxylate bypass in connection with the expression of the polycistronic ace operon. The structural organization of this operon is presented, including the 3 structural genes coding respectively for malate synthase (aceB), isocitrate lyase (aceA) and isocitrate dehydrogenase kinase/phosphatase (aceK), and the surrounding genes iclR and metA. In addition, the differential expression of genes aceB, aceA, and aceK has been tested both in vivo in a minicell system and in vitro in a plasmid-directed transcription-translation coupled system. Moreover, the codon usage and adaptation to transfer RNA frequencies during translation of the corresponding messenger RNAs have been measured.

Acetates↗

WWW-query: an on-line retrieval system for biological sequence banks.

We have developed a World Wide Web (WWW) version of the sequence retrieval system Query: WWW-Query. This server allows to query nucleotide sequence banks in the EMBL/GenBank/DDBJ formats and protein sequence banks in the NBRF/PIR format. WWW-Query includes all the features of the on-line sequences browsers already available: possibility to build complex queries, integration of cross-references with different data banks, and access to the functional zones of biological interest. It also provides original services not available elsewhere: introduction of the notion of re-usable sequence lists, integration of dedicated helper applications for visualizing alignments and phylogenetic trees and links with multivariate methods for studying codon usage or for complementing phylogenies.

Amino Acid Sequence↗

Reconstruction and expression of the autolytic gene from Clostridium acetobutylicum ATCC 824 in Escherichia coli.

The complete lyc gene encoding the autolytic lysozyme of Clostridium acetobutylicum ATCC 824 was reconstructed from two overlapping DNA fragments and cloned into a suitable plasmid enabling Escherichia coli to produce this lytic enzyme under the control of the lac promoter. A polypeptide with an apparent M(r) of 35,000, corresponding to that predicted from the nucleotide sequence, was observed by maxicell analysis of whole-cell extracts of E. coli harboring the clostridial gene. The enzyme yield was shown to depend on the pH of the culture medium, since the protein was unstable at alkaline pH. The expression of the lyc gene was not increased by using the E. coli strong promoter, lpp-lac, probably due to the limit imposed by the extreme differences in codon usage. Although the LYC lysozyme does not contain a cleavable signal peptide, most of the protein was found in the periplasmic fraction of E. coli suggesting that this enzyme was secreted through a specific mechanism, as already observed for other autolysins.

Bacteriolysis↗