Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Mutational analysis of the HIS4 translational initiator region in Saccharomyces cerevisiae.

We have mutated various features of the 5' noncoding region of the HIS4 mRNA in light of established Saccharomyces cerevisiae and mammalian consensus translational initiator regions. Our analysis indicates that insertion mutations that introduce G + C-rich sequences in the leader, particularly those that result in stable stem-loop structures in the 5' noncoding region of the HIS4 message, severely affect translation initiation. Mutations that alter the length of the HIS4 leader from 115 to 39 nucleotides had no effect on expression, and sequence context changes both 5' and 3' to the HIS4 AUG start codon resulted in no more than a twofold decrease of expression. Changing the normal context at HIS4 5'-AAUAAUGG-3' to the optimal sequence context proposed for mammalian initiator regions 5'-CACCAUGG-3' did not result in stimulation of HIS4 expression. These studies, in conjunction with comparative and genetic studies in S. cerevisiae, support a general mechanism of initiation of protein synthesis as proposed by the ribosomal scanning model.

Alleles↗

Two genes encoding an endoglucanase and a cellulose-binding protein are clustered and co-regulated by a TTA codon in Streptomyces halstedii JM8.

Streptomyces halstedii JM8 Cel2 is an endoglucanase of 28 kDa that is first produced as a protein of 42 kDa (p42) and is later processed at its C-terminus. Cel2 displays optimal activity towards CM-cellulose at pH6 and 50 degrees C and shows no activity against crystalline cellulose or xylan. The N-terminus of p42 shares similarity with cellulases included in family 12 of the beta-glycanases and the C-terminus shares similarity with bacterial cellulose-binding domains included in family II. This latter domain enables the precursor to bind so tightly to Avicel that it can only be eluted by boiling in 10% (w/v) SDS. Another open reading frame (ORF) situated 216 bp downstream from the p42 ORF encodes a protein of 40 kDa (p40) that does not have any clear hydrolytic activity against cellulosic or xylanosic compounds, but shows high affinity for Avicel (crystalline cellulose). The p40 protein is processed in old cultures to give a protein of 35 kDa that does not bind to Avicel. Translation of both ORFs is impaired in Streptomyces coelicolor bldA mutants, suggesting that a TTA codon situated at the fourth position of the first ORF is responsible for this regulation. S1 nuclease protection experiments demonstrate that both ORFs are co-transcribed.

Amino Acid Sequence↗

Endosymbiotic origin and codon bias of the nuclear gene for chloroplast glyceraldehyde-3-phosphate dehydrogenase from maize.

The nuclei of plant cells harbor genes for two types of glyceraldehyde-3-phosphate dehydrogenases (GAPDH) displaying a sequence divergence corresponding to the prokaryote/eukaryote separation. This strongly supports the endosymbiotic theory of chloroplast evolution and in particular the gene transfer hypothesis suggesting that the gene for the chloroplast enzyme, initially located in the genome of the endosymbiotic chloroplast progenitor, was transferred during the course of evolution into the nuclear genome of the endosymbiotic host. Codon usage in the gene for chloroplast GAPDH of maize is radically different from that employed by present-day chloroplasts and from that of the cytosolic (glycolytic) enzyme from the same cell. This reveals the presence of subcellular selective pressures which appear to be involved in the optimization of gene expression in the economically important graminaceous monocots.

Amino Acid Sequence↗

Genetic code and optimal resistance to the effects of mutations.

This paper deals with the notion of resistance of the genetic code to the effects of mutations. We measure the resistance of a group of t codons as the number of pairs of those which differ from each other in only one of their three bases. We find for each value of t the maximum possible value of the resistance and we describe some groups of codons giving this value. Important examples of such configurations are found in the genetic code, among these are the groups of synonymous codons, as observed elsewhere, and the cluster of codons which have an hydrophobic amino acid for translation.

Base Sequence↗

Optimally parsing a sequence into different classes based on multiple types of evidence.

We consider the problem of parsing a sequence into different classes of subsequences. Two common examples are finding the exons and introns in genomic sequences and identifying the secondary structure domains of protein sequences. In each case there are various types of evidence that are relevant to the classification, but none are completely reliable, so we expect some weighted average of all the evidence to provide improved classifications. For example, in the problem of identifying coding regions in genomic DNA, the combined use of evidence such as codon bias and splice junction patterns can give more reliable predictions than either type of evidence alone. We show three main results: 1. For a given weighting of the evidence a dynamic programming algorithm returns the optimal parse and any number of sub-optimal parses. 2. For a given weighting of the evidence a dynamic programming algorithm determines the probability of the optimal parse and any number of sub-optimal parses under a natural Boltzmann-Gibbs distribution over the set of possible parses. 3. Given a set of sequences with known correct parses, a dynamic programming algorithm allows one to apply gradient descent to obtain the weights that maximize the probability of the correct parses of these sequences.

Algorithms↗

Many combinations of amino acid sequences in a conserved region of the D1 protein satisfy photosystem II function.

The putative de helix of the D1 protein is located at the acceptor side of photosystem II (PS II) and serves as an indispensable part of a niche that binds the secondary plastoquinone QB. Combinatorial mutagenesis was applied to a stretch of four residues in a highly conserved region of this putative helix in order to reveal amino acid combinations that are able to support PS II function. An obligate photoheterotrophic mutant of the cyanobacterium Synechocystis sp. PCC 6803, missing four residues (delta YFGR254-7) in the de helix, was transformed with a D1-coding sequence carrying fully degenerate combinations of codons at the site of the deletion. Upon selection for photoautotrophy, 25 mutants with functional PS II were isolated. All mutants showed different codon combinations at positions 254 to 257; none was identical to the wild-type sequence, and none of the conserved residues was found to be mandatory for PS II function. However, 24 of the mutants contained Tyr of Phe at position 254 while at the other three positions many different amino acid combinations could be functionally accommodated. Most sequences maintained an amphiphilic arrangement of the helix that may align Tyr254 facing the QB binding pocket. This residue is proposed to be functionally analogous to Phe216 of the L subunit in purple bacteria which contributes to binding of QB. Most of the PS II properties were similar in the mutants compared to wild-type. Noticeable modifications in the mutants concerned the semiquinone equilibrium of electron transfer between QA and QB, and the affinity of PS II inhibitors. Differential effects on the semiquinone equilibrium were observed between two distinct quinones occupying the QB site (plastoquinone versus 2,5-dichloro-p-benzo- quinone), implying that residues in this domain are involved, directly or indirectly, with different binding determinants of the quinones. Even though many different combinations of amino acids in positions 254 to 257 of the D1 protein may satisfy the primary function of PS II, complex requirements need to be combined for optimized performance of the QB binding niche.

Amino Acid Sequence↗

Local activation and inactivation of thyroid hormones: the deiodinase family.

Tissue-specific activation and inactivation of ligands of nuclear receptors which belong to the steroid retinoid-thyroid hormone superfamily of transcription factors represents an important principle of development- and tissue-specific local modulation of hormone action. Recently, several enzyme families have been identified which act as 'guardians of the gate' of ligand-activated transcription modulation. Three monodeiodinase isoenzymes which are involved in activation the 'prohormone' L-thyroxine (T4), the main secretory product of the thyroid gland, have been identified, characterized, and cloned. Both, type I and type II 5'-deiodinase generate the thyromimetically active hormone 3,3',5-triiodothyronine (T3) by reductive deiodination of the phenolic ring of T4. Inactivation of T4 and its product T3 occurs by deiodination of iodothyronines at the tyrosyl ring. This reaction is catalyzed both the type III 5-deiodinase and also by the type I enzyme, which has a broader substrate specificity. The three deiodinases appear to constitute a newly discovered family of selenocysteine-containing proteins and the presence of selenocysteine in the protein is critical for enzyme activity. Whereas the selenoenzyme characteristics of the type I and type III deiodinases are definitively established some controversy still exists for the type II 5'-deiodinase in mammals. The mRNA probably encoding the type II 5'-deiodinase subunit is markedly longer than those of the two other deiodinases and its selenocysteine-insertion element is located more than 5 kB downstream of the UGA-codon in the 3'-untranslated region. The three deiodinase isoenzymes show a distinct development- and tissue-specific pattern of expression, operate at individual optimal substrate levels, are differently regulated and modulated by hormones, cytokines, signaling pathways, natural factors, and pharmaceuticals. Whereas circulating T3 mainly originates from hepatic production via the type I 5'-deiodinase, the local cellular thyroid hormone concentration in various tissues including the central nervous system is controlled by complex para-, auto-, and intracrine interactions of all three deiodinases. Local thyroid hormone availability is further modulated by conjugation reactions of the phenolic 4'-OH-group of iodothyronines, which also inactivate the thyroid hormones.

Animals↗

Rapid and direct detection of the most frequent Mediterranean beta-thalassemic mutations by multiplex allele-specific enzymatic amplification.

A rapid nonradioactive method for the diagnosis of the most frequent Mediterranean beta-thalassemic mutations is described based on a multiplex allele-specific polymerase chain reaction (PCR). This method allows direct detection of normal or mutated alleles on genomic DNA. We have used this approach to detect the most frequent Mediterranean mutations: IVS-1 nt 110 (G----A) and 39 nonsense (C----T). For each mutation three allele-specific oligonucleotides were used: one common upstream primer and two downstream primers differing in their terminal 3' nucleotide (one specific for the normal allele and one for the mutant allele). For each sample two PCR reactions were performed in parallel using in one case IVS-1 nt 110 and codon 39 normal primers and in the second case using the corresponding mutated primers. In both cases the different PCR fragments were visualized. After optimization these primers directed only amplification of their complementary allele. A single blind study was performed on the DNA of 18 individuals who were homozygous or heterozygous for these mutations. In comparison with a parallel investigation, using oligonucleotide probes, all the results were unambiguous. This diagnosis method, which is rapid, easy, direct, and inexpensive, allows the screening of a population group, including heterozygotes, which is required from an epidemiological and anthropological point of view. It could be extended to the large series screening of haplotypes before targeted diagnosis of various genetic diseases.

Algeria↗

[Cloning of the gene for thermostable Thermus aquaticus YT1 DNA polymerase and its expression in Escherichia coli].

Using the phasmid vector pSL5, the genomic DNA fragment of T. aquaticus YT1 which contained the thermostable DNA polymerase (Taq-polymerase) gene was cloned. The BglII fragment of this genome locus was subcloned in the BamHI site of the pUC19 plasmid. To optimize the Taq-polymerase gene expression in E. coli cells, the gene was cloned in the correct reading frame regarding the initiation ATG codon of the pPR-TGATG-1 expression vector. The gene expression in this vector was controlled by the phage lambda PR promoter and the temperature-sensitive phage lambda repressor. We used PCR to amplify the short 5'-end fragment of the Taq-polymerase gene coding for the part into which an artificial SacI site was introduced. This site has been used for cloning the PCR product into the pPR-TGATG-1 vector, and the missing gene part was cloned into the KpnI site of the PCR product from the natural cloned gene. The cells of the E. coli PVG-A1 strain, which was obtained in the end, expressed efficiently the Taq-polymerase gene at the nonpermissive temperature. The content of the recombinant Taq-polymerase in the cells was about 1-2% of total proteins. The purified nearly homogeneous Taq-polymerase amplified efficiently in the PCR DNA fragments up to 5.5 kb long and was useful in DNA sequencing the by Sanger method. The half-life of the purified Taq polymerase was about 60 min at 95 degrees C, it was active for at least 65 standard PCR circles. The specific activity of recombinant enzyme preparations was about 180-200,000 units per mg of protein. The E. coli PVG-A1 strain enables one to isolate up to 500,000 units of purified enzyme from 2 l of bacterial culture.

Bacteriophage lambda↗

Efficient translation initiation is required for replication of bovine viral diarrhea virus subgenomic replicons.

An internal ribosome entry site (IRES) mediates translation initiation of bovine viral diarrhea virus (BVDV) RNA. Studies have suggested that a portion of the N(pro) open reading frame (ORF) is required, although its exact function has not been defined. Here we show that a subgenomic (sg) BVDV RNA in which the NS3 ORF is preceded only by the 5' nontranslated region did not replicate to detectable levels following transfection. However, RNA synthesis and cytopathic effects were observed following serial passage in the presence of a noncytopathic helper virus. Five sg clones derived from the passaged virus contained an identical, silent substitution near the beginning of the NS3 coding sequence (G400U), as well as additional mutations. Four of the reconstructed mutant RNAs replicated in transfected cells, and in vitro translation showed increased levels of NS3 for the mutant RNAs compared to that of wild-type (wt) MetNS3. To more precisely dissect the role of these mutations, we constructed two sg derivatives: ad3.10, which contains only the G400U mutation, and ad3.7, with silent substitutions designed to minimize RNA secondary structure downstream of the initiator AUG. Both RNAs replicated and were translated in vitro to similar levels. Moreover, ad3.7 and ad3.10, but not wt MetNS3, formed toeprints downstream of the initiator AUG codon in an assay for detecting the binding of 40S ribosomal subunits and 43S ribosomal complexes to the IRES. These results suggest that a lack of stable RNA secondary structure(s), rather than a specific RNA sequence, immediately downstream of the initiator AUG is important for optimal translation initiation of pestivirus RNAs.

Amino Acid Sequence↗

Use of mitogenomic information in teleostean molecular phylogenetics: a tree-based exploration under the maximum-parsimony optimality criterion.

We explored the phylogenetic utility and limits of the individual and concatenated mitochondrial genes for reconstructing the higher-level relationships of teleosts, using the complete (or nearly complete) mitochondrial DNA sequences of eight teleosts (including three newly determined sequences), whose relative phylogenetic positions were noncontroversial. Maximum-parsimony analyses of the nucleotide and amino acid sequences of 13 protein-coding genes from the above eight teleosts, plus two outgroups (bichir and shark), indicated that all of the individual protein-coding genes, with the exception of ND5, failed to recover the expected phylogeny, although unambiguously aligned sequences from 22 concatenated transfer RNA (tRNA) genes (stem regions only) recovered the expected phylogeny successfully with moderate statistical support. The phylogenetic performance of the 13 protein-coding genes in recovering the expected phylogeny was roughly classified into five groups, viz. very good (ND5, ND4, COIII, COI), good (COII, cyt b), medium (ND3, ND2), poor (ND1, ATPase 6), and very poor (ND4L, ND6, ATPase 8). Although the universality of this observation was unclear, analysis of successive concatenation of the 13 protein-coding genes in the same ranking order revealed that the combined data sets comprising nucleotide sequences from the several top-ranked protein-coding genes (no 3rd codon positions) plus the 22 concatenated tRNA genes (stem regions only) best recovered the expected phylogeny, with all internal branches being supported by bootstrap values >90%. We conclude that judicious choice of mitochondrial genes and appropriate data weighting, in conjunction with purposeful taxonomic sampling, are prerequisites for resolving higher-level relationships in teleosts under the maximum-parsimony optimality criterion.

Animals↗

High guanine-cytosine content is not an adaptation to high temperature: a comparative analysis amongst prokaryotes.

The causes of the variation between genomes in their guanine (G) and cytosine (C) content is one of the central issues in evolutionary genomics. The thermal adaptation hypothesis conjectures that, as G:C pairs in DNA are more thermally stable than adenonine:thymine pairs, high GC content may he a selective response to high temperature. A compilation of data on genomic GC content and optimal growth temperature for numerous prokaryotes failed to demonstrate the predicted correlation. By contrast, the GC content of Structural RNAs is higher at high temperatures. The issue that we address here is whether more freely evolving sites in exons (i.e. codonic third positions) evolve in the same manner as genomic DNA as a whole, Showing no correlated response, or like structural RNAs showing a strong correlation. The latter pattern would provide strong support for the thermal adaptation hypothesis, as the variation in GC content between orthologous genes is typically most profoundly seen at codon third sites (GC3). Simple analysis of completely sequenced prokaryotic genomes shows that GC3, but not genomic GC, is higher on average in thermophilic species. This demonstrates, if nothing else, that the results from the two measures cannot be presumed to be the same. A proper analysis, however, requires phylogenetic control. Here, therefore, we report the results of a comparative analysis of GC composition and optimal growth temperature for over 100 prokaryotes. Comparative analysis fails to show, in either Archea or Eubacteria, any hint of connection between optimal growth temperature and GC content in the genome as a whole, in protein-coding regions or, more crucially at GC. Conversely, comparable analysis confirms that GC content of structural RNA is strongly correlated with optimal temperature. Against the expectations of the thermal adaptation hypothesis, within prokaryotes GC content in protein-coding genies, even at relatively freely evolving sites, cannot be considered an adaptation to the thermal environment.

Adaptation, Physiological↗

The human cytomegalovirus UL97 protein is a protein kinase that autophosphorylates on serines and threonines.

The product of the human cytomegalovirus (CMV) UL97 gene, which controls ganciclovir phosphorylation in virus-infected cells, is homologous to known protein kinases but diverges from them at a number of positions that are functionally important. To investigate UL97, we raised an antibody against it and overexpressed it in baculovirus-infected insect cells. Recombinant baculovirus expressing full-length UL97 directed the phosphorylation of ganciclovir in insect cells, which was abolished by a four-codon deletion that confers ganciclovir resistance to CMV. When incubated with [gamma-32P]ATP, full-length UL97 was phosphorylated on serine and threonine residues. Phosphorylation was severely impaired by a point mutation that alters lysine-355 in a motif that aligns with subdomain II of protein kinases. However, phosphorylation was impaired much less severely by the four-codon deletion. A UL97 fusion protein expressed from recombinant baculovirus was purified to near homogeneity. It too was phosphorylated upon incubation with [gamma-32P]ATP in vitro. This phosphorylation, which was abolished by the lysine 355 mutation, was optimal at high NaCl and high pH. The activity required either Mn2+ or Mg2+, with a preference for Mn2+, and utilized either ATP or GTP as a phosphate donor, with Kms of 2 and 4 microM, respectively. The phosphorylation rate was first order with protein concentration, consistent with autophosphorylation. These data strongly argue that UL97 is a serine/threonine protein kinase that autophosphorylates and suggest that the four-codon deletion affects its substrate specificity.

Animals↗

Purification and characterization of recombinant cytochrome P450TYR expressed at high levels in Escherichia coli.

The multifunctional tyrosine N-hydroxylase, cytochrome P450TYR (CYP79), from Sorghum bicolor catalyzing the conversion of tyrosine to p-hydroxyphenyl-acetaldoxime in the biosynthesis of the cyanogenic glucoside dhurrin, has been expressed in Escherichia coli using the isopropyl-beta-D-thiogalactopyranoside-inducible vector pSP19g10L, containing the cDNA encoding CYP79. The expression construct was optimized by reducing the length of the N-terminal hydrophobic core of the signal sequence of cytochrome P450TYR and by exchanging the first eight codons with the first eight codons of bovine P45017 alpha. The highest yielding construct provided 200-500 nmol P450TYR/liter cell culture. The recombinant P450TYR was gently and efficiently extracted from E. coli spheroblasts by temperature-induced phase partitioning of Triton X-114 in the presence of 30% glycerol and isolated by DEAE and reactive red chromatography. In reconstitution experiments using saturating amounts of sorghum NADPH-cytochrome P450 reductase, the Km and turnover rate for isolated recombinant P450TYR was 0.22 +/- 0.06 mM and 49.2 +/- 3.8 min-1, respectively, whereas a turnover rate as high as 350 min-1, was obtained using E. coli membranes. Addition of 3 mM glutathione stimulated the activity of reconstituted P450TYR and of sorghum microsomes although the effect was highly variable. Phenylalanine, the precursor of several cyanogenic glucosides, gave a type I binding spectrum, but was not metabolized by P450TYR, demonstrating the high substrate specificity of this P450. Administration of radioactively labeled p-hydroxyphenylacetaldoxime to E. coli cells, showed E. coli metabolized p-hydroxyphenylacetaldoxime independent of the expression of P450TYR.

Amino Acid Sequence↗

The production, purification, and bioactivity of recombinant bovine trophoblast protein-1 (bovine trophoblast interferon).

Bovine trophoblast protein-1 (bTP-1) is a 172-amino acid interferon- alpha that has a role in maternal recognition of pregnancy in cattle. Here we describe production of bTP-1 by recombinant procedures in Escherichia coli. A bTP-1 gene was constructed which lacked the codons representing the signal sequence and provided a Met initiation codon ahead of the TGT codon encoding Cys1 of the mature protein. This construct was placed under the control of the Trp promoter within the expression vector pTrp2. Expression occurred optimally in E. coli D112 in the absence of tryptophan and in the presence of 0.5% acid-hydrolyzed casein (casamino acids) when 0.5 mM indole acetic acid was included in the medium. The bTP-1 was deposited in inclusion bodies and accounted for as much as 27% of the total cellular protein. The inclusion bodies were isolated by differential centrifugation and washed. The bTP-1 was solubilized by use of guanidinium-HCI and 2-mercaptoethanol and allowed to renature in air. Final purification was achieved by anion exchange chromatography on DEAE-cellulose. The yield of purified product, which had an antiviral activity greater than 10(8) international reference units/mg, was approximately 20 mg/liter. The recombinant bTP-1 was relatively stable to freeze-thawing and frozen storage, and could induce the production of an acidic protein of 70,000 mol wt in cultured explants of endometrium prepared from ewes on day 13 of the estrous cycle. The latter protein is a characteristic product of interferon-alpha action on uterine tissue.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

The level and landscape of optimization in the origin of the genetic code.

We consider a model of the origin of genetic code organization incorporating the biosynthetic relationships between amino acids and their physicochemical properties. We study the behavior of the genetic code in the set of codes subject both to biosynthetic constraints and to the constraint that the biosynthetic classes of amino acids must occupy only their own codon domain, as observed in the genetic code. Therefore, this set contains the smallest number of elements ever analyzed in similar studies. Under these conditions and if, as predicted by physicochemical postulates, the amino acid properties played a fundamental role in genetic code organization, it can be expected that the code must display an extremely high level of optimization. This prediction is not supported by our analysis, which indicates, for instance, a minimization percentage of only 80%. These observations can therefore be more easily explained by the coevolution theory of genetic code origin, which postulates a role that is important but not fundamental for the amino acid properties in the structuring of the code. We have also investigated the shape of the optimization landscape that might have arisen during genetic code origin. Here, too, the results seem to favor the coevolution theory because, for instance, the fact that only a few amino acid exchanges would have been sufficient to transform the genetic code (which is not a local minimum) into a much better optimized code, and that such exchanges did not actually take place, seems to suggest that, for instance, the reduction of translation errors was not the main adaptive theme structuring the genetic code.

Algorithms↗

Effects of various amino acid 256 mutations on sarcoplasmic/endoplasmic reticulum Ca2+ ATPase function and their role in the cellular adaptive response to thapsigargin.

Upon direct selection of mammalian cells for resistance to thapsigargin (TG), a potent inhibitor of the sarcoplasmic/endoplasmic reticulum Ca2+ transport ATPase (SERCA), the ATPase can acquire specific mutations at amino acid position 256 (aa256). In particular, Phe256 --> Leu and Phe256 --> Ser substitutions can occur upon TG selection, with each substitution resulting in a SERCA that is 4- to 5-fold resistant to TG inhibition (M. Yu et al., J. Biol. Chem. 273, 3542-3546, 1998). We have now identified a third substitution, i.e., Phe256 --> Val, that occurs when the Chinese hamster lung fibroblast cell line DC-3F is selected for TG resistance. Although the Phe256 --> Val substitution at codon 256 results in a SERCA whose enzymological properties in terms of Ca2+ transport and ATP hydrolysis are essentially similar to that of wild-type (wt) SERCA, the mutant enzyme is more than 40-fold resistant to TG inhibition. To analyze further the role of aa256 in TG-SERCA interactions, mutational analysis of this particular residue was also carried out. Of all the mutations introduced, only the Phe256 --> Glu substitution interferes with expression of the ATPase. The Phe256 --> Arg substitution does not interfere with SERCA expression, but the resulting enzyme is totally inactive. In terms of sensitivity of the various mutants to TG, maximal reduction in the ATPase's affinity for TG occurs with amino acid substitutions containing branched side chains, i.e. with the Phe256 --> Val, Phe256 --> Ile, and Phe256 --> Thr mutants. Since a corresponding Phe is conserved in the Na+, K+-ATPase which is not sensitive to TG, our findings suggest that this amino acid provides stabilization of the stalk segment with respect to the membrane interface, thereby optimizing specific interactions of TG with neighboring S3 residues (L. Zhong and G. Inesi, J. Biol. Chem. 273, 12994-12998, 1998). It is likely that a relatively high frequency of codon 256 mutations favor the aa256 mutants as a specific adaptive response to TG selection.

ATP Binding Cassette Transporter, Subfamily B↗

Expression of polypeptides of human immunodeficiency virus-1 reverse transcriptase in Escherichia coli.

We have prepared a plasmid, pRC-RT, for expression of HXB2 HIV-1 reverse transcriptase (RT) in Escherichia coli (Becerra et al., Biochemistry 30, 11707-11719, 1991). Here we describe the optimization of RT overexpression and its purification. In pRC-RT, the precise RT coding region of HXB2 proviral DNA is flanked by start and stop codons, and expression is driven by the phage lambda pL promoter in a temperature-inducible system. The 64,484-Da RT polypeptide (termed p66) is expressed as approximately 10% of total cell protein after 2 h of induction, and the RT is readily solubilized and purified free of DNA Pol I and to near homogeneity as a homodimer of p66 or as a heterodimer of p66 and p51, resembling the natural enzyme. After achieving appropriate expression of the full-length p66 RT, we next created vectors to express multiple individual segments of the p66 polypeptide. These segments are: a 51,000-Da peptide, representing C-terminal truncation of p66, and several peptides representing consecutive N-terminal, central, and C-terminal segments of p66. The latter peptide, corresponding to the RNase H domain of RT, has been purified in large quantities and is currently under study for solution of its structure by NMR. This peptide is devoid of enzyme activity and of substrate-binding capacity, but exists in solution as a folded globular protein with structure resembling that of E. coli ribonuclease H and that of a similar HIV-1 RT RNase H domain peptide examined by X-ray crystallography (Becerra et al., FEBS Lett. 270, 67-80, 1990). Various other RT peptides described here should prove to be similarly useful for structural studies, as well as other approaches.

Amino Acid Sequence↗