Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

The complete mitochondrial sequence of Tarsius bancanus: evidence for an extensive nucleotide compositional plasticity of primate mitochondrial DNA.

Inconsistencies between phylogenetic interpretations obtained from independent sources of molecular data occasionally hamper the recovery of the true evolutionary history of certain taxa. One prominent example concerns the primate infraordinal relationships. Phylogenetic analyses based on nuclear DNA sequences traditionally represent Tarsius as a sister group to anthropoids. In contrast, mitochondrial DNA (mtDNA) data only marginally support this affiliation or even exclude Tarsius from primates. Two possible scenarios might cause this conflict: a period of adaptive molecular evolution or a shift in the nucleotide composition of higher primate mtDNAs through directional mutation pressure. To test these options, the entire mt genome of Tarsius bancanus was sequenced and compared with mtDNA of representatives of all major primate groups and mammals. Phylogenetic reconstructions at both the amino acid (AA) and DNA level of the protein-coding genes led to faulty tree topologies depending on the algorithms used for reconstruction. We propose that these artifactual affiliations rather reflect the nucleotide compositional similarity than phylogenetic relatedness and favor the directional mutation pressure hypothesis because: (1) the overall nucleotide composition changes dramatically on the lineage leading to higher primates at both silent and nonsilent sites, and (2) a highly significant correlation exists between codon usage and the nucleotide composition at the third, silent codon position. Comparisons of mt genes with mt pseudogenes that presumably transferred to the nucleus before the directional mutation pressure took place indicate that the ancestral DNA composition is retained in the relatively fossilized mtDNA-like sequences, and that the directed acceleration of the substitution rate in higher primates is restricted to mtDNA.

Animals↗

Nucleotide sequence and characteristics of the gene for L-lactate dehydrogenase of Thermus caldophilus GK24 and the deduced amino-acid sequence of the enzyme.

The gene for L-lactate dehydrogenase (LDH) (EC 1.1.1.27) of Thermus caldophilus GK24 was cloned in Escherichia coli using synthetic oligonucleotides as hybridization probes. The nucleotide sequence of the cloned DNA was determined. The primary structure of the LDH was deduced from the nucleotide sequence. The deduced amino acid sequence agreed with the NH2-terminal and COOH-terminal sequences previously reported and the determined amino acid sequences of the peptides obtained from trypsin-digested T. caldophilus LDH. The LDH comprised 310 amino acid residues and its molecular mass was determined to be 32,808. On alignment of the whole amino acid sequences, the T. caldophilus LDH showed about 40% identity with the Bacillus stearothermophilus, Lactobacillus casei and dogfish muscle LDHs. The T. caldophilus LDH gene was expressed with the E. coli lac promoter in E. coli, which resulted in the production of the thermophilic LDH. The gene for the T. caldophilus LDH showed more than 40% identity with those for the human and mouse muscle LDHs on alignment of the whole nucleotide sequences. The G + C content of the coding region for the T. caldophilus LDH was 74.1%, which was higher than that of the chromosomal DNA (67.2%). The G + C contents in the first, second and third positions of the codons used were 77.7%, 48.1% and 95.5% respectively. The high G + C content in the third base caused extremely non-random codon usage in the LDH gene. About half (48.7%) the codons in the LDH gene started with G, and hence there were relatively high contents of Val, Ala, Glu and Gly in the LDH. The contents of Pro, Arg, Ala and Gly, which have high G + C contents in their codons, were also high. Rare codons with U or A as the third base were sometimes used to avoid the TCGA sequence, the recognition site for the restriction endonuclease, TaqI. Two TCGA sequences were found only in the sequence of CTCGAG (XhoI site) in the sequenced region of the T. caldophilus DNA. There were three segments with similar sequences in the two 5' non-coding regions, probably the promoter and ribosome-binding regions, of the genes for the T. caldophilus LDH and the Thermus thermophilus 3-isopropylmalate dehydrogenase.

Amino Acid Sequence↗

Expression of the strA-strB streptomycin resistance genes in Pseudomonas syringae and Xanthomonas campestris and characterization of IS6100 in X. campestris.

Expression of the strA-strB streptomycin resistance (SMr) genes was examined in Pseudomonas syringae pv. syringae and Xanthomonas campestris pv. vesicatoria. The strA-strB genes in P. syringae and X. campestris were encoded on elements closely related to Tn5393 from Erwinia amylovora and designated Tn5393a and Tn5393b, respectively. The putative recombination site (res) and resolvase-repressor (tnpR) genes of Tn5393 from E. amylovora, P syringae, and X. campestris were identical; however, IS6100 mapped within tnpR in X. campestris, and IS1133 was previously located downstream of tnpR in E. amylovora (C.-S Chiou and A. L. Jones, J. Bacteriol. 175:732-740, 1993). Transcriptional fusions (strA-strB::uidA) indicated that a strong promoter sequence was located within res in Tn5393a. Expression from this promoter sequence was reduced when the tnpR gene was present in cis position relative to the promoter. In X. campestris pv. vesicatoria, analysis of promoter activity with transcriptional fusions indicated that IS6100 increased the expression of strA-strB. Analysis of codon usage patterns and percent G+C in the third codon position indicated that IS6100 could have originated in a gram-negative bacterium. The data obtained in the present study help explain differences observed in the levels of SMr expressed by three genera which share common genes for resistance. Furthermore, the widespread dissemination of Tn5393 and derivatives in phytopathogenic prokaryotes confirms the importance of these bacteria as reservoirs of antibiotic resistance in the environment.

Base Sequence↗

Characterization of the rfc region of Shigella flexneri.

The O antigen of the Shigella flexneri lipopolysaccharide (LPS) is an important virulence determinant and immunogen. We have isolated S. flexneri mutants which produce a semi-rough LPS by using an O-antigen-specific phage, Sf6c. Western immunoblotting was used to show that the LPS produced by the semi-rough mutants contained only one O-antigen repeat unit. Thus, the mutants are deficient in production of the O-antigen polymerase and were termed rfc mutants. Complementation experiments were used to locate the rfc adjacent to the rfb genes on plasmid clones previously isolated and containing this region (D. F. Macpherson, R. Morona, D. W. Beger, K.-C. Cheah, and P. A. Manning, Mol. Microbiol 5:1491-1499, 1991). A combination of deletions and subcloning analysis located the rfc gene as spanning a 2-kb region. Insertion of a kanamycin resistance cartridge into a SalI site in this region inactivated the rfc gene. The DNA sequence of the rfc region was determined. An open reading frame spanning the SalI site was identified and encodes a protein with a predicted molecular mass of 43.7 kDa. The predicted protein is highly hydrophobic and showed little sequence homology with any other protein. Comparison of its hydropathy plot with that of other Rfc proteins from Salmonella enterica (typhimurium) and Salmonella enterica (muenchen) revealed that the profiles were similar and that the proteins have 12 or more potential membrane-spanning segments. A comparison of the S. flexneri rfc gene and protein product with other rfc and rfc-like proteins revealed that they have a similarly low percentage of G + C content and have similar codon usage, and all have a high percentage of rare codons. An attempt to identify the S. flexneri Rfc protein was unsuccessful, although proteins encoded upstream and downstream of the rfc gene could be identified. Examination of the distribution of rare or minor codons in the rfc gene revealed that it has several minor codons within the first 25 amino acids. This is in contrast to the upstream gene rfbG, which also has a high percentage of rare codons but whose gene product could be detected. The positioning of the rare codons in the rfc gene may restrict translation and suggests that minor isoaccepting tRNA species may be involved in translational regulation of rfc expression. The low percentage of G + C content of rfc genes may be a consequence of the selection pressure to maintain this form of control.

Amino Acid Sequence↗

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing↗

The compositional distribution of coding sequences and DNA molecules in humans and murids.

The compositional distributions of coding sequences and DNA molecules (in the 50-100-kb range) are remarkably narrower in murids (rat and mouse) compared to humans (as well as to all other mammals explored so far). In murids, both distributions begin at higher and end at lower GC values. A comparison of homologous coding sequences from murids and humans revealed that their different compositional distributions are due to differences in GC levels in all three codon positions, particularly of genes located at both ends of the distribution. In turn, these differences are responsible for differences in both codon usage and amino acids. When GC levels at first + second codon positions and third codon positions, respectively, of murid genes are plotted against corresponding GC levels of homologous human genes, linear relationships (with very high correlation coefficients and slopes of about 0.78 and 0.60, respectively) are found. This indicates a conservation of the order of GC levels in homologous genes from humans and murids. (The same comparison for mouse and rat genes indicates a conservation of GC levels of homologous genes.) A similar linear relationship was observed when plotting GC levels of corresponding DNA fractions (as obtained by density gradient centrifugation in the presence of a sequence-specific ligand) from mouse and human. These findings indicate that orderly compositional changes affecting not only coding sequences but also noncoding sequences took place since the divergence of murids. Such directional fixations of mutations point to the existence of selective pressures affecting the genome as a whole.

Amino Acid Sequence↗

Sequence diversity and molecular evolution of the merozoite surface antigen 2 of Plasmodium falciparum.

Eleven new alleles of the Plasmodium falciparum merozoite surface antigen 2 (MSA2) from Papua New Guinea were analyzed by direct sequencing of polymerase chain reaction (PCR) products. We have used the sequence information to trace the molecular evolution of MSA2. The repeats of ten alleles belonging to the 3D7 allelic family differed considerably in size, nucleotide sequence, and repeat copy number. In the repeat region of these new alleles, codon usage was extremely biased with an exclusive use of NNT codons. Another new allele sequenced belonged to the FC27 family and confirmed the family-specific conserved structure of 96 and 36 bp repeats. In order to assess sequence microheterogeneity within samples defined as the same genotype by restriction fragment length polymorphism (RFLP), we have analyzed single-strand conformation polymorphism (SSCP) of different samples of the most frequent allele (D10 of the FC27 family) in the study population. No sequence heterogeneity could be detected within the repeat region. Based on analysis of the repeat regions in both allelic families, we discuss the hypothesis of a different evolutionary strategy being represented by each of the allelic families. Kew words: Merozoite surface antigen 2 - Nucleotide sequence comparisons - Molecular evolution

Amino Acid Sequence↗

Gene expression, amino acid conservation, and hydrophobicity are the main factors shaping codon preferences in Mycobacterium tuberculosis and Mycobacterium leprae.

Mycobacterium tuberculosis and Mycobacterium leprae are the ethiological agents of tuberculosis and leprosy, respectively. After performing extensive comparisons between genes from these two GC-rich bacterial species, we were able to construct a set of 275 homologous genes. Since these two bacterial species also have a very low growth rate, translational selection could not be so determinant in their codon preferences as it is in other fast-growing bacteria. Indeed, principal-components analysis of codon usage from this set of homologous genes revealed that the codon choices in M. tuberculosis and M. leprae are correlated not only with compositional constraints and translational selection, but also with the degree of amino acid conservation and the hydrophobicity of the encoded proteins. Finally, significant correlations were found between GC3 and synonymous distances as well as between synonymous and nonsynonymous distances.

Amino Acid Sequence↗

The plastid genome of the critically endangered Valeriana trinervis (= Centranthus trinervis) and insights from comparison with other Valeriana plastomes (Caprifoliaceae).

The first complete plastid genome of the critically endangered species Valeriana trinervis was sequenced, assembled and compared with other published Valeriana plastomes. In this study, we assembled the plastid genome of the critically endangered, endemic species Valeriana trinervis (= Centranthus trinervis) and compare it with all published plastomes of Valeriana. We found not only differences in the inverted repeats boundaries, in the type and abundance of repeats, but also similarities in codon usage and microsatellite numbers. We detected non-canonical start codons in several genes and identified variation in several regions that could be useful for phylogenetic and phylogeographic studies. The phylogenetic tree inference based on both full plastomes and coding sequence data indicated that V. trinervis is sister to all Eurasian Valeriana accessions confirming the phylogenetic position recently investigated. This is the first plastome available for a species of the Mediterranean clade of Valeriana previously known as Centranthus, and it adds further data to understand the evolution and diversification of this systematically debated genus.

Genome, Plastid↗

Nucleotide sequence of the mitochondrial structural gene for subunit 9 of yeast ATPase complex.

We have determined the nucleotide sequence of a segment of Saccharomyces mtDNA that contains the structural gene for one of the subunits (the dicyclohexylcarbodiimide-binding protein) of the mitochondrial ATPase complex. The sequence fits the known amino acid sequence of this protein with the exception of one amino acid. Codon usage is biased in favor of A + T-rich codons. On both sides of the gene, the nucleotide sequence contains less than 4% (mol/mol) G + C for at least 180 nucleotides; these A + T sequences show no evidence of internal repetition. The gene and all the A + T-rich sequence preceding the gene are present in a 12S RNA that is the major transcript of this segment of mtDNA. The nature of the sequences responsible for binding ribosomes to mitochondrial mRNA and for termination of RNA synthesis is considered.

Adenosine Triphosphatases↗

Structure and regulation of the anthranilate synthase genes in Pseudomonas aeruginosa: I. Sequence of trpG encoding the glutamine amidotransferase subunit.

We have determined the DNA sequence of the distal 148 codons of trpE and all of trpG in Pseudomonas aeruginosa. These genes encode, respectively, the large and small (glutamine amidotransferase) subunits of anthranilate synthase, the first enzyme in the tryptophan synthetic pathway. The sequenced region of trpE is homologous with the distal portion of E. coli and Bacillus subtilis trpE, whereas the trpG sequence is homologous to the glutamine amidotransferase subunit genes of a number of bacterial and fungal anthranilate synthases. The two coding sequences overlap by 23 bp. Codon usage in these Pseudomonas genes shows a marked preference for codons ending in G or C, thereby resembling that of trpB, trpA, and several other chromosomal loci from this species and others with a high G + C content in their DNA. The deduced amino acid sequence for the P. aeruginosa trpG gene product differs to a surprising extent from the directly determined amino acid sequence of the glutamine amidotransferase subunit of P. putida anthranilate synthase (Kawamura et al. 1978). This suggests that these two proteins are encoded by loci that duplicated much earlier in the phylogeny of these organisms but have recently assumed the same function. We have also determined 490 bp of DNA sequence distal to trpG but have not ascertained the function of this segment, though it is rich in dyad symmetries.

Amino Acid Sequence↗

Cloning of the Zymomonas mobilis structural gene encoding alcohol dehydrogenase I (adhA): sequence comparison and expression in Escherichia coli.

Zymomonas mobilis ferments sugars to produce ethanol with two biochemically distinct isoenzymes of alcohol dehydrogenase. The adhA gene encoding alcohol dehydrogenase I has now been sequenced and compared with the adhB gene, which encodes the second isoenzyme. The deduced amino acid sequences for these gene products exhibited no apparent homology. Alcohol dehydrogenase I contained 337 amino acids, with a subunit molecular weight of 36,096. Based on comparisons of primary amino acid sequences, this enzyme belongs to the family of zinc alcohol dehydrogenases which have been described primarily in eucaryotes. Nearly all of the 22 strictly conserved amino acids in this group were also conserved in Z. mobilis alcohol dehydrogenase I. Alcohol dehydrogenase I is an abundant protein, although adhA lacked many of the features previously reported in four other highly expressed genes from Z. mobilis. Codon usage in adhA is not highly biased and includes many codons which were unused by pdc, adhB, gap, and pgk. The ribosomal binding region of adhA lacked the canonical Shine-Dalgarno sequence found in the other highly expressed genes from Z. mobilis. Although these features may facilitate the expression of high enzyme levels, they do not appear to be essential for the expression of Z. mobilis adhA.

Alcohol Dehydrogenase↗

Archaeal grpE: transcription in two different morphologic stages of Methanosarcina mazei and comparison with dnaK and dnaJ.

Transcription of the heat shock gene grpE was studied in two different morphologic stages of the archaeon Methanosarcina mazei S-6 that differ in resistance to physical and chemical traumas: single cells and packets. While single cells are directly exposed to environmental changes, such as temperature elevations, cells in packets are surrounded by intercellular and peripheral material that keeps them together in a globular structure which can reach several millimeters in diameter. grpE transcript levels determined by Northern (RNA) blotting peaked after a 15-min heat shock in single cells. In contrast, the highest transcript levels in packets were observed after the longest heat shock tested, 60 min. The same response profiles were demonstrated by primer extension experiments and S1 nuclease analysis. A comparison of the grpE response to heat shock with those of dnaK and dnaJ showed that the grpE transcript level was the most increased, closely followed by that of the dnaK transcript, with that of the dnaJ gene being the least augmented. Transcription of grpE started at the same site under normal and heat shock temperatures, and the transcript was consistently approximately 700 bases long. Codon usage patterns revealed that the three archaeal genes use most codons and have the same codon preference for 61% of the amino acids.

Bacterial Proteins↗

Assembly and comparative analysis of the mitochondrial genome of Pleione yunnanensis: genome structure and evolutionary insights.

BACKGROUND: Pleione yunnanensis a terrestrial or semi-epiphytic herbaceous plant belonging to the Orchidaceae family, is valued for both its medicinal uses and ornamental appeal. Although its chloroplast genomes have been sequenced, its complete mt genome had not previously been resolved, limiting genetic and evolutionary studies of the species. RESULTS: In this work, we assembled and characterized the first complete mt genome of P. yunnanensis, revealing a structurally complex, multibranched system composed of 14 circular-mapping molecules totaling 468,176 bp with a GC content of 44.32%. The genome encodes 44 annotated genes, including 28 protein-coding genes (PCGs), 15 tRNAs, and one rRNA. The multibranched architecture provides new evidence supporting the dynamic and recombinational nature of plant mt genomes. Repeat analysis uncovered 29 simple sequence repeats (SSRs), 19 tandem repeats, and 118 dispersed repeats, indicating a comparatively lower repeat abundance than that found in closely related orchids with similar mt genome sizes. Codon-usage profiling of PCGs showed a marked bias toward A/T-ending codons. Prediction of RNA editing sites identified 4,708 putative edits across mitochondrial PCGs. Most mitochondrial genes displayed Ka/Ks ratios close to 1.0, suggesting relaxed selective constraints or lineage-specific evolutionary patterns rather than strong positive selection. Moreover, we detected 69 chloroplast-derived homologous fragments, including 15 intact genes, suggesting ongoing plastid-mitochondrial DNA transfer. Phylogenetic reconstruction and collinearity comparisons demonstrated that P. yunnanensis clustered closely with Dendrobium species, including D. amplum and D. hancockii, within the Orchidaceae clade. CONCLUSIONS: This study provides the first complete mt genome of P. yunnanensis, providing a foundational genomic resource for the genus Pleione. The results not only improve our understanding of mt genome structure and evolution in Orchidaceae, but also offer valuable molecular evidence for phylogenetic inference, germplasm identification, and conservation of this endangered medicinal species.

Orchidaceae↗

Heuristic approach to deriving models for gene finding.

Computer methods of accurate gene finding in DNA sequences require models of protein coding and non-coding regions derived either from experimentally validated training sets or from large amounts of anonymous DNA sequence. Here we propose a new, heuristic method producing fairly accurate inhomogeneous Markov models of protein coding regions. The new method needs such a small amount of DNA sequence data that the model can be built 'on the fly' by a web server for any DNA sequence >400 nt. Tests on 10 complete bacterial genomes performed with the GeneMark.hmm program demonstrated the ability of the new models to detect 93.1% of annotated genes on average, while models built by traditional training predict an average of 93.9% of genes. Models built by the heuristic approach could be used to find genes in small fragments of anonymous prokaryotic genomes and in genomes of organelles, viruses, phages and plasmids, as well as in highly inhomogeneous genomes where adjustment of models to local DNA composition is needed. The heuristic method also gives an insight into the mechanism of codon usage pattern evolution.

Codon↗

Increased levels of glycine tRNA associated with collagen synthesis.

Analysis of codon usage for chick Type I collagen indicates that 89% of glycine codons are GGU/C. Since collagens are one-third glycine, chick Type I collagen synthesis should require large amounts of tRNAGly with the anticodon GCC. Earlier chromatographic studies of chick tRNA had indicated that connective tissues showed altered tRNAGly isoacceptor profiles [P. J. Christner and J. Rosenbloom (1976) Arch. Biochem. Biophys. 172, 399-409; H. J. Drabkin and L. N. Lukens (1978) J. Biol. Chem. 253, 6233-6241]. We have therefore used both two-dimensional gel electrophoresis and hybridization analysis to investigate whether collagen synthesis in chick connective tissues is associated with expression of a novel tRNAGly. Liver and calvaria tRNAs produced qualitatively similar patterns when separated on 2-D gels. Northern blots of 2-D-separated tRNAs from liver and calvaria, when hybridized to genes for vertebrate tRNAGly isoacceptors with GCC or UCC anticodons, showed hybridization to the same tRNAs in both tissues. Quantitation of tRNA species by dot blot hybridization indicated an increase in levels of the tRNAGly isoacceptor with anticodon GCC. Tissues synthesizing Type I collagen had a two- to threefold increase in this tRNA while tissues synthesizing Type II collagen showed a more modest increase. We conclude that elevated tRNAGly levels associated with collagen synthesis are due to increased amounts of the same isoacceptor which is the major tRNAGly in other tissues.

Animals↗

Characterization of two divergent beta-tubulin genes from Colletotrichum graminicola.

We have cloned and sequenced two beta-tubulin genes, TUB1 and TUB2, from the phytopathogenic fungus, Colletotrichum graminicola. The nucleotide sequences of the coding regions of the two genes are only 72.8% homologous. This divergence is reflected in the deduced amino acid (aa) sequences which differ at 94 aa residues. Comparison with the aa sequences of other fungal beta-tubulins indicates that the C. graminicola TUB2 gene encodes a conserved isotype, whereas the C. graminicola TUB1 product is highly divergent. Both genes contain six identically placed introns and the position of each intron is conserved in other fungal beta-tubulin genes. Also typical of other fungal beta-tubulin genes, there is a pronounced bias in codon usage in the C. graminicola TUB2 gene; there is a lesser codon bias in TUB1 from C. graminicola. Both C. graminicola beta-tubulin genes are transcribed and yield similar sized messages.

Amino Acid Sequence↗

Cloning and sequence of several alpha 2u-globulin cDNAs.

We describe a simple cloning procedure for alpha 2u-globulin that requires neither enrichment of mRNA for cloning nor purification of a specific probe for screening recombinant colonies. Total adult male liver poly(A)+RNA was used as template for cloning, and the subsequent recombinant colonies were screened by comparing hybridization to radioactive cDNA probes prepared from hepatic male and female mRNA, respectively. Almost all of the selected "male-specific" clones were later shown to contain alpha 2u-globulin sequences. This cloned alpha 2u-globulin cDNA has been shown to specifically hybridize to male rat liver RNA, which, when isolated and translated in vitro, codes for a 21,000-dalton protein (pro-alpha 2u-globulin) immunologically identical to alpha 2u-globulin. When translation occurs in the presence of pancreatic microsomes this in vitro synthesized pro-alpha 2u-globulin is processed to the 19,000-dalton mature form of alpha 2u-globulin. The nucleotide sequence of the alpha 2u-globulin cDNA has been determined, thus elucidating the complete amino acid sequence of alpha 2u-globulin and most of the hydrophobic "leader" sequence of pro-alpha 2u-globulin. The amino acid sequence deduced from the cDNA is in agreement with the partial sequence that we previously determined by sequential Edman degradation of the purified protein. alpha 2u-Globulin cDNA clones contain within the 3'-untranslated region one or both of the two putative polyadenylylation/transcription termination sites (A-A-T-A-A-A and A-A-T-T-A-A-A). Either of these can be used, generating alpha 2u-globulin mRNA species of two lengths. A codon usage analysis of the cDNA showed that, although all six leucine codons are used for the 14 leucine residues in mature alpha 2u-globulin, the seven leucines in the partial leader sequence reported are all encoded by the same codon, CTG. The primary amino acid sequence contains a unique Asn-Gly-Ser sequence, likely to be in beta-turn conformation, as the probable site of glycosylation for this glycoprotein.

Alpha-Globulins↗