Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Complete nucleotide sequence of Tn1721: gene organization and a novel gene product with features of a chemotaxis protein.

The complete 11,139-nucleotide sequence of transposon Tn1721 has been determined. It contains three 38-bp inverted repeats, and (in this order) a new orfI, a resolution site (res), genes encoding resolvase (tnpR), transposase (tnpA), tetracycline-resistance (TcR) repressor (tetR), TcR (tetA) and a truncated transposase gene (tnpA'). The modulator origin of Tn1721 from at least three separate sources is supported by the distinctive codon usages of orfI, tnpR/tnpA and tetR/tetA, and by sequence similarities with Tn501 (tnpR/tnpA) and RP1 (tetR/tetA). The ORFI-encoded 56-kDa polypeptide exhibits features of a methyl-accepting chemotaxis protein (MCP) with a conserved signal domain and a potential transmembrane domain; this polypeptide cross-reacts with anti-MCP antiserum. Like chemotaxis genes, orfI is transcribed from a sigma 28-like promoter. The overexpressed orfI gene product interferes with MCP-dependent chemotaxis suggesting that it completes for soluble transducer protein(s) in the cell. The potential selective advantage of this novel transposon-borne gene is discussed.

Amino Acid Sequence↗

Overexpression of the Thermus aquaticus B malate dehydrogenase-encoding gene in Escherichia coli.

Expression of the Thermus aquaticus B malate dehydrogenase (MDH)-encoding gene (mdh), cloned in Escherichia coli, was initially at a relatively low level (0.1% of soluble cell protein) and was effected by read-through from the tac promoter in the plasmid vector used. An enhancement in expression to 0.4% of soluble cell protein was achieved by shortening the intervening sequence between the promoter and the translation start codon of mdh. An NdeI restriction site (5'-CAT-ATG-3') was engineered in the shortened fragment, which also changed the start codon from GTG to ATG. This resulted in an eightfold increase in expression, to 3.2% of soluble cell protein. Expression was further increased by subcloning the mdh gene via the engineered NdeI site, into two plasmid expression vectors, one carrying the E. coli trpP promoter and the other the E. coli mdhP promoter. In both these expression systems, 40-50% of the soluble cell protein was T. aquaticus MDH. This suggests that expression of the cloned T. aquaticus mdh in E. coli is enhanced predominantly by the optimisation of transcription and translation initiation signals. Moreover, the base composition of the coding region and the pattern of codon usage dictated by it appear to have little effect on expression. Heat treatment of the cell extract at 85 degrees C further effected purification of T. aquaticus MDH to over 80% of the soluble cell protein. The MDHs purified to homogeneity from the high-expression clones were identical with the MDH isolated from T. aquaticus B cells with respect to all measured parameters.

Base Sequence↗

Organization and sequence of the gene encoding the human acrosin-trypsin inhibitor (HUSI-II).

A complete cDNA encoding the acrosin-trypsin inhibitor, HUSI-II, was used as a probe to isolate genomic clones from a human placenta library. Three clones which cover the entire HUSI-II gene were isolated and characterized. The exon-intron organization of the gene was determined and found to be identical to other known Kazal-type inhibitor-encoding genes. The striking similarity in the amino acid sequences which was found previously in HUSI-II and glycoprotein hormone beta-subunits, is neither reflected in codon usage nor in the exon-intron arrangement of the genes. A 1.8-kb segment 5' of the gene was sequenced. The analysis of this sequence showed that HUSI-II contains a G + C-rich region upstream from the transcription start point (tsp) which fulfills the criteria for a CpG island. Furthermore, in the first intron, a potential glucocorticoid-responsive element was found as a half-palindrome flanked by two CACCC elements. Determination of the tsp by S1 mapping revealed that HUSI-II has multiple tsp. Genomic Southern hybridization was used to show that HUSI-II is a single-copy gene. The localization of the gene to chromosome 4 was determined by hybridization of a 5' genomic fragment to the DNA of a panel of somatic hybrids between human and rodent cells.

Acrosin↗

Cloning, characterization and functional expression of an endoglucanase-encoding gene from the phytopathogenic fungus Macrophomina phaseolina.

An endoglucanase-encoding clone (egl2) was isolated from the phytopathogenic soilborne deuteromycete fungus Macrophomina phaseolina (Mp). Clones were obtained from a cDNA library by functional expression in Escherichia coli. The egl2 clone hybridized to a 1.3-kb mRNA. Expression is induced by carboxymethylcellulose (CMC) and repressed by glucose. The deduced amino acid (aa) sequence revealed strong similarity to the egl3 from Trichoderma reesei (Tr) (72% for identical residues and 81% with conservative substitution over a span of 324 aa). The Mp egl2 lacks the cellulose-binding domain and linker region found in the Tr egl3. Different codon usage between the two fungi resulted in a much shorter span of nucleotide homology. The Egl2 protein cleaves cellodextrins with continguous beta, 1-4 linkages of four and larger, and shows activity against CMC and birchwood xylan.

Amino Acid Sequence↗

A human pancreatic secretory trypsin inhibitor presenting a hypervariable highly constrained epitope via monovalent phagemid display.

Hypervariable gene banks displaying ligands which can be used for affinity optimisation are valuable resources for examining shape space. They have added value if the ligand is small, if there is extensive information on its tertiary structure and if the variable region is highly constrained. These features would be expected to stabilise complexes by reducing the dissociation constants and to facilitate their use as 'lead substances' for the development of synthetic mimetics. The synthesis and characterisation of such phagemid-display banks is described here, in which the variable region is a 7-amino acid (aa) (pSKAN8-HyB/C) or 8-aa (pSKAN8-HyA) extended peptide held between two disulfide bridges at the exposed tip of the human pancreatic secretory trypsin inhibitor (PSTI). A phagemid pSKAN8 was created which contains a fusion between the PSTI and M13 pIII protein-coding genes. Cassettes containing the sequences (NNK)8 [HyA], (NNK)7 [HyB] or (NNK)6GTT [Hy-C] (where K = G or T) were used to randomize the aa coding region in the trypsin-inhibitory loop (aa 17 to 23) of PSTI. Some 31 million individual clones were generated in a mutS Escherichia coli strain kept as frozen cell stocks. Analysis of controls which had not undergone selection showed very low levels of deletion. The quality of the hypervariable region and bias of codon usage was quantified by DNA sequencing. It was estimated from SDS-PAGE that hybrid protein was represented statistically at a frequency of one molecule per two phagemid particles. The functionality and reproducibility of the system was demonstrated by trypsin-binding of the original vector and in selecting novel chymotrypsin inhibitors from the banks.

Amino Acid Sequence↗

Synthesis of a modified gene encoding human ornithine transcarbamylase for expression in mammalian mitochondrial and universal translation systems: a novel approach towards correction of a genetic defect.

The mitochondrial (MT) genome is a potential means of gene delivery to human cells for therapeutic expression. As a first step towards this, we have synthesized a gene coding for mature human ornithine transcarbamylase (OTC) by recursive PCR using 18 oligodeoxyribonucleotides, each 70-80 nucleotides in length, using codons which should allow translation in accordance with both mammalian mt and universal codon usage. Flanking mt DNA sequences were incorporated which are designed to facilitate site-specific cloning into the mt genome. Expression of this human gene in Escherichia coli leads to an immunoreactive OTC product of the correct size and N-terminal amino-acid sequence, but which forms inclusion bodies and lacks enzymatic activity.

Amino Acid Sequence↗

Actin-encoding cDNAs and gene expression during the intermolt cycle of the Bermuda land crab Gecarcinus lateralis.

Two actin-encoding cDNAs (act1 and act2) from Gecarcinus lateralis have been sequenced or partially sequenced and the corresponding proteins deduced. The act1 cDNA has a complete ORF; the act2 cDNA lacks most of the 5' end of the coding region. The nucleotide (nt) sequences of both clones are very similar to act sequences of many organisms, the most closely related being from another arthropod, the silkmoth Bombyx mori. The proteins Act1 and Act2 are more similar to vertebrate cytoplasmic actin isoforms (beta-actins) than to vertebrate muscle actins (alpha-actins); they are also more similar to animal actins than to those of fungi or plants. Codon usage is strongly biased toward C or G in the third position. The deduced number of amino acid (aa) residues and calculated Mr for Act1 are 376 aa and 41.94 kDa, respectively. The deduced aa sequence of Act1 is very similar to those of muscle actins of B. mori and Drosophila melanogaster. Southern blots indicated seven to eleven act genes in the crab genome. Northern blots probed with a segment from the 3' UTR of act1 showed a single band of approx. 1.6 kb in poly(A)+ mRNAs from epidermis, limb bud or claw muscle and in total RNAs from ovary and gill, and two bands of approx. 1.6 and 1.8 kb in total RNA from midgut gland. Western blots of one-dimensional gels of proteins from the four layers of the exoskeleton, epidermis, limb buds and claw muscle were probed with a monoclonal Ab against chicken gizzard actin; tissue- and stage-specific changes in actin content were observed. The presence of several isoforms, and differences in their number and occurrence at various stages of the intermolt cycle, were detected on Western blots of two-dimensional gels.

Actins↗

Cloning, sequencing and analysis of the ggh-A gene encoding a 1,4-beta-D-glucan glucohydrolase from Microbispora bispora.

The ggh-A gene, encoding a 1,4-beta-D-glucan glucohydrolase/beta-glucosidase, of Microbispora bispora (Mb) was subcloned and expressed from a 4.0-kb XhoI DNA fragment. The nucleotide sequence of this fragment was determined. Analysis of the sequence revealed one open reading frame (ORF) which encodes a 986-amino-acid (aa) protein with a calculated molecular weight of 107,510. The ggh-A ORF has features typical of an actinomycete gene including high GC content (70.5%) and corresponding biased codon usage. Comparison of the aa sequence of the Mb 1,4-beta-D-glucan glucohydrolase (Mbggh-A) with other glycosidases reveals high overall homology to several beta-glucosidases and a 1,4-beta-D-glucan glucohydrolase belonging to the glycosyl hydrolase family 3. The aa sequence alignments of Mbggh-A and beta-glucosidases show that the active site region potentially involves two Asp residues. The aa sequence homology studies revealed a potential two-domain structure for Mbggh-A and other beta-glucosidases. Furthermore, Mbggh-A has localized homology to a cellulose-binding domain present in some xylanases. This report is significant, as, to date, 1,4-beta-D-glucan glucohydrolases have rarely been reported, though they are assumed to have a critical role in cellulolysis.

Actinomycetales↗

The restriction-modification system of Pasteurella haemolytica is a member of a new family of type I enzymes.

Genes encoding the type I restriction-modification (R-M) system of the bovine pathogen, Pasteurella haemolytica, have been identified immediately downstream of a locus that encodes a transcriptional activator of P. haemolytica leukotoxin expression. Type I enzymes are encoded by three genes called hsdM, hsdS and hsdR, and have fallen into three groups, called Ia, Ib and Ic. HsdS provides a sequence recognition function which in concert with HsdM forms an active methyltransferase (MTase). Inclusion of the HsdR subunit in the complex creates an active restriction endonuclease (ENase) capable of cleaving unmethylated target DNA. The P. haemolytica hsdMSR genes were mapped using transposon Tn10d-Cam insertions, and bacteriophage restriction and modification assays in Escherichia coli. We determined the nucleotide sequences of hsdM, hsdS and hsdR, and observed that the deduced amino acid (aa) sequences were very similar to predicted R-M subunits in the respiratory pathogen, Haemophilus influenzae. Phylogenetic comparisons of all known Hsd aa sequences placed the P. haemolytica and H. influenzae proteins into a new group which we labeled the Type Id R-M family. Expression of the P. haemolytica R-M genes in E. coli was inefficient and is likely to be a consequence of the unusual codon usage in P. haemolytica genes.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of a Rickettsia tsutsugamushi 22 kDa antigen containing B- and T-cell epitopes.

The identification of Rickettsia tsutsugamushi T-cell epitopes is necessary for the characterization of the protective immune response (of which the T-cell response is essential) against scrub typhus rickettsiae. A T-helper cell line derived from R. tsutsugamushi (Karp strain) immune mice reacted with rickettsial protein antigens eluted from the 18-35 kDa region of polyacrylamide gels. Within this region is a 22 kDa protein which is reactive with immune serum. The gene encoding the 22 kDa scrub typhus antigen (sta22) was cloned and expressed in Escherichia coli. Nucleotide sequence analysis of the sta22 gene revealed a potential open reading frame (ORF) in the sta22 sequence encoding a 22 kDa protein. A recombinant 22 kDa protein synthesized in E. coli maxicells was reactive with anti-rickettsial antibodies. The codon usage of the adenine and thymine rich sta22 sequence was similar to other previously sequenced R. tsutsugamushi genes. Computer analysis of the deduced amino acid sequence suggested that the Sta22 protein has several amphipathic regions which may be potential T-cell epitopes. The recombinant Sta22 protein eluted from polyacrylamide gels induced a strong proliferative response from the scrub typhus rickettsiae reactive T-cell line. Recognition of the R. tsutsugamushi Sta22 polypeptide by both cellular and humoral immune mechanisms implicates this antigen as one of potential importance in vaccine development.

Amino Acid Sequence↗

Human arylsulfatase B: MOPAC cloning, nucleotide sequence of a full-length cDNA, and regions of amino acid identity with arylsulfatases A and C.

cDNAs encoding the human lysosomal hydrolase, arylsulfatase B (ASB; N-acetylgalactosamine-4-sulfatase, EC 3.1.6.1), were isolated from a hepatoma cell cDNA library using an ASB-specific oligonucleotide generated by the MOPAC (mixed oligonucleotide primed amplification of cDNA) technique. To facilitate cDNA cloning, human ASB was purified to apparent homogeneity and a total of 112 amino acid residues were microsequenced from the N-terminus and four internal tryptic peptides of the 47-kDa subunit. Based on the ASB N-terminal amino acid sequence, two oligonucleotide mixtures containing inosines to reduce the mixture complexity were constructed and used as primers to amplify an ASB-specific product from human placental cDNA by the polymerase chain reaction. DNA sequencing of this MOPAC product demonstrated colinearity with 21 N-terminal ASB amino acids. Based on this sequence and on codon usage for the adjacent conserved amino acids in human arylsulfatases A and C, a unique 66-mer was synthesized and used to screen a human hepatoma cell cDNA library. Four putative positive cDNA clones were isolated, and the largest insert (pASB-1) was sequenced in both orientations. The 1834-bp pASB-1 insert had a 1278-bp open reading frame encoding 425 amino acids that was colinear with 85 microsequenced amino acids of the purified enzyme, demonstrating its authenticity. Using the pASB-1 cDNA as a probe, a full-length cDNA clone, pASB-4, was isolated from a human testes library and sequenced in both orientations. pASB-4 had a 2811-bp insert containing a 559-bp 5' untranslated sequence, a 1602-bp open reading frame encoding 533 amino acids (six potential N-glycosylation sites), a 641-bp 3' untranslated sequence, and a 9-bp poly(A) tract. Comparison of the predicted amino acid sequences of arylsulfatases A, B, and C revealed regions of identity, particularly in their N-termini.

Amino Acid Sequence↗

A transcribed gene in an intron of the human factor VIII gene.

We have identified a CpG island contained within the largest factor VIII intron. This island is associated with a 1.8-kb transcript and, unlike factor VIII, is produced abundantly in a wide variety of cell types. The nested gene is oriented in a direction opposite to that of factor VIII and contains no intervening sequences. A cDNA of 1739 bases was isolated from a human liver library and found to have a GC-rich, long open reading frame. Two computer-assisted methods (Fickett TESTCODE and Staden-McLachlan codon usage) predict that the gene codes for a protein. Two other copies of this gene are located within 1.1 Mb of the factor VIII gene. Northern blot analysis of RNA isolated from hemophilia patients deleted for factor VIII sequences has shown that both the intron gene and at least one other copy of the gene are transcribed. A homologous, transcribed sequence is also present in mice.

Amino Acid Sequence↗

Novel muteins of human tumor necrosis factor alpha.

For chemical synthesis of a gene coding for human tumor necrosis factor alpha (TNF-alpha), DNA sequence predicted by the amino acid sequence of human TNF molecule was prepared. Codons were chosen according to the codon usage in Escherichia coli (E. coli). The 490 bp gene was assembled by enzymic ligation of 42 oligonucleotides and was cloned into a vector (pKK223-3) for high expression of active TNF-alpha in E. coli. With use of site-directed mutagenesis on this DNA, five different muteins of TNF-alpha were synthesized. TNF-M1 and TNF-M4 have deletions of His-73 and Gln-102, respectively. These deletions didn't cause loss of the cytotoxic activity against L929 cells. TNF-M5, which has a substitution of Asp-10 to Arg, had the similar cytotoxic activity to that of TNF-alpha. The cytotoxic spectra against several tumor cells were not changed by this substitution. TNF-M3 has an amino acid substitution of Glu-116 to His which occupies this position in human TNF-beta. This substitution didn't change the cytotoxicity. In addition, evidence was presented that the change of the carboxyl terminal residue doesn't always influence the cytotoxic activity of TNF-alpha. Many different muteins were also isolated by random mutagenesis with hydroxylamine-HCl. One of the muteins, which carries a mutation of His-15 to Tyr, lost the cytotoxic activity almost completely.

Amino Acid Sequence↗

The actin gene family in the oriental fruit fly Bactrocera dorsalis. Muscle specific actins.

The actin protein is a critical protein in eukaryotic cells. Four actin genes, constituting what appear to be a set of muscle specific actin genes, have been isolated from the genome of the oriental fruit fly Bactrocera dorsalis. DNA sequences have been determined for the coding as well as 3' and 5' flanking regions for each of these genes. These genes have also been characterized in terms of RNA expression patterns, and comparisons have been made to actin genes from other species. Consistent with other actins, there is a high degree of amino acid sequence conservation in the coding regions of these genes. However, even within the coding regions codon usage patterns in the oriental fruit fly are quite different from some other well characterized species. In addition, the DNA sequences in the intermediate 3' and 5' flanking regions exhibit virtually no detectable sequence homology both within and between species. In terms of introns, three of the four actin genes from the oriental fruit fly described here have a single intervening sequence. Two of these genes share the same intron position with the two muscle specific actin genes act79B and act88F from Drosophila melanogaster and with one muscle specific actin gene CcA1 from the Mediterranean fruit fly, Ceratitis capitata. Another gene from the oriental fruit fly shares the same intron position as the muscle specific actin gene act57B from D. melanogaster. Such conservation of intron positioning between species is highly unusual among previously characterized actin genes. Using unique sequences found in the 3' untranslated regions, gene specific probes have also been constructed. These have been used to detect the expression patterns of individual genes in a temporal and spatial manner. Each of the four genes examined here show differential patterns of expression. The patterns indicate that all four genes are most likely to encode muscle specific actins.

Actins↗

Metazoan OXPHOS gene families: evolutionary forces at the level of mitochondrial and nuclear genomes.

Mitochondrial and nuclear DNAs contribute to encode the whole mitochondrial protein complement. The two genomes possess highly divergent features and properties, but the forces influencing their evolution, even if different, require strong coordination. The gene content of mitochondrial genome in all Metazoa is in a frozen state with only few exceptions and thus mitochondrial genome plasticity especially concerns some molecular features, i.e. base composition, codon usage, evolutionary rates. In contrast the high plasticity of nuclear genomes is particularly evident at the macroscopic level, since its redundancy represents the main feature able to introduce genetic material for evolutionary innovations. In this context, genes involved in oxidative phosphorylation (OXPHOS) represent a classical example of the different evolutionary behaviour of mitochondrial and nuclear genomes. The simple DNA sequence of Cytochrome c oxidase I (encoded by the mitochondrial genome) seems to be able to distinguish intra- and inter-species relations between organisms (DNA Barcode). Some OXPHOS subunits (cytochrome c, subunit c of ATP synthase and MLRQ) are encoded by several nuclear duplicated genes which still represent the trace of an ancient segmental/genome duplication event at the origin of vertebrates.

Animals↗

Relation between mRNA expression and sequence information in Desulfovibrio vulgaris: combinatorial contributions of upstream regulatory motifs and coding sequence features to variations in mRNA abundance.

The context-dependent expression of genes is the core for biological activities, and significant attention has been given to identification of various factors contributing to gene expression at genomic scale. However, so far this type of analysis has been focused either on relation between mRNA expression and non-coding sequence features such as upstream regulatory motifs or on correlation between mRNA abundance and non-random features in coding sequences (e.g., codon usage and amino acid usage). In this study multiple regression analyses of the mRNA abundance and all sequence information in Desulfovibrio vulgaris were performed, with the goal to investigate how much coding and non-coding sequence features contribute to the variations in mRNA expression, and in what manner they act together. Using the AlignACE program, 442 over-represented motifs were identified from the upstream 100bp region of 293 genes located in the known regulons. Regression of mRNA expression data against the measures of coding and non-coding sequence features indicated that 54.1% of the variations in mRNA abundance can be explained by the presence of upstream motifs, while coding sequences alone contribute to 29.7% of the variations in mRNA abundance. Interestingly, most of contribution from coding sequences is overlapping with that from upstream motifs; thereby a total of 60.3% of the variations in mRNA abundance can be explained when coding and non-coding information was included. This result demonstrates that upstream regulatory motifs and coding sequence information contribute to the overall mRNA expression in a combinatorial rather than an additive manner.

Base Sequence↗

Functional insights from the molecular modelling of a novel two-component system.

Two-component systems (TCSs) are the major signalling pathway in bacteria and represent potential drug targets. Among the 11 paired TCS proteins present in Mycobacterium tuberculosis H37Rv, the histidine kinases (HKs) Rv0600c (HK1) and Rv0601c (HK2) are annotated to phosphorylate one response regulator (RR) Rv0602c (TcrA). We wanted to establish the sequence-structure-function relationship to elucidate the mechanism of phosphotransfer using in silico methods. Sequence alignments and codon usage analysis showed that the two domains encoded by a single gene in homologous HKs have been separated into individual open-reading frames in M. tuberculosis. This is the first example where two incomplete HKs are involved in phosphorylating a single RR. The model shows that HK2 is a unique histidine phosphotransfer (HPt)-mono-domain protein, not found as lone protein in other bacteria. The secondary structure of HKs was confirmed using "far-UV" circular dichroism study of purified proteins. We propose that HK1 phosphorylates HK2 at the conserved H131 and the phosphoryl group is then transferred to D73 of TcrA.

Amino Acid Sequence↗

Phylogenetic analyses of penicillia based on partial calmodulin gene sequences.

Partial sequences (about 600 nucleotides) of the calmodulin gene were used for the phylogenetic studies on Eupenicillium, Talaromyces and Penicillium. This region is from the 3rd base of the codon for the 9th amino acid Gln to the 3rd base of the codon for the 122th amino acid Val, flanking parts of the 2nd and 5th exons with complete sequences of two exons and three introns. Seventy-six isolates of 56 taxa of penicillia were involved. The nucleotide sequences with and without introns were analyzed respectively using the neighbor-joining (NJ) and maximum parsimony (MP) methods. The cluster analysis on relative synonymous codon usage (RSCU) of each sequence was also carried out. The fact that species of penicillia belong to the two subfamilies of the Trichocomaceae proposed by Malloch based on traditional methods is supported by our molecular data, whereas, the development of asci and patterns of penicilli show little phylogenetic information. Nine groups in the lineage of Eupenicillium and two in that of Talaromyces were recognized in our studies. In addition to the teleomorph-holomorph-anamorph evolutionary model of penicillia suggested by LoBuglio et al., and Pitt, we proposed that a mutation bias of holomorphs/anamorphs with or without selection is another evolutionary path of these organisms.

Base Sequence↗