Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Distinct patterns of evolution between respiratory syncytial virus subgroups A and B from New Zealand isolates collected over thirty-seven years.

Respiratory syncytial virus (RSV) is the most important cause of viral lower respiratory tract infections in infants and children worldwide. In New Zealand, infants with RSV disease are hospitalized at a higher rate than other industrialized countries, without a proportionate increase in known risk factors. The molecular epidemiology of RSV in New Zealand has never been described. Therefore, we analyzed viral attachment glycoprotein (G) gene sequences from 106 RSV subgroup A isolates collected in New Zealand between 1967 and 2003, and 38 subgroup B viruses collected between 1984 and 2004. Subgroup A and B sequences were aligned separately, and compared to sequences of viruses isolated from other countries during a similar period. Genotyping and clustering analyses showed RSV in New Zealand is similar and temporally related to viruses found in other countries. By quantifying temporal clustering, we found subgroup B viruses clustered more strongly than subgroup A viruses. RSV B sequences displayed more variability in stop codon usage and predicted protein length, and had a higher degree of predicted O-glycosylation site changes than RSV A. The mutation rate calculated for the RSV B G gene was significantly higher than for RSV A. Together, these data reveal that RSV subgroups exhibit different patterns of evolution, with subgroup B viruses evolving faster than A.

Biological Evolution↗

Vitellogenin motifs conserved in nematodes and vertebrates.

Caenorhabditis elegans vitellogenins are encoded by a family of six genes, one of which, vit-5, has been previously sequenced and shown to be surprisingly closely related to the vertebrate vitellogenin genes. Here we report an alignment of the amino acid sequences of vitellogenins from frog and chicken with those from three C. elegans genes: vit-5 and two newly sequenced genes, vit-2 and vit-6. The four introns of vit-6 are all in different places from the four introns of vit-5, but three of these eight positions are identical or close to intron locations in the vertebrate vitellogenin genes. The encoded polypeptides have diverged from one another sufficiently to allow us to draw some conclusions about conserved positions. Many cysteine residues have been conserved, suggesting that vitellogenin structure has been maintained over a long evolutionary distance and is dependent upon disulfide bonds. In addition, a 20-residue segment shows conservation between the vertebrate and the nematode vitellogenins. This sequence may play a highly conserved role in vitellogenesis, such as specific recognition by oocytes. On the whole, however, selection may be acting more strongly on amino acid composition and codon usage than on amino acid sequence, as might be expected for abundant storage proteins: The amino acid compositions of vit-2, vit-5, and vit-6 products are remarkably similar, despite the fact that the sequence of the vit-2 protein is only 22% and 50% identical to the sequences of vit-6 and vit-5 proteins, respectively.

Amino Acid Sequence↗

DNA sequences of yeast H3 and H4 histone genes from two non-allelic gene sets encode identical H3 and H4 proteins.

The complete DNA sequences of two loci encoding H3 and H4 histones in Saccharomyces cerevisiae have been determined. Each locus contains one H3 and one H4 gene. The genes at each locus are divergently transcribed and the coding sequences are separated by 646 base-pairs at one locus and 676 base-pairs at the other. The H3 genes code for identical histone H3 proteins and the H4 genes code for identical histone H4 proteins. The yeast proteins differ from histones H3 and H4 of calf by 15 and 8 amino acid substitutions, respectively, and these differences are largely confined to the carboxy-terminal halves of the proteins. The genes demonstrate a bias in synonymous codon usage similar to that noted for other yeast genes. This bias is confined to the coding sequences of the genes and is specific for the reading frame encoding the proteins. The coding sequence of each gene is flanked on both sides by DNA with an A + T content of 70 to 80%. Possible regulatory sequences are located relative to the 5' and 3'-termini of the histone H3 and H4 RNA transcripts.

Base Sequence↗

Site dependent time optimization of protein synthesis with special regard to accuracy.

The efficiency of protein synthesis is determined by its rate, accuracy, and energy consumption. With the energy consumption fixed, we optimize the system with respect to time and accuracy. Using an analytic model for a simple system and computer simulations for more complex systems, where also the possibility of errors is included, we demonstrate how different parts of the messenger RNA influence the protein production rate differently. The first part of the coding sequence is of major importance, since the availability of empty initiation sites is crucial, and queuing back to that region may interfere with initiation. The elongation rate at different positions depends on codon usage, on the concentrations of substrate and co-factors, and on the kinetic rate constants, including those of the proofreading branch(es). Ribosomal proofreading is a time consuming process and by allowing for more errors in the beginning of a protein, it is possible to increase the production rate of that protein. We calculate the mean translation time per functioning protein for various translation accuracies, and discuss the different strategies open to living cells.

Algorithms↗

On-line tools for sequence retrieval and multivariate statistics in molecular biology.

We have developed a World-Wide Web server for browsing sequence collections structured under the ACNUC format and for performing multivariate analyses on sequences. General collections (like GenBank or EMBL), as well as specialized data banks (like Hovergen and NRSub) can be accessed. This system allows complex queries to be constructed, and the result of each query, represented by a list of sequences, is stored on the server. It is then possible to reuse this list to compute multivariate analyses on the sequences. Two examples of applications are shown. The first one consists in a study of codon usage with correspondence analysis on all the protein genes of Haemophilus influenzae Rd. This study allows the highly expressed genes and the integral membrane proteins of this organism to be identified. The second one consists in an ordering of 70 aligned protein sequences of growth hormone with principal coordinate analysis. With this method, we are able to re-establish the patterns of relationships between the sequences previously determined with tree building programs.

Algorithms↗

Putidaredoxin reductase and putidaredoxin. Cloning, sequence determination, and heterologous expression of the proteins.

The oxidation of camphor by cytochrome P-450cam requires the participation of a flavoprotein, putidaredoxin reductase, and an iron-sulfur protein, putidaredoxin, to mediate the transfer of electrons from NADH to P-450 for oxygen activation. A 2.2-kilobase pair BamHI-StuI fragment from whole cell DNA of camphor-grown Pseudomonas putida has been cloned and sequenced. Translation of the sequence revealed two open reading frames that could code for putidaredoxin reductase and putidaredoxin. In the case of putidaredoxin, the translated sequence matched the published sequence (Tanaka, M., Haniu, M., Yasunobu, K. T., Dus, K., and Gunsalus, I. C. (1974) J. Biol. Chem. 249, 3689-3701) with the exception of one amino acid. Codon usage in these proteins, like the proteins of other Pseudomonads, is strongly biased to G + C in the third nucleotide. A potential transcription termination site was found 3' to the putidaredoxin coding region. The "FAD-binding" amino acid consensus sequence, present in other flavoproteins, was found in putidaredoxin reductase beginning at residue 11 and a second occurrence of this sequence was found beginning with amino acid 156. The second sequence could represent the NAD-binding site. The regions encoding putidaredoxin reductase and putidaredoxin were subcloned and independently expressed in Escherichia coli at the level of 0.4 and 4.8 mg of enzymatically active protein/g wet weight of cells, respectively. Site-directed mutagenesis was used to change the rare start codon, GTG, of putidaredoxin reductase to ATG which resulted in an 18-fold increase in the level of expression of this protein to 7.4 mg/g wet weight of cells. The construction of these two clones, which express these important proteins, will facilitate studies of their interaction with each other and with P-450cam.

Amino Acid Sequence↗

Compositional heterogeneity of the Escherichia coli genome: a role for VSP repair?

E. coli genes that contain a high frequency of the tetranucleotide CTAG are also rich in the tetramers CTTG, CCTA, CCAA, TTGG, TAGG, and CAAG (group-I tetramers). Conversely, E. coli genes lacking CTAG are rich in the tetranucleotides CCTG, CCAG, CTGG, and CAGG (group-II tetramers). These two gene samples differ also in codon usage, amino acid composition, frequency of Dcm sites, and contrast vocabularies. Group-I tetramers have in common that they are depleted by very-short-patch repair (VSP), while group-II tetramers are favored by VSP activity. The VSP system repairs G:T mismatches to G:C, thereby increasing the overall G+C content of the genome; for this reason the CTAG-rich sample has a lower G+C content than the CTAG-poor sample. This compositional heterogeneity can be tentatively explained by a low level of VSP activity on the CTAG-rich sample. A negative correlation is found between the frequency of group-I tetramers and the level of gene expression, as measured by the Codon Adaptation Index (CAI). A possible link between the rate of VSP activity and the level of gene expression is considered.

Base Sequence↗

Isolation and sequencing of a new beta-galactosidase-encoding archaebacterial gene.

The gene lacS coding for a beta-galactosidase (beta Gal; EC 3.2.1.23) has been cloned from the thermoacidophilic archaebacterium Sulfolobus solfataricus, strain MT-4. It encodes a polypeptide chain of 489 amino acids (aa) (56,764 Da) in good agreement with the value directly measured for the enzyme (60 +/- 2 kDa per subunit). The aa composition of the enzyme and, in particular, its peculiarly low cysteine content (one Cys per subunit) has been confirmed; at the same time, it has been observed that the very low G + C content of the S. solfataricus genome strongly influences the codon usage preferences in the lacS sequence. There appears to be no evident similarity between this and the Escherichia coli lacZ sequence, thus suggesting that the two enzymes have analogous function, but are not homologous. By comparison with the published sequences of archaebacterial promoters, terminators and ribosome-binding sites, potential regulatory sites have been identified in the flanking regions of the S. solfataricus lacS gene.

Amino Acid Sequence↗

Four synonymous genes encode calmodulin in the teleost fish, medaka (Oryzias latipes): conservation of the multigene one-protein principle.

We cloned four distinct calmodulin (CaM)-encoding cDNAs from a small teleost fish, medaka (Oryzias latipes). The deduced amino acid (aa) sequences were exactly the same in these four genes and identical to the aa sequence of mammalian CaM, because of synonymous codon usages. The four cDNAs from medaka, termed CaM-A, -B, -C and -D, corresponded to mRNAs of 1.8, 1.4, 2.5 and 1.8 kb, respectively, in Northern blot analysis. Our results demonstrated that the 'multigene one-protein' principle of CaM synthesis is applicable to medaka, as well as to mammals whose CaM is encoded by at least three different genes.

Amino Acid Sequence↗

RNA secondary structure and compensatory evolution.

The classic concept of epistatic fitness interactions between genes has been extended to study interactions within gene regions, especially between nucleotides that are important in maintaining pre-mRNA/mRNA secondary structures. It is shown that the majority of linkage disequilibria found within the Drosophila Adh gene are likely to be caused by epistatic selection operating on RNA secondary structures. A recently proposed method of RNA secondary structure prediction based on DNA sequence comparisons is reviewed and applied to several types of RNAs, including tRNA, rRNA, and mRNA. The patterns of covariation in these RNAs are analyzed based on Kimura's compensatory evolution model. The results suggest that this model describes the substitution process in the pairing regions (helices) of RNA secondary structures well when the helices are evolutionarily conserved and thermodynamically stable, but fails in some other cases. Epistatic selection maintaining pre-mRNA/mRNA secondary structures is compared to weak selective forces that determine features such as base composition and synonymous codon usage. The relationships among these forces and their relative strengths are addressed. Finally, our mutagenesis experiments using the Drosophila Adh locus are reviewed. These experiments analyze long-range compensatory interactions between the 5' and 3' ends of Adh mRNA, the different constraints on secondary structures in introns and exons, and the possible role of secondary structures in RNA splicing.

Alcohol Dehydrogenase↗

Completion of the sequence of a cetacean morbillivirus and comparative analysis of the complete genome sequences of four morbilliviruses.

The gene encoding the large (L) protein and the genome termini of the dolphin strain of cetacean morbillivirus (CeMV) were sequenced. The CeMV genome is 15702 nucleotides long and has been compared with other available morbillivirus genome sequences in regards to the "rule of six" and the "phase" of any particular nucleotide, defined as its position within a given hexamer, which here is defined as a group of six nucleotides starting from the 3' end of the genomic RNA. With exception of the position of the start of the F gene, the phase of the transcription start sites of each gene is strictly conserved between the morbilliviruses, but each gene is in a different phase. The lengths of gene transcripts differ between viruses by multiples of six nucleotides with exception of the M and F transcripts. The differences between the various morbilliviruses result from deletions or insertions of multiples of six nucleotides in the 3' and 5' UTRs of the different viral genes. The four bases were distributed non-randomly over the six positions in the hexamer boxes. However, the distribution patterns of each of the four bases indicated that multiples of three were more prevalent than those of six nucleotides. This reflected the positions of nucleotides in codons and codon usage in the reading frames. The L protein of CeMV was found to be 2183 amino acids in length and similar to that of MV and RPV. The CeMV L protein sequence was found to be equidistant between those of the CDV/PDV and MV/RPV subgroups of the morbilliviruses. This concurs with the analyses carried out on the other structural proteins.

3' Untranslated Regions↗

The mitochondrial genome of the primary screwworm fly Cochliomyia hominivorax (Diptera: Calliphoridae).

The complete sequence of the mitochondrial genome of the screwworm Cochliomyia hominivorax was determined. This genome is 16,022 bp in size and corresponds to a typical Brachycera mtDNA. A Serine start codon for COI and incomplete termination codons for COII, NADH 5 and NADH 4 genes were described. The nucleotide composition of C. hominivorax mtDNA is 77% AT-rich, reflected in the predominance of AT-rich codons in protein-coding genes. Non-optimal codon usage was commonly observed in C. hominivorax mitochondrial genes. Phylogenetic analysis distributed the Acalypterate species as a monophyletic group and assembled the C. hominivorax (Calyptratae) and the Acalyptratae in a typical Brachycera cluster. The identification of diagnostic restriction sites on the sequenced mitochondrial genome and the correlation with previous RFLP analysis are discussed.

Animals↗

Isolation of cDNAs encoding 6-phosphogluconate dehydrogenase and glucose-6-phosphate dehydrogenase from the mediterranean fruit fly Ceratitis capitata: correlating genetic and physical maps of chromosome 5.

We have isolated and determined the nucleotide sequences for cDNA clones encoding glucose-6-phosphate dehydrogenase (G6PD) and 6-phosphogluconate dehydrogenase (6PGD) from the medfly Ceratitis capitata. The derived amino acid sequences for G6PD and 6PGD are presented and compared with G6PDs and 6PGDs from other species. The codon usage of the cDNA clones has little bias with the notable exceptions of arginine, glycine and leucine. The chromosomal location of the genes for 6PGD and G6PD were determined by in situ hybridization to salivary gland polytene chromosomes. This localization orients a genetic map of enzymatic loci and illustrates a remarkable similarity in the intra chromosomal order of homologous genes between Drosophila melanogaster and medfly.

Amino Acid Sequence↗

[Amplification, cloning and sequence analysis of spider dragline silk cDNA].

Spider dragline silk is synthesized in special gland named major ampulate (MA) gland. The MA glands were dissected from the abdomen of the spiders Nephila clavata and the total RNA was extracted by the TRIZOL. The cDNA of dragline silk was amplificated by RT-PCR (reverse transcription polymerase chain reaction), multiplex PCR and cloned. PCR identification, restriction analysis and DNA sequence analysis were carried out to verify the recombinant plasmids. The codon usage frequencies of the cloned cDNA were added up, and the predicted amino acid sequence was compared with Spidroin2 of Nephila clavipes. Predicted secondary structure of the predicted amino-acid sequence was analysized by DNAStar software. All results showed that the cloned cDNA we got (GenBank Accession No. AF441245) was the very fragment of spider dragline silk Spidroin2 cDNA.

Amino Acid Sequence↗

The control region of the F plasmid transfer operon: DNA sequence of the traJ and traY genes and characterisation of the traY leads to Z promoter.

The complete nucleotide sequence of the F plasmid transfer genes traJ and traY, together with the promoter-proximal region of the traA gene has been determined. The traJ reading frame has been confirmed by sequencing the traJ90 amber mutant allele. The predicted amino acid sequence of the TraJ protein shows that this outer-membrane protein lacks a signal sequence. The pattern of codon usage within the traJ gene is different from that of genes for abundant outer-membrane proteins and is closer to that of genes that are expressed at relatively low levels. We have located the traY leads to Z operon promoter by in vitro run-off transcription experiments and have developed in vivo assays for the activity of the promoter by fusing it to galactokinase and kanamycin-resistance genes.

Bacterial Proteins↗

Cloning, sequencing and biochemical characterization of xylose isomerase from Thermoanaerobacterium saccharolyticum strain B6A-RI.

The xylose isomerase gene from Thermoanaerobacterium saccharolyticum strain B6A-RI was cloned by complementation using Escherichia coli xyl-5 mutant strain HB101. One positive clone was detected and the recombinant plasmid, pZX16, was isolated. The clone contained the vector pUC18 and an insert fragment of 4.5 kb. The cloned xylose isomerase gene (xylA) was expressed constitutively in E. coli. The gene contained one open reading frame (ORF) of 1317 bp, which corresponds to 439 amino acid residues. The molecular mass of the gene product was calculated to be 50474 Da from the deduced amino acid sequence. A putative promoter region (Pribnow box), TATAATATATAAT, which repeated twice at the -10 region in E. coli, was found 25 bp upstream of the ribosomal binding site. The deduced amino acid sequence of T. saccharolyticum strain B6A-RI xylose isomerase exhibited very high homology to those from Thermoanaerobacterium thermosulfurigenes 4B (formerly Clostridium thermosulfurogenes 4B) and Thermoanaerobacter ethanolicus 39E (formerly Clostridium thermohydrosulfuricum 39E). Codon usage in xynA, xynB and xylA showed a clear propensity for AT-containing isocodons. The native molecular mass of the purified recombinant thermostable xylose isomerase was 200 kDa, and the enzyme was a tetramer comprised of identical subunits. The apparent temperature and pH optima for activity of the cloned xylose isomerase were 80 degrees C and 7.0 to 7.5, respectively.

Aldose-Ketose Isomerases↗

The CDC25 "Start" gene of Saccharomyces cerevisiae: sequencing of the active C-terminal fragment and regional homologies with rhodopsin and cytochrome P450.

The CDC25 Start gene whose product appears to be required for traversing the Go phase of the cell cycle in Saccharomyces cerevisiae, has been previously cloned (J. Daniel and G. Simchen (1986), Curr Genet 10:643-646). By nucleotide sequencing of an active subclone, we found that only a region of the gene that codes for the C-terminal portion of the CDC25 protein was required for full suppression of the cdc25 mutation. The codon usage in this region indicates a poor translation of the transcript compared to genes encoding abundant proteins. The derived CDC25 protein fragment contains two regions of homology, one with the rhodopsin family, the other with the cytochrome P450 family. Strikingly, these two regions of homology are adjacent on the CDC25 protein. In view of the likely involvement of the CDC25 protein in the regulation of adenylate cyclase activity, a working hypothesis is proposed that accounts for the observed homologies.

Adenylyl Cyclases↗

Yeast omnipotent supressor SUP1 (SUP45): nucleotide sequence of the wildtype and a mutant gene.

The primary structures of the yeast recessive omnipotent suppressor gene SUP1 (SUP45) and one of its mutant alleles (sup1-ts36) was determined. The gene codes for a protein of 49 kD. The mutant protein differs from the wildtype form in one amino acid residue (Ser instead of Leu) in the N-terminal part. The codon usage differs significantly from that of yeast ribosomal protein genes. However, an upstream element resembling a conserved oligonucleotide in the region 5' to ribosomal protein genes in S. cerevisiae has been found. A DNA probe internal to the SUP1 gene does not exhibit detectable homology to genomic DNA neither from higher eucaryotes nor from eu- or archaebacteria. The hypothetical function of this protein in control of translational fidelity is discussed.

Amino Acid Sequence↗