Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

The Caenorhabditis elegans unc-93 gene encodes a putative transmembrane protein that regulates muscle contraction.

unc-93 is one of a set of five interacting genes involved in the regulation or coordination of muscle contraction in Caenorhabditis elegans. Rare altered-function alleles of unc-93 result in sluggish movement and a characteristic "rubber band" uncoordinated phenotype. By contrast, null alleles cause no visibly abnormal phenotype, presumably as a consequence of the functional redundancy of unc-93. To understand better the role of unc-93 in regulating muscle contraction, we have cloned and molecularly characterized this gene. We isolated transposon-insertion alleles and used them to identify the region of DNA encoding the unc-93 protein. Two unc-93 proteins differing at their NH2 termini are potentially encoded by transcripts that differ at their 5' ends. The putative unc-93 proteins are 700 and 705 amino acids in length and have two distinct regions: the NH2 terminal portion of 240 or 245 amino acids is extremely hydrophilic, whereas the rest of the protein has multiple potential membrane-spanning domains. The unc-93 transcripts are low in abundance and the unc-93 gene displays weak codon usage bias, suggesting that the unc-93 protein is relatively rare. The unc-93 protein has no sequence similarity to proteins listed in current data-bases. Thus, unc-93 is likely to encode a novel membrane-associated muscle protein. We discuss possible roles for the unc-93 protein either as a component of an ion transport system involved in excitation-contraction coupling in muscle or in coordinating muscle contraction between muscle cells by affecting the functioning of gap junctions.

Alleles↗

Neighboring-nucleotide effects on the rates of germ-line single-base-pair substitution in human genes.

The spectrum of single-base-pair substitutions logged in The Human Gene Mutation Database (HGMD), comprising 7,271 different lesions in the coding regions of 547 different human genes, was analyzed for nearest-neighbor effects on relative mutation rates. Owing to its retrospective nature, HGMD allows mutation rates to be estimated only in relative terms. Therefore, a novel methodology was devised in order to obtain these estimates in iterative fashion, correcting, at the same time, for the confounding effects of differential codon usage and for the fact that different types of amino acid replacement come to clinical attention with different probabilities. Over and above the hypermutability of CpG dinucleotides, reflected in transition rates five times the base mutation rate, only a subtle and locally confined influence of the surrounding DNA sequence on relative single-base-pair substitution rates was observed, which extended no farther than 2 bp from the substitution site. A disparity between the two DNA strands was evidenced by the fact that, when substitution rates were estimated conditional on the 5' and 3' flanking nucleotides, a significant rate difference emerged for 10 of 96 possible pairs of complementary substitutional events. Mutational bias, favoring substitutions toward flanking bases, a phenomenon reminiscent of misalignment mutagenesis, was apparent and exhibited both directionality and reading-frame sensitivity. No specific preponderance of repeat-sequence motifs was observed in the vicinity of nucleotide substitutions, but a moderate correlation between the relative mutability and thermodynamic stability of DNA triplets emerged, suggesting either inefficient DNA replication in regions of high stability or the transient stabilization of misaligned intermediates.

Base Pairing↗

Rev-independent expression of synthetic gag-pol genes of human immunodeficiency virus type 1 and simian immunodeficiency virus: implications for the safety of lentiviral vectors.

The safety of lentiviral vectors for clinical applications is still a major concern. The gag-pol expression plasmids and the lentiviral vectors used in previous studies contain homologous regions, which constitute a risk for recombination events. Synthetic gag-pol genes of human immunodeficiency virus type 1 (HIV-1) and simian immunodeficiency virus (SIV) were therefore constructed, in which the codon usage was optimized for expression in human cells without altering the amino acid sequences. The synthetic gag-pol genes allowed efficient expression of these genes in the absence of Rev and the 5' untranslated leader region. Both the HIV-1 and the SIV synthetic gag-pol expression plasmids could mediate transduction of an SIV vector into nondividing human cells with titers of about 10(6) transducing units/ml. Similar titers were obtained with a four-plasmid vector-packaging system based on HIV-1. Using a biological assay, homologous recombination events between the synthetic gag-pol expression plasmids and an SIV vector were undetectable and in comparison with a previously used gag-pol expression plasmid at least approximately 100-fold less frequent. By eliminating regions of homology and sequences involved in packaging, synthetic gag-pol genes should improve the safety profile of lentiviral vectors.

5' Untranslated Regions↗

Subtype-specific patterns in HIV Type 1 reverse transcriptase and protease in Oyo State, Nigeria: implications for drug resistance and host response.

As the use of antiretroviral therapy becomes more widespread across Africa, it is imperative to characterize baseline molecular variability and subtype-specific peculiarities of drug targets in non-subtype B HIV-1 infection. We sequenced and analyzed 35 reverse transcriptase (RT) and 43 protease (PR) sequences from 50 therapy-naive HIV-1-infected Nigerians. Phylogenetic analyses of RT revealed that the predominant viruses were CRF02_AG (57%), subtype G (26%), and CRF06_cpx (11%). Six of 35 (17%) individuals harbored primary mutations for RT inhibitors, including M41L, V118I, Y188H, P236L, and Y318F, and curiously three of the six were infected with CRF06_cpx. Therefore, CRF06_cpx drug-naive individuals had significantly more drug resistance mutations than the other subtypes (p = 0.011). By combining data on quasisynonymous codon bias with the influence of the differential genetic cost of mutations, we were able to predict some mutations, which are likely to predominate by subtype, under drug pressure. Some subtype-specific polymorphisms occurred within epitopes for HLA B7 and B35 in the RT, and HLA A2 and A*6802 in PR, at positions implicated in immune evasion. Balanced polymorphism was also observed at predicted serine-threonine phosphorylation sites in the RT of subtype G viruses. The subtype-specific codon usage and polymorphisms observed suggest the involvement of differential pathways for drug resistance and host-driven viral evolution in HIV-1 CRF02_AG, subtype G, and CRF06_cpx, compared to subtype B. Subtype-specific responses to HIV therapy may have significant consequences for efforts to provide effective therapy to the populations infected with these HIV-1 subtypes.

Amino Acid Sequence↗

Recognition of genes in human DNA sequences.

A new approach to computer-assisted gene recognition in higher eukaryote DNA is suggested. It allows one to use not only linear functions for scoring structures, but all functions satisfying natural monotonicity conditions. The algorithm constructs the set of structures guaranteed to contain an optimal structure for every function. So, it uncouples the time-consuming step of generation of this set from the fast step of structure scoring, thus making it simple to experiment with different functions. One particular scoring function, taking into account only codon usage and positional nucleotide frequencies of the splicing sites, has been implemented in the Genome Recognition and Exon Assembly Tool program, and has been tested on an independent sample of human genes, yielding 88% sensitivity and 79% specificity.

Algorithms↗

Identification of a Balb/c mouse pro alpha 1(I) procollagen gene: evidence for insertions or deletions in gene coding sequences.

We report the first isolation and identification of a mouse genomic fragment encoding amino acid sequences for the pro alpha 1(I) chain of type I procollagen. The DNA sequence of eight coding sequences is presented; five of these are 54 bp and three are 108 bp in length. Together these specify 198 amino acids which are 94% homologous to the corresponding bovine pro alpha 1(I) chain protein sequences. Each of the eight coding sequences is flanked by appropriate splice-junction sequences that exhibit considerable sequence complementarity to the rat small nuclear U1a RNA. In the 198 codons examined in this mouse genomic clone, the preferred codons for glycine and alanine are GGU (46/67) and GCU (23/30), respectively. This is in contrast to the codon usage reported for the chicken pro alpha 1(I) cDNA clone (Fuller and Boedtker, 1981). The examined coding sequences exhibit considerable nucleotide homology in both end-to-end and in staggered alignments. Based on an analysis of this homology data, a model is presented for the generation of 108-bp coding sequences from 54-bp units by two successive homologous recombinational events within coding sequences. Alternatively, the 108-bp units may have arisen by precise deletions of an intervening sequence between 54-bp coding sequences. Evidence supporting this is provided by a comparison of pro alpha 1(I) and pro alpha 2(I) genes. In the mouse pro alpha 1(I) gene amino acids 856-891 are encoded in a 108-bp unit; the chicken pro alpha 2(I) gene these residues are encoded in two 54-bp coding sequences. In addition, the coding sequences for nearly 50% of the alpha domain are condensed in the pro alpha 1(I) gene into a region approximately one half the size occupied by the comparable sequences in the pro alpha 2(I) gene.

Amino Acid Sequence↗

Construction and identification by partial nucleotide sequence analysis of bovine casein and beta-lactoglobulin cDNA clones.

Double stranded (DS) DNA molecules obtained by reverse transcription of a partially purified lactating bovine mammary gland mRNA fraction were cloned into pBR322. Restriction maps for four recombinants were constructed and partial nucleotide sequence analysis of these revealed coding sequences corresponding to alpha s1-, beta-, and kappa-casein and beta-lactoglobulin. The specific single-stranded (SS) cDNAs representing each of these species were identified and their nucleotide lengths estimated. Evidence is presented that these are essentially full-length transcripts of the major mRNA species. On this basis, the cDNA clones range in size from 50% for beta-casein to about 95% for alpha s1-casein in comparison with their respective mRNAs. The DNA sequence spanning all eight phosphoserine residues in alpha s1-casein is presented. These data, together with other serine codon usage data, indicate that the mammary gland phosphoseryl tRNA does not play a role in the incorporation of serine phosphate residues during casein synthesis. The observation that the nucleotide sequences for the serine phosphate cluster in bovine alpha s1-and rat beta-casein exhibit close homology supports the suggestion that these regions have evolved from a common primordial sequence.

Animals↗

The Tc2 transposon of Caenorhabditis elegans has the structure of a self-regulated element.

We have analyzed the sequence of the Tc2 transposon of the nematode Caenorhabditis elegans. The Tc2 element is 2,074 bp in length and has perfect inverted terminal repeats of 24 bp. The structure of this element suggests that it may have the capacity to code for a transposase protein and/or for regulatory functions. Three large reading frames on one strand exhibit nonrandom codon usage and may represent exons. The first open coding region is preceded by a potential CAAT box, TATA box, and consensus heat shock sequence. In addition to its inverted terminal repeats, Tc2 has an unusual structural feature: subterminal degenerate direct repeats that are arranged in an irregular overlapping pattern. We have also examined the insertion sites of two Tc2 elements previously identified as the cause of restriction fragment length polymorphisms. Both insertions generated a target site duplication of 2 bp. One element had inserted inside the inverted terminal repeat of another transposon, splitting it into two unequal parts.

Animals↗

In vitro and in silico cloning of Xenopus laevis SOD2 cDNA and its phylogenetic analysis.

By using the methodology of both wet and dry biology (i.e., RT-PCR and cycle sequencing, and biocomputational technology, respectively) and the data obtained through the Genome Projects, we have cloned Xenopus laevis SOD2 (MnSOD) cDNA and determined its nucleotide sequence. These data and the deduced protein primary structure were compared with all the other SOD2 nucleotide and amino acid sequences from eukaryotes and prokaryotes, published in public databases. The analysis was performed by using both Clustal W, a well known and widely used program for sequence analysis, and AntiClustAl, a new algorithm recently created and implemented by our group. Our results demonstrate a very high conservation of the enzyme amino acid sequence during evolution, which proves a close structure-function relationship. This is to be expected for very ancient molecules endowed with critical biological functions, performed through a specific structural organization. The nucleotide sequence conservation is less pronounced: this too was foreseeable, due to neutral mutations and to the species-specific codon usage. The data obtained by using AntiClustAl are comparable with those produced with Clustal W, which validates this algorithm as an important new tool for biocomputational analysis. Finally, it is noteworthy that evolutionary trees, drawn by using all the available data on SOD2 nucleotide sequences and amino acid and either Clustal W or AntiClustAl, are comparable to those obtained through phylogenetic analysis based on fossil records.

Amino Acid Sequence↗

Immunogenicity testing of a novel engineered HIV-1 envelope gp140 DNA vaccine construct.

DNA vaccines expressing the envelope (env) of the human immunodeficiency virus type 1 (HIV-1) have been relatively ineffective at generating strong immune responses. In this study, we described the development of a recombinant plasmid DNA (pEK2P-B) expressing an engineered codon-optimized envelope gp140 gene of primary (nonrecombinant) HIV-1 subtype B isolate 6101. Codon usage and RNA optimization of HIV-1 structural genes has been shown to increase protein expression in vitro as well as in the context of DNA vaccines in vivo. To further increase the expression, a synthetic IgE leader with kozak sequences were fused into the env gene. The cytoplasmic tail of the gene was also truncated to prevent recycling. The expression of env by the recombinant pEK2P-B was evaluated using T7 coupled transcription/translation. The construct demonstrated high expression of the HIV-1 env gene in eukaryotic cells as demonstrated in transfected 293-T and RD cells. Immunogenicity of pEK2P-B was evaluated in mice using IFN-gamma ELISpot assay, and the construct was found to be highly immunogenic and crossreactive with HIV-1 clade C env peptides. Three immunodominant peptides were also mapped out. Furthermore, by performing a CFSE flow cytometry-based proliferation assay, 2.4 and 1.5% proliferation was observed in CD4+, CD8+, and CCR+ memory T cells, respectively. Therefore, this engineered synthetic optimized env DNA vaccine may be useful in DNA vaccine and other studies of HIV-1 immunogenicity.

AIDS Vaccines↗

MEGA: Molecular Evolutionary Genetics Analysis software for microcomputers.

A computer program package called MEGA has been developed for estimating evolutionary distances, reconstructing phylogenetic trees and computing basic statistical quantities from molecular data. It is written in C++ and is intended to be used on IBM and IBM-compatible personal computers. In this program, various methods for estimating evolutionary distances from nucleotide and amino acid sequence data, three different methods of phylogenetic inference (UPGMA, neighbor-joining and maximum parsimony) and two statistical tests of topological differences are included. For the maximum parsimony method, new algorithms of branch-and-bound and heuristic searches are implemented. In addition, MEGA computes statistical quantities such as nucleotide and amino acid frequencies, transition/transversion biases, codon frequencies (codon usage tables), and the number of variable sites in specified segments in nucleotide and amino acid sequences. Advanced on-screen sequence data and phylogenetic-tree editors facilitate publication-quality outputs with a wide range of printers. Integrated and interactive designs, on-line context-sensitive helps, and a text-file editor make MEGA easy to use.

Algorithms↗

A set of Macintosh computer programs for the design and analysis of synthetic genes.

Computer programs that can be used for the design of synthetic genes and that are run on an Apple Macintosh computer are described. These programs determine nucleic acid sequences encoding amino acid sequences. They select DNA sequences based on codon usage as specified by the user, and determine the placement of base changes that can be used to create restriction enzyme sites without altering the amino acid sequence. A new algorithm for finding restriction sites by translating the restriction endonuclease target sequence in all three reading frames and then searching the given peptide or protein amino acid sequence with these short restriction enzyme peptide sequences is described. Examples are given for the creation of synthetic DNA sequences for the bovine prethrombin-2 and ribonuclease A genes.

Algorithms↗

Genome inhomogeneity is determined mainly by WW and SS dinucleotides.

According to the hypothesis of the modular structure of DNA, genomes consist of modules of various nature which may differ in statistical characteristics. Statistical analysis helps in revealing the differences in statistical characteristics and predicting the modular structure. In this connection the question about the contribution of each word of length l (l-tuple) to the inhomogeneity of genetic text arises. The notion of stationary (i.e. relatively evenly distributed over a genome) versus non-stationary l-tuples has been introduced previously. In this paper, the dinucleotide distributions for all long sequences from GenBank were analyzed and it was shown that non-stationary dinucleotides are closely associated with polyW and polyS tracts (W denotes 'weak' nucleotides A or T, while S stands for the 'strong' nucleotides G or C). Thus, genome inhomogeneity is shown to be determined mainly by AA, TT, GG, CC, AT, TA, GC and CG dinucleotides. It has been demonstrated that neither 'codon usage' nor the 'isochore model' can account for this phenomenon.

Algorithms↗

Preference of simple sequence repeats in coding and non-coding regions of Arabidopsis thaliana.

MOTIVATION: Simple sequence repeats or microsatellites have been found abundantly in many genomes. However, the significance of distribution preference has not been completely understood. Completion of the Arabidopsis genome sequencing allows us to better understand and characterize microsatellites. RESULTS: Microsatellite distribution was more abundant in 5'-flanking regions of genes compared with that expected in the whole genome, with an over-representation of AG and AAG repeats; there were clear differences from distributions in 3'-flanks and coding fractions, where triplet frequencies evidently corresponded to codon usage. We identified 1140 full-length genes that contained at least one locus of AG or AAG repeats in their upstream sequences, and whose functional characteristics were significantly associated with the repeats. This observation indicates that selective pressure markedly differed in the three transcribed regions, with positive selection of AG and AAG repeats in 5'-flanks close to those genes whose products are preferentially involved in transcription.

Algorithms↗

TIP: protein backtranslation aided by genetic algorithms.

UNLABELLED: Several applications require the backtranslation of a protein sequence into a nucleic acid sequence. The degeneracy of the genetic code makes this process ambiguous; moreover, not every translation is equally viable. The usual answer is to mimic the codon usage of the target species; however, this does not capture all the relevant features of the 'genomic styles' from different taxa. The program TIP ' Traducción Inversa de Proteínas') applies genetic algorithms to improve the backtranslation, by minimizing the difference of some coding statistics with respect to their average value in the target. AVAILABILITY: http://www.cmm.uchile.cl/genoma/tip/

Algorithms↗

Molecular analysis of surface-associated enzymes of Porphyromonas gingivalis.

There is now increasing evidence that surface-associated enzymes, previously considered to be involved in intermediary metabolism or virulence, play a role in physiological reactions such as signal transduction, transport systems, and metabolic processes. Herein we report the molecular aspects of two such enzymes, the cysteine proteinase gingivain and NAD-dependent glutamate dehydrogenase of Porphyromonas gingivalis. The gdh gene comprises an open reading frame of 1,335 base pairs that encodes a 49,000-M(r) protein of 445 amino acids. The gdh gene showed high homology (78.3%) with that of Clostridium symbiosum. Optimal codons accounted for 35.9% of the total codon usage, indicating high expression of this enzyme. These data are currently being used to carry out targeted mutagenesis, which was established here for gingivain. Conditions for targeted mutagenesis within the histidine domain of the catalytic site of gingivain using Tn 4351 was successfully achieved. Consequently, the catalytic functions, such as gingivain's capacity to hydrolyze the synthetic substrate alpha-benzoyl-arginine-4-nitroanilide, were disrupted.

Amino Acid Sequence↗

Complete genome sequence of enterohemorrhagic Escherichia coli O157:H7 and genomic comparison with a laboratory strain K-12.

Escherichia coli O157:H7 is a major food-borne infectious pathogen that causes diarrhea, hemorrhagic colitis, and hemolytic uremic syndrome. Here we report the complete chromosome sequence of an O157:H7 strain isolated from the Sakai outbreak, and the results of genomic comparison with a benign laboratory strain, K-12 MG1655. The chromosome is 5.5 Mb in size, 859 Kb larger than that of K-12. We identified a 4.1-Mb sequence highly conserved between the two strains, which may represent the fundamental backbone of the E. coli chromosome. The remaining 1.4-Mb sequence comprises of O157:H7-specific sequences, most of which are horizontally transferred foreign DNAs. The predominant roles of bacteriophages in the emergence of O157:H7 is evident by the presence of 24 prophages and prophage-like elements that occupy more than half of the O157:H7-specific sequences. The O157:H7 chromosome encodes 1632 proteins and 20 tRNAs that are not present in K-12. Among these, at least 131 proteins are assumed to have virulence-related functions. Genome-wide codon usage analysis suggested that the O157:H7-specific tRNAs are involved in the efficient expression of the strain-specific genes. A complete set of the genes specific to O157:H7 presented here sheds new insight into the pathogenicity and the physiology of O157:H7, and will open a way to fully understand the molecular mechanisms underlying the O157:H7 infection.

Bacterial Proteins↗

Distribution of repetitive sequences on the leading and lagging strands of the Escherichia coli genome: comparative study of Long Direct Repeat (LDR) sequences.

In the present study, we developed a method for detecting sequences whose similarity to a target sequence is statistically significant and we examined the distribution of these sequences in the E. coli K-12 genome. Target sequences examined are as follows: (i) short repeat: Crossover hot-spot instigator (Chi) sequence, replication termination (Ter) sequence, and DnaA binding sequence (DnaA box); (ii) potential stem-loop structure repeats: palindromic unit (PU), boxC sequences, and intergenic repeat unit (IRU); (iii) potential RNA coding repeats: rRNAs, PAIR, TRIP, and QUAD; and (iv) potential protein coding repeats: insertion elements (ISs) and Long Direct Repeats (LDRs). We also examined the distribution of these sequences on leading and lagging strands. We obtained another four statistically significant LDR sequences with more than 187 bp matched to LDR-A near the LDR loci, suggesting that these regions might be used as high recombination hot spots for LDR. Adaptation of individual LDRs to E. coli genome is also discussed on the basis of codon usage.

Bacterial Proteins↗