Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon optimality”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

The role of cytochrome c-550 as studied through reverse genetics and mutant characterization in Synechocystis sp. PCC 6803.

The gene coding for cytochrome c-550 in Synechocystis sp. PCC 6803 was cloned based on the N-terminal sequence of the mature polypeptide. Using the most probable translation start codon, the gene is expected to code for 160 amino acid residues. This includes a cleavable N-terminal leader sequence of 25 residues. This leader sequence has an Arg-Asn-Arg sequence immediately before the cleavage site; this is characteristic for transit peptides in prokaryotes. Comparison of this sequence with the leader sequence of the photosystem II-associated extrinsic 33-kDa protein from the same cyanobacterium showed an identity of 13 out of 25 residues. These results suggest that after synthesis of the apoprotein, cytochrome c-550 is transported into the thylakoid lumen. Using the cloned gene, insertion and deletion mutants of Synechocystis sp. PCC 6803 were constructed. In the absence of cytochrome c-550, both mutants were capable of photoautotrophic growth but at a significantly reduced rate. Atrazine bindng and Western blot analysis showed that these mutants on a per-chlorophyll basis contained 53-67% of the amount of photosystem II as compared with wild type. The photosystem II-specific oxygen-evolving activity at saturating light intensity was reduced to about 40% of that in the wild type strain. Taken together, these results indicate that the cytochrome c-550 is transported into the thylakoid lumen and contributes to optimal functional stability of photosystem II in cyanobacteria. This supports our biochemical evidence that cytochrome c-550 is associated with the lumenal side of photosystem II as one of the extrinsic proteins enhancing oxygen evolution (Shen, J.-R., Ikeuchi, M., and Inoue, Y. (1992) FEBS Lett. 301, 145-149; Shen, J.-R., and Inoue, Y. (1993) Biochemistry 32, 1825-1832). Based on these results, the gene for cytochrome c-550 was named psbV. The possible evolutionary relationship among extrinsic proteins of the photosystem II donor side is discussed.

Amino Acid Sequence↗

The involvement of the anticodon adjacent modified nucleoside N-(9-(BETA-D-ribofuranosyl) purine-6-ylcarbamoyl)-threonine in the biological function of E. coli tRNAile.

tRNAile was isolated from E. coli Cp 79 (leu-, arg-, thr-, his-, thiamin-, RCrel) which had been grown on a sub-optimal concentration of thr and was found to contain an average of 50% less N-[9-(beta-D-ribofuranosyl)- purin-6-ylcarbamoyl]threonine, t6Ado, than tRNAile from cells grown on an optimum concentration of thr and containing a normal complement of t6Ado. The two tRNA's were identical in their ability to be aminoacylated, to accept the 3'-terminal dinucleotide, and to form an ile-tRNAile-Tu-GTP complex. In contrast, the t6Ado-deficient-tRNA was significantly less efficient in binding to ribosomes compared to the normal tRNA. This difference was seen in the binding of deacylated tRNA and in the nonenzymatic and enzymatic binding of ile-tRNA, all in response to poly AUC. The t6Ado-deficient ile-tRNA demonstrated no binding at Mg2+ concentrations less than or equal to 10 mM, while the normal ile-tRNA bound at low Mg2+ concentrations. Tetracycline had the same effect on the normal as on the t6Ado-deficient ile-tRNA binding. As a control, the binding of phe-tRNA (which does not contain t6Ado) from normal and thr-starved cells in response to poly U was identical. It was concluded that t6Ado is required for proper codon-anticodon interaction.

Amino Acyl-tRNA Synthetases↗

Recognition of Li Fraumeni syndrome at diagnosis of a locally advanced extremity rhabdomyosarcoma.

A contemporaneous presentation of a second breast cancer in a mother and an extremity rhabdomyosarcoma (RMS) in her daughter led to the diagnosis of the Li Fraumeni syndrome (LFS). Although the association between LFS and RMS in young patients is well recognised 1 there are no guidelines as to how this knowledge should influence the optimal management of these patients. After reviewing the literature about the natural history of the LFS 2, the incidence of second malignancy (SMN) in RMS survivors 3-6 and the management of extremity RMS 7-9, we are concerned that contemporary RMS treatment, combining non-mutilating surgery with chemoradiotherapy, may be associated with an excessive SMN risk in LFS patients with advanced RMS. We question whether treatment should be individualised and, where possible and acceptable to the family, measures such as amputation should be the considered to attain local control for LFS patients with RMS as this will avoid the need for local radiotherapy without compromising long-term function and quality of life 10.

Adrenal Cortex Neoplasms↗

Analysis of Fc gamma RIII (CD16) membrane expression and association with CD3 zeta and Fc epsilon RI-gamma by site-directed mutation.

Two genes encode Fc gamma RIII (CD16), a low affinity FcR for IgG. CD16-I is expressed as a phosphatidylinositol glycan-anchored membrane glycoprotein on neutrophils, whereas CD16-II is a transmembrane-linked glycoprotein on NK cells. Membrane anchoring is determined by codon 203. Site-directed mutation of codon 203 and transient expression of these cDNA in COS-7 cells indicated that Phe, Ile, Leu, and Val permit transmembrane expression, whereas Ser, Thr, Tyr, Asn, Gly, Ala, Asp and Lys enable phosphatidylinositol-glycan attachment. Thus, the involvement of amino acid 203 in membrane anchoring cannot be explained simply on the basis of size, charge, or polarity of the amino acid side groups at this site. Efficient expression of CD16-II in COS-7 cells requires co-transfection with either CD3 zeta or Fc epsilon RI-gamma. Truncation of the cytoplasmic segment of CD16 failed to affect association with CD3 zeta. CD3 zeta and Fc epsilon RI-gamma with truncated cytoplasmic segments were also able to facilitate membrane expression of CD16-II, implicating the transmembrane segments as the interaction site between CD16-II and CD3 zeta or Fc epsilon RI-gamma. Prior studies have suggested that the acidic residue in the CD3 zeta transmembrane segment may be important for the association of CD3 zeta complexes. Although site-directed mutation of CD3 zeta-Asp36 to Glu, Leu, or Val retained the ability to permit membrane expression of CD16-II, quantitatively the wild-type CD3 zeta-Asp36 provided optimal levels of expression, consistent with conservation of this amino acid in mouse and human CD3 zeta.

Antigens, CD↗

Analysis of high throughput protein expression in Escherichia coli.

The ability to efficiently produce hundreds of proteins in parallel is the most basic requirement of many aspects of proteomics. Overcoming the technical and financial barriers associated with high throughput protein production is essential for the development of an experimental platform to query and browse the protein content of a cell (e.g. protein and antibody arrays). Proteins are inherently different one from another in their physicochemical properties; therefore, no single protocol can be expected to successfully express most of the proteins. Instead of optimizing a protocol to express a specific protein, we used sequence analysis tools to estimate the probability of a specific protein to be expressed successfully using a given protocol, thereby avoiding a priori proteins with a low success probability. A set of 547 proteins, to be used for antibody production and selection, was expressed in Escherichia coli using a high throughput protein production pipeline. Protein properties derived from sequence alone were correlated to successful expression, and general guidelines are given to increase the efficiency of similar pipelines. A second set of 68 proteins was expressed to investigate the link between successful protein expression and inclusion body formation. More proteins were expressed in inclusion bodies; however, the formation of inclusion bodies was not a requirement for successful expression.

Cloning, Molecular↗

Regulatory role of the conserved stem-loop structure at the 5' end of collagen alpha1(I) mRNA.

Three fibrillar collagen mRNAs, alpha1(I), alpha2(I), and alpha1(III), are coordinately upregulated in the activated hepatic stellate cell (hsc) in liver fibrosis. These three mRNAs contain sequences surrounding the start codon that can be folded into a stem-loop structure. We investigated the role of this stem-loop structure in expression of collagen alpha1(I) reporter mRNAs in hsc's and fibroblasts. The stem-loop dramatically decreases accumulation of mRNAs in quiescent hsc's and to a lesser extent in activated hsc's and fibroblasts. The stem-loop decreases mRNA stability in fibroblasts. In activated hsc's and fibroblasts, a protein complex binds to the stem-loop, and this binding requires the presence of a 7mG cap on the RNA. Placing the 3' untranslated region (UTR) of collagen alpha1(I) mRNA in a reporter mRNA containing this stem-loop further increases the steady-state level in activated hsc's. This 3' UTR binds alphaCP, a protein implicated in increasing stability of collagen alpha1(I) mRNA in activated hsc's (B. Stefanovic, C. Hellerbrand, M. Holcik, M. Briendl, S. A. Liebhaber, and D. A. Brenner, Mol. Cell. Biol. 17:5201-5209, 1997). A set of protein complexes assembles on the 7mG capped stem-loop RNA, and a 120-kDa protein is specifically cross-linked to this structure. Thus, collagen alpha1(I) mRNA is regulated by a complex interaction between the 5' stem-loop and the 3' UTR, which may optimize collagen production in activated hsc's.

3T3 Cells↗

Assessment of sequence-based p53 gene analysis in human breast cancer: messenger RNA in comparison with genomic DNA targets.

The high prevalence of p53 mutations in human cancers and the suggestion from several groups that the presence or absence of p53 mutations might have both prognostic and therapeutic consequences point to the importance of optimal methods for p53 determination. Several strategies exploring this have been described, based either on mRNA or genomic DNA as a template. However, no comparative study on the reliability of the two templates has been performed. The principal aim of this study was to study the concordance of RNA- and DNA-based direct sequencing methods in detecting p53 mutations in breast tumors. In 100 tumors, 22 mutations were detected by both methods. Furthermore, one stop mutation, two splice-site mutations, and one intron alteration were found only by genomic sequencing. In addition, the comparative study suggests that cells with missense mutations have increased steady-state concentrations of p53-specific mRNA, in contrast to cells with a gene encoding a truncated protein.

Aneuploidy↗

Emergence of zidovudine and multidrug-resistance mutations in the HIV-1 reverse transcriptase gene in therapy-naive patients receiving stavudine plus didanosine combination therapy. STADI Group.

OBJECTIVE: Assessment of genotypic changes in the reverse transcriptase gene of HIV-1 occurring in antiretroviral naive patients treated by stavudine plus didanosine combination therapy. METHODS: Sequence analysis (codons 1-230) was performed after amplification of the reverse transcriptase gene from plasma samples collected at baseline and at the end of treatment from 39 previously treatment-naive patients treated for 24-48 weeks. RESULTS: At baseline, mutations associated with zidovudine resistance were detected in plasma from two patients: Asp67Asn/Lys219Gln and Leu210Trp. Among the 39 subjects, 18 (46%) developed mutations: one developed the Val75Thr/Ala mutation, four (10%) developed a Gln151Met multidrug-resistance mutation (MDR), associated in one of them with the Phe77Leu and the Phe116Tyr MDR mutations and 14 (36%) developed one or more zidovudine-specific mutations (Met41Leu, Asp67Asn, Lys70Arg, Leu210Trp, Thr215Tyr/Phe). The development of a Met41Leu zidovudine-specific mutation was associated with the development of a Gln151Met mutation in one patient. Other reverse transcriptase mutations known to confer resistance to nucleoside analogues were not detected. At inclusion, there was no statistical difference in HIV-1 load between patients who developed resistance mutations and those who did not. RNA HIV-1 load decrease was higher (P = 0.05) in patients who maintained a wild-type reverse transcriptase genotype (-2.22 log10 copies/ml) than in patients who developed resistance mutations (-1.14 log10 copies/ml). CONCLUSION: Stavudine/didanosine combination therapy is associated with emergence of zidovudine-related resistance or MDR mutations in naive patients. These findings should be considered when optimizing salvage therapy for patients who have received a treatment including stavudine/didanosine combination.

Adult↗

Improved green fluorescent protein by molecular evolution using DNA shuffling.

Green fluorescent protein (GFP) has rapidly become a widely used reporter of gene regulation. However, for many organisms, particularly eukaryotes, a stronger whole cell fluorescence signal is desirable. We constructed a synthetic GFP gene with improved codon usage and performed recursive cycles of DNA shuffling followed by screening for the brightest E. coli colonies. A visual screen using UV light, rather than FACS selection, was used to avoid red-shifting the excitation maximum. After 3 cycles of DNA shuffling, a mutant was obtained with a whole cell fluorescence signal that was 45-fold greater than a standard, the commercially available Clontech plasmid pGFP. The expression level in E. coli was unaltered at about 75% of total protein. The emission and excitation maxima were also unchanged. Whereas in E. coli most of the wildtype GFP ends up in inclusion bodies, unable to activate its chromophore, most of the mutant protein is soluble and active. Three amino acid mutations appear to guide the mutant protein into the native folding pathway rather than toward aggregation. Expressed in Chinese Hamster Ovary (CHO) cells, this shuffled GFP mutant showed a 42-fold improvement over wildtype GFP sequence, and is easily detected with UV light in a wide range of assays. The results demonstrate how molecular evolution can solve a complex practical problem without needing to first identify which process is limiting. DNA shuffling can be combined with screening of a moderate number of mutants. We envision that the combination of DNA shuffling and high throughput screening will be a powerful tool for the optimization of many commercially important enzymes for which selections do not exist.

Animals↗

Dimethyl sulfoxide-mediated primer Tm reduction: a method for analyzing the role of renaturation temperature in the polymerase chain reaction.

We report a method for optimizing the specificity of product formation in the polymerase chain reaction (PCR). This technique is based on the use of dimethyl sulfoxide (DMSO) and takes into account primer Tm. The reduction in Tm by DMSO is directly correlated with renaturation temperature such that a DMSO gradient reflects a temperature gradient. We use this relationship to show that optimum product formation usually occurs at or within several degrees of the midpoint Tm of a given primer pair. We illustrate these correlations using three examples deriving PCR products from a human cDNA library, representing the casein kinase II alpha and beta subunits as well as the 5' untranslated region for the beta subunit. By following product formation as a function of renaturation temperature, we postulate rules for cycle design based on primer Tm. Implications for the use of degenerate primers are discussed.

Base Sequence↗

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals↗

Genetics and molecular biology of haemophilias A and B.

The development of rapid procedures for the characterization of mutations is advancing the knowledge of the molecular biology of the haemophilias and transforming the strategies for the diagnoses required for genetic counselling. In haemophilia B more than 300 mutants have been fully characterized. These comprise complete and partial deletions, rare insertions, and 'point' mutations. The latter may impair transcription (promoter mutations), RNA processing (splicing mutations) and translation (frameshifts and stop codons) or cause single amino acid (aa) changes. Eighty-four residues are involved in the 105 presumed detrimental aa substitutions reported so far and these are usually conserved in the factor IX homologues (factors VII, X and protein C) and/or the factor IX of different mammalian species. There are clear correlations between the mutation and clinical features. In addition mutations causing gross physical or functional loss of coding information appear to predispose to the development of antibodies against therapeutic factor IX. Hotspots of mutations have been identified and are usually associated with CpG sequences. In haemophilia A the size and complexity of the factor VIII gene has hindered the analysis of mutants. Most of the studies published so far have analysed only a small fraction of the essential region of the factor VIII gene and this led to the repeated observation of specific types of mutation. The recent development of a rapid method to analyse RNA splicing and the whole coding region of the factor VIII gene should unblock this situation. With regard to genetic counselling, the direct detection of gene defects has increased the proportion of haemophilia B families that can be helped from 60% to virtually 100% and similar expectations may now be formulated for haemophilia A. In the UK a national database of haemophilia B mutations is being constructed to optimize genetic counselling. This should offer a model for a similar development in haemophilia A.

DNA Mutational Analysis↗

Hamster polyomavirus-derived virus-like particles are able to transfer in vitro encapsidated plasmid DNA to mammalian cells.

The authentic major capsid protein 1 (VP1) of hamster polyomavirus (HaPyV) consists of 384 amino acid (aa) residues (42 kDa). Expression from an additional in-frame initiation codon located upstream from the authentic VP1 open reading frame (at position -4) might result in the synthesis of a 388 aa-long, amino-terminally extended VP1 (aa -4 to aa 384; VP1(ext)). In a plasmid-mediated Drosophila Schneider (S2) cell expression system, both VP1 derivatives as well as a VP1(ext) variant with an amino acid exchange of the authentic Met1Gly (VP1(ext-M1)) were expressed to a similar high level. Although all three proteins were detected in nuclear as well as cytoplasmic fractions, formation of virus-like particles (VLPs) was observed exclusively in the nucleus as confirmed by negative staining electron microscopy. The use of a tryptophan promoter-driven Escherichia coli expression system resulted in the efficient synthesis of VP1 and VP1(ext) and formation of VLPs. In addition, establishment of an in vitro disassembly/reassembly system allowed the encapsidation of plasmid DNA into VLPs. Encapsidated DNA was found to be protected against the action of DNase I. Mammalian COS-7 and CHO cells were transfected with HaPyV-VP1-VLPs carrying a plasmid encoding enhanced green fluorescent protein (eGFP). In both cell lines eGFP expression was detected indicating successful transfer of the plasmid into the cells, though at a still low level. Cesium chloride gradient centrifugation allowed the separation of VLPs with encapsidated DNA from "empty" VLPs, which might be useful for further optimization of transfection. Therefore, heterologously expressed HaPyV-VP1 may represent a promising alternative carrier for foreign DNA in gene transfer applications.

Amino Acid Sequence↗

Protein- and mRNA-based phenotype-genotype correlations in DMD/BMD with point mutations and molecular basis for BMD with nonsense and frameshift mutations in the DMD gene.

Straightforward detectable Duchenne muscular dystrophy (DMD) gene rearrangements, such as deletions or duplications involving an entire exon or more, are involved in about 70% of dystrophinopathies. In the remaining 30% a variety of point mutations or "small" mutations are suspected. Due to their diversity and to the large size and complexity of the DMD gene, these point mutations are difficult to detect. To overcome this diagnostic issue, we developed and optimized a routine muscle biopsy-based diagnostic strategy. The mutation detection rate is almost as high as 100% and mutations were identified in all patients for whom the diagnosis of DMD and Becker muscular dystrophy (BMD) was clinically suspected and further supported by the detection on Western blot of quantitative and/or qualitative dystrophin protein abnormalities. Here we report a total of 124 small mutations including 11 nonsense and frameshift mutations detected in BMD patients. In addition to a comprehensive assessment of muscular phenotypes that takes into account consequences of mutations on the expression of the dystrophin mRNA and protein, we provide and discuss genomic, mRNA, and protein data that pinpoint molecular mechanisms underlying BMD phenotypes associated with nonsense and frameshift mutations.

Adolescent↗

Optimized procedure for renaturation of recombinant human bone morphogenetic protein-2 at high protein concentration.

The human gene encoding the mature form of bone morphogenetic protein-2 (hBMP-2), a dimeric disulfide-bonded protein of the cystine knot growth factor family, was expressed in recombinant Escherichia coli using a temperature-inducible expression system. The recombinant protein was produced in the form of cytoplasmic inclusion bodies and the effect of different variables on the renaturation of rhBMP-2 was investigated. In particular, variables such as pH, redox conditions, protein concentration, temperature, the presence of different types of aggregation suppressors, and host cell contaminants were studied with respect to their effect on aggregation during refolding and on the final renaturation yield of rhBMP-2. It is shown that the renaturation yield is particularly sensitive to pH, temperature, protein concentration, and the presence of aggregation suppressors. In contrast, little effect of the redox conditions and the ionic strength on the renaturation yield was observed, as equal yields were obtained in a broad range of reduced to oxidized glutathione ratios and concentrations of NaCl, respectively. The aggregation suppressor 2-(cyclohexylamino)ethanesulfonic acid (CHES) proved to be superior with respect to the final renaturation yield, although, in comparison to the more common arginine, it was less efficient in preventing aggregation of rhBMP-2 during refolding. Detergent washing of inclusion bodies was sufficient, as further purification of rhBMP-2 prior to refolding was without effect on the final renaturation yield. An increase in the concentration of renatured rhBMP-2 was achieved by a pulsed refolding procedure by which up to a total amount of 2.1 mg mL(-1) rhBMP-2 could be transferred in seven pulses into the renaturation buffer with an overall refolding yield of 38%, corresponding to 0.8 mg mL(-1) renatured dimeric rhBMP-2. Furthermore, a simplified purification procedure is presented that also includes freeze-drying for long-term storage of biologically active rhBMP-2. Finally, it is shown that the appearance of rhBMP-2 variants could be avoided by using a host strain overexpressing rare codon tRNAs.

Bone Morphogenetic Protein 2↗

Assessment of protein coding measures.

A number of methods for recognizing protein coding genes in DNA sequence have been published over the last 13 years, and new, more comprehensive algorithms, drawing on the repertoire of existing techniques, continue to be developed. To optimize continued development, it is valuable to systematically review and evaluate published techniques. At the core of most gene recognition algorithms is one or more coding measures--functions which produce, given any sample window of sequence, a number or vector intended to measure the degree to which a sample sequence resembles a window of 'typical' exonic DNA. In this paper we review and synthesize the underlying coding measures from published algorithms. A standardized benchmark is described, and each of the measures is evaluated according to this benchmark. Our main conclusion is that a very simple and obvious measure--counting oligomers--is more effective than any of the more sophisticated measures. Different measures contain different information. However there is a great deal of redundancy in the current suite of measures. We show that in future development of gene recognition algorithms, attention can probably be limited to six of the twenty or so measures proposed to date.

Algorithms↗

C-->U editing of neurofibromatosis 1 mRNA occurs in tumors that express both the type II transcript and apobec-1, the catalytic subunit of the apolipoprotein B mRNA-editing enzyme.

C-->U RNA editing of neurofibromatosis 1 (NF1) mRNA changes an arginine (CGA) to a UGA translational stop codon, predicted to result in translational termination of the edited mRNA. Previous studies demonstrated varying degrees of C-->U RNA editing in peripheral nerve-sheath tumor samples (PNSTs) from patients with NF1, but the basis for this heterogeneity was unexplained. In addition, the role, if any, of apobec-1, the catalytic deaminase that mediates C-->U editing of mammalian apolipoprotein B (apoB) RNA, was unresolved. We have examined these questions in PNSTs from patients with NF1 and demonstrate that a subset (8/34) manifest C-->U editing of RNA. Two distinguishing characteristics were found in the PNSTs that demonstrated editing of NF1 RNA. First, these tumors express apobec-1 mRNA, the first demonstration, in humans, of its expression beyond the luminal gastrointestinal tract. Second, PNSTs with C-->U editing of RNA manifest increased proportions of an alternatively spliced exon, 23A, downstream of the edited base. C-->U editing of RNA in these PNSTs was observed preferentially in transcripts containing exon 23A. These findings were complemented by in vitro studies using synthetic RNA templates incubated in the presence of recombinant apobec-1, which again confirmed preferential editing of transcripts containing exon 23A. Finally, adenovirus-mediated transfection of HepG2 cells revealed induction of editing of apoB RNA, along with preferential editing of NF1 transcripts containing exon 23A. Taken together, the data support the hypothesis that C-->U RNA editing of the NF1 transcript occurs both in a subset of PNSTs and in an alternatively spliced form containing a downstream exon, presumably an optimal configuration for enzymatic deamination by apobec-1.

APOBEC-1 Deaminase↗

Inferring parameters shaping amino acid usage in prokaryotic genomes via Bayesian MCMC methods.

Molar content of guanine plus cytosine (G + C) and optimal growth temperature (OGT) are main factors characterizing the frequency distribution of amino acids in prokaryotes. Previous work, using multivariate exploratory methods, has emphasized ascertainment of biological factors underlying variability between genomes, but the strength of each identified factor on amino acid content has not been quantified. We combine the flexibility of the phylogenetic mixed model (PMM) with the power of Bayesian inference via Markov Chain Monte Carlo (MCMC) methods, to obtain a novel evolutionary picture of amino acid usage in prokaryotic genomes. We implement a Bayesian PMM which incorporates the feature that evolutionary history makes observed data interdependent. As in previous studies with PMM, we present a variance partition; however, attention is also given to the posterior distribution of "systematic effects" that may shed light about the relative importance of and relationships between evolutionary forces acting at the genomic level. In particular, we analyzed influences of G + C, OGT, and respiratory metabolism. Estimates of G + C effects were significant for amino acids coded by G + C or molar content of adenine plus thymine (A + T) in first and second bases. OGT had an important effect on 12 amino acids, probably reflecting complex patterns of protein modifications, to cope with varying environments. The effect of respiratory metabolism was less clear, probably due to the already reported association of G + C with aerobic metabolism. A "heritability" parameter was always high and significant, reinforcing the importance of accommodating phylogenetic relationships in these analyses. "Heritable" component correlations displayed a pattern that tended to cluster "pure" G + C (A + T) in first and second codon positions, suggesting an inherited departure from linear regression on G + C.

Amino Acids↗