Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Extracting phylogenetic information from whole-genome sequencing projects: the lactic acid bacteria as a test case.

The availability of an ever increasing number of complete genome sequences of diverse prokaryotic taxa has led to the introduction of novel approaches to infer phylogenetic relationships among bacteria. In the present study the sequences of the 16S rRNA gene and nine housekeeping genes were compared with the fraction of shared putative orthologous protein-encoding genes, conservation of gene order, dinucleotide relative abundance and codon usage among 11 genomes of species belonging to the lactic acid bacteria. In general there is a good correlation between the results obtained with various approaches, although it is clear that there is a stronger phylogenetic signal in some data sets than in others, and that different parameters have different taxonomic resolutions. It appears that trees based on different kinds of information derived from whole-genome sequencing projects do not provide much additional information about the phylogenetic relationships among bacterial taxa compared to more traditional alignment-based methods. Nevertheless, it is expected that the study of these novel forms of information will have its value in taxonomy, to determine which genes are shared, when genes or sets of genes were lost in evolutionary history, to detect the presence of horizontally transferred genes and/or confirm or enhance the phylogenetic signal derived from traditional methods. Although these conclusions are based on a relatively small data set, they are largely in agreement with other studies and it is anticipated that similar trends will be observed when comparing other genomes.

Bacterial Proteins↗

Synonymous genetic polymorphisms within Brazilian human immunodeficiency virus Type 1 subtypes may influence mutational routes to drug resistance.

BACKGROUND: Most published data on antiretroviral-drug resistance is generated from in vitro or in vivo studies of subtype B virus. However, this subtype is associated with <10% of HIV infections worldwide, and it is essential to explore subtype-specific determinants of drug resistance. One potential cause of the differences between subtypes is the synonymous codon usage at key resistance positions. METHODS: We investigated the nucleotide sequences at drug resistance-related sites, for all major Brazilian subtypes (B, C, and F1) of human immunodeficiency virus type 1 (HIV-1) group M. RESULTS: We identified a change at positions 151 and 210 of the reverse-transcriptase region in subtype F1, such that the emergence of these key nucleoside/nucleotide analogue resistance mutations required an extra nucleotide change in subtype F1, compared with subtypes B and C. The clinical significance of position 210 was confirmed within a large Brazilian database, in which we identified a lower prevalence of the L210W mutation in subtype F1 virus, compared with subtype B virus, in patients matched for thymidine-analogue experience. An inverse relationship between the L210W and K70R mutations was also observed. CONCLUSIONS: The findings of the present study illustrate an important mechanism by which a subtype may determine genetic routes to resistance, with implications for treatment strategies for populations infected with HIV-1 subtype F.

Base Sequence↗

Complete nucleotide sequence and gene rearrangement of the mitochondrial genome of the Japanese pond frog Rana nigromaculata.

In this study, we determined the complete nucleotide sequence of the mitochondrial genome of the Japanese pond frog Rana nigromaculata. The length of the sequence of the frog was 17,804 bp, though this was not absolute due to length variation caused by differing numbers of repetitive units in the control regions of individual frogs. The gene content, base composition, and codon usage of the Japanese pond frog conformed to those of typical vertebrate patterns. However, the comparison of gene organization between three amphibian species (Rana, Xenopus and caecilian) provided evidence that the gene arrangement of Rana differs by four tRNA gene positions from that of Xenopus or caecilian, a common gene arrangement in vertebrates. These gene rearrangements are presumed to have occurred by the tandem duplication of a gene region followed by multiple deletions of redundant genes. It is probable that the rearrangements start and end at tRNA genes involved in the initial production of a tandemly duplicated gene region. Putative secondary structures for the 22 tRNAs and the origin of the L-strand replication (OL) are described. Evolutionary relationships were estimated from the concatenated sequences of the 12 proteins encoded in the H-strand of mtDNA among 37 vertebrate species. A quartet-puzzling tree showed that three amphibian species form a monophyletic clade and that the caecilian is a sister group of the monophyletic Anura.

Amino Acid Sequence↗

Cloning, sequence and expression of a beta-tubulin-encoding gene in the homobasidiomycete Schizophyllum commune.

The beta-tubulin (beta Tub)-encoding gene (tub-2) of Schizophyllum commune is the first tubulin gene isolated, cloned and sequenced from higher filamentous fungi (homobasidiomycetes). The S. commune tub-2 gene is organized into nine exons and eight introns. The introns vary from 48 to 107 nt in length, and are distributed throughout the gene. The tub-2 exons code for a protein of 445 amino acids (aa), which shows great homology with beta Tubs of filamentous ascomycetes, plants, and animals, but less homology with yeasts. The codon usage of tub-2 from S. commune is biased, as it is in most beta Tub-encoding genes of filamentous fungi. The S. commune beta Tub shows a conserved aa sequence in the C-terminal domain, which is suggested to interact with microtubule-associated proteins in animals. In contrast, the S. commune beta Tub deviates from most known beta Tubs by having a Cys165 residue, which might be significant for the insensitivity of S. commune haploid strains to the antimicrotubule drug, benomyl. In tub-2 of different haploid strains, sequence polymorphisms occur in the 5' and 3' flanking regions. The expression of tub-2 is high in young mycelium, which has a high number of extending apical cells, but decreases with the aging of the mycelium. No significant difference in the hybridization signal intensity for the tub-2 transcripts was recorded either during intercellular nuclear migration at early mating, or in mycelia with a mutation in the B mating-type gene.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Evolution of 4-coumarate:coenzyme A ligase (4CL) gene and divergence of Larix (Pinaceae).

The evolutionary dynamics of the 4CL gene encoding 4-coumarate:coenzyme A ligase was investigated in the genus Larix (Pinaceae) by comparing copy number, GC content and codon usage, sequence divergence, and phylogenetic analysis. All 4CL clones of Larix formed a strongly supported monophyletic group, in which two robust clades (4clA and 4clB) derived from an ancient gene duplication event in the common ancestor of Larix were identified. Further gene duplication in the 4clA clade gave rise to two subclades 4clA(1) and 4clA(2). Frequent duplication/deletion appears to be a common evolutionary phenomenon in the 4CL gene family and paralogous genes differ greatly in their evolution rate. The existence of L. speciosa in subclades 4clA(1) and 4clA(2) suggests that this species may represent a primitive form of Larix or the closest relative of the common ancestor of the Eurasian Sect. Multiserialis. In addition, cpDNA and nrDNA ITS analyses support the hypothesis of an early separation of Larix into a North American and a Eurasian clade, which is congruent with the results of previous allozyme and very recent AFLP analyses. The unexpected close relationship between North American larches and the short-bracted species L. gmelinii in East Asia, based on the 4CL gene tree, may stem from lineage sorting.

Base Composition↗

Comparison of the amino acid sequences of the transacylase components of branched chain oxoacid dehydrogenase of Pseudomonas putida, and the pyruvate and 2-oxoglutarate dehydrogenases of Escherichia coli.

The nucleotide sequence of bkdB, the structural gene for E2b, the transacylase component of branched-chain-oxoacid dehydrogenase of Pseudomonas putida has been determined and translated into its amino acid sequence. The start of bkdB was identified from the N-terminal sequence of E2b isolated from branched-chain-oxoacid dehydrogenase of the closely related species, P. aeruginosa. The reading frame was composed of 65.5% G + C with 82.3% of the codons ending in G or C. There was no intergenic space between bkdA2 and bkdB. No codons requiring minor tRNAs were utilized and the codon bias index indicated a preferential codon usage. The bkdB gene encoded 423 amino acids although the N-terminal methionine was absent from E2b prepared from P. aeruginosa. The relative molecular mass of the encoded protein was 45,134 (45,003 minus methionine) vs 47,000 obtained by SDS/polyacrylamide gel electrophoresis. There was a single lipoyl domain in E2b compared to three lipoyl domains in E2p, and one domain in E2o, the transacylases of pyruvate and 2-oxoglutarate dehydrogenases of Escherichia coli respectively. There was significant similarity between the lipoyl domain of E2b and of E2p and E2o as well as between the E1-E2 binding domains of E2b, E2p and E2o. There was no similarity between the E3 binding domain of E2b to E2p and E2o which may reflect the uniqueness of the E3 component of branched-chain-oxoacid dehydrogenase of P. putida. The conclusions drawn from these comparisons are that the transacylases of prokaryotic pyruvate, 2-oxoglutarate and branched-chain-oxoacid dehydrogenases descended from a common ancestral protein probably at about the same time.

3-Methyl-2-Oxobutanoate Dehydrogenase (Lipoamide)↗

The histone H1 genes of the dipteran insect, Chironomus thummi, fall under two divergent classes and encode proteins with distinct intranuclear distribution and potentially different functions.

Four histone H1 genes of the midge, Chironomus thummi piger, and three H1 genes of the subspecies C. thummi thummi have been cloned and assigned to the four different H1 proteins from C. thummi larvae. Together with an earlier cloned H1 gene from C. thummi thummi [Hankeln, T. & Schmidt, E. R. (1991) Chromosoma 101, 25-31], these genes probably constitute the complete complement of H1 genes in both subspecies. They were found to fall under two classes that differ remarkably in their gene copy numbers, genomic organization, structure of flanking sequences, codon usage, and expression during embryonic development, and that encode H1 proteins of divergent structure. Histone H1 I-1 contains an inserted sequence, KAPKAPKAPKSPKAE in C. thummi piger, and KAPKAPKSPKAE in C. thummi thummi, that is lacking in the other H1 variants, H1 II-1, H1 II-2, and H1 III-1. In the immediate neighbourhood to the inserted sequence, a substitution in the H1 I-1 protein sequence dramatically enhances the potential to form a reversed turn. In early development, H1 I-1 is expressed at a higher rate than the other H1 genes. The transcripts have a size of about 1 kb; in addition, the H1 I-1 gene exhibited two minor transcripts of about 2.5 and > 3 kb size in middle blastoderm that are possibly polyadenylated. Together with our earlier finding that histone H1 I-1 is found in a limited number of polytene chromosome bands whereas the other H1 histones are uniformly distributed in chromatin, these results intimate functional differences between the two classes of H1 genes and their products.

Amino Acid Sequence↗

Polyhedrin structure.

Polyhedrin has evolved two highly specialized functions. Firstly, it forms a protective crystal around the virus and secondly it resists solubilization except under strongly alkaline conditions similar to those found in the insect midgut. Both of these properties allow the virus to remain viable for many years outside the insect host. Although polyhedrin and granulin can vary by about 50% in amino acid sequence, many of their structural features are highly conserved, reflecting the similar function and biochemical properties of these proteins. By comparing the sequences, domains within the gene which evolve rapidly have been identified and molecular phylogenies have been proposed. Information on predicted secondary structure has also been obtained and some insight into the possible role of codon usage in baculovirus function has also been gained from the sequence information. In addition to the conserved structural properties of the polyhedrin protein, there is a conserved regulatory process which results in the synthesis of massive amounts of polyhedrin. This process is probably governed by a virus-specific RNA polymerase. A potential regulatory signal shared by all these genes has been identified upstream from the polyhedrin gene. A number of additional granulin and lepidopteran polyhedrin sequences will certainly be forthcoming because of the ease with which these genes are identified by cross-hybridization with available related probes. However, of special interest will be sequences from dipteran and hymenopteran polyhedrins which will add greatly to our understanding of the constraints governing polyhedrin structure and diversity. Another logical step in their study will be to examine polyhedrin quaternary structure utilizing X-ray crystallography. Additional areas of future emphasis will probably focus on the regulation of polyhedrin synthesis. Elucidation of the regulatory signals governing transcription of these genes are of prime interest as are complementary studies on the characterization of the RNA polymerase which transcribes these genes.

Amino Acid Sequence↗

Comparison of two human ovarian carcinoma cell lines (A2780/CP70 and MCAS) that are equally resistant to platinum, but differ at codon 118 of the ERCC1 gene.

ERCC1 is an essential gene within the nucleotide excision repair process. We studied two human ovarian carcinoma cell lines for cisplatin resistance, which differed with respect to ERCC1. The A2780/CP70 cell line has been extensively studied previously, and has the wild-type ERCC1 sequence. The MCAS cell line has a recently described ERCC1 polymorphism at codon 118, which is associated with an approximate 50% reduction in codon usage. These cells did not differ with respect to p53 sequence nor p53 mRNA induction following cisplatin exposure. The induction of ERCC1 mRNA was markedly reduced in MCAS cells as compared to A2780/CP70 cells. At the IC50 cisplatin dose for each cell line, MCAS cells were less proficient at cisplatin-DNA adduct repair than A2780/CP70 cells. In absolute terms, A2780/CP70 cells repaired 3-fold as much adduct (2.7 pg/microgram DNA over 6 h vs 0.86 pg/microgram DNA); and when expressed in terms of the maximal DNA adduct load, A2780/CP70 cells repaired 50% more adduct than MCAS cells. MCAS cells had increased cytosolic inactivation of drug at the IC50 dose level, which has been previously suggested to be a compensatory cellular response for reduced DNA repair capacity. These data suggest the possibility that this specific ERCC1 polymorphism, may be associated with reduced DNA repair capacity in human ovarian cancer cells. This association may be effected through a reduction in peak production of ERCC1 mRNA, and a consequent reduction in the translation of ERCC1 mRNA into protein.

Antineoplastic Agents↗

Partition of unit-copy miniplasmids to daughter cells. III. The DNA sequence and functional organization of the P1 partition region.

The boundaries of the P1 par (plasmid partition) region of the unit-copy plasmid P1 were defined to within 2.7 X 10(3) base-pairs of DNA. The DNA sequence of the region revealed two large open reading frames that could encode proteins of Mr 44,000 and Mr 38,000. Both would be read in the same direction. The first open reading frame corresponds to the par A gene, the Mr 44,000 protein product of which was shown to be trans acting and essential for partition. The second open reading frame (parB) follows closely and may be cotranscribed with par A. The codon usage frequency for parB is consistent with its producing a protein product. The ParB protein was identified in cell extracts as a product with an apparent Mr of 45,000, suggesting that it behaves anomolously on gel electrophoresis. Following parB is the incB region, an incompatibility determinant thought to be the cis acting site that constitutes the putative attachment point on the DNA for the cellular partition apparatus. Subcloning of this site showed it to consist of a maximum of 174 base-pairs. The incB sequence is highly A + T-rich and contains a 20 base-pair inverted repeat. Another A + T-rich inverted repeat of similar size but different sequence is found between the putative parA promoter and the ribosome initiation sequence at the start of the parA open reading frame and may be involved in the autoregulation of ParA synthesis. The par region appears to contain a functional analog of the centromere of eukaryotic chromosomes. It is responsible for ensuring that newly replicated plasmids are properly distributed to daughter cells during cell division of its Escherichia coli host.

Bacterial Proteins↗

Toward more efficient protein expression: keep the message simple.

Optimization of gene coding-sequence, including preferred codon usage and removal of cryptic splice sites and mRNA-destabilizing motifs, has been shown to improve recombinant protein production of different proteins. Here, we present data to show that gene optimization can also be used to improve the production of a complex macromolecule, namely an antibody. When applied to the heavy and light chain genes of our model antibody, we found that greater numbers of high-producing transfectants as well as increased levels of protein production were observed (approximately 1.5-fold). In this test model, production was improved even though the antibody has previously been demonstrated to give high expression in stably transfected cells (up to 5 g/L in bioreactors). Because the parental heavy chain sequence contained introns, and the process of gene optimization is most efficiently performed on sequences without introns, we demonstrated that removal of introns in the coding sequence had no effect on the quantity of antibody produced. All constructs were evaluated using Lonza's glutamine synthetase gene expression vectors in Chinese hamster ovary cells. Our findings suggest that significant improvements in product yields can be achieved by gene optimization, which may facilitate the processing and translation of gene transcripts.

Animals↗

[Molecular evolution of MHC DQA genes. II. Phylogenetic analysis based on nucleotide substitution and SCU bias].

Phylogenetics of 23 alleles at MHC DQA loci in 7 mammalian species was studied based on their nucleotide (NT) substitution and synonymous codon usage (SCU) bias. (1) It was demonstrated that the NT substitution rates are 1.0 x 10(-9) NT/site/yr for exon2 and 1.3 x 10(-9) NT/site/yr for exon2-4 in a large time scale, which is similar to other nuclear genes, while for mouse and rat the rates are nearly twice as high as above mentioned. (2) The DQA locus diversity and their interallelic diversity developed long after the radiation of mammalian 80Mya (million years ago). The bovine counterpart, of, and with the same recent ancestor of ovine DQA2, remains to be discovered. HLA-DQA2 locus split from HLA-DQA1 ancestor at the time between 12 approximately 20 Mya while allele diversity of HLA-DQA1 emerged and developed from 24 Mya to less than 1 Mya. (3) The phylogenetic trees based on SCU divergence reflect the phylogenetics of MHC DQA genes quite well generally in a new respect and reveal that HLA-DQA2 has a distinctive SCU bias different from all other MHC DQA locianalyzed. It indicates that SCU statistics plays an important and unique role in phylogenetic analysis of orthologous genes. The method to estimate the SCU divergence and SCU similarity was improved in this research.

Animals↗

Variation suggestive of horizontal gene transfer at a lipopolysaccharide (lps) biosynthetic locus in Xanthomonas oryzae pv. oryzae, the bacterial leaf blight pathogen of rice.

BACKGROUND: In animal pathogenic bacteria, horizontal gene transfer events (HGT) have been frequently observed in genomic regions that encode functions involved in biosynthesis of the outer membrane located lipopolysaccharide (LPS). As a result, different strains of the same pathogen can have substantially different lps biosynthetic gene clusters. Since LPS is highly antigenic, the variation at lps loci is attributed to be of advantage in evading the host immune system. Although LPS has been suggested as a potentiator of plant defense responses, interstrain variation at lps biosynthetic gene clusters has not been reported for any plant pathogenic bacterium. RESULTS: We report here the complete sequence of a 12.2 kb virulence locus of Xanthomonas oryzae pv. oryzae (Xoo) encoding six genes whose products are homologous to functions involved in LPS biosynthesis and transport. All six open reading frames (ORFs) have atypical G+C content and altered codon usage, which are the hallmarks of genomic islands that are acquired by horizontal gene transfer. The lps locus is flanked by highly conserved genes, metB and etfA, respectively encoding cystathionine gamma lyase and electron transport flavoprotein. Interestingly, two different sets of lps genes are present at this locus in the plant pathogens, Xanthomonas campestris pv. campestris (Xcc) and Xanthomonas axonopodis pv. citri (Xac). The genomic island is present in a number of Xoo strains from India and other Asian countries but is not present in two strains, one from India (BXO8) and another from Nepal (Nepal624) as well as the closely related rice pathogen, Xanthomonas oryzae pv. oryzicola (Xoor). TAIL-PCR analysis indicates that sequences related to Xac are present at the lps locus in both BXO8 and Nepal624. The Xoor strain has a hybrid lps gene cluster, with sequences at the metB and etfA ends, being most closely related to sequences from Xac and the tomato pathogen, Pseudomonas syringae pv. tomato respectively. CONCLUSION: This is the first report of hypervariation at an lps locus between different strains of a plant pathogenic bacterium. Our results indicate that multiple HGT events have occurred at this locus in the xanthomonad group of plant pathogens.

Base Sequence↗

Incorporating a TEV cleavage site reduces the solubility of nine recombinant mouse proteins.

Failure to express soluble proteins in bacteria is mainly attributed to the properties of the target protein itself, as well as the choice of the vector, the purification tag and the linker between the tag and protein, and codon usage. The expression of proteins with fusion tags to facilitate subsequent purification steps is a widely used procedure in the production of recombinant proteins. However, the additional residues can affect the properties of the protein; therefore, it is often desirable to remove the tag after purification. This is usually done by engineering a cleavage site between the tag and the encoded protein that is recognised by a site-specific protease, such as the one from tobacco etch virus (TEV). In this study, we investigated the effect of four different tags on the bacterial expression and solubility of nine mouse proteins. Two of the four engineered constructs contained hexahistidine tags with either a long or short linker. The other two constructs contained a TEV cleavage site engineered into the linker region. Our data show that inclusion of the TEV recognition site directly downstream of the recombination site of the Invitrogen Gateway vector resulted in a loss of solubility of the nine mouse proteins. Our work suggests that one needs to be very careful when making modifications to expression vectors and combining different affinity and fusion tags and cleavage sites.

Amino Acid Sequence↗

The mitochondrial genomes of the human hookworms, Ancylostoma duodenale and Necator americanus (Nematoda: Secernentea).

The complete mitochondrial genome sequences were determined for two species of human hookworms, Ancylostoma duodenale (13,721 bp) and Necator americanus (13,604 bp). The circular hookworm genomes are amongst the smallest reported to date for any metazoan organism. Their relatively small size relates mainly to a reduced length in the AT-rich region. Both hookworm genomes encode 12 protein, two ribosomal RNA and 22 transfer RNA genes, but lack the ATP synthetase subunit 8 gene, which is consistent with three other species of Secernentea studied to date. All genes are transcribed in the same direction and have a nucleotide composition high in A and T, but low in G and C. The AT bias had a significant effect on both the codon usage pattern and amino acid composition of proteins. For both hookworm species, genes were arranged in the same order as for Caenorhabditis elegans, except for the presence of a non-coding region between genes nad3 and nad5. In A. duodenale, this non-coding region is predicted to form a stem-and-loop structure which is not present in N. americanus. The mitochondrial genome structure for both hookworms differs from Ascaris suum only in the location of the AT-rich region, whereas there are substantial differences when compared with Onchocerca volvulus, including four gene or gene-block translocations and the positions of some transfer RNA genes and the AT-rich region. Based on genome organisation and amino acid sequence identity, A. duodenale and N. americanus were more closely related to C. elegans than to A. suum or O. volvulus (all secernentean nematodes), consistent with a previous phylogenetic study using ribosomal DNA sequence data. Determination of the complete mitochondrial genome sequences for two human hookworms (the first members of the order Strongylida ever sequenced) provides a foundation for studying the systematics, population genetics and ecology of these and other nematodes of socio-economic importance.

Amino Acid Sequence↗

Structure and organization of the mitochondrial genome of the canine heartworm, Dirofilaria immitis.

This study determined the complete mitochondrial (mt) genome sequence of the canine heartworm, Dirofilaria immitis, and compared its structure, organization and other characteristics with Onchocerca volvulus and other secernentean nematodes. The D. immitis mt genome is 13814 bp in size and contains 36 of the 37 genes typical of metazoan organisms, and lacks the ATP synthetase subunit 8 gene. All of the genes are transcribed in the same direction. For the entire genome, the nucleotide contents are approximately 55% (T), approximately 19% (each for A and G) and approximately 7% (C), which is very similar to those of the protein-coding genes. In the latter genes, most (approximately 69%) third codon positions have a T, but rarely (approximately 1-9%) have an A or a C. The C content (8-12%) is higher at the first and second codon positions compared with the third position (approximately 1%). These nucleotide biases have a significant effect on the codon usage patterns and, thus, on the amino acid composition of the proteins. The mt genome organization of D. immitis is essentially the same as that of O. volvulus, but is distinctly different from other secernentean nematodes sequenced thus far. Irrespective of transpositions of transfer RNA (trn) genes and the non-coding, AT-rich region, there are 4 gene- or gene block-translocations between the mt genome of D. immitis and those of Caenorhabditis elegans, Ascaris suum and the 2 human hookworms, Ancylostoma duodenale and Necator americanus. For D. immitis, the 22 trn genes have secondary structures typical of other secernentean nematodes, and possess a TV-replacement loop instead of a TpsiC arm and loop. Like O. volvulus, the mt trnK and trnP of D. immitis use the anticodons CUU and AGG, whereas in other nematodes, UUU and UGG are employed, respectively. Also, the secondary structures of the 2 ribosomal RNA (rrn) genes are similar to the models for other nematodes. Overall, the availability of the complete D. immitis mt genome sequence provides a resource for future studies of the comparative mt genomics and of the population genetics and/or phylogeny of parasitic nematodes.

Amino Acid Sequence↗

Score-based prediction of genomic islands in prokaryotic genomes using hidden Markov models.

BACKGROUND: Horizontal gene transfer (HGT) is considered a strong evolutionary force shaping the content of microbial genomes in a substantial manner. It is the difference in speed enabling the rapid adaptation to changing environmental demands that distinguishes HGT from gene genesis, duplications or mutations. For a precise characterization, algorithms are needed that identify transfer events with high reliability. Frequently, the transferred pieces of DNA have a considerable length, comprise several genes and are called genomic islands (GIs) or more specifically pathogenicity or symbiotic islands. RESULTS: We have implemented the program SIGI-HMM that predicts GIs and the putative donor of each individual alien gene. It is based on the analysis of codon usage (CU) of each individual gene of a genome under study. CU of each gene is compared against a carefully selected set of CU tables representing microbial donors or highly expressed genes. Multiple tests are used to identify putatively alien genes, to predict putative donors and to mask putatively highly expressed genes. Thus, we determine the states and emission probabilities of an inhomogeneous hidden Markov model working on gene level. For the transition probabilities, we draw upon classical test theory with the intention of integrating a sensitivity controller in a consistent manner. SIGI-HMM was written in JAVA and is publicly available. It accepts as input any file created according to the EMBL-format.It generates output in the common GFF format readable for genome browsers. Benchmark tests showed that the output of SIGI-HMM is in agreement with known findings. Its predictions were both consistent with annotated GIs and with predictions generated by different methods. CONCLUSION: SIGI-HMM is a sensitive tool for the identification of GIs in microbial genomes. It allows to interactively analyze genomes in detail and to generate or to test hypotheses about the origin of acquired genes.

Algorithms↗

An aromatic-dependent mutant of the fish pathogen Aeromonas salmonicida is attenuated in fish and is effective as a live vaccine against the salmonid disease furunculosis.

Aeromonas salmonicida is the etiological agent of furunculosis in salmonid fish. The disease is responsible for severe economic losses in intensively cultured salmon and trout. Bacterin vaccines provide inadequate protection against infection. We have constructed an aromatic-dependent mutant of A. salmonicida in order to investigate the possibility of an effective live-attenuated vaccine. The aroA gene of A. salmonicida was cloned in Escherichia coli, and the nucleotide sequence was determined. The codon usage pattern of aroA was found to be quite distinct from that of the vapA gene coding for the surface array protein layer (A layer). The aroA gene was inactivated by inserting a fragment expressing kanamycin resistance within the coding sequence. The aroA::Kar mutation was introduced into the chromosome of virulent A. salmonicida 644Rb and 640V2 by allele replacement by using a suicide plasmid delivery system. The aroA mutation did not revert at a detectable frequency (< 10(-11). The mutation resulted in attenuation when bacteria were injected intramuscularly into Atlantic salmon (Salmo salar L.). Introduction of the wild-type aroA gene into the A. salmonicida mutants on a broad-host-range plasmid restored virulence. A. salmonicida mutant 644Rb aroA::Kar persisted in the kidney of brown trout (Salmo trutta L.) for 12 days at 10 degrees C. Vaccination of brown trout with 10(7) CFU of A. salmonicida 644Rb aroA by intraperitoneal injection resulted in a 253-fold increase in the 50% lethal dose (LD50) compared with unvaccinated controls challenged with a virulent clinical isolate 9 weeks later. A second vaccination after 6 weeks increased the LD50 by a further 16-fold.

4-Aminobenzoic Acid↗