Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “convergent evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Analysis of HLA class II haplotypes in the Cayapa Indians of Ecuador: a novel DRB1 allele reveals evidence for convergent evolution and balancing selection at position 86.

PCR amplification, oligonucleotide probe typing, and sequencing were used to analyze the HLA class II loci (DRB1, DQA1, DQB1, and DPB1) of an isolated South Amerindian tribe. Here we report HLA class II variation, including the identification of a new DRB1 allele, several novel DR/DQ haplotypes, and an unusual distribution of DPB1 alleles, among the Cayapa Indians (N = 100) of Ecuador. A general reduction of HLA class II allelic variation in the Cayapa is consistent with a population bottle-neck during the colonization of the Americas. The new Cayapa DRB1 allele, DRB1*08042, which arose by a G-->T point mutation in the parental DRB1*0802, contains a novel Val codon (GTT) at position 86. The generation of DRB1*08042 (Val-86) from DRB1*0802 (Gly-86) in the Cayapa, by a different mechanism than the (GT-->TG) change in the creation of DRB1*08041 (Val-86) from DRB1*0802 in Africa, implicates selection in the convergent evolution of position 86 DR beta variants. The DRB1*08042 allele has not been found in > 1,800 Amerindian haplotypes and thus presumably arose after the Cayapa separated from other South American Amerindians. Selection pressure for increased haplotype diversity can be inferred in the generation and maintenance of three new DRB1*08042 haplotypes and several novel DR/DQ haplotypes in this population. The DPB1 allelic distribution in the Cayapa is also extraordinary, with two alleles, DPB1*1401, a very rare allele in North American Amerindian populations, and DPB1*0402, the most common Amerindian DPB1 allele, constituting 89% of the Cayapa DPB1.(ABSTRACT TRUNCATED AT 250 WORDS)

Alleles↗

Genetic flexibility in the convergent evolution of hermaphroditism in Caenorhabditis nematodes.

The self-fertile hermaphrodites of C. elegans and C. briggsae evolved from female ancestors by acquiring limited spermatogenesis. Initiation of C. elegans hermaphrodite spermatogenesis requires germline translational repression of the female-promoting gene tra-2, which allows derepression of the three male-promoting fem genes. Cessation of hermaphrodite spermatogenesis requires fem-3 translational repression. We show that C. briggsae requires neither fem-2 nor fem-3 for hermaphrodite development, and that XO Cb-fem-2/3 animals are transformed into hermaphrodites, not females as in C. elegans. Exhaustive screens for Cb-tra-2 suppressors identified another 75 fem-like mutants, but all are self-fertile hermaphrodites rather than females. Control of hermaphrodite spermatogenesis therefore acts downstream of the fem genes in C. briggsae. The outwardly similar hermaphrodites of C. elegans and C. briggsae thus achieve self-fertility via intervention at different points in the core sex determination pathway. These findings are consistent with convergent evolution of hermaphroditism, which is marked by considerable developmental genetic flexibility.

Animals↗

Threonine aldolase and alanine racemase: novel examples of convergent evolution in the superfamily of vitamin B6-dependent enzymes.

Vitamin B(6)-dependent enzymes may be grouped into five evolutionarily unrelated families, each having a different fold. Within fold type I enzymes, L-threonine aldolase (L-TA) and fungal alanine racemase (AlaRac) belong to a subgroup of structurally and mechanistically closely related proteins, which specialised during evolution to perform different functions. In a previous study, a comparison of the catalytic properties and active site structures of these enzymes suggested that they have a catalytic apparatus with the same basic features. Recently, recombinant D-threonine aldolases (D-TAs) from two bacterial organisms have been characterised, their predicted amino acid sequences showing no significant similarities to any of the known B(6) enzymes. In the present work, a comparative structural analysis suggests that D-TA has an alpha/beta barrel fold and therefore is a fold type III B(6) enzyme, as eukaryotic ornithine decarboxylase (ODC) and bacterial AlaRac. The presence of both TA and AlaRac in two distinct evolutionary unrelated families represents a novel and interesting example of convergent evolution. The independent emergence of the same catalytic properties in families characterised by completely different folds may have not been determined by chance, but by the similar structural features required to catalyse pyridoxal phosphate-dependent aldolase and racemase reactions.

Alanine Racemase↗

Phylogenetic analysis of triple gene block viruses based on the TGB 1 homolog gene indicates a convergent evolution.

The complete nucleotide sequence of the triple gene block one (TGB 1) of cymbidium mosaic potexvirus (CymMV) was compared to those from other potex-, carla-, furo- and hordeiviruses. Seven conserved motifs in the TGB 1, including the ATP-GTP binding domain (P-Loop) consensus GXXGKTSTS, were found in all four virus genera. We propose that all TGBV can be classified into phylogenetic clusters based on their TGB 1 homolog genes. These clusters can be further delineated to form subgroups. The first cluster comprises the potexviruses which are further subdivided into three subgroups; BaMV, FMV, PlaMV and PapMV (subgroup Ia); CymMV, PAMV, NMV, SMYEaV and WC1MV (subgroup Ib) and PVX (subgroup Ic). The second cluster comprises carlaviruses with a dual subgrouping; CVB, LSV, PVM, PMV and ASPV (subgroup IIa) and LVX (subgroup IIb). The third cluster carries the most diverse of TGBV comprising furoviruses PCV, PMTV and BSBV (subgroup IIIa) and hordeiviruses PSLV, BSMV and LRSV (subgroup IIIb). The phylogenetic relationships of triple gene block viruses (TGBV) based on the TGB 1 homolog gene indicates a convergent evolution.

Amino Acid Sequence↗

A case of convergent evolution of nucleic acid binding modules.

Divergent evolution can explain how many proteins containing structurally similar domains, which perform a variety of related functions, have evolved from a relatively small number of modules or protein domains. However, it cannot explain how protein domains with similar, but distinguishable, functions and similar, but distinguishable, structures have evolved. Examples of this are the RNA-binding protein containing the RNA-binding domain (RBD), and a newly established protein group, the cold-shock domain (CSD) protein family. Both protein domains contain conserved RNP motifs on similar single-stranded nucleic acid-binding surfaces. Apart from the RNP motifs, which have a similar function, the two families show little similarity in topology or amino acid sequence. This can be considered an interesting example of convergent evolution at the molecular level. Previously, a beta-sheet surface was found to interact with RNA in non-homologous proteins from yeast, phage and man, revealing that this mode of RNA binding may be a widely recurring theme.

Amino Acid Sequence↗

Multiple independent origins of Shigella clones of Escherichia coli and convergent evolution of many of their characteristics.

The evolutionary relationships of 46 Shigella strains representing each of the serotypes belonging to the four traditional Shigella species (subgroups), Dysenteriae, Flexneri, Boydii, and Sonnei, were determined by sequencing of eight housekeeping genes in four regions of the chromosome. Analysis revealed a very similar evolutionary pattern for each region. Three clusters of strains were identified, each including strains from different subgroups. Cluster 1 contains the majority of Boydii and Dysenteriae strains (B1-4, B6, B8, B10, B14, and B18; and D3-7, D9, and D11-13) plus Flexneri 6 and 6A. Cluster 2 contains seven Boydii strains (B5, B7, B9, B11, B15, B16, and B17) and Dysenteriae 2. Cluster 3 contains one Boydii strain (B12) and the Flexneri serotypes 1-5 strains. Sonnei and three Dysenteriae strains (D1, D8, and D10) are outside of the three main clusters but, nonetheless, are clearly within Escherichia coli. Boydii 13 was found to be distantly related to E. coli. Shigella strains, like the other pathogenic forms of E. coli, do not have a single evolutionary origin, indicating convergent evolution of Shigella phenotypic properties. We estimate the three main Shigella clusters to have evolved within the last 35,000 to 270,000 years, suggesting that shigellosis was one of the early infectious diseases of humans.

Base Sequence↗

Convergent evolution of perenniality in rice and sorghum.

Annual and perennial habit are two major strategies by which grasses adapt to seasonal environmental change, and these distinguish cultivated cereals from their wild relatives. Rhizomatousness, a key trait contributing to perenniality, was investigated by using an F(2) population from a cross between cultivated rice (Oryza sativa) and its wild relative, Oryza longistaminata. Molecular mapping based on a complete simple sequence-repeat map revealed two dominant-complementary genes controlling rhizomatousness. Rhz3 was mapped to the interval between markers OSR16 [1.3 centimorgans (cM)] and OSR13 (8.1 cM) on rice chromosome 4 and Rhz2 located between RM119 (2.2 cM) and RM273 (7.4 cM) on chromosome 3. Comparative mapping indicated that each gene closely corresponds to major quantitative trait loci (QTLs) controlling rhizomatousness in Sorghum propinquum, a wild relative of cultivated sorghum. Correspondence of these genes in rice and sorghum, which diverged from a common ancestor approximately 50 million years ago, suggests that the two genes may be key regulators of rhizome development in many Poaceae. Many additional QTLs affecting abundance of rhizomes in O. longistaminata were identified, most of which also corresponded to the locations of S. propinquum QTLs. Convergent evolution of independent mutations at, in some cases, corresponding genes may have been responsible for the evolution of annual cereals from perennial wild grasses. DNA markers closely linked to Rhz2 and Rhz3 will facilitate cloning of the genes, which may contribute significantly to our understanding of grass evolution, advance opportunities to develop perennial cereals, and offer insights into environmentally benign weed-control strategies.

Biological Evolution↗

The structure of Leishmania mexicana ICP provides evidence for convergent evolution of cysteine peptidase inhibitors.

Clan CA, family C1 cysteine peptidases (CPs) are important virulence factors and drug targets in parasites that cause neglected diseases. Natural CP inhibitors of the I42 family, known as ICP, occur in some protozoa and bacterial pathogens but are absent from metazoa. They are active against both parasite and mammalian CPs, despite having no sequence similarity with other classes of CP inhibitor. Recent data suggest that Leishmania mexicana ICP plays an important role in host-parasite interactions. We have now solved the structure of ICP from L. mexicana by NMR and shown that it adopts a type of immunoglobulin-like fold not previously reported in lower eukaryotes or bacteria. The structure places three loops containing highly conserved residues at one end of the molecule, one loop being highly mobile. Interaction studies with CPs confirm the importance of these loops for the interaction between ICP and CPs and suggest the mechanism of inhibition. Structure-guided mutagenesis of ICP has revealed that residues in the mobile loop are critical for CP inhibition. Data-driven docking models support the importance of the loops in the ICP-CP interaction. This study provides structural evidence for the convergent evolution from an immunoglobulin fold of CP inhibitors with a cystatin-like mechanism.

Amino Acid Sequence↗

A hypothesis for the HLA-B27 immune dysregulation in spondyloarthropathy: contributions from enteric organisms, B27 structure, peptides bound by B27, and convergent evolution.

Several human rheumatic diseases occur predominantly in persons who carry the histocompatibility (HLA) class I allele B27. They have also been related to Gram-negative enteric microorganisms. In addition, the recent recovery of peptides bound to B27 has allowed an understanding of the structural requirements for their binding. Using the accumulated data base of protein sequences, we have tested a series of hypotheses. First, we have asked whether the primary amino acid sequence of the hypervariable regions of HLA-B27 shares short sequences with the proteins of Gram-negative enteric bacteria. The data demonstrate that, unique among the HLA-B molecules, the hypervariable regions of HLA-B27 unexpectedly share short peptide sequences with proteins from these bacteria. Second, we have asked whether the enteric proteins tend to satisfy the structural requirements for peptide binding to B27 in those regions of the sequence shared with B27. This hypothesis also tends to be true, especially in an allelically variable part of the B27 sequence which is predicted to bind B27 if it were to be presented as a free peptide. We conclude that HLA-B27 and enteric Gram-negative bacteria have undergone a previously unappreciated form of convergent evolution which may be important in the process leading to these rheumatic diseases. Moreover, the regions of the enteric bacterial proteins which are contiguous with the short sequences shared with B27 tend to have structures which are also predicted to bind B27. These observations suggest a mechanism for autoimmunity and lead to the prediction that the B27-associated diseases are mediated by a subset of T-cell receptors, B27, and the peptides bound by B27.

Amino Acid Sequence↗

The structure and organization of lamprin genes: multiple-copy genes with alternative splicing and convergent evolution with insect structural proteins.

Lamprin is a unique structural protein which forms the extracellular matrix of several cartilaginous structures found in the lamprey. Lamprin is noncollagenous in nature but shows sequence similarities to elastins and to insect structural proteins. Here, we characterize the structure and organization of lamprin genes, demonstrating the presence of multiple similar but not identical copies of the lamprin gene in the genome of the lamprey. In at least one species of lamprey, Lampetra richardsoni, the multiple gene copies are arranged in tandem in the genome in a head-to-tail orientation. Lamprin genes from Petromyzon marinus contain either seven or eight exons, with exon 4 being alternatively spliced in all genes, resulting in a total of six different lamprin transcripts. All exon junctions are of class 1,1. An unusual feature of the lamprin gene structure is the distribution of the 3' untranslated region sequence among multiple exons. A TATA box and cap sequence have been identified in upstream sequences in close proximity to the transcription start site, but no CAAT box could be identified. Sequence and gene structure comparisons between lamprins, elastins, and insect structural proteins suggest that the regions of sequence similarity are the result of a process of convergent evolution.

Alternative Splicing↗

Convergent evolution of similar enzymatic function on different protein folds: the hexokinase, ribokinase, and galactokinase families of sugar kinases.

Kinases that catalyze phosphorylation of sugars, called here sugar kinases, can be divided into at least three distinct nonhomologous families. The first is the hexokinase family, which contains many prokaryotic and eukaryotic sugar kinases with diverse specificities, including a new member, rhamnokinase from Salmonella typhimurium. The three-dimensional structure of hexokinase is known and can be used to build models of functionally important regions of other kinases in this family. The second is the ribokinase family, of unknown three-dimensional structure, and comprises pro- and eukaryotic ribokinases, bacterial fructokinases, the minor 6-phosphofructokinase 2 from Escherichia coli, 6-phosphotagatokinase, 1-phosphofructokinase, and, possibly, inosine-guanosine kinase. The third family, also of unknown three-dimensional structure, contains several bacterial and yeast galactokinases and eukaryotic mevalonate and phosphomevalonate kinases and may have a substrate binding region in common with homoserine kinases. Each of the three families of sugar kinases appears to have a distinct three-dimensional fold, since conserved sequence patterns are strikingly different for the three families. Yet each catalyzes chemically equivalent reactions on similar or identical substrates. The enzymatic function of sugar phosphorylation appears to have evolved independently on the three distinct structural frameworks, by convergent evolution. In addition, evolutionary trees reveal that (1) fructokinase specificity has evolved independently in both the hexokinase and ribokinase families and (2) glucose specificity has evolved independently in different branches of the hexokinase family. These are examples of independent Darwinian adaptation of a structure to the same substrate at different evolutionary times. The flexible combination of active sites and three-dimensional folds observed in nature can be exploited by protein engineers in designing and optimizing enzymatic function.

Amino Acid Sequence↗

Convergent evolution of receptors for protein import into mitochondria.

BACKGROUND: Mitochondria evolved from intracellular bacterial symbionts. Establishing mitochondria as organelles required a molecular machine to import proteins across the mitochondrial outer membrane. This machinery, the TOM complex, is composed of at least seven component parts, and its creation and evolution represented a sizeable challenge. Although there is good evidence that a core TOM complex, composed of three subunits, was established in the protomitochondria, we suggest that the receptor component of the TOM complex arose later in the evolution of this machine. RESULTS: We have solved by nuclear magnetic resonance the structure of the presequence binding receptor from the TOM complex of the plant Arabidopsis thaliana. The protein fold suggests that this protein, AtTom20, belongs to the tetratricopeptide repeat (TPR) superfamily, but it is unusual in that it contains insertions lengthening the helices of each TPR motif. Peptide titrations map the presequence binding site to a groove of the concave surface of the receptor. In vitro functional assays and peptide titrations suggest that the plant Tom20 is functionally equivalent to fungal and animal Tom20s. CONCLUSIONS: Comparison of the sequence and structure of Tom20 from plants and animals suggests that these two presequence binding receptors evolved from two distinct ancestral genes following the split of the animal and plant lineages. The need to bind equivalent mitochondrial targeting sequences and to make similar interactions within an equivalent protein translocation machine has driven the convergent evolution of two distinct proteins to a common structure and function.

Amino Acid Sequence↗

Convergent evolution of strigiform and caprimulgiform dark-activity is supported by phylogenetic analysis using the arylalkylamine N-acetyltransferase (Aanat) gene.

Alternative hypotheses propose the sister order of owls (Strigiformes) to be either day-active raptors (Falconiformes) or dark-active nightjars and allies (Caprimulgiformes). In an effort to identify molecular characters distinguishing between these hypotheses we examined a gene, arylalkylamine N-acetyltransferase (Aanat), potentially associated with the evolution of avian dark-activity. Partial Aanat coding sequences, and two introns, were obtained from the genomic DNA of 16 species: Strigiformes (four species), Falconiformes (four species), Caprimulgiformes (five species), with outgroups: Ciconiiformes (one species), Passeriformes (one species), and Apterygiformes (one species). Phylogenetic trees derived from aligned, evolutionarily conserved Aanat regions did not consistently recover clades corresponding to orders Strigiformes and Falconiformes but did place a caprimulgiform clade more distant from the strigiform and falconiform species than the latter two groups are to each other. This finding was supported by spectral analysis. The taxonomic distribution of seven intronic indels is consistent with the Aanat derived phylogenetic trees and supports conventional family-level groupings within both Strigiformes and Caprimulgiformes. The phylogenetic analyses also indicate that Caprimulgiformes is a polyphyletic grouping. In conclusion the data support, but do not conclusively prove, the proposal that Falconiformes is the sister order to Strigiformes and therefore, that the dark-activity characteristic of Strigiformes and Caprimulgiformes arose by convergent evolution.

Animals↗

Photoenergetics of octopus rhodopsin. Convergent evolution of biological photon counters?

The enthalpy changes associated with each of the major steps in the photoconversion of octopus rhodopsin have been measured by direct photocalorimetry. Formation of the primary photoproduct (bathorhodopsin) involves energy uptake of about 130 kJ/mol, corresponding to storage of over 50% of the exciting photon energy, and is comparable to the energy storage previously observed in bovine rhodopsin. Subsequent intermediates involve the step-wise dissipation of this energy to give the physiological end-product (acid metarhodopsin) at a level only slightly above the parent rhodopsin. No significant differences in energetics are observed between rhodopsin in microvilli membrane suspensions or detergent dispersions. Use of different buffer systems in the calorimetric experiments shows that conversion of rhodopsin to acid metarhodopsin involves no light-induced protonation change, whereas alkali metarhodopsin photoproduction occurs with the release of one proton per molecule and an additional enthalpy increase of about 50 kJ/mol. Van't Hoff analysis of the effect of temperature on the reversible metarhodopsin equilibrium gives an enthalpy for the acid----alkali transition consistent with this calorimetric result, and the proton release is confirmed by direct observation of light-induced pH changes. Acid-base titration of metarhodopsin yields an apparent pK of 9.5 for this transition, though the pH profile deviates slightly from ideal titration behaviour. We suggest that a high energy primary photoproduct is an obligatory feature of efficient biological photo-detectors, as opposed to photon energy transducers, and that the similarity at this stage between cephalopod and vertebrate rhodopsins represents either convergent evolution at the molecular level or strong conservation of a crucial functional characteristic.

Animals↗

Convergent evolution of antifreeze glycoproteins in Antarctic notothenioid fish and Arctic cod.

Antarctic notothenioid fishes and several northern cods are phylogenetically distant (in different orders and superorders), yet produce near-identical antifreeze glycoproteins (AFGPs) to survive in their respective freezing environments. AFGPs in both fishes are made as a family of discretely sized polymers composed of a simple glycotripeptide monomeric repeat. Characterizations of the AFGP genes from notothenioids and the Arctic cod show that their AFGPs are both encoded by a family of polyprotein genes, with each gene encoding multiple AFGP molecules linked in tandem by small cleavable spacers. Despite these apparent similarities, detailed analyses of the AFGP gene sequences and substructures provide strong evidence that AFGPs in these two polar fishes in fact evolved independently. First, although Antarctic notothenioid AFGP genes have been shown to originate from a pancreatic trypsinogen, Arctic cod AFGP genes share no sequence identity with the trypsinogen gene, indicating trypsinogen is not the progenitor. Second, the AFGP genes of the two fish have different intron-exon organizations and different spacer sequences and, thus, different processing of the polyprotein precursors, consistent with separate genomic origins. Third, the repetitive AFGP tripeptide (Thr-Ala/Pro-Ala) coding sequences are drastically different in the two groups of genes, suggesting that they arose from duplications of two distinct, short ancestral sequences with a different permutation of three codons for the same tripeptide. The molecular evidence for separate ancestry is supported by morphological, paleontological, and paleoclimatic evidence, which collectively indicate that these two polar fishes evolved their respective AFGPs separately and thus arrived at the same AFGPs through convergent evolution.

Amino Acid Sequence↗

Conservation of structure and cold-regulation of RNA-binding proteins in cyanobacteria: probable convergent evolution with eukaryotic glycine-rich RNA-binding proteins.

The rbp gene family of the cyanobacterium Anabaena variabilis strain M3 consists of eight members that encode small RNA-binding proteins containing a single RNA recognition motif (RRM). Similar genes are found in the genomes of Synechocystis sp. PCC6803, Helicobacter pylori and Treponema pallidum, but are absent from the other completely sequenced prokaryotic genomes. The expression of the rbp genes of Anabaena is induced by low temperature, with the exception of the rbpD gene. We found four stretches of conserved sequences in the 5'-untranslated region of the cyanobacterial rbp genes that are known to be induced by low temperature. The cold-regulated Rbp proteins contain a short C-terminal glycine-rich domain. In this respect, these proteins are similar to plant and mammalian glycine-rich RNA-binding proteins (GRPs), which also contain a single RRM domain with a C-terminal glycine-rich domain and are highly expressed at low temperature. Detailed phylogenetic analysis showed, however, that the cyanobacterial Rbp proteins and the eukaryotic GRPs do not belong to a single lineage, but that the glycine-rich domains are likely to have been added independently. The cold-regulation of both types of proteins is also likely to have evolved independently. Furthermore, the chloroplast RNA-binding proteins are not likely to have originated from the Rbp proteins of endosymbiont cyanobacterium, but are supposed to have diverged from the GRPs. These results suggest that the cyanobacterial Rbp proteins and the eukaryotic GRPs are similar in both structure and regulation, but that this apparent similarity has resulted from convergent evolution.

Amino Acid Sequence↗

Convergent evolution of a 2-methylbutyryl-CoA dehydrogenase from isovaleryl-CoA dehydrogenase in Solanum tuberosum.

The potato cDNAs Solanum tuberosum isovaleryl-CoA dehydrogenases 1 and 2 (St-IVD1 and St-IVD2) encode proteins that are 84% identical to each other and 65 and 64% identical to human IVD, respectively. St-IVD2 protein was previously partially purified from potato tubers and confirmed to be an IVD. The function of St-IVD1 is unknown. In these experiments, both proteins were expressed in Escherichia coli and purified as intact homotetramers. The substrate preference profile of the St-IVD2 protein was similar to that of human IVD. However, recombinant St-IVD1 had maximal activity with 2-methylbutyryl-CoA, which in humans is dehydrogenated by short/branched-chain acyl-CoA dehydrogenase (SBCAD). Whereas molecular modeling predicts that the 2-methylbutyryl-CoA dehydrogenase (2MBCD) and IVD substrate binding pockets are nearly identical, 2MBCD has amino acid substitutions at five residues that are invariant among all of the known and putative IVDs. Site-directed mutagenesis was used to match the human IVD active site with that of potato 2MBCD. The resulting mutant IVD had detectable activity with 2-methylbutyryl-CoA and no activity with isovaleryl-CoA. The 2MBCD active site was compared with that of human SBCAD using molecular modeling. Residues Met-361 and Ala-365 of 2MBCD appear to partially substitute for the function of Tyr-380 in human SBCAD, binding the methyl branch linked to C2 of 2-methylbutyryl-CoA, whereas residues Val-88, Val-92, and Val-96 appear to bind the distal C4 methyl group. The presence of a 2MBCD in potato that is highly homologous to IVD is an example of convergent evolution within the acyl-CoA dehydrogenase family, leading to the independent occurrence of two enzymes (SBCAD and 2MBCD) specific for 2-methylbutyryl-CoA.

Binding Sites↗

The genes and enzymes for the catabolism of galactitol, D-tagatose, and related carbohydrates in Klebsiella oxytoca M5a1 and other enteric bacteria display convergent evolution.

Enteric bacteria (Enteriobacteriaceae) carry on their single chromosome about 4000 genes that all strains have in common (referred to here as "obligatory genes"), and up to 1300 "facultative" genes that vary from strain to strain and from species to species. In closely related species, obligatory and facultative genes are orthologous genes that are found at similar loci. We have analyzed a set of facultative genes involved in the degradation of the carbohydrates galactitol, D-tagatose, D-galactosamine and N-acetyl-galactosamine in various pathogenic and non-pathogenic strains of these bacteria. The four carbohydrates are transported into the cell by phosphotransferase (PTS) uptake systems, and are metabolized by closely related or even identical catabolic enzymes via pathways that share several intermediates. In about 60% of Escherichia coli strains the genes for galactitol degradation map to a gat operon at 46.8 min. In strains of Salmonella enterica, Klebsiella pneumoniae and K. oxytoca, the corresponding gat genes, although orthologous to their E. coli counterparts, are found at 70.7 min, clustered in a regulon together with three tag genes for the degradation of D-tagatose, an isomer of D-fructose. In contrast, in all the E. coli strains tested, this chromosomal site was found to be occupied by an aga/kba gene cluster for the degradation of D-galactosamine and N-acetyl-galactosamine. The aga/kba and the tag genes were paralogous either to the gat cluster or to the fru genes for degradation of D-fructose. Finally, in more then 90% of strains of both Klebsiella species, and in about 5% of the E. coli strains, two operons were found at 46.8 min that comprise paralogous genes for catabolism of the isomers D-arabinitol (genes atl or dal) and ribitol (genes rtl or rbt). In these strains gat genes were invariably absent from this location, and they were totally absent in S. enterica. These results strongly indicate that these various gene clusters and metabolic pathways have been subject to convergent evolution among the Enterobacteriaceae. This apparently involved recent horizontal gene transfer and recombination events, as indicated by major chromosomal rearrangements found in their immediate vicinity.

Acetylgalactosamine↗