Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “diverged evolution”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Functional elements of the ribosomal protein L7a (rpL7a) gene promoter region and their conservation between mammals and birds.

The transcriptional initiation sites of the chicken ribosomal protein L7a (rpL7a) gene have been determined and found to occur at three consecutive cytidine residues at the start of a polypyrimidine tract of 8 base pairs (bp). A comparative analysis of the 5' upstream regions of the mouse, human and chicken rpL7a genes identified two sequence elements (Box A and Box B) conserved over the 600 million years of divergent evolution that separate mammals and birds. Only Box A (nts - 56 to - 39) and Box B (nts - 25 to - 4) sequences were detected to bind nuclear factors from mouse nuclear extracts in an analysis of the mouse rpL7a 5' upstream sequence. Box A and Box B bind different nuclear factors and the factor binding to mouse Box A and mouse Box B sequences could be effectively competed by corresponding homologous sequences from the human and chicken rpL7a promoters. These results indicate that elements of the rpL7a promoter region are conserved between mammals and birds. An in vivo analysis of the mouse rpL7a 5' upstream sequence required for efficient transcription identified the 5' border of the minimal promoter region as lying between nts - 50 and - 56. Constructs containing 56 bp of 5' upstream DNA and the first 25 bp rpL7a exon were very efficiently transcribed indicating that sequences within the first intron are not required for gene expression. No sequence similarity was detected between the rpL7a promoter elements and described promoter elements of other eukaryotic ribosomal protein genes.

Animals↗

M.phi 3TII: a new monospecific DNA (cytosine-C5) methyltransferase with pronounced amino acid sequence similarity to a family of adenine-N6-DNA-methyltransferases.

The temperate B.subtilis phages phi 3T and rho 11s code, in addition to the multispecific DNA (cytosine-C5) methyltransferases (C5-MTases) M.phi 3TI and M.rho 11sI, which were previously characterized, for the identical monospecific C5-MTases M.phi 3TII and M.rho 11sII. These enzymes modify the C to TCGA sites, a novel target specificity among C5-MTases. The primary sequence of M.phi 3TII (326 amino acids) shows all conserved motifs typical of the building plan of C5-MTases. The degree of relatedness between M.phi 3TII and all other mono- or multispecific C5-MTases ranges from 30-40% amino acid identity. Particularly M.phi 3TII does not show pronounced similarity to M.phi 3TI indicating that both MTase genes were not generated from one another but were acquired independently by the phage. The amino terminal part of the M.phi 3TII (preceding the variable region 'V'), which predominantly constitutes the catalytic domain of the enzyme, exhibits pronounced sequence similarity to the amino termini of a family of A-N6-MTases, which--like M.Taql--recognize the general sequence TNNA. This suggests that recently described similarities in the general three dimensional organization of C5- and A-N6-MTases imply divergent evolution of these enzymes originating from a common molecular ancestor.

Amino Acid Sequence↗

M.phi 3TII: a new monospecific DNA (cytosine-C5) methyltransferase with pronounced amino acid sequence similarity to a family of adenine-N6-DNA-methyltransferases.

The temperate B.subtilis phages phi 3T and rho 11s code, in addition to the multispecific DNA (cytosine-C5) methyltransferases (C5-MTases) M. phi 3TI and M. rho 11sI, which were previously characterized, for the identical monospecific C5-MTases M. phi 3TII and M. rho 11sII. These enzymes modify the C of TCGA sites, a novel target specificity among C5-MTases. The primary sequence of M. phi 3TII (326 amino acids) shows all conserved motifs typical of the building plan of C5-MTases. The degree of relatedness between M. phi 3TII and all other mono- or multispecific C5-MTases ranges from 30-40% amino acid identity. Particularly M. phi 3TII does not show pronounced similarity to M. phi 3TI indicating that both MTase genes were not generated from one another but were acquired independently by the phage. The amino terminal part of the M. phi 3TII (preceding the variable region 'V'), which predominantly constitutes the catalytic domain of the enzyme, exhibits pronounced sequence similarity to the amino termini of a family of A-N6-MTases, which--like M.TaqI--recognize the general sequence TNNA. This suggests that recently described similarities in the general three dimensional organization of C5- and A-N6-MTases imply divergent evolution of these enzymes originating from a common molecular ancestor.

Amino Acid Sequence↗

Cloning and analysis of the four genes coding for Bpu10I restriction-modification enzymes.

The Bpu 10I R-M system from Bacillus pumilus 10, which recognizes the asymmetric 5'-CCTNAGC sequence, has been cloned, sequenced and expressed in Escherichia coli . The system comprises four adjacent, similarly oriented genes encoding two m5C MTases and two subunits of Bpu 10I ENase (34.5 and 34 kDa). Both bpu10IR genes either in cis or trans are needed for the manifestation of R. Bpu 10I activity. Subunits of R. Bpu 10I, purified to apparent homogeneity, are both required for cleavage activity. This heterosubunit structure distinguishes the Bpu 10I restriction endonuclease from all other type II restriction enzymes described previously. The subunits reveal 25% amino acid identity. Significant similarity was also identified between a 43 amino acid region of R. Dde I and one of the regions of higher identity shared between the Bpu 10I subunits, a region that could possibly include the catalytic/Mg2+binding center. The similarity between Bpu 10I and Dde I MTases is not limited to the conserved motifs (CM) typical for m5C MTases. It extends into the variable region that lies between CMs VIII and IX. Duplication of a progenitor gene, encoding an enzyme recognizing a symmetric nucleotide sequence, followed by concerted divergent evolution, may provide a possible scenario leading to the emergence of the Bpu 10I ENase, which recognizes an overall asymmetric sequence and cleaves within it symmetrically.

Amino Acid Sequence↗

SUPFAM--a database of potential protein superfamily relationships derived by comparing sequence-based and structure-based families: implications for structural genomics and function annotation in genomes.

Members of a superfamily of proteins could result from divergent evolution of homologues with insignificant similarity in the amino acid sequences. A superfamily relationship is detected commonly after the three-dimensional structures of the proteins are determined using X-ray analysis or NMR. The SUPFAM database described here relates two homologous protein families in a multiple sequence alignment database of either known or unknown structure. The present release (1.1), which is the first version of the SUPFAM database, has been derived by analysing Pfam, which is one of the commonly used databases of multiple sequence alignments of homologous proteins. The first step in establishing SUPFAM is to relate Pfam families with the families in PALI, which is an alignment database of homologous proteins of known structure that is derived largely from SCOP. The second step involves relating Pfam families which could not be associated reliably with a protein superfamily of known structure. The profile matching procedure, IMPALA, has been used in these steps. The first step resulted in identification of 1280 Pfam families (out of 2697, i.e. 47%) which are related, either by close homologous connection to a SCOP family or by distant relationship to a SCOP family, potentially forming new superfamily connections. Using the profiles of 1417 Pfam families with apparently no structural information, an all-against-all comparison involving a sequence-profile match using IMPALA resulted in clustering of 67 homologous protein families of Pfam into 28 potential new superfamilies. Expansion of groups of related proteins of yet unknown structural information, as proposed in SUPFAM, should help in identifying 'priority proteins' for structure determination in structural genomics initiatives to expand the coverage of structural information in the protein sequence space. For example, we could assign 858 distinct Pfam domains in 2203 of the gene products in the genome of Mycobacterium tubercolosis. Fifty-one of these Pfam families of unknown structure could be clustered into 17 potentially new superfamilies forming good targets for structural genomics. SUPFAM database can be accessed at http://pauling.mbu.iisc.ernet.in/~supfam.

Animals↗

Prediction of lipid posttranslational modifications and localization signals from protein sequences: big-Pi, NMT and PTS1.

Many posttranslational modifications (N-myristoylation or glycosylphosphatidylinositol (GPI) lipid anchoring) and localization signals (the peroxisomal targeting signal PTS1) are encoded in short, partly compositionally biased regions at the N- or C-terminus of the protein sequence. These sequence signals are not well defined in terms of amino acid type preferences but they have significant interpositional correlations. Although the number of verified protein examples is small, the quantification of several physical conditions necessary for productive protein binding with the enzyme complexes executing the respective transformations can lead to predictors that recognize the signals from the amino acid sequence of queries alone. Taxon-specific prediction functions are required due to the divergent evolution of the active complexes. The big-Pi tool for the prediction of the C-terminal signal for GPI lipid anchor attachment is available for metazoan, protozoan and plant sequences. The myristoyl transferase (NMT) predictor recognizes glycine N-myristoylation sites (at the N-terminus and for fragments after processing) of higher eukaryotes (including their viruses) and fungi. The PTS1 signal predictor finds proteins with a C-terminus appropriate for peroxisomal import (for metazoa and fungi). Guidelines for application of the three WWW-based predictors (http://mendel.imp.univie.ac.at/) and for the interpretation of their output are described.

Acyltransferases↗

Comprehensive splice-site analysis using comparative genomics.

We have collected over half a million splice sites from five species-Homo sapiens, Mus musculus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-and classified them into four subtypes: U2-type GT-AG and GC-AG and U12-type GT-AG and AT-AC. We have also found new examples of rare splice-site categories, such as U12-type introns without canonical borders, and U2-dependent AT-AC introns. The splice-site sequences and several tools to explore them are available on a public website (SpliceRack). For the U12-type introns, we find several features conserved across species, as well as a clustering of these introns on genes. Using the information content of the splice-site motifs, and the phylogenetic distance between them, we identify: (i) a higher degree of conservation in the exonic portion of the U2-type splice sites in more complex organisms; (ii) conservation of exonic nucleotides for U12-type splice sites; (iii) divergent evolution of C.elegans 3' splice sites (3'ss) and (iv) distinct evolutionary histories of 5' and 3'ss. Our study proves that the identification of broad patterns in naturally-occurring splice sites, through the analysis of genomic datasets, provides mechanistic and evolutionary insights into pre-mRNA splicing.

Animals↗

Up-to-dating of complete sequenced DNA data of Hansenula wingei yeast mitochondria.

To update sequenced data, we determined the 5' and 3' termini of yeast Hansenula wingei (Pichia canadensis) mitochondrial (mt) large subunit ribosomal RNA (LSU) which is encoded in the mt genome. The 5' end position was mapped downstream from a putative transcription starting site which is homologous to a Saccharomyces cerevisiae mitochondrial promoter sequence. This suggests that the primary transcript of LSU is processed from 5' end and then mature transcript is formed. This processing is different from that of S. cerevisiae mt LSU in which processing on its 5' end does not occur. Based on the sequence data of H. wingei mt LSU, we constructed its secondary structure, and compared it with those of the other fungal organisms. Conserved regions of H. wingei LSU were identified and used for subsequent phylogenetic analysis. In genome structure and gene content, H. wingei mt genome has several characteristics similar to those in filamentous fungi, but the phylogenetic analysis indicates closer kinship to yeast S. cerevisiae. This agrees with previous non-sequencing phylogenies and suggests that extraordinary rearrangements have occurred in yeast mt genomes during divergent evolution.

Base Sequence↗

Purification and characterization of alpha-macroglobulin and ovomacroglobulin of the green turtle (Chelonia mydas japonica).

The plasma alpha-macroglobulin and egg white ovomacroglobulin were purified from the sea turtle, Chelonia mydas japonica, and their structural and functional properties were studied with the aim of clarifying the degree of evolutional divergence of two homologous proteins specific to different tissues of the same animal. The concentration of alpha-macroglobulin in green turtle plasma was about 4 mg/ml. The protein was purified from the plasma by precipitation with polyethylene glycol 6000, followed by zinc chelate chromatography and gel chromatography on Sepharose CL-6B. The concentration of ovomacroglobulin in green turtle egg white was about 0.4 mg/ml. Ovomacroglobulin was purified by gel chromatography on Sepharose CL-6B. The two proteins had similar molecular weights and amino acid compositions, and both inhibited proteinases such as trypsin, chymotrypsin, papain, and thermolysin. The amino terminal sequences of the two proteins were homologous to each other but higher homologies were found between the ovomacroglobulin of turtle and chicken, and between the serum macroglobulins of the same animals. The functional difference between turtle alpha-macroglobulin and ovomacroglobulin became clear when they were treated with methylamine, which is known to destroy the inhibitory activity of human alpha 2-macroglobulin by splitting internal thiolester bonds. The inhibitory activity of the turtle plasma protein was completely destroyed by methylamine but that of ovomacroglobulin was only partially affected. The number of sulfhydryl groups as titrated with 5,5'-dithiobis(2-nitrobenzoate) before and after treatment with proteinases or methylamine was different for the two proteins. The amount of radioactive methylamine that was incorporated was also different between the two proteins. The two proteins purified in this study had no immunological cross-reactivity.

Amino Acid Sequence↗

Conformational changes of alpha-macroglobulin and ovomacroglobulin from the green turtle (Chelonia mydas japonica).

Green turtle plasma alpha-macroglobulin and ovomacroglobulin underwent conformational changes when they were treated with proteinases or methylamine. Their conformational changes were studied by HPLC gel chromatography, circular dichroism, and electron microscopy. The Stokes radii of native green turtle alpha-macroglobulin and ovomacroglobulin were estimated to be 84.3 +/- 0.5 A, and 93.0 +/- 0.5 A, respectively, by means of an HPLC experiment. After reaction with methylamine or proteinases, the Stokes radius of alpha-macroglobulin changed to 83.0 +/- 0.5 A or 85.4 +/- 0.5 A, respectively, and that of ovomacroglobulin to 93.0 +/- 0.5 A or 87.1 +/- 0.5 A. The circular dichroic spectra of native alpha-macroglobulin and ovomacroglobulin exhibited a negative band at around 215 nm, indicating the presence of beta-structure. Reaction of the two macroglobulins with methylamine resulted in a slight decrease in the ellipticity and reaction with proteinases led to a slight increase. The electron micrographic images of native alpha-macroglobulin and ovomacroglobulin can be described as deformed rings for the former and rugby balls for the latter. A common characteristic feature of the two molecules was that the central parts of the molecules were only thinly occupied by subunit. After reaction of macroglobulins with proteinases, the void spaces became partially filled and their overall shape more rectangular. Methylamine treatment caused a structural change only in alpha-macroglobulin but not in ovomacroglobulin. The difference in the susceptibility of the macroglobulins to methylamine was taken as an indication of evolutional divergence of the two homologous proteins within the last 300 million years.

Animals↗

RbcX can function as a rubisco chaperonin, but is non-essential in Synechococcus PCC7942.

In most cyanobacteria, the gene rbcX is co-transcribed with the rbcL and rbcS genes that code for the large and small subunits of ribulose-1,5-bisphosphate carboxylase/oxygenase (Rubisco). Previous co-expression studies in Escherichia coli of cyanobacterial Rubisco and RbcX have identified a chaperonin-like function for RbcX. The organization of the rbcLXS operon has, to a certain extent, precluded definitive gene function studies of rbcX in cyanobacteria. In Synechococcus PCC7942, however, rbcX is located >100 kb away from the rbcLS operon, providing an opportunity to examine the role of RbcX by insertional inactivation without interference from the Rubisco genes. Fully segregated Synechococcus PCC7942 DeltarbcX::KmR mutants were readily obtained that showed no perturbations in growth rate or Rubisco content and activity. Low amounts of rbcX transcript were detected in Synechococcus PCC7942; however, a sensitive antibody raised against purified RbcX failed to detect RbcX expression in cells exposed to different stress treatments. In contrast, co-expression studies of Rubisco assembly in E. coli showed that RbcX from Synechococcus PCC7942 and PCC7002 are functionally interchangeable and can stimulate assembly of the PCC7942 and PCC7002 Rubisco subunits. Our results indicate that Rubisco folding and assembly in Synechococcus PCC7942 may have evolved to be independent of RbcX function, apparently in contrast to other beta-cyanobacteria. We speculate that divergent evolution of the RbcL sequence may have relaxed a requirement for RbcX function in Synechococcus PCC7942 and propose a new approach for definitively isolating RbcX function in other beta-cyanobacteria.

Amino Acid Sequence↗

An analysis of simultaneous variation in protein structures.

The simultaneous substitution of pairs of buried amino acid side chains during divergent evolution has been examined in a set of protein families with known crystal structures. A weak signal is found that shows that amino acid pairs near in space in the folded structure preferentially undergo substitution in a compensatory way. Three different physicochemical types of covariation 'signals' were then examined separately, with consideration given to the evolutionary distance at which different types of compensation occur. Where the compensatory covariation tends towards retaining the combined residue volumes, the signal is significant only at very low evolutionary distances. Where the covariation compensates for changes in the hydrogen bonding, the signal is strongest at intermediate evolutionary distances. Covariations that compensate for charge variations appeared with equal strength at all the evolutionary distances examined. A recipe is suggested for using the weak covariation signal to assemble the predicted secondary structural elements, where the evolutionary distance, covariation type and weighting are considered together with the tertiary structural context (interior or surface) of the residues being examined.

Computer Simulation↗

Recognition of analogous and homologous protein folds--assessment of prediction success and associated alignment accuracy using empirical substitution matrices.

Fold recognition methods aim to use the information in the known protein structures (the targets) to identify that the sequence of a protein of unknown structure (the probe) will adopt a known fold. This paper highlights that the structural similarities sought by these methods can be divided into two types: remote homologues and analogues. Homologues are the result of divergent evolution and often share a common function. We define remote homologues as those that are not easily detectable by sequence comparison methods alone. Analogues do not have a common ancestor and generally do not have a common function. Several sets of empirical matrices for residue substitution, secondary structure conservation and residue accessibility conservation have previously been derived from aligned pairs of remote homologues and analogues (Russell et al., J. Mol. Biol., 1997, 269, 423-439). Here a method for fold recognition, FOLDFIT, is introduced that uses these matrices to match the sequences, secondary structures and residue accessibilities of the probe and target. The approach is evaluated on distinct datasets of analogous and remotely homologous folds. The accuracy of FOLDFIT with the different matrices on the two datasets is contrasted to results from another fold recognition method (THREADER) and to searches using mutation matrices in the absence of any structural information. FOLDFIT identifies at top rank 12 out of 18 remotely homologous folds and five out of nine analogous folds. The average alignment accuracies for residue and secondary structure equivalencing are much higher for homologous folds (residue approximately 42%, secondary structure approximately 78%) than for analogues folds (approximately 12%, approximately 47%). Sequence searches alone can be successful for several homologues in the testing sets but nearly always fail for the analogues. These results suggest that the recognition of analogous and remotely homologous folds should be assessed separately. This study has implications for the development and comparative evaluation of fold recognition algorithms.

Evolution, Molecular↗

Natural instability of Agrobacterium vitis Ti plasmid due to unusual duplication of a 2.3-kb DNA fragment.

The octopine/cucumopine (o/c) Ti plasmids of Agrobacterium vitis carry two T regions, TA and TB. The TA region resembles the octopine TL region. The TB region contains the auxin synthesis genes TB-iaaM and TB-iaaH and the cucumopine synthesis gene cus. Within the group of o/c isolates, strains 2608 and 2641 are closely related. However, 2641 lacks the TB region. The restriction maps of pTi2608 and pTi2641 were established and showed that the TB deletion resulted from intramolecular recombination between two directly repeated sequences separated by 66 kb in pTi2608. The 2,294-bp repeated sequence lacks inverted repeats and does not duplicate its target site, indicating that it is not a classical bacterial insertion sequence (IS element). It was therefore called an RSAv element (repeated sequence of A. vitis). The RSAv element carries two open reading frames: ORF234 is homologous to the traR gene of the A. tumefaciens nopaline Ti plasmid pTiC58; ORF488 is homologous to the sucrose phosphorylase gene of Leuconostoc mesenteroides and the glucosyl transferase A gene of Streptococcus mutans. The RSAv repeat starts precisely at the start codon of ORF488 and ends two base pairs 3' of the stop codon of ORF234. The structural organization of the RSAv element suggests that the amplification event did not result from a random amplification process. A study of the distribution of the two RSAv copies (RSAv-1 and RSAv-2) in pTi2608 and other o/c isolates indicates that the ancestor o/c Ti plasmid contained only RSAv-1 and that this sequence was duplicated at one point during the divergent evolution of the o/c Ti plasmids.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Development of mouse and rat oocytes in chimeric reaggregated ovaries after interspecific exchange of somatic and germ cell components.

The germ cell and somatic cell compartments of newborn rat and mouse ovaries, which contain only primordial stage follicles, were completely exchanged and reaggregated to produce xenogeneic chimeric ovaries. The reaggregated ovaries were grafted beneath the renal capsules of ovariectomized SCID mice to develop for periods up to 21 days. Xenogeneic follicles developed with essentially normal morphological characteristics. Both rat and mouse oocytes with species-specific characteristics grew within follicles that were composed of somatic cells exclusively of the alternative species. Rat oocytes grown in mouse follicles became competent to resume meiosis, and progressed to metaphase II when they were removed from follicles and cultured. In addition, mouse oocytes grown in rat follicles underwent fertilization and preimplantation development in vitro, and developed to term after embryos were transferred to pseudopregnant mouse foster mothers. Therefore, despite an estimated 11 million years of divergent evolution, oocytes and somatic cells of rat and mouse ovaries can be exchanged and can produce functional oocytes. It is concluded that factors involved in oocyte-somatic cell interactions necessary to support oocyte development and appropriate differentiation of the oocyte-associated granulosa cells are conserved between rats and mice. Moreover, although granulosa cells play important roles in oocyte development, the development of species-specific characteristics of oocytes occurs without apparent modification by a xenogeneic follicular environment.

Animals↗

Pathways of ubiquitin conjugation.

The covalent attachment of the polypeptide ubiquitin to proteins marks them for degradation by the ubiquitin/26S proteasome-dependent degradation pathway. This pathway functions in regulating many fundamental processes required for cell viability. Phylogenetic analysis of ubiquitin sequences reveals greater variability among lower eukaryotes and defines essential residues, many of which are conserved among the three ubiquitin-like proteins known to undergo parallel ligation pathways. The hierarchical design of the ubiquitin conjugation mechanism provides great flexibility for the divergent evolution of new functions mediated by this posttranslational modification. Within this hierarchy, a single ubiquitin-activating enzyme provides charged intermediates to multiple targeting pathways defined by cognate ubiquitin carrier protein (E2)/ligase (E3) pairs. Sequence analysis of E2 isozymes shows that the E2 superfamily is composed of distinct function-specific families. The apparent lack of E2/E3 specificity suggested in the literature results from the presence of multiple isozymes within many E2 families and erroneous family assignments based on incomplete data sets. Other apparent inconsistencies are explained by interfamily sequence relationships among some E2 isoforms.

Amino Acid Sequence↗

P450 superfamily: update on new sequences, gene mapping, accession numbers and nomenclature.

We provide here a list of 481 P450 genes and 22 pseudogenes, plus all accession numbers that have been reported as of October 18, 1995. These genes have been described in 85 eukaryote (including vertebrates, invertebrates, fungi, and plants) and 20 prokaryote species. Of 74 gene families so far described, 14 families exist in all mammals examined to date. These 14 families comprise 26 mammalian subfamilies, of which 20 and 15 have been mapped in the human genome and the mouse genome, respectively. Each subfamily usually represents a cluster of tightly linked genes widely scattered throughout the genome, but there are exceptions. Interestingly, the CYP51 family has been found in mammals, filamentous fungi and yeast, and plants-attesting to the fact that this P450 gene family is very ancient. One functional CYP51 gene and two processed pseudogenes, which are the first examples of intronless pseudogenes within the P450 superfamily, have been mapped to three different human chromosomes. This revision supersedes the four previous updates in which a nomenclature system, based on divergent evolution of the superfamily, has been described. For the gene, we recommend that the italicized root symbol "CYP' for human ("Cyp' for mouse and Drosophila), representing "cytochrome P450', be followed by an Arabic number denoting the family, a letter designating the subfamily (when two or more exist), and an Arabic numeral representing the individual gene within the subfamily. A hyphen is no longer recommended in mouse gene nomenclature. "P' ("ps' in mouse and Drosophila) after the gene number denotes a pseudogene; "X' after the gene number means its use has been discontinued. If a gene is the sole member of a family, the subfamily letter and gene number would be helpful but need not be included. The human nomenclature system should be used for all species other than mouse and Drosophila. The cDNAs, mRNAs and enzymes in all species (including mouse) should include all capital letters, and without italics or hyphens. This nomenclature system is similar to that proposed in our previous updates.

Alleles↗

The UDP glycosyltransferase gene superfamily: recommended nomenclature update based on evolutionary divergence.

This review represents an update of the nomenclature system for the UDP glucuronosyltransferase gene superfamily, which is based on divergent evolution. Since the previous review in 1991, sequences of many related UDP glycosyltransferases from lower organisms have appeared in the database, which expand our database considerably. At latest count, in animals, yeast, plants and bacteria there are 110 distinct cDNAs/genes whose protein products all contain a characteristic 'signature sequence' and, thus, are regarded as members of the same superfamily. Comparison of a relatedness tree of proteins leads to the definition of 33 families. It should be emphasized that at least six cloned UDP-GlcNAc N-acetylglucosaminyltransferases are not sufficiently homologous to be included as members of this superfamily and may represent an example of convergent evolution. For naming each gene, it is recommended that the root symbol UGT for human (Ugt for mouse and Drosophila), denoting 'UDP glycosyltransferase,' be followed by an Arabic number representing the family, a letter designating the subfamily, and an Arabic numeral denoting the individual gene within the family or subfamily, e.g. 'human UGT2B4' and 'mouse Ugt2b5'. We recommend the name 'UDP glycosyltransferase' because many of the proteins do not preferentially use UDP glucuronic acid, or their nucleotide sugar preference is unknown. Whereas the gene is italicized, the corresponding cDNA, transcript, protein and enzyme activity should be written with upper-case letters and without italics, e.g. 'human or mouse UGT1A1.' The UGT1 gene (spanning > 500 kb) contains at least 12 promoters/first exons, which can be spliced and joined with common exons 2 through 5, leading to different N-terminal halves but identical C-terminal halves of the gene products; in this scheme each first exon is regarded as a distinct gene (e.g. UGT1A1, UGT1A2, ... UGT1A12). When an orthologous gene between species cannot be identified with certainty, as occurs in the UGT2B subfamily, sequential naming of the genes is being carried out chronologically as they become characterized. We suggest that the Human Gene Nomenclature Guidelines (http://www.gene.acl.ac.uk/nomenclature/guidelines.html++ +) be used for all species other than the mouse and Drosophila. Thirty published human UGT1A1 mutant alleles responsible for clinical hyperbilirubinemias are listed herein, and given numbers following an asterisk (e.g. UGT1A1*30) consistent with the Human Gene Nomenclature Guidelines. It is anticipated that this UGT gene nomenclature system will require updating on a regular basis.

Amino Acid Sequence↗