Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,675 records · Page 93Linked to original sources

Comparative analysis of the catalytic domain of hemorrhagic and non-hemorrhagic snake venom metallopeptidases using bioinformatic tools.

Snake venom metalloproteases (SVMPs) are a set of interesting enzymes that are one of the major components of venom affecting hemostasis. A great challenge since their discovery has been to find molecular features responsible for their hemorrhagic potency and many attempts have been made without any consistent result. Here we describe a series of comparisons between the catalytic domains of hemorrhagic and non-hemorrhagic SVMPs made with the help of bioinformatics. These involved sequence and structure-based multiple alignments, phylogenetic reconstruction, predicted physical and chemical properties, motif scanning and structural analyses. Although hemorrhagic activity seems to be complex, involving multiple factors, we found some molecular characteristics that may influence the toxic effects. Among these findings, it was possible to use a molecular surface feature to subdivide the P-I class in hemorrhagic and non-hemorrhagic SVMPs. It was also possible to suggest a role for the conserved Asp148 and Ser176 residues in the stabilization of the active site.

Animals↗

BCM Search Launcher--an integrated interface to molecular biology data base search and analysis services available on the World Wide Web.

The BCM Search Launcher is an integrated set of World Wide Web (WWW) pages that organize molecular biology-related search and analysis services available on the WWW by function, and provide a single point of entry for related searches. The Protein Sequence Search Page, for example, provides a single sequence entry form for submitting sequences to WWW servers that offer remote access to a variety of different protein sequence search tools, including BLAST, FASTA, Smith-Waterman, BEAUTY, PROSITE, and BLOCKS searches. Other Launch pages provide access to (1) nucleic acid sequence searches, (2) multiple and pair-wise sequence alignments, (3) gene feature searches, (4) protein secondary structure prediction, and (5) miscellaneous sequence utilities (e.g., six-frame translation). The BCM Search Launcher also provides a mechanism to extend the utility of other WWW services by adding supplementary hypertext links to results returned by remote servers. For example, links to the NCBI's Entrez data base and to the Sequence Retrieval System (SRS) are added to search results returned by the NCBI's WWW BLAST server. These links provide easy access to auxiliary information, such as Medline abstracts, that can be extremely helpful when analyzing BLAST data base hits. For new or infrequent users of sequence data base search tools, we have preset the default search parameters to provide the most informative first-pass sequence analysis possible. We have also developed a batch client interface for Unix and Macintosh computers that allows multiple input sequences to be searched automatically as a background task, with the results returned as individual HTML documents directly to the user's system. The BCM Search Launcher and batch client are available on the WWW at URL http:@gc.bcm.tmc.edu:8088/search-launcher.html.

Animals↗

Secondary structural predictions for the clostridial neurotoxins.

The primary structures of a family of ten clostridial neurotoxins have recently been deduced yet little information is presently available concerning their secondary or tertiary structures. Because the overall similarity percentage of multiply aligned sequences is high, the secondary structures of these metalloendopeptidases are also expected to be conserved. The neural net program, PHD (Rost and Sander, Proc. Natl. Acad. Sci. USA 90:7558-7562, 1993), predicted that the secondary structures of the neurotoxins were indeed conserved in both single and multiple sequence modes of analysis. Predictions for the amounts of helical, extended, and loop states from the single sequence analyses were consistent with previously published data from circular dichroism studies on some of these neurotoxins. In the single analysis mode, only the aligned regions were predicted to show conservation of the three-state structure. In contrast, the multiple sequence analysis predicted that a conserved state (variable loops) also exists in non-aligned regions. Alignments with the primary structure of the prototypic metalloendopeptidase thermolysin showed that about 25% of the residues within this enzyme are similar to those in the neurotoxins. A comparison of thermolysin's known secondary structure with the predictions from this study showed that about 80% of thermolysin's residues could be structurally aligned with those in the neurotoxins. These predictions provide the necessary framework to build a homologous low-resolution tertiary structure of the neurotoxin active site that will be essential in the development of synthetic inhibitors.

Amino Acid Sequence↗

Genetic organization of the streptokinase region of the Streptococcus equisimilis H46A chromosome.

The complete nucleotide sequences of four genes and one open reading frame (ORF1) adjacent to the streptokinase gene, skc, from Streptococcus equisimilis H46A were determined. These genes are encoded on the opposite DNA strand to skc and are arranged as follows: dexB-abc-lrp-skc-ORF1-rel. The dexB gene, coding for an alpha-glucosidase (M(r) 61,733), and abc, encoding an ABC transporter (M(r) 42,080), are similar to the dexB and msmK genes, respectively, from the multiple sugar metabolism operon of S. mutans. The lrp gene specifies a leucine-rich protein (M(r) 32,302) that has a leucine-zipper motif at its C-terminus. The function of the Lrp protein is not known but appeared to be detrimental when overexpressed in Escherichia coli. Although lrp appears not to be an essential gene, as judged by plasmid insertion mutagenesis, it is conserved in all streptococcal strains carrying a streptokinase gene. The rel gene showed significant homology to the E. coli relA and spoT genes involved in the stringent response to amino acid deprivation. Multiple alignment of the amino acid sequences of Rel (M(r) 83,913), RelA and SpoT revealed 59.4% homology of the primary structures. Northern hybridization analyses of the genes in the skc region showed skc to be transcribed most abundantly. In addition to transcripts for skc, monocistronic mRNAs were detected for all three genes divergently transcribed from skc. Although there was also some read-through transcription from lrp into abc, and from abc into dexB, the transcription pattern suggests a high degree of transcriptional and functional independence not only of skc but also abc and dexB. Prominent structural features in intergenic regions included a static DNA bending locus located upstream and a putative bidirectional transcription terminator downstream of skc.

Amino Acid Sequence↗

Polyester synthases: natural catalysts for plastics.

Polyhydroxyalkanoates (PHAs) are biopolyesters composed of hydroxy fatty acids, which represent a complex class of storage polyesters. They are synthesized by a wide range of different Gram-positive and Gram-negative bacteria, as well as by some Archaea, and are deposited as insoluble cytoplasmic inclusions. Polyester synthases are the key enzymes of polyester biosynthesis and catalyse the conversion of (R)-hydroxyacyl-CoA thioesters to polyesters with the concomitant release of CoA. These soluble enzymes turn into amphipathic enzymes upon covalent catalysis of polyester-chain formation. A self-assembly process is initiated resulting in the formation of insoluble cytoplasmic inclusions with a phospholipid monolayer and covalently attached polyester synthases at the surface. Surface-attached polyester synthases show a marked increase in enzyme activity. These polyester synthases have only recently been biochemically characterized. An overview of these recent findings is provided. At present, 59 polyester synthase structural genes from 45 different bacteria have been cloned and the nucleotide sequences have been obtained. The multiple alignment of the primary structures of these polyester synthases show an overall identity of 8-96% with only eight strictly conserved amino acid residues. Polyester synthases can been assigned to four classes based on their substrate specificity and subunit composition. The current knowledge on the organization of the polyester synthase genes, and other genes encoding proteins related to PHA metabolism, is compiled. In addition, the primary structures of the 59 PHA synthases are aligned and analysed with respect to highly conserved amino acids, and biochemical features of polyester synthases are described. The proposed catalytic mechanism based on similarities to alpha/beta-hydrolases and mutational analysis is discussed. Different threading algorithms suggest that polyester synthases belong to the alpha/beta-hydrolase superfamily, with a conserved cysteine residue as catalytic nucleophile. This review provides a survey of the known biochemical features of these unique enzymes and their proposed catalytic mechanism.

Acyltransferases↗

The major opsin in bees (Insecta: Hymenoptera): A promising nuclear gene for higher level phylogenetics.

We report the phylogenetic utility of the nuclear gene encoding the long-wavelength opsin (LW Rh) for tribes of bees. Aligned nucleotide sequences were examined in multiple taxa from the four tribes comprising the corbiculate bees within the subfamily Apinae. Phylogenetic analyses of sequence variation in a 502-bp fragment (approx 40% of the coding region) strongly supported the monophyly of each of the four tribes, which are well established from previous studies of morphology and DNA. Trees estimated from parsimony and maximum likelihood analyses of LW Rh sequences show a strongly supported relationship between the tribes Meliponini and Bombini, a relationship that has been found uniformly in studies of other genes (28S, 16S, and cytochrome b). All of the tribal clades as well as relationships among the tribes are supported by high bootstrap values, suggesting the utility of LW Rh in estimating tribal and subfamily rank for these bees. The sequences exhibit minimal base composition bias. Both 1st + 2nd and 3rd position sites provide information for estimating a reliable tree topology. These results suggest that LW Rh, which has not been reported previously in studies of organismal phylogenetics, could provide important new data from the nuclear genome for phylogeny reconstruction.

Animals↗

Adaptations of the helix-grip fold for ligand binding and catalysis in the START domain superfamily.

With a protein structure comparison, an iterative database search with sequence profiles, and a multiple-alignment analysis, we show that two domains with the helix-grip fold, the star-related lipid-transfer (START) domain of the MLN64 protein and the birch allergen, are homologous. They define a large, previously underappreciated superfamily that we call the START superfamily. In addition to the classical START domains that are primarily involved in eukaryotic signaling mediated by lipid binding and the birch antigen family that consists of plant proteins implicated in stress/pathogen response, the START superfamily includes bacterial polyketide cyclases/aromatases (e.g., TcmN and WhiE VI) and two families of previously uncharacterized proteins. The identification of this domain provides a structural prediction of an important class of enzymes involved in polyketide antibiotic synthesis and allows the prediction of their active site. It is predicted that all START domains contain a similar ligand-binding pocket. Modifications of this pocket determine the ligand-binding specificity and may also be the basis for at least two distinct enzymatic activities, those of a cyclase/aromatase and an RNase. Thus, the START domain superfamily is a rare case of the adaptation of a protein fold with a conserved ligand-binding mode for both a broad variety of catalytic activities and noncatalytic regulatory functions. Proteins 2001;43:134-144.

Allergens↗

Repeating sequence homologies in the p36 target protein of retroviral protein kinases and lipocortin, the p37 inhibitor of phospholipase A2.

Although considerable information has emerged on the molecular properties of the p36 target protein its function as well as the possible implications of its tyrosine phosphorylation have remained elusive. Here we show that all sequence segments of p36 published so far can be aligned by homology along the complete sequence of lipocortin, which has been reported recently. This alignment extends beyond multiple Geisow motifs, thought to indicate a sequence principle implicated in Ca2+ and/or lipid binding. While the latter properties are already established for p36 one may expect them also for lipocortin, an inhibitor of phospholipase A2 activity. Certain implications of these results are discussed.

Annexins↗

Genetic aspects of aromatic amino acid biosynthesis in Lactococcus lactis.

Polymerase chain reaction (PCR) primers designed from a multiple alignment of predicted amino acid sequences from bacterial aroA genes were used to amplify a fragment of Lactococcus lactis DNA. An 8 kb fragment was then cloned from a lambda library and the DNA sequence of a 4.4 kb region determined. This region was found to contain the genes tyrA, aroA, aroK, and pheA, which are involved in aromatic amino acid biosynthesis and folate metabolism. TyrA has been shown to be secreted and AroK also has a signal sequence, suggesting that these proteins have a secondary function, possibly in the transport of amino acids. The aroA gene from L. lactis has been shown to complement an E. coli mutant strain deficient in this gene. The arrangement of genes involved in aromatic amino acid biosynthesis in L. lactis appears to differ from that in other organisms.

3-Phosphoshikimate 1-Carboxyvinyltransferase↗

Proposed classification of the bipartite-genomed raspberry bushy dwarf idaeovirus, with tripartite-genomed viruses in the family Bromoviridae.

Raspberry bushy dwarf virus (RBDV) has an unusual combination of properties and has been classified as the sole member of a new plant virus genus, for which the name idaeovirus has been proposed. Particles of RBDV resemble those of ilarviruses (family Bromoviridae) in appearance and in being transmitted in association with pollen. RBDV has two genomic RNA species, RNA-1 (5,449 nt) and RNA-2 (2,231 nt). The particles also contain RNA-3 (946 nt), a subgenomic monocistronic coat protein mRNA which is derived from the 3' end of the bicistronic RNA-2. The single 190 K protein encoded by RNA-1 contains methyltransferase, helicase and polymerase domains. Evolutionary distance data obtained from multiple alignments of the amino acid sequence of the RBDV 190 K protein and corresponding proteins with replicative function from other plant viruses suggest that the closest affinities of RBDV are with the tripartite genomed viruses in the family Bromoviridae. We propose that the genus idaeovirus be included in the family Bromoviridae.

Cluster Analysis↗

An unusual seed-specific 3-ketoacyl-ACP synthase associated with the biosynthesis of petroselinic acid in coriander.

Petroselinic acid (18:1 delta6) is the major component of the seed oil of Umbelliferae species such as coriander (Coriandrum sativum) as well as Araliaceae and Garryaceae species. This unusual fatty acid is synthesized in plastids by the delta4 desaturation of palmitoyl-acyl carrier protein (16:0-ACP) and subsequent elongation of delta4-hexadecenoyl (16:1 delta4)-ACP. To characterize the enzymatic nature of the elongation reaction, an in vitro assay was developed with 16:1 delta4-ACP and 16:0-ACP as substrates. Extracts from developing coriander seeds elongated 16:1 delta4-ACP in a competitive assay at rates ten-fold greater than that with 16:0-ACP. In contrast, extracts from castor seeds, which do not synthesize petroselinic acid, displayed a strong preference for the elongation of 16:0-ACP rather than 16:1 A4-ACP. In addition, the elongation of 16:1 A4-ACP and 16:0-ACP by coriander seed extracts was strongly inhibited by cerulenin at concentrations as low as 10 microM. This finding suggested that the elongation of 16:1 A4-ACP and 16:0-ACP in coriander seed is catalyzed by a 3-ketoacyl-ACP synthase (KAS) 1-type enzyme(s), rather than a KAS II-type activity that is typically associated with 16:0-ACP elongation. Consistent with this, a cDNA for a diverged form of KAS I was isolated from a cDNA library prepared from developing coriander seed. Using a variety of heterologous probing techniques, no KAS II-type cDNAs could be identified in this library. Multiple alignment of KAS amino acid sequences indicated that, although the polypeptide corresponding to the coriander cDNA is more closely related to KAS I. its active site motif deviates from those found in both KAS I and KAS II enzymes. Also suggestive of a possible role in petroselinic acid synthesis, antibodies raised to the recombinant protein recognize an abundant 45 kDa polypeptide in coriander endosperm that is not detected in coriander leaves. These antibodies also recognize a major band of similar size in developing seeds of English ivy (Hedera helix), an Araliaceae species that also accumulates petroselinic acid in a seed-specific manner.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

Identification of a gadd45beta 3' enhancer that mediates SMAD3- and SMAD4-dependent transcriptional induction by transforming growth factor beta.

GADD45beta regulates cell growth, differentiation, and cell death following cellular exposure to diverse stimuli, including DNA damage and transforming growth factor-beta (TGFbeta). We examined how cells transduce the TGFbeta signal from the cell surface to the gadd45beta genomic locus and describe how GADD45beta contributes to TGFbeta biology. Following an alignment of gadd45beta genomic sequences from multiple organisms, we discovered a novel TGFbeta-responsive enhancer encompassing the third intron of the gadd45beta gene. Using three different experimental approaches, we found that SMAD3 and SMAD4, but not SMAD2, mediate transcription from this enhancer. Three lines of evidence support our conclusions. First, overexpression of SMAD3 and SMAD4 activated the transcriptional activity from this enhancer. Second, silencing of SMAD protein levels using short interfering RNAs revealed that TGFbeta-induced activation of the endogenous gadd45beta gene required SMAD3 and SMAD4 but not SMAD2. In contrast, we found that the regulation of plasminogen activator inhibitor type I depended upon all three SMAD proteins. Last, SMAD3 and SMAD4 reconstitution in SMAD-deficient cancer cells restored TGFbeta induction of gadd45beta. Finally, we assessed the function of GADD45beta within the TGFbeta response and found that GADD45beta-deficient cells arrested in G2 following TGFbeta treatment. These data support a role for SMAD3 and SMAD4 in activating gadd45beta through its third intron to facilitate G2 progression following TGFbeta treatment.

Animals↗

Comparative analysis of docking motifs in MAP-kinases and nuclear receptors.

Nuclear receptor (NR) agonists induce activation of mitogen-activated protein kinases (MAPK) through an yet unknown rapid non-genomic mechanism. Vice versa, NR are targets for phosphorylation by MAPK. By multiple alignment of the amino acid sequences and comparative analysis of the secondary and tertiary structures we identified four peptides in MAPK with similarity to bona fide protein-protein-interaction motifs in NR. In both molecule species, these motifs mediate selective docking to dimerization partners, coregulators or phosphoacceptors. We therefore propose that similar motifs may direct the site-specific association of NR with MAPK. Based on mutual allosteric interactions within a kinase-receptor complex, we discuss a novel principle how NR-agonists may regulate kinase activity and thus expression of hormone-dependent genes.

Amino Acid Motifs↗

ADAPTSITE: detecting natural selection at single amino acid sites.

UNLABELLED: ADAPTSITE is a program package for detecting natural selection at single amino acid sites, using a multiple alignment of protein-coding sequences for a given phylogenetic tree. The program infers ancestral codons at all interior nodes, and computes the total numbers of synonymous (c(S)) and nonsynonymous (c(N)) substitutions as well as the average numbers of synonymous (s(S)) and nonsynonymous (s(N)) sites for each codon site. The probabilities of occurrence of synonymous and nonsynonymous substitutions are approximated by s(S) / (s(S) + s(N)) and s(N) / (s(S) + s(N)), respectively. The null hypothesis of selective neutrality is tested for each codon site, assuming a binomial distribution for the probability of obtaining c(S) and c(N). AVAILABILITY: ADAPTSITE is available free of charge at the World-Wide Web sites http://mep.bio.psu.edu/adaptivevol.html and http://www.cib.nig.ac.jp/dda/yossuzuk/welcome.html. The package includes the source code written in C, binary files for UNIX operating systems, manual, and example files.

Algorithms↗

A method for detecting positive selection at single amino acid sites.

A method was developed for detecting the selective force at single amino acid sites given a multiple alignment of protein-coding sequences. The phylogenetic tree was reconstructed using the number of synonymous substitutions. Then, the neutrality was tested for each codon site using the numbers of synonymous and nonsynonymous changes throughout the phylogenetic tree. Computer simulation showed that this method accurately estimated the numbers of synonymous and nonsynonymous substitutions per site, as long as the substitution number on each branch was relatively small. The false-positive rate for detecting the selective force was generally low. On the other hand, the true-positive rate for detecting the selective force depended on the parameter values. Within the range of parameter values used in the simulation, the true-positive rate increased as the strength of the selective force and the total branch length (namely the total number of synonymous substitutions per site) in the phylogenetic tree increased. In particular, with the relative rate of nonsynonymous substitutions to synonymous substitutions being 5.0, most of the positively selected codon sites were correctly detected when the total branch length in the phylogenetic tree was > or = 2.5. When this method was applied to the human leukocyte antigen (HLA) gene, which included antigen recognition sites (ARSs), positive selection was detected mainly on ARSs. This finding confirmed the effectiveness of the present method with actual data. Moreover, two amino acid sites were newly identified as positively selected in non-ARSs. The three-dimensional structure of the HLA molecule indicated that these sites might be involved in antigen recognition. Positively selected amino acid sites were also identified in the envelope protein of human immunodeficiency virus and the influenza virus hemagglutinin protein. This method may be helpful for predicting functions of amino acid sites in proteins, especially in the present situation, in which sequence data are accumulating at an enormous speed.

Amino Acid Substitution↗

Analysis of the levels of conservation of the J domain among the various types of DnaJ-like proteins.

DnaJ-like proteins are defined by the presence of an approximately 73 amino acid region termed the J domain. This region bears similarity to the initial 73 amino acids of the Escherichia coli protein DnaJ. Although the structures of the J domains of E coli DnaJ and human heat shock protein 40 have been solved using nuclear magnetic resonance, no detailed analysis of the amino acid conservation among the J domains of the various DnaJ-like proteins has yet been attempted. A multiple alignment of 223 J domain sequences was performed, and the levels of amino acid conservation at each position were established. It was found that the levels of sequence conservation were particularly high in 'true' DnaJ homologues (ie, those that share full domain conservation with DnaJ) and decreased substantially in those J domains in DnaJ-like proteins that contained no additional similarity to DnaJ outside their J domain. Residues were also identified that could be important for stabilizing the J domain and for mediating the interaction with heat shock protein 70.

Amino Acid Sequence↗

Combining bioinformatics and phylogenetics to identify large sets of single-copy orthologous genes (COSII) for comparative, evolutionary and systematic studies: a test case in the euasterid plant clade.

We report herein the application of a set of algorithms to identify a large number (2869) of single-copy orthologs (COSII), which are shared by most, if not all, euasterid plant species as well as the model species Arabidopsis. Alignments of the orthologous sequences across multiple species enabled the design of "universal PCR primers," which can be used to amplify the corresponding orthologs from a broad range of taxa, including those lacking any sequence databases. Functional annotation revealed that these conserved, single-copy orthologs encode a higher-than-expected frequency of proteins transported and utilized in organelles and a paucity of proteins associated with cell walls, protein kinases, transcription factors, and signal transduction. The enabling power of this new ortholog resource was demonstrated in phylogenetic studies, as well as in comparative mapping across the plant families tomato (family Solanaceae) and coffee (family Rubiaceae). The combined results of these studies provide compelling evidence that (1) the ancestral species that gave rise to the core euasterid families Solanaceae and Rubiaceae had a basic chromosome number of x=11 or 12.2) No whole-genome duplication event (i.e., polyploidization) occurred immediately prior to or after the radiation of either Solanaceae or Rubiaceae as has been recently suggested.

Algorithms↗

[Cloning and identification of the priming glycosyltransferase gene involved in exopolysaccharide 139A biosynthesis in Streptomyces].

Recently in our laboratory, Streptomyces sp. 139 has been identified to produce a new exopolysaccharide designated EPS 139A that shows anti-rheumatic arthritis activity. The strategy of studying EPS 139A biosynthesis is to clone the key gene in the EPS biosynthesis pathway, i.e. the priming glycosyltransferase gene catalyzing the first step of nucleotide sugar transfer. Degenerate primers-based PCR approach was adopted to isolate the putative priming glycosyltransferase gene in Streptomyces sp. 139. According to the genes encoding the priming glycosyltransferases that have been identified in several microorganisms, a multiple alignment of the amino acid sequences of these genes was used to identify regions conserved between all genes. To clone the priming glycosyltransferase gene in Streptomyces sp. 139, degenerate primers were designed from these conserved regions taking into account information on Streptomyces codon usage to amplify an internal DNA fragment of this gene. A distinctive PCR product with the expected size of 0.3 kb was amplified from Streptomyces sp. 139 total genomic DNA. Sequence analysis showed that it is part of a putative priming glycosyltransferase gene and contains the predicted conserved domain B. To isolate the complete priming glycosyltransferase gene, a Streptomyces sp. 139 genomic library was constructed in the E. coli--Streptomyces shuttle vector pOJ446. Using the 0.3 kb PCR product of priming glycosyltransferase gene as a probe, 17 positive colonies were isolated by colony hybridization. A 4.0 kb BamHI fragment from all positive cosmids that hybridized to this probe was sequenced, which revealed the complete priming glycosyltransferase gene. The priming glycosyltransferase gene ste5 (GenBank under accession number AY131229) most likely begins with GTG, preceded by a probable ribosome binding site (RBS), GGGGA. It encodes a 492-amino-acid protein with molecular weight of 54 kDa and isoelectric point of 10.6. The G + C content of ste5 is 73%, close to the average of G + C content (74%) for Streptomyces. Moreover, the preference usage of G or C as third base of codons are found in the ste5, which is in accordance with the Streptomyces codon usage. A BlastP search showed that the C-terminal region of Ste5 shows highly homology with a number of priming glycosyltransferases from many different organisms. Ste5 contains two putative catalytic residues, Glu and Asp (residues 423 and 474) with a spacing of approximately 50 amino acids that conserved in various beta-glycosyltransferases. Moreover, the C-terminal one third of Ste5 contains three domains, A, B and C that is reported to be common to glycosyltransferases. By hydrophilicity plot prediction, the N-terminal two thirds of Ste5 exhibits 5 putative transmembrane domains. To investigate the involvement of the identified polysaccharide gene cluster in EPS 139A biosynthesis, the gene ste5 encoding priming glycosyltransferase was insertionally disrupted by a single-crossover homologous recombination event. A 0.85 kb internal fragment of ste5 was cloned into vector pKC1139 to yield pLY5015 that was transduced into Streptomyces sp. 139. Correct integration in Streptomyces LY1001 ste5- mutant strain was confirmed by Southern hybridization. After fermentation, no EPS 139A could be detected in the cultures of ste5- mutant strain Streptomyces LY1001. Therefore, the gene ste5 identified in this work is involved in the synthesis of the Streptomyces sp. 139 EPS.

Amino Acid Sequence↗