Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Structural and functional characterization of Mycobacterium tuberculosis CmtR, a PbII/CdII-sensing SmtB/ArsR metalloregulatory repressor.

The SmtB/ArsR family of prokaryotic metalloregulators are winged-helix transcriptional repressors that collectively provide resistance to a wide range of both biologically required and toxic heavy-metal ions. CmtR is a recently described Cd(II)/Pb(II) regulator expressed in Mycobacterium tuberculosis that is structurally distinct from the well-characterized SmtB/ArsR Cd(II)/Pb(II) sensor, Staphylococcus aureus plasmid pI258-encoded CadC. From functional analyses and a multiple sequence alignment of CmtR paralogs, M. tuberculosis CmtR is proposed to bind Pb(II) and Cd(II) via coordination by Cys57, Cys61, and Cys102 [Cavet et al. (2003) J. Biol. Chem. 278, 44560-44566]. We establish here that both wild-type and C102S CmtR are homodimers and bind Cd(II) and Pb(II) via formation of cysteine thiolate-rich coordination bonds. UV-vis optical spectroscopy, (113)Cd NMR spectroscopy (delta = 480 ppm), and (111m)Cd perturbed angular correlation (PAC) spectroscopy suggest two or three thiolate donors in the wild-type protein. Cys57 and Cys61 anchor the coordination complex, while Cys102 plays only an accessory role in stabilizing the metal chelate in the free protein because C102S CmtR binds Cd(II) and Zn(II) with only approximately 10-20-fold lower affinity relative to wild-type CmtR but approximately 100-1000-fold lower for Pb(II). Quantitative investigation of CmtR-cmt O/P binding equilibria using fluorescence anisotropy, however, reveals that Cys102 functions as a key allosteric metal ligand, because substitution of Cys102 abrogates disassembly of oligomeric CmtR-cmt O/P oligomeric complexes. The implications of these findings on the evolution of distinct metal-sensing sites in a family of homologous proteins are discussed.

Amino Acid Sequence↗

The periplasmic domains of Escherichia coli HflKC oligomerize through right-handed coiled-coil interactions.

The periplasmic domains of the Escherichia coli HflK and HflC were coexpressed and purified. The two polypeptides copurified in a 1:1 ratio, as determined by quantitative amino acid analysis. Circular dichroism studies showed the complex to have substantial helical/coiled-coil content that melted with midpoints in the range of 26-29 degrees C depending upon the concentration, implying a reversible oligomerization. The average molecular weight of the soluble HflKC determined by sedimentation equilibrium ultracentrifugation using a single-species model varied with rotor speed, providing further evidence of concentration-dependent oligomerization. The data were well-fit by models that specified a protomer to n-mer oligomerization, with the heterodimeric HflKC as the protomer and values of n between 7 and 10. Multiple-sequence alignments of both HflK and HflC revealed regions near the C termini to contain 11-residue hendecad repeats, indicative of right-handed coiled coils, with characteristic small residues in the a, f, h, and j positions. To test the importance of the small size of these positions, two residues in the HflC domain, Ala-262 in a f position and Gly-268 in an a position, were mutated to isoleucine. The HflKC:A262I mutant complex showed lower helicity than the wild type, and its melting was less concentration-dependent. During purification of HflKC:G268I, the mutated HflC subunit precipitated, leaving a preparation of the pure peripheral HflK domain. This polypeptide behaved as a monomer in sedimentation equilibrium experiments and showed low helicity, implying that the protein conformation is largely dependent upon heteromeric subunit interactions. These results demonstrate the importance of right-handed coiled-coil interactions in the oligomerization of HflKC, and a model entailing the formation of a right-handed helical barrel is proposed.

Alanine↗

A monofunctional and thermostable prephenate dehydratase from the archaeon Methanocaldococcus jannaschii.

Prephenate dehydratase (PDT) is an important but poorly characterized enzyme that is involved in the production of L-phenylalanine. Multiple-sequence alignments and a phylogenetic tree suggest that the PDT family has a common structural fold. On the basis of its sequence, the PDT from the extreme thermophile Methanocaldococcus jannaschii (MjPDT) was chosen as a promising representative of this family for pursuing structural and functional studies. The corresponding pheA gene was cloned and expressed in Escherichia coli. It encodes a monofunctional and thermostable enzyme with an N-terminal catalytic domain and a C-terminal regulatory ACT domain. Biophysical characterization suggests a dimeric (62 kDa) protein with mixed alpha/beta secondary structure elements. MjPDT unfolds in a two-state manner (Tm = 94 degrees C), and its free energy of unfolding [DeltaGU(H2O)] is 32.0 kcal/mol. The purified enzyme catalyzes the conversion of prephenate to phenylpyruvate according to Michaelis-Menten kinetics (kcat = 12.3 s-1 and Km = 22 microM at 30 degrees C), and its activity is pH-independent over the range of pH 5-10. It is feedback-inhibited by L-phenylalanine (Ki = 0.5 microM), but not by L-tyrosine or L-tryptophan. Comparison of its activation parameters (DeltaH(++)= 15 kcal/mol and DeltaS(++)= -3 cal mol-1 K-1) with those for the spontaneous reaction (DeltaH(++) = 17 kcal/mol and DeltaS(++)= -28 cal mol-1 K-1) suggests that MjPDT functions largely as an entropy trap. By providing a highly preorganized microenvironment for the dehydration-decarboxylation sequence, the enzyme may avoid the extensive solvent reorganization that accompanies formation of the carbocationic intermediate in the uncatalyzed reaction.

Amino Acid Sequence↗

Photoaffinity labeling with UMP of lysine 992 of carbamyl phosphate synthetase from Escherichia coli allows identification of the binding site for the pyrimidine inhibitor.

UMP is a highly specific reagent for photoaffinity labeling of the allosteric inhibitor site of carbamyl phosphate synthetase (CPS) from Escherichia coli and has been found to be photoincorporated in the COOH-terminal domain of the large subunit [Rubio et al. (1991) Biochemistry 30, 1068-1075]. In the present work we identify lysine 992 as the residue that is covalently attached to UMP. This identification is based on two lines of evidence. First, [14C]UMP is found to be incorporated between residues 939 and 1006, as shown by peptide mapping and by mass estimates of [14C]UMP-peptides generated by chemical and enzymatic cleavage of CPS. Secondly, we have purified two radioactive peptides derived exclusively from those enzyme molecules (approximately 5% of the total enzyme) that had incorporated [14C]-UMP. Edman analyses show the sequences of the labeled peptides (989)LVNXVHEGRPHIQD and (989)LVNXVHE to be overlapping. Since neither a phenylthiohydantoin (Pth) derivative (in cycle 4) nor any radioactivity is released from the membrane during sequencing, we can conclude that Lys992 and [14C]-UMP form a covalent adduct that remains bound to the membrane. Formation of this adduct agrees with all of the evidence and with the finding that UMP labeling prevents trypsin cleavage at Lys992. Lysine 992 is invariant in those CPSs that are inhibited by UMP, and is located 30 residues upstream of the site whose phosphorylation in hamster CAD reduces inhibition of CAD by UTP. Multiple sequence alignment of the residues surrounding Lys992 of the E. coli enzyme and the corresponding residues of the yeast and animal enzymes supports the existence of a uridine nucleotide binding fold in this region of the protein. We conclude that sequence changes in the binding fold provide a structural basis for the different regulatory properties found among CPSs I, II, and III.

Affinity Labels↗

Single amino acid substitutions disrupt tetramer formation in the dihydroneopterin aldolase enzyme of Pneumocystis carinii.

In the opportunistic pathogen Pneumocystis carinii, dihydroneopterin aldolase function is expressed as the N-terminal portion of the multifunctional folic acid synthesis protein (Fas). This region encompasses two domains, FasA and FasB, which are 27% amino acid identical. FasA and FasB also share significant amino acid sequence similarity with bacterial dihydroneopterin aldolases. In the present study, this enzyme function has been overproduced as an independent monofunctional activity in Escherichia coli. Recombinant FasAB-Met23 (amino acids 23-290 of the predicted open reading frame) was purified and shown to contain dihydroneopterin aldolase activity. The native FasAB-Met23 is a tetramer of the 30-kDa subunit, demonstrating characteristics of an associating-dissociating equilibrium system in which only the multimeric form of the enzyme is active. Multiple sequence alignment of FasA and FasB with other dihydroneopterin aldolases highlights only three positions where the amino acid is invariable between all the predicted proteins. The role of these conserved amino acid residues in enzyme function was investigated using site-directed mutagenesis. Mutant FasAB-Met23 species were overproduced and purified to near homogeneity. Three FasA domain mutants and two FasB domain mutants had little or no detectable dihydroneopterin aldolase activity, implicating both FasA and FasB in the catalytic mechanism. We show that each mutant protein containing an inactivating amino acid substitution has lost its ability to form stable tetramers.

Aldehyde-Lyases↗

Effects of mutations in M4 of the gastric H+,K+-ATPase on inhibition kinetics of SCH28080.

The effects of site-directed mutagenesis were used to explore the role of residues in M4 on the apparent Ki of a selective, K+-competitive inhibitor of the gastric H+,K+ ATPase, SCH28080. A double transfection expression system is described, utilizing HEK293 cells and separate plasmids encoding the alpha and beta subunits of the H+,K+-ATPase. The wild-type enzyme gave specific activity (micromoles of Pi per hour per milligram of expressed H+,K+-ATPase protein), apparent Km for ammonium (a K+ surrogate), and apparent Ki for SCH28080 equal to the H+, K+-ATPase purified from hog gastric mucosa. Amino acids in the M4 transmembrane segment of the alpha subunit were selected from, and substituted with, the nonconserved residues in M4 of the Na+, K+-ATPase, which is insensitive to SCH28080. Most of the mutations produced competent enzyme with similar Km,app values for NH4+ and Ki,app for SCH28080. SCH28080 affinity was decreased 2-fold in M330V and 9-fold in both M334I and V337I without significant effect on Km,app. Hence methionine 334 and valine 337 participate in binding but are not part of the NH4+ site. Methionine 330 may be at the periphery of the inhibitor site, which must have minimum dimensions of approximately 16 x 8 x 5 A and be accessible from the lumen in the E2-P conformation. Multiple sequence alignments place the membrane surface near arginine 328, suggesting that the side chains of methionine 334 and valine 337, on one side of the M4 helix, project into a binding cavity within the membrane domain.

Amino Acid Sequence↗

A versatile structural domain analysis server using profile weight matrices.

The WEB tool "AnDom" assigns to a given protein sequence all experimentally determined structural domains contained within it, including multidomain and large proteins. The server uses profile specific matrices from custom generated multiple sequence alignments of all known SCOP domains (SCOP version 1.50). Prediction time is short allowing numerous applications for structural genomics including investigation of complex eucaryotic protein families. The WWW server is at http://www.bork.embl-heidelberg.de/AnDom, and profiles can be downloaded at ftp.bork.embl-heidelberg.de/pub/users/ schmidt/AnDom.

Amino Acid Sequence↗

The footprint sorting problem.

Phylogenetic footprints are short pieces of noncoding DNA sequence in the vicinity of a gene that are conserved between evolutionary distant species. A seemingly simple problem is to sort footprints in their order along the genomes. It is complicated by the fact that not all footprints are collinear: they may cross each other. The problem thus becomes the identification of the crossing footprints, the sorting of the remaining collinear cliques, and finally the insertion of the noncollinear ones at "reasonable" positions. We show that solving the footprint sorting problem requires the solution of the "Minimum Weight Vertex Feedback Set Problem", which is known to be NP-complete and APX-hard. Nevertheless good approximations can be obtained for data sets of interest. The remaining steps of the sorting process are straightforward: computation of the transitive closure of an acyclic graph, linear extension of the resulting partial order, and finally sorting w.r.t. the linear extension. Alternatively, the footprint sorting problem can be rephrased as a combinatorial optimization problem for which approximate solutions can be obtained by means of general purpose heuristics. Footprint sortings obtained with different methods can be compared using a version of multiple sequence alignment that allows the identification of unambiguously ordered sublists. As an application we show that the rat has a slightly increased insertion/deletion rate in comparison to the mouse genome.

Journal Article↗

Architecture of P2Y nucleotide receptors: structural comparison based on sequence analysis, mutagenesis, and homology modeling.

Human P2Y receptors encompass at least eight subtypes of Class A G protein-coupled receptors (GPCRs), responding to adenine and/or uracil nucleotides. Using a BLAST search against the Homo sapiens subset of the SWISS-PROT and TrEMBL databases, we identified 68 proteins showing high similarity to P2Y receptors. To address the problem of low sequence identity between rhodopsin and the P2Y receptors, we performed a multiple-sequence alignment of the retrieved proteins and the template bovine rhodopsin, combining manual identification of the transmembrane domains (TMs) with automatic techniques. The resulting phylogenetic tree delineated two distinct subgroups of P2Y receptors: Gq-coupled subtypes (e.g., P2Y1) and those coupled to Gi (e.g., P2Y12). On the basis of sequence comparison we mutated three Tyr residues of the putative P2Y1 binding pocket to Ala and Phe and characterized pharmacologically the mutant receptors expressed in COS-7 cells. The mutation of Y306 (7.35, site of a cationic residue in P2Y12) or Y203 in the second extracellular loop selectively decreased the affinity of the agonist 2-MeSADP, and the Y306F mutation also reduced antagonist (MRS2179) affinity by 5-fold. The Y273A (6.48) mutation precluded the receptor activation without a major effect on the ligand-binding affinities, but the Y273F mutant receptor still activated G proteins with full agonist affinity. Thus, we have identified new recognition elements to further define the P2Y1 binding site and related these to other P2Y receptor subtypes. Following sequence-based secondary-structure prediction, we constructed complete models of all the human P2Y receptors by homology to rhodopsin. Ligand docking on P2Y1 and P2Y12 receptor models was guided by mutagenesis results, to identify the residues implicated in the binding process. Different sets of cationic residues in the two subgroups appeared to coordinate phosphate-bearing ligands. Within the P2Y1 subgroup these residues are R3.29, K/R6.55, and R7.39. Within the P2Y12 subgroup, the only residue in common with P2Y1 is R6.55, and the role of R3.29 in TM3 seems to be fulfilled by a Lys residue in EL2, whereas the R7.39 in TM7 seems to be substituted by K7.35. Thus, we have identified common and distinguishing features of P2Y receptor structure and have proposed modes of ligand binding for the two representative subtypes that already have well-developed ligands.

Amino Acid Sequence↗

Molecular docking-based study of vasopressin analogues modified at positions 2 and 3 with N-methylphenylalanine: influence on receptor-bound conformations and interactions with vasopressin and oxytocin receptors.

In this study, four cyclic vasopressin (CYFQNCPRG-NH(2), AVP) analogues substituted at positions 2 and 3 with four combinations of enantiomers of N-methylphenylalanine have been investigated. Three-dimensional structures of analogues have been formerly determined using NMR spectroscopy in dimethyl sulfoxide. Three-dimensional models of the vasopressin and oxytocin receptors were constructed by combining the multiple sequence alignment and the RD crystal structure as a template. The analogues have been docked into the receptor using the AutoDock program. The relaxation of the receptor-ligand complexes using energy minimization, followed by the constrained simulated annealing protocols (CSA), has been performed. The receptor-bound conformations of the investigated analogues have been proposed. We concluded that the N-methylated residues at positions 2 and 3 act as a structural restraint, determining the conformation of analogues, their location inside the receptor cavity, and mutual arrangement of the aromatic side chains. The conserved polar residues constitute the handles keeping the biologically active analogues inside the binding cavity. The Arg(8)-D(2.65) salt bridge might be responsible for analogue-selective binding in OTR and V1aR versus V2R, where the positively charged K(2.65) 100 is present at the equivalent position.

Amino Acid Sequence↗

Counting the zinc-proteins encoded in the human genome.

Metalloproteins are proteins capable of binding one or more metal ions, which may be required for their biological function, or for regulation of their activities or for structural purposes. Genome sequencing projects have provided a huge number of protein primary sequences, but, even though several different elaborate analyses and annotations have been enabled by a rich and ever-increasing portfolio of bioinformatic tools, metal-binding properties remain difficult to predict as well as to investigate experimentally. Consequently, the present knowledge about metalloproteins is only partial. The present bioinformatic research proposes a strategy to answer the question of how many and which proteins encoded in the human genome may require zinc for their physiological function. This is achieved by a combination of approaches, which include: (i) searching in the proteome for the zinc-binding patterns that, on their turn, are obtained from all available X-ray data; (ii) using libraries of metal-binding protein domains based on multiple sequence alignments of known metalloproteins obtained from the Pfam database; and (iii) mining the annotations of human gene sequences, which are based on any type of information available. It is found that 1684 proteins in the human proteome are independently identified by all three approaches as zinc-proteins, 746 are identified by two, and 777 are identified by only one method. By assuming that all proteins identified by at least two approaches are truly zinc-binding and inspecting the proteins identified by a single method, it can be proposed that ca. 2800 human proteins are potentially zinc-binding in vivo, corresponding to 10% of the human proteome, with an uncertainty of 400 sequences. Available functional information suggests that the large majority of human zinc-binding proteins are involved in the regulation of gene expression. The most abundant class of zinc-binding proteins in humans is that of zinc-fingers, with Cys4 and Cys2His2 being the most common types of coordination environment.

Computational Biology↗

Ferric hydroxamate binding protein FhuD from Escherichia coli: mutants in conserved and non-conserved regions.

Uptake of iron complexes into the gram-negative bacterial cell requires highly specific outer membrane receptors and specific ATP-dependent (ATP-Binding-Cassette (ABC)) transport systems located in the inner membrane. The latter type of import system is characterized by a periplasmic binding protein (BP), integral membrane proteins, and membrane-associated ATP-hydrolyzing proteins. In gram-positive bacteria lacking the periplasmic space, the binding proteins are lipoproteins tethered to the cytoplasmic membrane. To date, there is little structural information about the components of ABC transport systems involved in iron complex transport. The recently determined structure of the Escherichia coli periplasmic ferric siderophore binding protein FhuD is unique for an ABC transport system (Clarke et al. 2000). Unlike other BP's, FhuD has two domains connected by a long alpha-helix. The ligand binds in a shallow pocket between the two domains. In vivo and in vitro analysis of single amino acid mutants of FhuD identified several residues that are important for proper functioning of the protein. In this study, the mutated residues were mapped to the protein structure to define special areas and specific amino acid residues in E. coli FhuD that are vital for correct protein function. A number of these important residues were localized in conserved regions according to a multiple sequence alignment of E. coli FhuD with other BP's that transport siderophores, heme, and vitamin B12. The alignment and structure prediction of these polypeptides indicate that they form a distinct family of periplasmic binding proteins.

Amino Acid Sequence↗

Structural investigations on human erythrocyte acylpeptide hydrolase by mass spectrometric procedures.

The complete primary structure of human erythrocyte acylpeptide hydrolase has been determined by using a combination of different mass spectrometric procedures and sequencing techniques. These data allowed us to correct the incomplete nucleotide sequence of the DNF15S2 locus on the short arm of human chromosome 3 at region 21, coding for the enzyme. The protein consists of 732 amino acid residues and is acetylated at the N-terminus. Alkylation experiments on the native enzyme demonstrated that all 17 cysteine residues present in the polypeptide chain are in reduced form. Multiple sequence alignment did not reveal striking similarity with proteases of known tertiary structure with the exception of members of the serine oligopeptidase family. Limited proteolysis experiments generated a C-terminal portion, containing all the catalytic triad elements responsible for proteolytic activity, and an N-terminal domain of unknown function, both still strongly associated in a completely active nicked form. The site of tryptic hydrolysis was identified as Arg193. The secondary structural organization of the protease domain of the enzyme is consistent with the alpha/beta hydrolase fold.

Alkylation↗

Genome organization in Arabidopsis thaliana: a survey for genes involved in isoprenoid and chlorophyll metabolism.

The isoprenoid biosynthetic pathway provides intermediates for the synthesis of a multitude of natural products which serve numerous biochemical functions in plants: sterols (isoprenoids with a C30 backbone) are essential components of membranes; carotenoids (C40) and chlorophylls (which contain a C20 isoprenoid side-chain) act as photosynthetic pigments; plastoquinone, phylloquinone and ubiquinone (all of which contain long isoprenoid side-chains) participate in electron transport chains; gibberellins (C20), brassinosteroids (C30) and abscisic acid (C15) are phytohormones derived from isoprenoid intermediates; prenylation of proteins (with C15 or C20 isoprenoid moieties) may mediate subcellular targeting and regulation of activity; and several monoterpenes (C10), sesquiterpenes (C15) and diterpenes (C20) have been demonstrated to be involved in plant defense. Here we present a comprehensive analysis of genes coding for enzymes involved in the metabolism of isoprenoid-derived compounds in Arabidopsis thaliana. By combining homology and sequence motif searches with knowledge regarding the phylogenetic distribution of pathways of isoprenoid metabolism across species, candidate genes for these pathways in A. thaliana were obtained. A detailed analysis of the vicinity of chromosome loci for genes of isoprenoid metabolism in A. thaliana provided evidence for the clustering of genes involved in common pathways. Multiple sequence alignments were used to estimate the number of genes in gene families and sequence relationship trees were utilized to classify their individual members. The integration of all these datasets allows the generation of a knowledge-based metabolic map of isoprenoid metabolic pathways in A. thaliana and provides a substantial improvement of the currently available gene annotation.

Abscisic Acid↗

Cloning and sequencing of coat protein gene of an Indian potato leaf roll virus (PLRV) isolate and its similarity with other members of Luteoviridae.

An Indian strain of potato leaf roll virus (PLRV) was purified to generate complementary DNA corresponding to the coat protein (CP) gene. Virus cDNA was synthesized from purified viral RNA using oligo (dT)-anchor primer and virus specific primers. The viral sequence encoding the coat protein was specifically amplified by polymerase chain reaction (PCR), using specific primers bordering the CP gene. The unique amplified product thus obtained was A-T cloned into the pGEM-T Easy vector and the authenticity of the cloned gene verified by dot blot hybridization and sequence analysis. Run-way-transcripts of the cloned CP gene could detect PLRV in tissue imprints and tissue dilution. The nucleotide sequences and the deduced amino acid sequences were compared with the other PLRV isolates and found to be 97-99% identical at both the nucleotide and amino acid sequence level of other isolates. Multiple sequence alignment of deduced amino acid sequences revealed considerable homology to other luteoviruses. A nuclear localization signal located close to the N-terminus of the CP gene was predicted. This is the first report of PLRV coat protein sequence from an Indian strain.

Amino Acid Sequence↗

Predicting differential antigen-antibody contact regions based on solvent accessibility.

A novel computational approach was examined for predicting epitopes from primary structures of the seven immunologically distinct botulinum neurotoxins (BoNT/A-G) and tetanus toxin (TeTX). An artificial neural network [Rost and Sander (1994), Proteins 20, 216] was used to estimate residue solvent accessibilities in multiple aligned sequences. A similar network trained to predict secondary structures was also used to examine this protein family, whose tertiary fold is presently unknown. The algorithm was validated by showing that it was 80% accurate in determining the secondary structure of avian egg-white lysozyme and that it correctly identified highly solvent-exposed residues that correspond to the major contact regions of lysozyme-antibody cocrystals. When sequences of the heavy (H) chains of TeTX and BoNT/A-G were analyzed, this algorithm predicted that the most highly exposed regions were clustered at the sequentially nonconserved N- and C-termini [Lebeda and Olson (1994), Proteins 20, 293]. The secondary structures and the remaining highly solvent-accessible regions were, in contrast, predicted to be conserved. In experiments reported by others, H-chain fragments that induced immunological protection against BoNT/A overlap with these predicted most highly exposed regions. It is also known that the C-terminal halves of the TeTX and BoNT/A H-chains interfere with holotoxin binding to ectoacceptors on nerve endings. Thus, the present results provide a theoretical framework for predicting the sites that could assist in the development of genetically engineered vaccines and that could interact with neurally located toxin ectoacceptors. Finally, because the most highly solvent-exposed regions were not well conserved, it is hypothesized that nonconserved, potential contact sites partially account for the existence of different dominant binding regions for type-specific neutralizing antibodies.

Algorithms↗

On the classification and evolution of protein modules.

Our efforts to classify the functional units of many proteins, the modules, are reviewed. The data from the sequencing projects for various model organisms are extremely helpful in deducing the evolution of proteins and modules. For example, a dramatic increase of modular proteins can be observed from yeast to C. elegans in accordance with new protein functions that had to be introduced in multicellular organisms. Our sequence characterization of modules relies on sensitive similarity search algorithms and the collection of multiple sequence alignments for each module. To trace the evolution of modules and to further automate the classification, we have developed a sequence and a module alerting system that checks newly arriving sequence data for the presence of already classified modules. Using these systems, we were able to identify an unexpected similarity between extracellular C1Q modules with bacterial proteins.

Amino Acid Sequence↗

Prokaryotic orthologues of mitochondrial alternative oxidase and plastid terminal oxidase.

The mitochondrial alternative oxidase (AOX) and the plastid terminal oxidase (PTOX) are two similar members of the membrane-bound diiron carboxylate group of proteins. AOX is a ubiquinol oxidase present in all higher plants, as well as some algae, fungi, and protists. It may serve to dampen reactive oxygen species generation by the respiratory electron transport chain. PTOX is a plastoquinol oxidase in plants and some algae. It is required in carotenoid biosynthesis and may represent the elusive oxidase in chlororespiration. Recently, prokaryotic orthologues of both AOX and PTOX proteins have appeared in sequence databases. These include PTOX orthologues present in four different cyanobacteria as well as an AOX orthologue in an alpha-proteobacterium. We used PCR, RT-PCR and northern analyses to confirm the presence and expression of the PTOX gene in Anabaena variabilis PCC 7120. An extensive phylogeny of newly found prokaryotic and eukaryotic AOX and PTOX proteins supports the idea that AOX and PTOX represent two distinct groups of proteins that diverged prior to the endosymbiotic events that gave rise to the eukaryotic organelles. Using multiple sequence alignment, we identified residues conserved in all AOX and PTOX proteins. We also provide a scheme to readily distinguish PTOX from AOX proteins based upon differences in amino acid sequence in motifs around the conserved iron-binding residues. Given the presence of PTOX in cyanobacteria, we suggest that this acronym now stand for plastoquinol terminal oxidase. Our results have implications for the photosynthetic and respiratory metabolism of these prokaryotes, as well as for the origin and evolution of eukaryotic AOX and PTOX proteins.

Amino Acid Sequence↗