Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

An L40C mutation converts the cysteine-sulfenic acid redox center in enterococcal NADH peroxidase to a disulfide.

Multiple sequence alignments including the enterococcal NADH peroxidase and NADH oxidase indicate that residues Ser38 and Cys42 align with the two cysteines of the redox-active disulfides found in glutathione reductase (GR), lipoamide dehydrogenase, mercuric reductase, and trypanothione reductase. In order to evaluate those structural determinants involved in the selection of the cysteine-sulfenic acid (Cys-SOH) redox centers found in the two peroxide reductases and the redox-active disulfides present in the GR class of disulfide reductases, NADH peroxidase residues Ser38, Phe39, Leu40, and Ser41 have been individually replaced with Cys. Both the F39C and L40C mutant peroxidases yield active-site disulfides involving the new Cys and the native Cys42; formation of the Cys39-Cys42 disulfide, however, precludes binding of the FAD coenzyme. In contrast, the L40C mutant contains tightly-bound FAD and has been analyzed by both kinetic and spectroscopic approaches. In addition, the L40C and S41C mutant structures have been determined at 2.1 and 2.0 A resolution, respectively, by X-ray crystallography. Formation of the Cys40-Cys42 disulfide bond requires a movement of Cys42-SG to a new position 5.9 A from the flavin-C(4a) position; this is consistent with the inability of the new disulfide to function as a redox center in concert with the flavin. Stereochemical constraints prohibit formation of the Cys41-Cys42 disulfide in the latter mutant.

Amino Acid Sequence

A proteolytically sensitive region common to several rat liver cytochromes P450: effect of cleavage on substrate binding.

Limited proteolysis of rat liver microsomes was used to probe the topography and structure of cytochrome P450 bound to the endoplasmic reticulum. Three cytochromes P450 from two families were examined. Monoclonal antibodies to cytochrome P450 forms 1A1, 2B1, and 2E1 were used to immunopurify these proteolyzed cytochromes P450 from microsomes from rats treated with 3-methylcholanthrene, phenobarbital, and acetone, respectively. Electrophoretic and immunoblot analysis of tryptic fragments revealed a highly sensitive cleavage site in all three cytochromes P450. N-Terminal sequencing was performed on the fragments after transfer onto poly(vinylidene difluoride) membranes and showed that this preferential cleavage site is at amino acid position 298 of P450 1A1, position 277 of P450 2B1, and position 278 of P450 2E1. Multiple sequence alignment revealed that these positions are at the amino terminal of a highly conserved region of these cytochromes P450. The important functional role implied by primary sequence conservation along with the proteolytic sensitivity at its amino terminal suggests that this region is a protein domain. Comparison with the known structure of the bacterial cytochrome P450cam predicts that this proteolytically sensitive site is within an interhelical turn region connected to the distal helix that partially encompasses the heme-containing active site. Substrate binding to the cleaved cytochromes P450 was examined in order to determine whether the newly added conformational freedom near the cleavage site functionally altered these cytochromes P450. Cleavage of P450 2B1 abolished benzphetamine binding, which indicates that the cleavage site contains an important structural determinant for binding this substrate. However, cleavage did not affect benzo[a]pyrene binding to P450 1A1.

Amino Acid Sequence

Identification of Cys-150 in the active site of phosphomannose isomerase from Candida albicans.

Candida albicans phosphomannose isomerase (PMI) (EC 5.3.1.8) has been recently cloned and overexpressed in Escherichia coli. The enzyme can be irreversibly inactivated by iodoacetate in 50 mM borate buffer, pH 9.0, in a time-dependent manner at a rate of 4.2 +/- 0.03 min-1 M-1. This inhibition can be prevented by the substrate mannose 6-phosphate with a Ks of 0.22 +/- 0.05 mM, slightly lower than its Km value. However, metals such as zinc and cadmium, which are reversible, competitive inhibitors for PMI, do not protect the enzyme against modification. The protein has been labeled by using [2-14C]iodoacetate, in the presence or absence of substrate, and the protein is fully inactivated when 1.0 thiol group is modified per molecule of enzyme. Tryptic maps of the modified protein have been produced. The protected peptide has been identified and sequenced, and the phenylthiohydantoin amino acids have been collected. The modified amino acid is Cys-150. This cysteine residue is conserved in mammalian and yeast phosphomannose isomerases, but not in bacterial species where it is replaced with asparagine. We therefore purified PMI from E. coli and showed that this enzyme is not sensitive to inactivation by iodoacetate. The iodoacetate is presumably inhibiting PMI by sterically blocking the mannose 6-phosphate binding site. Multiple sequence alignment procedures were used to try to identify potential ligands of the zinc atom that is essential for enzyme activity and thus to delineate the active site region.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

Photoaffinity labeling with UMP of lysine 992 of carbamyl phosphate synthetase from Escherichia coli allows identification of the binding site for the pyrimidine inhibitor.

UMP is a highly specific reagent for photoaffinity labeling of the allosteric inhibitor site of carbamyl phosphate synthetase (CPS) from Escherichia coli and has been found to be photoincorporated in the COOH-terminal domain of the large subunit [Rubio et al. (1991) Biochemistry 30, 1068-1075]. In the present work we identify lysine 992 as the residue that is covalently attached to UMP. This identification is based on two lines of evidence. First, [14C]UMP is found to be incorporated between residues 939 and 1006, as shown by peptide mapping and by mass estimates of [14C]UMP-peptides generated by chemical and enzymatic cleavage of CPS. Secondly, we have purified two radioactive peptides derived exclusively from those enzyme molecules (approximately 5% of the total enzyme) that had incorporated [14C]-UMP. Edman analyses show the sequences of the labeled peptides (989)LVNXVHEGRPHIQD and (989)LVNXVHE to be overlapping. Since neither a phenylthiohydantoin (Pth) derivative (in cycle 4) nor any radioactivity is released from the membrane during sequencing, we can conclude that Lys992 and [14C]-UMP form a covalent adduct that remains bound to the membrane. Formation of this adduct agrees with all of the evidence and with the finding that UMP labeling prevents trypsin cleavage at Lys992. Lysine 992 is invariant in those CPSs that are inhibited by UMP, and is located 30 residues upstream of the site whose phosphorylation in hamster CAD reduces inhibition of CAD by UTP. Multiple sequence alignment of the residues surrounding Lys992 of the E. coli enzyme and the corresponding residues of the yeast and animal enzymes supports the existence of a uridine nucleotide binding fold in this region of the protein. We conclude that sequence changes in the binding fold provide a structural basis for the different regulatory properties found among CPSs I, II, and III.

Affinity Labels

Single amino acid substitutions disrupt tetramer formation in the dihydroneopterin aldolase enzyme of Pneumocystis carinii.

In the opportunistic pathogen Pneumocystis carinii, dihydroneopterin aldolase function is expressed as the N-terminal portion of the multifunctional folic acid synthesis protein (Fas). This region encompasses two domains, FasA and FasB, which are 27% amino acid identical. FasA and FasB also share significant amino acid sequence similarity with bacterial dihydroneopterin aldolases. In the present study, this enzyme function has been overproduced as an independent monofunctional activity in Escherichia coli. Recombinant FasAB-Met23 (amino acids 23-290 of the predicted open reading frame) was purified and shown to contain dihydroneopterin aldolase activity. The native FasAB-Met23 is a tetramer of the 30-kDa subunit, demonstrating characteristics of an associating-dissociating equilibrium system in which only the multimeric form of the enzyme is active. Multiple sequence alignment of FasA and FasB with other dihydroneopterin aldolases highlights only three positions where the amino acid is invariable between all the predicted proteins. The role of these conserved amino acid residues in enzyme function was investigated using site-directed mutagenesis. Mutant FasAB-Met23 species were overproduced and purified to near homogeneity. Three FasA domain mutants and two FasB domain mutants had little or no detectable dihydroneopterin aldolase activity, implicating both FasA and FasB in the catalytic mechanism. We show that each mutant protein containing an inactivating amino acid substitution has lost its ability to form stable tetramers.

Aldehyde-Lyases

Effects of mutations in M4 of the gastric H+,K+-ATPase on inhibition kinetics of SCH28080.

The effects of site-directed mutagenesis were used to explore the role of residues in M4 on the apparent Ki of a selective, K+-competitive inhibitor of the gastric H+,K+ ATPase, SCH28080. A double transfection expression system is described, utilizing HEK293 cells and separate plasmids encoding the alpha and beta subunits of the H+,K+-ATPase. The wild-type enzyme gave specific activity (micromoles of Pi per hour per milligram of expressed H+,K+-ATPase protein), apparent Km for ammonium (a K+ surrogate), and apparent Ki for SCH28080 equal to the H+, K+-ATPase purified from hog gastric mucosa. Amino acids in the M4 transmembrane segment of the alpha subunit were selected from, and substituted with, the nonconserved residues in M4 of the Na+, K+-ATPase, which is insensitive to SCH28080. Most of the mutations produced competent enzyme with similar Km,app values for NH4+ and Ki,app for SCH28080. SCH28080 affinity was decreased 2-fold in M330V and 9-fold in both M334I and V337I without significant effect on Km,app. Hence methionine 334 and valine 337 participate in binding but are not part of the NH4+ site. Methionine 330 may be at the periphery of the inhibitor site, which must have minimum dimensions of approximately 16 x 8 x 5 A and be accessible from the lumen in the E2-P conformation. Multiple sequence alignments place the membrane surface near arginine 328, suggesting that the side chains of methionine 334 and valine 337, on one side of the M4 helix, project into a binding cavity within the membrane domain.

Amino Acid Sequence

Structural investigations on human erythrocyte acylpeptide hydrolase by mass spectrometric procedures.

The complete primary structure of human erythrocyte acylpeptide hydrolase has been determined by using a combination of different mass spectrometric procedures and sequencing techniques. These data allowed us to correct the incomplete nucleotide sequence of the DNF15S2 locus on the short arm of human chromosome 3 at region 21, coding for the enzyme. The protein consists of 732 amino acid residues and is acetylated at the N-terminus. Alkylation experiments on the native enzyme demonstrated that all 17 cysteine residues present in the polypeptide chain are in reduced form. Multiple sequence alignment did not reveal striking similarity with proteases of known tertiary structure with the exception of members of the serine oligopeptidase family. Limited proteolysis experiments generated a C-terminal portion, containing all the catalytic triad elements responsible for proteolytic activity, and an N-terminal domain of unknown function, both still strongly associated in a completely active nicked form. The site of tryptic hydrolysis was identified as Arg193. The secondary structural organization of the protease domain of the enzyme is consistent with the alpha/beta hydrolase fold.

Alkylation

Predicting differential antigen-antibody contact regions based on solvent accessibility.

A novel computational approach was examined for predicting epitopes from primary structures of the seven immunologically distinct botulinum neurotoxins (BoNT/A-G) and tetanus toxin (TeTX). An artificial neural network [Rost and Sander (1994), Proteins 20, 216] was used to estimate residue solvent accessibilities in multiple aligned sequences. A similar network trained to predict secondary structures was also used to examine this protein family, whose tertiary fold is presently unknown. The algorithm was validated by showing that it was 80% accurate in determining the secondary structure of avian egg-white lysozyme and that it correctly identified highly solvent-exposed residues that correspond to the major contact regions of lysozyme-antibody cocrystals. When sequences of the heavy (H) chains of TeTX and BoNT/A-G were analyzed, this algorithm predicted that the most highly exposed regions were clustered at the sequentially nonconserved N- and C-termini [Lebeda and Olson (1994), Proteins 20, 293]. The secondary structures and the remaining highly solvent-accessible regions were, in contrast, predicted to be conserved. In experiments reported by others, H-chain fragments that induced immunological protection against BoNT/A overlap with these predicted most highly exposed regions. It is also known that the C-terminal halves of the TeTX and BoNT/A H-chains interfere with holotoxin binding to ectoacceptors on nerve endings. Thus, the present results provide a theoretical framework for predicting the sites that could assist in the development of genetically engineered vaccines and that could interact with neurally located toxin ectoacceptors. Finally, because the most highly solvent-exposed regions were not well conserved, it is hypothesized that nonconserved, potential contact sites partially account for the existence of different dominant binding regions for type-specific neutralizing antibodies.

Algorithms

On the classification and evolution of protein modules.

Our efforts to classify the functional units of many proteins, the modules, are reviewed. The data from the sequencing projects for various model organisms are extremely helpful in deducing the evolution of proteins and modules. For example, a dramatic increase of modular proteins can be observed from yeast to C. elegans in accordance with new protein functions that had to be introduced in multicellular organisms. Our sequence characterization of modules relies on sensitive similarity search algorithms and the collection of multiple sequence alignments for each module. To trace the evolution of modules and to further automate the classification, we have developed a sequence and a module alerting system that checks newly arriving sequence data for the presence of already classified modules. Using these systems, we were able to identify an unexpected similarity between extracellular C1Q modules with bacterial proteins.

Amino Acid Sequence

An investigation of the role of Glu-842, Glu-844 and His-846 in the function of the cytoplasmic domain of the epidermal growth factor receptor.

Activation of several protein kinases is mediated, at least in part, by phosphorylation of conserved Thr or Tyr residues located in a variable loop region, near the active site. In certain kinases, this activation loop also controls access of peptide substrates to the active site. In the corresponding region of the epidermal growth factor (EGF) receptor, a potential phosphorylation site, Tyr-845, does not appear to have a major regulatory role. In order to find out whether this variable loop can modulate the peptide phosphorylation and self-phosphorylation activities of the EGF receptor kinase, we investigated the role of residues around Tyr-845, using site-directed mutagenesis. Multiple sequence alignment showed that residues Glu-842, Glu-844 and His-846 are conserved or nearly conserved in eight members of the EGF receptor family. Mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala were expressed in the baculovirus/insect cell system, purified to near-homogeneity and characterized with respect to their peptide phosphorylation and self-phosphorylation activities. All three mutants were active, and these changes did not affect ATP binding directly. However, all mutations increased the Km(app.) for peptide substrates and MnATP in peptide phosphorylation reactions. The Vmax. for the phosphorylation of peptide RREELQDDYEDD was unaltered, but the Vmax. for self-phosphorylation (with variable [MnATP]) decreased 4-, 2- and 7-fold for mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala respectively, compared with the wild-type. These results suggest that binding of this peptide restored an optimal conformation at the active site that might be impaired by the mutations. A study of the dependence of initial rates of self-phosphorylation on cytoplasmic domain concentration showed that the order of reaction increased with the progress of self-phosphorylation. Both pre-phosphorylation and high concentrations of ammonium sulphate restored maximal or near-maximal levels of self-phosphorylation in the mutants, possibly through compensating conformational changes. A plausible homology model, based on the cyclic AMP-dependent protein kinase catalytic subunit, accommodated the sequence Glu-841-Glu-Lys-Glu as an insertion in the peptide binding loop at the edge of the active site cleft. The model suggests that Glu-844 and His-846 may participate in H-bonding interactions, thus stabilizing the active site region, while Glu-842 does not appear to interact with regions of the catalytic core.

Amino Acid Sequence

Prediction from sequence comparisons of residues of factor H involved in the interaction with complement component C3b.

The amino acid sequence of the region of bovine factor H containing the C3b binding site has been derived from sequencing overlapping cDNA clones. A cDNA sequence encoding 669 amino acids was obtained. Like human and mouse factor H the sequence can be arranged into a number of internally homologous units (CPs), each of which is about 60 amino acids long and is based on a framework of four conserved cysteine residues. Bovine factor H is of the same molecular mass as human and mouse factor H, and is therefore likely to be composed of 20 contiguous CPs. Comparisons with human and mouse factor H indicate that the partial bovine sequence encodes CPs 2-12 inclusive of bovine factor H. Bovine factor H binds to human ammonia-treated C3 (causing thiolester cleavage) [C3(NH3)] and promotes the cleavage of human C3(NH3) in the presence of bovine factor I. Other studies indicate that CPs 2-5 of human factor H encompass the C3b binding and factor I cofactor activity site. Multiple sequence alignments of human factor H, mouse factor H (which also interacts with human C3b) and bovine factor H with CP modules whose structures have been determined experimentally, have been used to predict residues in the hypervariable loops of CPs 2-5 and to identify residues of potential importance in human C3 binding and factor I cofactor activity. Leu-17 and Gly-20 of CP 2, Ser-17, Ala-19, Glu-21, Asp-23 and Glu-25 of CP 3 and Lys-18 of CP 4 are all conserved between the three species. It may be that CPs 3 and 4 interact with C3(NH3) directly, whilst CPs 2 and 5 maintain the correct orientation for CPs 3 and 4 to interact.

Amino Acid Sequence

Identification of sequence similarity between 60 kDa and 70 kDa molecular chaperones: evidence for a common evolutionary background?

Recent findings support the premise that chaperonins (60 kDa stress-proteins) and alpha-subunits of F-type ATPases (alpha-ATPase) are evolutionary related protein families. Two-dimensional gel patterns of synthesized proteins in unstressed and heat-shocked embryonic Drosophila melanogaster SL2 cells revealed that antibodies raised against the alpha-subunit of the F1-ATPase complex from rat liver recognize an inducible p71 member of the 70 kDa stress-responsive protein family. Molecular recognition of this stress-responsive 70 kDa protein by antibodies raised against the F1-ATPase alpha-subunit suggests the possibility of partial sequence similarity within these ATP-binding protein families. A multiple sequence alignment between alpha-ATPases and 60 kDa and 70 kDa molecular chaperones is presented. Statistical evaluation of sequence similarity reveals a significant degree of sequence conservation within the three protein families. The finding suggests a common evolutionary origin for the ATPases and molecular chaperone protein families of 60 kDa and 70 kDa, despite the lack of obvious structural resemblance between them.

Amino Acid Sequence

Comparative anatomy of the aldo-keto reductase superfamily.

The aldo-keto reductases metabolize a wide range of substrates and are potential drug targets. This protein superfamily includes aldose reductases, aldehyde reductases, hydroxysteroid dehydrogenases and dihydrodiol dehydrogenases. By combining multiple sequence alignments with known three-dimensional structures and the results of site-directed mutagenesis studies, we have developed a structure/function analysis of this superfamily. Our studies suggest that the (alpha/beta)8-barrel fold provides a common scaffold for an NAD(P)(H)-dependent catalytic activity, with substrate specificity determined by variation of loops on the C-terminal side of the barrel. All the aldo-keto reductases are dependent on nicotinamide cofactors for catalysis and retain a similar cofactor binding site, even among proteins with less than 30% amino acid sequence identity. Likewise, the aldo-keto reductase active site is highly conserved. However, our alignments indicate that variation ofa single residue in the active site may alter the reaction mechanism from carbonyl oxidoreduction to carbon-carbon double-bond reduction, as in the 3-oxo-5beta-steroid 4-dehydrogenases (Delta4-3-ketosteroid 5beta-reductases) of the superfamily. Comparison of the proposed substrate binding pocket suggests residues 54 and 118, near the active site, as possible discriminators between sugar and steroid substrates. In addition, sequence alignment and subsequent homology modelling of mouse liver 17beta-hydroxysteroid dehydrogenase and rat ovary 20alpha-hydroxysteroid dehydrogenase indicate that three loops on the C-terminal side of the barrel play potential roles in determining the positional and stereo-specificity of the hydroxysteroid dehydrogenases. Finally, we propose that the aldo-keto reductase superfamily may represent an example of divergent evolution from an ancestral multifunctional oxidoreductase and an example of convergent evolution to the same active-site constellation as the short-chain dehydrogenase/reductase superfamily.

Alcohol Oxidoreductases

Cloning and sequencing of four new mammalian monocarboxylate transporter (MCT) homologues confirms the existence of a transporter family with an ancient past.

Measurement of monocarboxylate transport kinetics in a range of cell types has provided strong circumstantial evidence for a family of monocarboxylate transporters (MCTs). Two mammalian MCT isoforms (MCT1 and MCT2) and a chicken isoform (REMP or MCT3) have already been cloned, sequenced and expressed, and another MCT-like sequence (XPCT) has been identified. Here we report the identification of new human MCT homologues in the database of expression sequence tags and the cloning and sequencing of four new full-length MCT-like sequences from human cDNA libraries, which we have denoted MCT3, MCT4, MCT5 and MCT6. Northern blotting revealed a unique tissue distribution for the expression of mRNA for each of the seven putative MCT isoforms (MCT1-MCT6 and XPCT). All sequences were predicted to have 12 transmembrane (TM) helical domains with a large intracellular loop between TM6 and TM7. Multiple sequence alignments showed identities ranging from 20% to 55%, with the greatest conservation in the predicted TM regions and more variation in the C-terminal than the N-terminal region. Searching of additional sequence databases identified candidate MCT homologues from the yeast Saccharomyces cerevisiae, the nematode worm Caenorhabditis elegans and the archaebacterium Sulfolobus solfataricus. Together these sequences constitute a new family of transporters with some strongly conserved sequence motifs, the possible functions of which are discussed.

Amino Acid Sequence

The profilin multigene family of maize: differential expression of three isoforms.

Profilin is a small (12-15 kDa) actin- and phospholipid-binding protein previously known only from studies on animals and lower eukaryotes but recently identified as a birch pollen allergen. Here we have identified and characterized three members of the profilin multigene family from the plant Zea mays. Two cDNAs isolated from a maize pollen library (ZmPRO 1 and ZmPRO 3) each have a single, large open reading frame encoding a putative polypeptide 131 amino acids long with a predicted molecular weight of approximately 14 kDa. A third maize pollen cDNA (ZmPRO 2) has two in-frame translation initiation codons. Use of the first ATG would result in a polypeptide 137 amino acids long with a molecular weight of 14.8 kDa. The three maize profilins are highly homologous to each other (> 90% nucleotide and amino acid sequence identity) as well as other plant profilins but show far less similarity (30-40% amino acid sequence identity) to animal and lower eukaryote profilins. Multiple sequence alignments indicate that only nine residues are shared by all eukaryotic profilins examined. However, limited comparisons reveal domains in the NH2 and COOH termini that have a high degree of similarity suggesting functional conservation. The maize gene family size is estimated to contain three to six members based on Southern blot experiments with gene-specific and coding region probes. Northern blot analysis demonstrates that the three maize profilin cDNAs characterized here are utilized in a tissue-specific manner and are anther or pollen specific.

Actins

Equus caballus gelsolin--cDNA sequence and protein structural implications.

We have generated and characterized the cDNA from equine smooth muscle that encodes gelsolin, an actin-modulating protein. Overlapping cDNA clones synthesized by the reverse transcriptase/polymerase chain reaction and clones isolated from a horse genomic library provided the complete primary structure for the intracellular isoform of gelsolin, while cDNA complemented with protein sequence data produced the full-length primary transcript of the gelsolin isoform found circulating in equine plasma. The deduced amino acid sequences of the intracellular and secreted versions of equine gelsolin infer polypeptides of 731 and 755 residues with apparent molecular masses of 80.7 kDa and 83.2 kDa, respectively. Multiple sequence alignment analysis of equine, human, porcine, and murine orthologs of gelsolin demonstrates prominent similarities among all of these proteins, with the horse and human molecules exhibiting the largest degree of likeness with respect to polypeptide length and overall sequence composition. Both horse and human plasma gelsolins are comprised of 755 amino acids with 94% of the residues identical, while the degree of sequence identity in the shorter (731 residues) cytoplasmic gelsolins is 95%. Analysis of the sequences and structures of the six related domains that comprise gelsolin emphasizes the strong correlation that exists between primary structural conservation among mammalian gelsolins and maintenance of the three-dimensional domain fold characteristic of members of this protein family.

Amino Acid Sequence

Subunit organization of the abalone Haliotis tuberculata hemocyanin type 2 (HtH2), and the cDNA sequence encoding its functional units d, e, f, g and h.

We have developed a HPLC procedure to isolate the two different hemocyanin types (HtH1 and HtH2) of the European abalone Haliotis tuberculata. On the basis of limited proteolytic cleavage, two-dimensional immunoelectrophoresis, PAGE, N-terminal protein sequencing and cDNA sequencing, we have identified eight different 40-60-kDa functional units (FUs) in HtH2, termed HtH2-a to HtH2-h, and determined their linear arrangement within the elongated 400-kDa subunit. From a Haliotis cDNA library, we have isolated and sequenced a cDNA clone which encodes the five C-terminal FUs d, e, f, g and h of HtH2. As shown by multiple sequence alignments, defg of HtH2 correspond structurally to defg from Octopus dofleini hemocyanin. HtH2-e is the first FU of a gastropod hemocyanin to be sequenced. The new Haliotis hemocyanin sequences are compared to their counterparts in Octopus, Helix pomatia and HtH1 (from the latter, the sequences of FU-f, FU-g and FU-h have recently been determined) and discussed in relation to the recent 2.3 A X-ray structure of FU-g from Octopus hemocyanin and the 15 A three-dimensional reconstruction of the Megathura crenulata hemocyanin didecamer from electron micrographs. This data allows, for the first time, an insight into the evolution of the two functionally different hemocyanin isoforms found in marine gastropods. It appears that they evolved several hundred million years ago within the Prosobranchia, after separation of the latter from the branch leading to the Pulmonata. Moreover, as a structural explanation for the inefficiency of the type 1 hemocyanin to form multidecamers in vivo, the additional N-glycosylation sites in HtH1 compared to HtH2 are discussed.

Amino Acid Sequence

Overview on the sub-grouping of the crustacean hyperglycemic hormone family.

The Crustacean hyperglycemic hormones (CHHs) are an ever extending family of crustacean hormones mainly involved in carbohydrate metabolism, molt and reproduction. In this paper, we drew together 32 available CHH sequences, and applied the techniques of multiple sequence alignment, motif searching and amino acid conservation analysis to the characterization of the molecules independently of their biological function. The analysis clearly showed that the proteins clustered into two groups (CHH and VIH). Amino acid conservation analysis also subdivided the VIH group into sequences involved in reproduction (RIH) or in molt (MIH). Motif searching identified five motifs in each group of mature hormones. Motifs A2 and A3 were conserved in all sequences while motifs A1 and A1' were specific of the CHH and VIH groups respectively. This approach demonstrated the S. gregaria ion transport peptides as true members of the CHH group. The two main groups, CHH and VIH, are also discussed in terms of functional homogeneity.

Animals