Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Angiotensin I-converting enzyme transition state stabilization by HIS1089: evidence for a catalytic mechanism distinct from other gluzincin metalloproteinases.

Angiotensin (Ang) I-converting enzyme (ACE) is a member of the gluzincin family of zinc metalloproteinases that contains two homologous catalytic domains. Both the N- and C-terminal domains are peptidyl-dipeptidases that catalyze Ang II formation and bradykinin degradation. Multiple sequence alignment was used to predict His(1089) as the catalytic residue in human ACE C-domain that, by analogy with the prototypical gluzincin, thermolysin, stabilizes the scissile carbonyl bond through a hydrogen bond during transition state binding. Site-directed mutagenesis was used to change His(1089) to Ala or Leu. At pH 7.5, with Ang I as substrate, k(cat)/K(m) values for these Ala and Leu mutants were 430 and 4,000-fold lower, respectively, compared with wild-type enzyme and were mainly due to a decrease in catalytic rate (k(cat)) with minor effects on ground state substrate binding (K(m)). A 120,000-fold decrease in the binding of lisinopril, a proposed transition state mimic, was also observed with the His(1089) --> Ala mutation. ACE C-domain-dependent cleavage of AcAFAA showed a pH optimum of 8.2. H1089A has a pH optimum of 5.5 with no pH dependence of its catalytic activity in the range 6.5-10.5, indicating that the His(1089) side chain allows ACE to function as an alkaline peptidyl-dipeptidase. Since transition state mutants of other gluzincins show pH optima shifts toward the alkaline, this effect of His(1089) on the ACE pH optimum and its ability to influence transition state binding of the sulfhydryl inhibitor captopril indicate that the catalytic mechanism of ACE is distinct from that of other gluzincins.

Amino Acid Sequence↗

Mapping the hyaluronan-binding site on the link module from human tumor necrosis factor-stimulated gene-6 by site-directed mutagenesis.

Link modules are hyaluronan-binding domains found in extracellular proteins involved in matrix assembly, development, and immune cell migration. Previously we have expressed the Link module from the inflammation-associated protein tumor necrosis factor-stimulated gene-6 (TSG-6) and determined its tertiary structure in solution. Here we generated 21 Link module mutants, and these were analyzed by nuclear magnetic resonance spectroscopy and a hyaluronan-binding assay. The individual mutation of five amino acids, which form a cluster on one face of the Link module, caused large reductions in functional activity but did not affect the Link module fold. This ligand-binding site in TSG-6 is similar to that determined previously for the hyaluronan receptor, CD44, suggesting that the location of the interaction surfaces may also be conserved in other Link module-containing proteins. Analysis of the sequences of TSG-6 and CD44 indicates that the molecular details of their association with hyaluronan are likely to be significantly different. This comparison identifies key sequence positions that may be important in mediating hyaluronan binding, across the Link module superfamily. The use of multiple sequence alignment and molecular modeling allowed the prediction of functional residues in link protein, and this approach can be extended to all members of the superfamily.

Amino Acid Sequence↗

Identification of the catalytic residues of bifunctional glycogen debranching enzyme.

Eukaryotic glycogen debranching enzyme (GDE) possesses two different catalytic activities (oligo-1,4-->1,4-glucantransferase/amylo-1,6-glucosidase) on a single polypeptide chain. To elucidate the structure-function relationship of GDE, the catalytic residues of yeast GDE were determined by site-directed mutagenesis. Asp-535, Glu-564, and Asp-670 on the N-terminal half and Asp-1086 and Asp-1147 on the C-terminal half were chosen by the multiple sequence alignment or the comparison of hydrophobic cluster architectures among related enzymes. The five mutant enzymes, D535N, E564Q, D670N, D1086N, and D1147N were constructed. The mutant enzymes showed the same purification profiles as that of wild-type enzyme on beta-CD-Sepharose-6B affinity chromatography. All the mutant enzymes possessed either transferase activity or glucosidase activity. Three mutants, D535N, E564Q, and D670N, lost transferase activity but retained glucosidase activity. In contrast, D1086N and D1147N lost glucosidase activity but retained transferase activity. Furthermore, the kinetic parameters of each mutant enzyme exhibiting either the glucosidase activity or transferase activity did not vary markedly from the activities exhibited by the wild-type enzyme. These results strongly indicate that the two activities of GDE, transferase and glucosidase, are independent and located at different sites on the polypeptide chain.

Amino Acid Sequence↗

Molecular modeling of the extracellular domain of the RET receptor tyrosine kinase reveals multiple cadherin-like domains and a calcium-binding site.

Using bioinformatic tools, mutagenesis, and binding studies, we have investigated the structural organization of the extracellular region of the RET receptor tyrosine kinase, a functional receptor for glial cell line-derived neurotrophic factor (GDNF). Multiple sequence alignments of seven vertebrate sequences and one invertebrate RET sequence delineated four distinct N-terminal domains, each of about 110 residues, containing many of the consensus motifs of the cadherin fold. Based on these alignments and the crystal structures of epithelial and neural cadherins, we have generated molecular models of each of the four cadherin-like domains in the extracellular region of human RET. The modeled structures represent realistic models from both energetic and geometrical points of view and are consistent with previous observations gathered from biochemical analyses of the effects of Hirschsprung's disease mutations affecting the folding and stability of the RET molecule, as well as our own site-directed mutagenesis studies of RET cadherin-like domain 1. We have also investigated the role of Ca(2+) in ligand binding by RET and found that Ca(2+) ions are required for RET binding to GDNF but not for GDNF binding to the GFRalpha1 co-receptor. In agreement with these results, RET, but not GFRalpha1, was found to bind Ca(2+) directly. Our results indicate that the overall architecture of the extracellular region of RET is more closely related to cadherins than previously thought. The models of the cadherin-like domains of human RET represent valuable tools with which to guide future site-directed mutagenesis studies aimed at identifying residues involved in ligand binding and receptor activation.

Amino Acid Sequence↗

Residues in internal repeats of the rice cation/H+ exchanger are involved in the transport and selection of cations.

In plants, the cation/H+ exchanger (CAX) translocates Ca2+ and other metal ions into vacuoles using the H+ gradient formed by H+-ATPase and H+-pyrophosphatase. Such exchangers carrying 11 transmembrane domains (TMs) have been isolated from plants, yeast, and bacteria. In this study, multiple sequence alignment of several CAXs revealed the presence of highly conserved 36-residue regions between TM3 and TM4 and between TM8 and TM9. These two repetitive motifs are designated repeats c-1 and c-2. Using site-directed mutagenesis, we generated 31 mutations in the repeats of the Oryza sativa CAX, which translocates Ca2+ and Mn2+. Mutant exchangers were expressed in a Saccharomyces cerevisiae strain that is sensitive to Ca2+ and Mn2+ because of the absence of vacuolar Ca2+-ATPase and the Ca2+/H+ exchanger. Mutant exchangers were classified into six classes according to their tolerance for Ca2+ and Mn2+. For example, the class III mutants had no tolerance for either ion, and the class IV mutants had tolerance only for Ca2+. The biochemical function of each residue was estimated. We investigated the membrane topology of the repeats using a method combining cysteine mutagenesis and sulfhydryl reagents. Our results suggest that repeat c-1 re-enters the membrane from the vacuolar luminal side and forms a solution-accessible region. Furthermore, several residues in repeats c-1 and c-2 were found to be conserved in animal Na+/Ca2+ exchangers. Finally, we suggest that these re-entrant repeats may form a vestibule or filter for cation selection.

Amino Acid Sequence↗

Distant structural homology leads to the functional characterization of an archaeal PIN domain as an exonuclease.

Genome sequencing projects have focused attention on the problem of discovering the functions of protein domains that are widely distributed throughout living species but which are, as yet, largely uncharacterized. One such example is the PIN domain, found in eukaryotes, bacteria, and Archaea, and with suggested roles in signaling, RNase editing, and/or nucleotide binding. The first reported crystal structure of a PIN domain (open reading frame PAE2754, derived from the crenarchaeon, Pyrobaculum aerophilum) has been determined to 2.5 A resolution and is presented here. Mapping conserved residues from a multiple sequence alignment onto the structure identifies a putative active site. The discovery of distant structural homology with several exonucleases, including T4 phage RNase H and flap endonuclease (FEN1), further suggests a likely function for PIN domains as Mg2+-dependent exonucleases, a hypothesis that we have confirmed in vitro. The tetrameric structure of PAE2754, with the active sites inside a tunnel, suggests a mechanism for selective cleavage of single-stranded overhangs or flap structures. These results indicate likely DNA or RNA editing roles for prokaryotic PIN domains, which are strikingly numerous in thermophiles, and in organisms such as Mycobacterium tuberculosis. They also support previous hypotheses that eukaryotic PIN domains participate in RNAi and nonsense-mediated RNA degradation.

Amino Acid Sequence↗

Molecular mechanism of dimerization of Bowman-Birk inhibitors. Pivotal role of ASP76 in the dimerzation.

Horsegram (Dolichos biflorus), a protein-rich leguminous pulse, is a crop native to Southeast Asia and tropical Africa. The seeds contain multiple forms of Bowman-Birk type inhibitors. The major inhibitor HGI-III, from the native seed with 76 amino acid residues exists as a dimer. The amino acid sequence of three isoforms of Bowman-Birk inhibitor from germinated horsegram, designated as HGGI-I, HGGI-II, and HGGI-III, have been obtained by sequential Edman analyses of the pyridylethylated inhibitors and peptides derived therefrom by enzymatic and chemical cleavage. The HGGIs are monomers, comprising of 66, 65, and 60 amino acid residues, respectively. HGGI-III from the germinated seed differs from the native seed inhibitor in the physiological deletion of a dodecapeptide at the amino terminus and a tetrapeptide, -SHDD, at the carboxyl terminus. The study of the state of association of HGI-III, by size-exclusion chromatography and SDS-PAGE in the presence of 1 mM ZnCl2, has revealed the role of charged interactions in the monomer <--> dimer equilibria. Chemical modification studies of Lys and Arg have confirmed the role of charge interactions in the above equilibria. These results support the premise that a unique interaction, which stabilizes the dimer, is the cause of self-association in the inhibitors. This interaction in HGI-III involves the epsilon-amino group of the Lys24 (P1 residue) at the first reactive site of one monomer and the carboxyl of an Asp86 at the carboxyl terminus of the second monomer. Identification of the role of these individual amino acids in the structure and stability of the dimer was accomplished by chemical modifications, multiple sequence alignment of legume Bowman-Birk inhibitors, and homology modeling. The state of association may also influence the physiological and functional role of these inhibitors.

Amino Acid Sequence↗

Hydralysins, a new category of beta-pore-forming toxins in cnidaria.

Cnidaria are venomous animals that produce diverse protein and polypeptide toxins, stored and delivered into the prey through the stinging cells, the nematocytes. These include pore-forming cytolytic toxins such as well studied actinoporins. In this work, we have shown that the non-nematocystic paralytic toxins, hydralysins, from the green hydra Chlorohydra viridissima comprise a highly diverse group of beta-pore-forming proteins, distinct from other cnidarian toxins but similar in activity and structure to bacterial and fungal toxins. Functional characterization of hydralysins reveals that as soluble monomers they are rich in beta-structure, as revealed by far UV circular dichroism and computational analysis. Hydralysins bind erythrocyte membranes and form discrete pores with an internal diameter of approximately 1.2 nm. The cytolytic effect of hydralysin is cell type-selective, suggesting a specific receptor that is not a phospholipid or carbohydrate. Multiple sequence alignment reveals that hydralysins share a set of conserved sequence motifs with known pore-forming toxins such as aerolysin, epsilon-toxin, alpha-toxin, and LSL and that these sequence motifs are found in and around the poreforming domains of the toxins. The importance of these sequence motifs is revealed by the cloning, expression, and mutagenesis of three hydralysin isoforms that strongly differ in their hemolytic and paralytic activities. The correlation between the paralytic and cytolytic activities of hydralysin suggests that both are a consequence of receptor-mediated pore formation. Hydralysins and their homologues exemplify the wide distribution of beta-pore formers in biology and provide a useful model for the study of their molecular mode of action.

Amino Acid Motifs↗

Evolutionarily conserved allosteric network in the Cys loop family of ligand-gated ion channels revealed by statistical covariance analyses.

The Cys loop family of ligand-gated ion channels mediate fast synaptic transmission for communication between neurons. They are allosteric proteins, in which binding of a neurotransmitter to its binding site in the extracellular amino-terminal domain triggers structural changes in distant transmembrane domains to open a channel for ion flow. Although the locations of binding site and channel gating machinery are well defined, the structural basis of the activation pathway coupling binding and channel opening remains to be determined. In this paper, by analyzing amino acid covariance in a multiple sequence alignment, we have identified an energetically interconnected network in the Cys loop family of ligand-gated ion channels. Statistical coupling and correlated mutational analyses along with clustering revealed a highly coupled cluster. Mapping the positions in the cluster onto a three-dimensional structural model demonstrated that these highly coupled positions form an interconnected network linking experimentally identified binding domains through the coupling region to the gating machinery. In addition, these highly coupled positions are also condensed in the transmembrane domains, which are a recent focus for the sites of action of many allosteric modulators. Thus, our results revealed a genetically interconnected network that potentially plays an important role in the allosteric activation and modulation of the Cys loop family of ligand-gated ion channels.

Allosteric Regulation↗

The solution structure of Escherichia coli Wzb reveals a novel substrate recognition mechanism of prokaryotic low molecular weight protein-tyrosine phosphatases.

Low molecular weight protein-tyrosine phosphatases (LMW-PTPs) are small enzymes that ubiquitously exist in various organisms and play important roles in many biological processes. In Escherichia coli, the LMW-PTP Wzb dephosphorylates the autokinase Wzc, and the Wzc/Wzb pair regulates colanic acid production. However, the substrate recognition mechanism of Wzb is still poorly understood thus far. To elucidate the molecular basis of the catalytic mechanism, we have determined the solution structure of Wzb at high resolution by NMR spectroscopy. The Wzb structure highly resembles that of the typical LMW-PTP fold, suggesting that Wzb may adopt a similar catalytic mechanism with other LMW-PTPs. Nevertheless, in comparison with eukaryotic LMW-PTPs, the absence of an aromatic amino acid at the bottom of the active site significantly alters the molecular surface and implicates Wzb may adopt a novel substrate recognition mechanism. Furthermore, a structure-based multiple sequence alignment suggests that a class of the prokaryotic LMW-PTPs may share a similar substrate recognition mechanism with Wzb. The current studies provide the structural basis for rational drug design against the pathogenic bacteria.

Amino Acid Sequence↗

Structural role for Tyr-104 in Escherichia coli isopentenyl-diphosphate isomerase: site-directed mutagenesis, enzymology, and protein crystallography.

Isopentenyl-diphosphate (IPP):dimethylallyl diphosphate isomerase is a key enzyme in the biosynthesis of isoprenoids. The mechanism of the isomerization reaction involves protonation of the unactivated carbon-carbon double bond in the substrate, but identity of the acidic moiety providing the proton is still not clear. Multiple sequence alignments and geometrical features observed in crystal structures of complexes with IPP isomerase suggest that Tyr-104 could play an important role during catalysis. A series of mutants was constructed by directed mutagenesis and characterized by enzymology. Crystallographic and thermal denaturation data for Y104A and Y104F mutants were obtained. Those data demonstrate the importance of residue Tyr-104 for proper folding of Escherichia coli type I IPP isomerase.

Carbon-Carbon Double Bond Isomerases↗

Triad3A regulates ubiquitination and proteasomal degradation of RIP1 following disruption of Hsp90 binding.

Toll-like receptors (TLRs) play a crucial role in innate immunity by recognizing microbial pathogens. Triad3A is an E3 ubiquitin-protein ligase that interacts with the Toll/interleukin-1 receptor domain of TLRs and promotes their proteolytic degradation. In the present study, we further investigated its activity on signaling molecules downstream of TLRs and tumor necrosis factor (TNF) receptor 1. Triad3A promoted down-regulation of two TIR domain-containing adapter proteins, TIRAP and TRIF, as well as a RIP1 but had no effect on other adapter molecules in either the TLRs or TNF-alpha signaling pathways. Multiple sequence alignment analysis suggested that RIP1 contains a TIR homologous domain, and mutation of amino acid residues in this domain identified three residues critical for its interaction with Triad3A. Moreover, Triad3A acted as a negative regulator in TNF-alpha signaling. Reduction of Triad3A expression by small interference RNAs rendered cells hyperresponsive to TNF-alpha stimulation. Conversely, overexpression of Triad3A in cells blocked TNF-alpha-induced cell activation. This negative regulation was effected independently of changes in the cellular protein level of RIP1. Further studies indicated that RIP1 formed a complex with Triad3A and heat shock protein 90 (Hsp90), which is a chaperone protein capable of maintaining the stability of its client proteins. Treatment of cells with geldanamycin to disrupt the Hsp90 complex led to proteasomal degradation of RIP1. Depletion of Triad3A by small interference RNA treatment inhibited geldanamycin-activated ubiquitination and proteolytic degradation of RIP1. These results suggest that Triad3A is an E3 ubiquitin-protein ligase to RIP1 and that Hsp90 and Triad3A cooperatively maintain the homeostasis of RIP1.

Adaptor Proteins, Vesicular Transport↗

A multidomain TIGR/olfactomedin protein family with conserved structural similarity in the N-terminal region and conserved motifs in the C-terminal region.

Based on the similarity between the TIGR (trabecular-meshwork inducible glucocorticoid response) (also known as myocilin) and olfactomedin protein families identified throughout the length of the TIGR protein, we have identified more distantly related proteins to determine the elements essential to the function/structure of the TIGR and olfactomedin proteins. Using a sequence walk method and the Shotgun program, we have identified a family including 31 olfactomedin domain-containing sequences. Multiple sequence alignments and secondary structure analyses were used to identify conserved sequence elements. Pairwise identity in the olfactomedin domain ranges from 8 to 64%, with an average pairwise identity of 24%. The N-terminal regions of the proteins fall into two subgroups, one including the TIGR and olfactomedin families and another group of apparently unrelated domains. The TIGR and olfactomedin sequences display conserved motifs including a residual leucine zipper region and maintain a similar secondary structure throughout the N-terminal region. The correlation between conserved elements and disease-associated mutations and apparent polymorphisms in human TIGR was also examined to evaluate the apparent importance of conserved residues to the function/structure of TIGR. Several residues have been identified as essential to the function and/or structure of the human TIGR protein based on their degree of conservation across the family and their implication in the pathogenesis of primary open-angle glaucoma. Additionally, we have identified a group of chitinase sequences containing several of the highly conserved motifs present in the C-terminal region of the olfactomedin domain-containing sequences.

Algorithms↗

Quantitative proteomics profiling of sarcomere associated proteins in limb and extraocular muscle allotypes.

The sarcomere is the major structural and functional unit of striated muscle. Approximately 65 different proteins have been associated with the sarcomere, and their exact composition defines the speed, endurance, and biology of each individual muscle. Past analyses relied heavily on electrophoretic and immunohistochemical techniques, which only allow the analysis of a small fraction of proteins at a time. Here we introduce a quantitative label-free, shotgun proteomics approach to differentially quantitate sarcomeric proteins from microgram quantities of muscle tissue in a fast and reliable manner by liquid chromatography and mass spectrometry. The high sequence similarity of some sarcomeric proteins poses a problem for shotgun proteomics because of limitations in subsequent database search algorithms in the exclusive assignment of peptides to specific isoforms. Therefore multiple sequence alignments were generated to improve the identification of isoform specific peptides. This methodology was used to compare the sarcomeric proteome of the extraocular muscle allotype to limb muscle. Extraocular muscles are a unique group of highly specialized muscles with distinct biochemical, physiological, and pathological properties. We were able to quantitate 40 sarcomeric proteins; although the basic sarcomeric proteins in extraocular muscle are similar to those in limb muscle, key proteins stabilizing the connection of the Z-bands to thin filaments and the costamere are augmented in extraocular muscle and may represent an adaptation to the eccentric contractions known to normally occur during eye movements. Furthermore, a number of changes are seen that closely relate to the unique nature of extraocular muscle.

Animals↗

Molecular modelling of mammalian CYP2B isoforms and their interaction with substrates, inhibitors and redox partners.

1. The construction of three-dimensional models of CYP2B isozymes from rat (CYP2B1), rabbit (CYP2B4) and man (CYP2B6), based on a multiple sequence alignment with CYP102, a unique eukaryotic-like bacterial P450 (in terms of possessing an NADPH-dependent FAD- and FMN-containing oxidoreductase redox partner) of known crystal structure, is reported. 2. The enzyme models described are shown to be consistent with experimental evidence from site-directed mutagenesis studies, antibody recognition sites and amino acid residues identified as being associated with redox partner interactions, together with the location of a key serine residue (Ser-128) likely to be involved in protein kinaseA-mediated phosphorylation. 3. A substantial number of known substrates and inhibitors of CYP2B isozymes are shown to fit the putative active sites of the enzyme models in agreement with their reported position of metabolism or mode of inhibition respectively. In particular, there is complementarity between the characteristic non-planar geometries of CYP2B substrates and key groups in the enzymes' active sites. 4. Molecular modelling of CYP2B isozymes appears to rationalize a number of the reported findings from quantitative structure-activity relationship investigations on series of CYP2B substrates and inhibitors.

Amino Acid Sequence↗

Calcium binding of transglutaminases: a 43Ca NMR study combined with surface polarity analysis.

Transglutaminases (TGases) form cross-links between glutamine and lysine side-chains of polypeptides in a Ca2+-dependent reaction. The structural basis of the Ca2+-effect is poorly defined. 43Ca NMR, surface polarity analysis combined with multiple sequence alignment and the construction of a new homology model of human tissue transglutaminase (tTGase) were used to obtain structural information about Ca2+ binding properties of factor XIII-A2, tTGase and TGase 3 (each of human origin). 43Ca NMR provided higher average dissociation constants titrating on a wide Ca2+-concentration scale than previous studies with equilibrium dialysis performed in shorter ranges. These results suggest the existence of low affinity Ca2+ binding sites on both FXIII-A and tTGase in addition to high affinity ones in accordance with our surface polarity analysis identifying high numbers of negatively charged clusters. Upon increasing the salt concentration or activating with thrombin, FXIII-A2 partially lost its original Ca2+ affinity; the NMR data suggested different mechanisms for the two activation processes. The NMR provided structural evidence of GTP-induced conformational changes on the tTGase molecule diminishing all of its Ca2+ binding sites. NMR data on the Ca2+ binding properties of the TGase 3 are presented here; it binds Ca2+ the most tightly, which is weakened after its proteolytic activation. The investigated TGases seem to have very symmetric Ca2+ binding sites and no EF-hand motifs.

Amino Acid Sequence↗

Study and prediction of secondary structure for membrane proteins.

In this paper we present a novel approach to membrane protein secondary structure prediction based on the statistical stepwise discriminant analysis method. A new aspect of our approach is the possibility to derive physical-chemical properties that may affect the formation of membrane protein secondary structure. The certain physical-chemical properties of protein chains can be used to clarify the formation of the secondary structure types under consideration. Another aspect of our approach is that the results of multiple sequence alignment, or the other kinds of sequence alignment, are not used in the frame of the method. Using our approach, we predicted the formation of three main secondary structure types (alpha-helix, beta-structure and coil) with high accuracy, that is Q(3) = 76%. Predicting the formation of alpha-helix and non-alpha-helix states we reached the accuracy which was measured as Q(2) = 86%. Also we have identified certain protein chain properties that affect the formation of membrane protein secondary structure. These protein properties include hydrophobic properties of amino acid residues, presence of Gly, Ala and Val amino acids, and the location of protein chain end.

Amino Acid Sequence↗

Sequence and hydropathy profile analysis of two classes of secondary transporters.

A structural class in the MemGen classification of membrane proteins is a set of evolutionary related proteins sharing a similar global fold. A structural class contains both closely related pairs of proteins for which homology is clear from sequence comparison and very distantly related pairs, for which it is not possible to establish homology based on sequence similarity alone. In the latter case the evolutionary link is based on hydropathy profile analysis. Here, we use these evolutionary related sets of proteins to analyze the relationship between E-values in BLAST searches, sequence similarities in multiple sequence alignments and structural similarities in hydropathy profile analyses. Two structural classes of secondary transporters termed ST[3], which includes the Ion Transporter (IT) superfamily and ST[4], which includes the DAACS family (TC# 2.A.23) were extracted from the NCBI protein database. ST[3] contains 2051 unique sequences distributed over 32 families and 59 subfamilies. ST[4] is a smaller class containing 399 unique sequences distributed over 2 families and 7 subfamilies. One subfamily in ST[4] contains a new class of binding protein dependent secondary transporters. Comparison of the averaged hydropathy profiles of the subfamilies in ST[3] and ST[4] revealed that the two classes represent different folds. Divergence of the sequences in ST[4] is much smaller than observed in ST[3], suggesting different constraints on the proteins during evolution. Analysis of the correlation between the evolutionary relationship of pairs of proteins in a class and the BLAST E-value revealed that: (i) the BLAST algorithm is unable to pick up the majority of the links between proteins in structural class ST[3], (ii) "low complexity filtering" and "composition based statistics" improve the specificity, but strongly reduce the sensitivity of BLAST searches for distantly related proteins, indicating that these filters are too stringent for the proteins analyzed, and (iii) the E-value cut-off, which may be used to evaluate evolutionary significance of a hit in a BLAST search is very different for the two structural classes of membrane proteins.

Algorithms↗