Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Analysis of the Structure of the PsbO Protein and its Implications.

The PsbO protein is a ubiquitous extrinsic subunit of Photosystem II (PS II), the water splitting enzyme of photosynthesis. A recently determined 3D X-ray structure of a cyanobacterial protein bound to PS II has given an opportunity to conduct complete analyses of its sequence and structural characteristics using bioinformatic methods. Multiple sequence alignments for the PsbO family are constructed and correlated with the cyanobacterial structure. We identify the most conserved regions of PsbO and the mapping of their positions within the structure indicates their functional roles especially in relation to interactions of this protein with the lumenal surface of PS II. Homologous models for eukaryotic PsbO were built in order to compare with the prokaryotic protein. We also explore structural homology between PsbO and other proteins for which 3D structures are known and determine its structural classification. These analyses contribute to the understanding of the function and evolutionary origin of the PS II manganese stabilising protein.

Journal Article↗

Arbitrarily primed PCR and sequencing of 16S rDNA for epidemiological typing and species identification of Burkholderia cepacia isolates from Swedish patients with cystic fibrosis reveal genetic heterogeneity.

To investigate whether arbitrarily primed (AP)-PCR and/or 16S rDNA sequencing could be used as rapid methods for epidemiological typing and species identification of clinical Burkholderia isolates from patients with cystic fibrosis (CF), a total of 39 clinical B. cepacia isolates, including 33 isolates from 14 CF patients, were fingerprinted. ERIC-2 primer was used for AP-PCR. The AP-PCR clustering analysis resulted in 14 different clusters at a 70% similarity level. The AP-PRC patterns were individual despite considerable similarities. To sequence rDNA, a broad-range PCR was applied. The PCR product included four variable loops (V8, V3, V4 and V9) of the 16S ribosomal small subunit RNA gene. The multiple sequence alignment produced 12 different patterns, 5 of them including more than one isolate. Heterogeneity of the bases in the V3 region, indicating the simultaneous presence of at least two different types of 16S rRNA genes in the same cell, was revealed in 10 isolates. Most of the CF patients were adults who had advanced disease at follow-up. Both the sequencing and the AP-PCR patterns revealed genetic heterogeneity of isolates between patients. According to the results obtained, AP-PCR could advantageously be used for epidemiological typing of Burkholderia, whereas partial species identification could effectively be obtained by sequencing of the V3 region of the 16S RNA gene.

Adolescent↗

Molecular determinants of complex formation between Clp/Hsp100 ATPases and the ClpP peptidase.

The Clp/Hsp100 ATPases are hexameric protein machines that catalyze the unfolding, disassembly and disaggregation of specific protein substrates in bacteria, plants and animals. Many family members also interact with peptidases to form ATP-dependent proteases. In Escherichia coli, for instance, the ClpXP protease is assembled from the ClpX ATPase and the ClpP peptidase. Here, we have used multiple sequence alignments to identify a tripeptide 'IGF' in E. coli ClpX that is essential for ClpP recognition. Mutations in this IGF sequence, which appears to be part of a surface loop, disrupt ClpXP complex formation and prevent protease function but have no effect on other ClpX activities. Homologous tripeptides are found only in a subset of Clp/Hsp100 ATPases and are a good predictor of family members that have a ClpP partner. Mapping of the IGF loop onto a homolog of known structure suggests a model for ClpX-ClpP docking.

ATP-Dependent Proteases↗

Natural-like function in artificial WW domains.

Protein sequences evolve through random mutagenesis with selection for optimal fitness. Cooperative folding into a stable tertiary structure is one aspect of fitness, but evolutionary selection ultimately operates on function, not on structure. In the accompanying paper, we proposed a model for the evolutionary constraint on a small protein interaction module (the WW domain) through application of the SCA, a statistical analysis of multiple sequence alignments. Construction of artificial protein sequences directed only by the SCA showed that the information extracted by this analysis is sufficient to engineer the WW fold at atomic resolution. Here, we demonstrate that these artificial WW sequences function like their natural counterparts, showing class-specific recognition of proline-containing target peptides. Consistent with SCA predictions, a distributed network of residues mediates functional specificity in WW domains. The ability to recapitulate natural-like function in designed sequences shows that a relatively small quantity of sequence information is sufficient to specify the global energetics of amino acid interactions.

Amino Acid Sequence↗

Evolutionary information for specifying a protein fold.

Classical studies show that for many proteins, the information required for specifying the tertiary structure is contained in the amino acid sequence. Here, we attempt to define the sequence rules for specifying a protein fold by computationally creating artificial protein sequences using only statistical information encoded in a multiple sequence alignment and no tertiary structure information. Experimental testing of libraries of artificial WW domain sequences shows that a simple statistical energy function capturing coevolution between amino acid residues is necessary and sufficient to specify sequences that fold into native structures. The artificial proteins show thermodynamic stabilities similar to natural WW domains, and structure determination of one artificial protein shows excellent agreement with the WW fold at atomic resolution. The relative simplicity of the information used for creating sequences suggests a marked reduction to the potential complexity of the protein-folding problem.

Algorithms↗

Snake venom disintegrins: novel dimeric disintegrins and structural diversification by disulphide bond engineering.

We report the isolation and amino acid sequences of six novel dimeric disintegrins from the venoms of Vipera lebetina obtusa (VLO), V. berus (VB), V. ammodytes (VA), Echis ocellatus (EO) and Echis multisquamatus (EMS). Disintegrins VLO4, VB7, VA6 and EO4 displayed the RGD motif and inhibited the adhesion of K562 cells, expressing the integrin alpha5beta1 to immobilized fibronectin. A second group of dimeric disintegrins (VLO5 and EO5) had MLD and VGD motifs in their subunits and blocked the adhesion of the alpha4beta1 integrin to vascular cell adhesion molecule 1 with high selectivity. On the other hand, disintegrin EMS11 inhibited both alpha5beta1 and alpha4beta1 integrins with almost the same degree of specificity. Comparison of the amino acid sequences of the dimeric disintegrins with those of other disintegrins by multiple-sequence alignment and phylogenetic analysis, in conjunction with current biochemical and genetic data, supports the view that the different disintegrin subfamilies evolved from a common ADAM (a disintegrin and metalloproteinase-like) scaffold and that structural diversification occurred through disulphide bond engineering.

Amino Acid Motifs↗

Role of conserved Asp293 of cytochrome P450 2C9 in substrate recognition and catalytic activity.

Human cytochrome P450 2C9 (CYP2C9) is important in the metabolism of non-steroidal anti-inflammatory compounds such as diclofenac, the antidiabetic agent tolbutamide and other clinically important drugs, many of which are weakly acidic. Multiple sequence alignment of CYPs identified CYP2C9 Asp(293) as corresponding to Asp(301) of CYP2D6, which has been suggested to play a role in the binding of basic substrates to the latter enzyme. Replacement of Asp(293) with Ala (D293A) decreased activity by more than 90%, and led to an approx. 3- to 10-fold increase in K (m) values for the three test substrates tolbutamide, dextromethorphan and diclofenac. Conservative replacement of the carboxyl side chain in a Glu (D293E) mutant produced no significant changes in K (m) values and slight increases in k (cat) values. Changes in regiospecificity were observed for both the Ala and Glu substitutions; low levels of both dextromethorphan O- and N-demethylation were observed in the D293A mutant, whereas increased preference for O-demethylation was observed for the D293E mutant. Expression of constructs coding for Asn (D293N) and Gln (D293Q) substitutions failed to form a P450 correctly. Our analysis suggests a structural role for the carboxyl side chain of Asp(293) in CYP2C9 substrate binding and catalysis. The conservation of an Asp residue in other CYP families in a position equivalent to Asp(293) indicates a common mechanism for maintaining the active-site architecture.

Amino Acid Sequence↗

Theoretical model of the three-dimensional structure of a sugar-binding protein from Pyrococcus horikoshii: structural analysis and sugar-binding simulations.

The three-dimensional structure of a sugar-binding protein from the thermophilic archaea Pyrococcus horikoshii has been predicted by a homology modelling procedure and investigated for its stability and its ability to bind different sugars. The model was created by using as templates the three-dimensional structures of a maltodextrin-binding protein from Pyrococcus furiosus, a trehalose-maltose-binding protein from Thermococcus litoralis and a maltodextrin-binding protein from Escherichia coli. According to the suggestions from the CASP (Critical Assessment of Structure Prediction) meetings, the homology modelling strategy was applied by assessing an accurate multiple sequence alignment, based on the high structural conservation in the family of ATP-binding cassette transporters to which all these proteins belong. The model has been deposited in the Protein Data Bank with the code 1R25. According to the origin of the protein, several characteristics in the organization of the secondary-structure elements and in the distribution of polar and non-polar amino acids are very similar to those of thermophilic proteins, compared with proteins from mesophilic organisms, and are analysed in detail. Finally, a simulation of the binding of several sugars in the binding site of this protein is presented, and interactions with amino acids are highlighted in detail.

Amino Acid Sequence↗

Crystal structure of levansucrase from the Gram-negative bacterium Gluconacetobacter diazotrophicus.

The endophytic Gram-negative bacterium Gluconacetobacter diazotrophicus SRT4 secretes a constitutively expressed levansucrase (LsdA, EC 2.4.1.10), which converts sucrose into fructooligosaccharides and levan. The enzyme is included in GH (glycoside hydrolase) family 68 of the sequence-based classification of glycosidases. The three-dimensional structure of LsdA has been determined by X-ray crystallography at a resolution of 2.5 A (1 A=0.1 nm). The structure was solved by molecular replacement using the homologous Bacillus subtilis (Bs) levansucrase (Protein Data Bank accession code 1OYG) as a search model. LsdA displays a five-bladed beta-propeller architecture, where the catalytic residues that are responsible for sucrose hydrolysis are perfectly superimposable with the equivalent residues of the Bs homologue. The comparison of both structures, the mutagenesis data and the analysis of GH68 family multiple sequences alignment show a strong conservation of the sucrose hydrolytic machinery among levansucrases and also a structural equivalence of the Bs levansucrase Ca2+-binding site to the LsdA Cys339-Cys395 disulphide bridge, suggesting similar fold-stabilizing roles. Despite the strong conservation of the sucrose-recognition site observed in LsdA, Bs levansucrase and GH32 family Thermotoga maritima invertase, structural differences appear around residues involved in the transfructosylation reaction.

Amino Acid Sequence↗

Probing the substrate binding site of Candida tenuis xylose reductase (AKR2B5) with site-directed mutagenesis.

Little is known about how substrates bind to CtXR (Candida tenuis xylose reductase; AKR2B5) and other members of the AKR (aldo-keto reductase) protein superfamily. Modelling of xylose into the active site of CtXR suggested that Trp23, Asp50 and Asn309 are the main components of pentose-specific substrate-binding recognition. Kinetic consequences of site-directed substitutions of these residues are reported. The mutants W23F and W23Y catalysed NADH-dependent reduction of xylose with only 4 and 1% of the wild-type efficiency (kcat/K(m)) respectively, but improved the wild-type selectivity for utilization of ketones, relative to xylose, by factors of 156 and 471 respectively. Comparison of multiple sequence alignment with reported specificities of AKR members emphasizes a conserved role of Trp23 in determining aldehyde-versus-ketone substrate selectivity. D50A showed 31 and 18% of the wild-type catalytic-centre activities for xylose reduction and xylitol oxidation respectively, consistent with a decrease in the rates of the chemical steps caused by the mutation, but no change in the apparent substrate binding constants and the pattern of substrate specificities. The 30-fold preference of the wild-type for D-galactose compared with 2-deoxy-D-galactose was lost completely in N309A and N309D mutants. Comparison of the 2.4 A (1 A=0.1 nm) X-ray crystal structure of mutant N309D bound to NAD+ with the previous structure of the wild-type holoenzyme reveals no major structural perturbations. The results suggest that replacement of Asn309 with alanine or aspartic acid disrupts the function of the original side chain in donating a hydrogen atom for bonding with the substrate C-2(R) hydroxy group, thus causing a loss of transition-state stabilization energy of 8-9 kJ/mol.

Aldehyde Reductase↗

An investigation of the role of Glu-842, Glu-844 and His-846 in the function of the cytoplasmic domain of the epidermal growth factor receptor.

Activation of several protein kinases is mediated, at least in part, by phosphorylation of conserved Thr or Tyr residues located in a variable loop region, near the active site. In certain kinases, this activation loop also controls access of peptide substrates to the active site. In the corresponding region of the epidermal growth factor (EGF) receptor, a potential phosphorylation site, Tyr-845, does not appear to have a major regulatory role. In order to find out whether this variable loop can modulate the peptide phosphorylation and self-phosphorylation activities of the EGF receptor kinase, we investigated the role of residues around Tyr-845, using site-directed mutagenesis. Multiple sequence alignment showed that residues Glu-842, Glu-844 and His-846 are conserved or nearly conserved in eight members of the EGF receptor family. Mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala were expressed in the baculovirus/insect cell system, purified to near-homogeneity and characterized with respect to their peptide phosphorylation and self-phosphorylation activities. All three mutants were active, and these changes did not affect ATP binding directly. However, all mutations increased the Km(app.) for peptide substrates and MnATP in peptide phosphorylation reactions. The Vmax. for the phosphorylation of peptide RREELQDDYEDD was unaltered, but the Vmax. for self-phosphorylation (with variable [MnATP]) decreased 4-, 2- and 7-fold for mutants Glu-842-->Ser, Glu-844-->Gln and His-846-->Ala respectively, compared with the wild-type. These results suggest that binding of this peptide restored an optimal conformation at the active site that might be impaired by the mutations. A study of the dependence of initial rates of self-phosphorylation on cytoplasmic domain concentration showed that the order of reaction increased with the progress of self-phosphorylation. Both pre-phosphorylation and high concentrations of ammonium sulphate restored maximal or near-maximal levels of self-phosphorylation in the mutants, possibly through compensating conformational changes. A plausible homology model, based on the cyclic AMP-dependent protein kinase catalytic subunit, accommodated the sequence Glu-841-Glu-Lys-Glu as an insertion in the peptide binding loop at the edge of the active site cleft. The model suggests that Glu-844 and His-846 may participate in H-bonding interactions, thus stabilizing the active site region, while Glu-842 does not appear to interact with regions of the catalytic core.

Amino Acid Sequence↗

Prediction from sequence comparisons of residues of factor H involved in the interaction with complement component C3b.

The amino acid sequence of the region of bovine factor H containing the C3b binding site has been derived from sequencing overlapping cDNA clones. A cDNA sequence encoding 669 amino acids was obtained. Like human and mouse factor H the sequence can be arranged into a number of internally homologous units (CPs), each of which is about 60 amino acids long and is based on a framework of four conserved cysteine residues. Bovine factor H is of the same molecular mass as human and mouse factor H, and is therefore likely to be composed of 20 contiguous CPs. Comparisons with human and mouse factor H indicate that the partial bovine sequence encodes CPs 2-12 inclusive of bovine factor H. Bovine factor H binds to human ammonia-treated C3 (causing thiolester cleavage) [C3(NH3)] and promotes the cleavage of human C3(NH3) in the presence of bovine factor I. Other studies indicate that CPs 2-5 of human factor H encompass the C3b binding and factor I cofactor activity site. Multiple sequence alignments of human factor H, mouse factor H (which also interacts with human C3b) and bovine factor H with CP modules whose structures have been determined experimentally, have been used to predict residues in the hypervariable loops of CPs 2-5 and to identify residues of potential importance in human C3 binding and factor I cofactor activity. Leu-17 and Gly-20 of CP 2, Ser-17, Ala-19, Glu-21, Asp-23 and Glu-25 of CP 3 and Lys-18 of CP 4 are all conserved between the three species. It may be that CPs 3 and 4 interact with C3(NH3) directly, whilst CPs 2 and 5 maintain the correct orientation for CPs 3 and 4 to interact.

Amino Acid Sequence↗

Identification of sequence similarity between 60 kDa and 70 kDa molecular chaperones: evidence for a common evolutionary background?

Recent findings support the premise that chaperonins (60 kDa stress-proteins) and alpha-subunits of F-type ATPases (alpha-ATPase) are evolutionary related protein families. Two-dimensional gel patterns of synthesized proteins in unstressed and heat-shocked embryonic Drosophila melanogaster SL2 cells revealed that antibodies raised against the alpha-subunit of the F1-ATPase complex from rat liver recognize an inducible p71 member of the 70 kDa stress-responsive protein family. Molecular recognition of this stress-responsive 70 kDa protein by antibodies raised against the F1-ATPase alpha-subunit suggests the possibility of partial sequence similarity within these ATP-binding protein families. A multiple sequence alignment between alpha-ATPases and 60 kDa and 70 kDa molecular chaperones is presented. Statistical evaluation of sequence similarity reveals a significant degree of sequence conservation within the three protein families. The finding suggests a common evolutionary origin for the ATPases and molecular chaperone protein families of 60 kDa and 70 kDa, despite the lack of obvious structural resemblance between them.

Amino Acid Sequence↗

Comparative anatomy of the aldo-keto reductase superfamily.

The aldo-keto reductases metabolize a wide range of substrates and are potential drug targets. This protein superfamily includes aldose reductases, aldehyde reductases, hydroxysteroid dehydrogenases and dihydrodiol dehydrogenases. By combining multiple sequence alignments with known three-dimensional structures and the results of site-directed mutagenesis studies, we have developed a structure/function analysis of this superfamily. Our studies suggest that the (alpha/beta)8-barrel fold provides a common scaffold for an NAD(P)(H)-dependent catalytic activity, with substrate specificity determined by variation of loops on the C-terminal side of the barrel. All the aldo-keto reductases are dependent on nicotinamide cofactors for catalysis and retain a similar cofactor binding site, even among proteins with less than 30% amino acid sequence identity. Likewise, the aldo-keto reductase active site is highly conserved. However, our alignments indicate that variation ofa single residue in the active site may alter the reaction mechanism from carbonyl oxidoreduction to carbon-carbon double-bond reduction, as in the 3-oxo-5beta-steroid 4-dehydrogenases (Delta4-3-ketosteroid 5beta-reductases) of the superfamily. Comparison of the proposed substrate binding pocket suggests residues 54 and 118, near the active site, as possible discriminators between sugar and steroid substrates. In addition, sequence alignment and subsequent homology modelling of mouse liver 17beta-hydroxysteroid dehydrogenase and rat ovary 20alpha-hydroxysteroid dehydrogenase indicate that three loops on the C-terminal side of the barrel play potential roles in determining the positional and stereo-specificity of the hydroxysteroid dehydrogenases. Finally, we propose that the aldo-keto reductase superfamily may represent an example of divergent evolution from an ancestral multifunctional oxidoreductase and an example of convergent evolution to the same active-site constellation as the short-chain dehydrogenase/reductase superfamily.

Alcohol Oxidoreductases↗

Cloning and sequencing of four new mammalian monocarboxylate transporter (MCT) homologues confirms the existence of a transporter family with an ancient past.

Measurement of monocarboxylate transport kinetics in a range of cell types has provided strong circumstantial evidence for a family of monocarboxylate transporters (MCTs). Two mammalian MCT isoforms (MCT1 and MCT2) and a chicken isoform (REMP or MCT3) have already been cloned, sequenced and expressed, and another MCT-like sequence (XPCT) has been identified. Here we report the identification of new human MCT homologues in the database of expression sequence tags and the cloning and sequencing of four new full-length MCT-like sequences from human cDNA libraries, which we have denoted MCT3, MCT4, MCT5 and MCT6. Northern blotting revealed a unique tissue distribution for the expression of mRNA for each of the seven putative MCT isoforms (MCT1-MCT6 and XPCT). All sequences were predicted to have 12 transmembrane (TM) helical domains with a large intracellular loop between TM6 and TM7. Multiple sequence alignments showed identities ranging from 20% to 55%, with the greatest conservation in the predicted TM regions and more variation in the C-terminal than the N-terminal region. Searching of additional sequence databases identified candidate MCT homologues from the yeast Saccharomyces cerevisiae, the nematode worm Caenorhabditis elegans and the archaebacterium Sulfolobus solfataricus. Together these sequences constitute a new family of transporters with some strongly conserved sequence motifs, the possible functions of which are discussed.

Amino Acid Sequence↗

The profilin multigene family of maize: differential expression of three isoforms.

Profilin is a small (12-15 kDa) actin- and phospholipid-binding protein previously known only from studies on animals and lower eukaryotes but recently identified as a birch pollen allergen. Here we have identified and characterized three members of the profilin multigene family from the plant Zea mays. Two cDNAs isolated from a maize pollen library (ZmPRO 1 and ZmPRO 3) each have a single, large open reading frame encoding a putative polypeptide 131 amino acids long with a predicted molecular weight of approximately 14 kDa. A third maize pollen cDNA (ZmPRO 2) has two in-frame translation initiation codons. Use of the first ATG would result in a polypeptide 137 amino acids long with a molecular weight of 14.8 kDa. The three maize profilins are highly homologous to each other (> 90% nucleotide and amino acid sequence identity) as well as other plant profilins but show far less similarity (30-40% amino acid sequence identity) to animal and lower eukaryote profilins. Multiple sequence alignments indicate that only nine residues are shared by all eukaryotic profilins examined. However, limited comparisons reveal domains in the NH2 and COOH termini that have a high degree of similarity suggesting functional conservation. The maize gene family size is estimated to contain three to six members based on Southern blot experiments with gene-specific and coding region probes. Northern blot analysis demonstrates that the three maize profilin cDNAs characterized here are utilized in a tissue-specific manner and are anther or pollen specific.

Actins↗

Gain-of-function and loss-of-function phenotypes of the protein phosphatase 2C HAB1 reveal its role as a negative regulator of abscisic acid signalling.

HAB1 was originally cloned on the basis of sequence homology to ABI1 and ABI2, and indeed, a multiple sequence alignment of 32 Arabidopsis protein phosphatases type-2C (PP2Cs) reveals a cluster composed by the four closely related proteins, ABI1, ABI2, HAB1 and At1g17550 (here named HAB2). Characterisation of transgenic plants harbouring a transcriptional fusion ProHAB1: green fluorescent protein (GFP) indicates that HAB1 is broadly expressed within the plant, including key target sites of abscisic acid (ABA) action as guard cells or seeds. The expression of the HAB1 mRNA in vegetative tissues is strongly upregulated in response to exogenous ABA. In this work, we show that constitutive expression of HAB1 in Arabidopsis under a cauliflower mosaic virus (CaMV) 35S promoter led to reduced ABA sensitivity both in seeds and vegetative tissues, compared to wild-type plants. Thus, in the field of ABA signalling, this work represents an example of a stable phenotype in planta after sustained overexpression of a PP2C genes. Additionally, a recessive T-DNA insertion mutant of HAB1 was analysed in this work, whereas previous studies of recessive alleles of PP2C genes were carried out with intragenic revertants of the abi1-1 and abi2-1 mutants that carry missense mutations in conserved regions of the PP2C domain. In the presence of exogenous ABA, hab1-1 mutant shows ABA-hypersensitive inhibition of seed germination; however, its transpiration rate was similar to that of wild-type plants. The ABA-hypersensitive phenotype of hab1-1 seeds together with the reduced ABA sensitivity of 35S:HAB1 plants are consistent with a role of HAB1 as a negative regulator of ABA signalling. Finally, these results provide new genetic evidence on the function of a PP2C in ABA signalling.

Abscisic Acid↗

cDNA sequences of the authentic keratins 8 and 18 in zebrafish.

From the zebrafish Danio rerio, we have cDNA cloned and sequenced a novel type II and a novel type I keratin, termed DreK8 and DreK18, respectively. We identified DreK8/18 as the true orthologs of the human keratin pair K8/18 as follows: (i) MALDI-MS assignment to the biochemically identified K8 and K18 candidates that are co-expressed in simple epithelia and absent in epidermal keratinocytes; (ii) multiple sequence alignments and phylogenetic tree analysis, showing that DreK8, within the phylogenetic tree of type II keratins, forms a highly bootstrap-supported branch together with K8 from goldfish and rainbow trout, whereas DreK18, within the phylogenetic tree of type I keratins, groups with the K18 sequences from all other vertebrates studied; (iii) presence of a conserved motif in the tail domain of DreK8 (VxKxxETxDGxxVSESSxV) that is typical for all hitherto sequenced K8 orthologs. Moreover, several zebrafish type II keratin sequences published by other authors have now been assigned to epidermal keratins, previously identified biochemically.

Amino Acid Sequence↗