Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Molecular evolution of SRP cycle components: functional implications.

Signal recognition particle (SRP) is a cytoplasmic ribonucleoprotein that targets a subset of nascent presecretory proteins to the endoplasmic reticulum membrane. We have considered the SRP cycle from the perspective of molecular evolution, using recently determined sequences of genes or cDNAs encoding homologs of SRP (7SL) RNA, the Srp54 protein (Srp54p), and the alpha subunit of the SRP receptor (SR alpha) from a broad spectrum of organisms, together with the remaining five polypeptides of mammalian SRP. Our analysis provides insight into the significance of structural variation in SRP RNA and identifies novel conserved motifs in protein components of this pathway. The lack of congruence between an established phylogenetic tree and size variation in 7SL homologs implies the occurrence of several independent events that eliminated more than half the sequence content of this RNA during bacterial evolution. The apparently non-essential structures are domain I, a tRNA-like element that is constant in archaea, varies in size among eucaryotes, and is generally missing in bacteria, and domain III, a tightly base-paired hairpin that is present in all eucaryotic and archeal SRP RNAs but is invariably absent in bacteria. Based on both structural and functional considerations, we propose that the conserved core of SRP consists minimally of the 54 kDa signal sequence-binding protein complexed with the loosely base-paired domain IV helix of SRP RNA, and is also likely to contain a homolog of the Srp68 protein. Comparative sequence analysis of the methionine-rich M domains from a diverse array of Srp54p homologs reveals an extended region of amino acid identity that resembles a recently identified RNA recognition motif. Multiple sequence alignment of the G domains of Srp54p and SR alpha homologs indicates that these two polypeptides exhibit significant similarity even outside the four GTPase consensus motifs, including a block of nine contiguous amino acids in a location analogous to the binding site of the guanine nucleotide dissociation stimulator (GDS) for E. coli EF-Tu. The conservation of this sequence, in combination with the results of earlier genetic and biochemical studies of the SRP cycle, leads us to hypothesize that a component of the Srp68/72p heterodimer serves as the GDS for both Srp54p and SR alpha. Using an iterative alignment procedure, we demonstrate similarity between Srp68p and sequence motifs conserved among GDS proteins for small Ras-related GTPases. The conservation of SRP cycle components in organisms from all three major branches of the phylogenetic tree suggests that this pathway for protein export is of ancient evolutionary origin.

Amino Acid Sequence

Homology modelling and protein engineering strategy of subtilases, the family of subtilisin-like serine proteinases.

Subtilases are members of the family of subtilisin-like serine proteases. Presently, greater than 50 subtilases are known, greater than 40 of which with their complete amino acid sequences. We have compared these sequences and the available three-dimensional structures (subtilisin BPN', subtilisin Carlsberg, thermitase and proteinase K). The mature enzymes contain up to 1775 residues, with N-terminal catalytic domains ranging from 268 to 511 residues, and signal and/or activation-peptides ranging from 27 to 280 residues. Several members contain C-terminal extensions, relative to the subtilisins, which display additional properties such as sequence repeats, processing sites and membrane anchor segments. Multiple sequence alignment of the N-terminal catalytic domains allows the definition of two main classes of subtilases. A structurally conserved framework of 191 core residues has been defined from a comparison of the four known three-dimensional structures. Eighteen of these core residues are highly conserved, nine of which are glycines. While the alpha-helix and beta-sheet secondary structure elements show considerable sequence homology, this is less so for peptide loops that connect the core secondary structure elements. These loops can vary in length by greater than 150 residues. While the core three-dimensional structure is conserved, insertions and deletions are preferentially confined to surface loops. From the known three-dimensional structures various predictions are made for the other subtilases concerning essential conserved residues, allowable amino acid substitutions, disulphide bonds, Ca(2+)-binding sites, substrate-binding site residues, ionic and aromatic interactions, proteolytically susceptible surface loops, etc. These predictions form a basis for protein engineering of members of the subtilase family, for which no three-dimensional structure is known.

Amino Acid Sequence

Secondary structure prediction for modelling by homology.

An improved method of secondary structure prediction has been developed to aid the modelling of proteins by homology. Selected data from four published algorithms are scaled and combined as a weighted mean to produce consensus algorithms. Each consensus algorithm is used to predict the secondary structure of a protein homologous to the target protein and of known structure. By comparison of the predictions to the known structure, accuracy values are calculated and a consensus algorithm chosen as the optimum combination of the composite data for prediction of the homologous protein. This customized algorithm is then used to predict the secondary structure of the unknown protein. In this manner the secondary structure prediction is initially tuned to the required protein family before prediction of the target protein. The method improves statistical secondary structure prediction and can be incorporated into more comprehensive systems such as those involving consensus prediction from multiple sequence alignments. Thirty one proteins from five families were used to compare the new method to that of Garnier, Osguthorpe and Robson (GOR) and sequence alignment. The improvement over GOR is naturally dependent on the similarity of the homologous protein, varying from a mean of 3% to 7% with increasing alignment significance score.

Algorithms

Protein fold refinement: building models from idealized folds using motif constraints and multiple sequence data.

A general solution to the problem of directly incorporating data from multiple sequence alignments into the construction of molecular models was approached through the calculation of an estimated pairwise distance based on conserved hydrophobicity. A scaling method was developed that allowed the required bulk geometric properties of the estimated pair-wise distances (mean and mean squared) to mimic those expected in a globular protein. These properties were maintained independently of the composition, length, number or degree of conservation of the original sequences. Despite being a poor estimate for individual distances were found to be compatible with the native structure and could be weighted highly. While the estimated distances provided a general drive towards hydrophobic packing, more specific structure (including secondary structures and motifs) were induced by regularization towards an ideal form. These constraints were used to refine an outline starting structure (derived only from secondary structure axes) towards a compact form that was sufficiently protein-like for side chains to be added with almost no further adjustment of the alpha-carbon positions. This process allows rough folds based on abstract representations of protein architecture to be rapidly converted to a form where they can be analysed by the growing number of methods designed to assess molecular models.

Chemical Phenomena

Sequence divergence analysis for the prediction of seven-helix membrane protein structures: II. A 3-D model of human rhodopsin.

A three-dimensional (3-D) model of the transmembrane domain of human rhodopsin was predicted from the sequence divergence analysis of 42 sequences of rhodopsins and visual pigments without a template. The prediction steps include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, their secondary structure and packing shape in a helix bundle, prediction of side-chain conformations and structure refinement. The identification of the retinal binding site was assisted by its known covalent linkage with K296. The structural features of the predicted 3-D model are in good agreement with a low resolution electron density map of bovine rhodopsin and with residues in contact with retinal as determined experimentally.

Amino Acid Sequence

A simple procedure for assigning a sequence motif with an obscure pattern: application to the basic/helix-loop-helix motif.

We have developed a simple method to assign a sequence motif with an obscure pattern. Given a multiple sequence alignment for a region of protein that is known or strongly believed to have the same secondary and tertiary structures, the quantification method by principal component analysis is designed to find the regions most likely to have the same structure in a protein outside of the original set. The potential of this newly developed method was evaluated with reference to the known basic/helix-loop-helix (bHLH) motifs, and its characteristics were discussed with four obscure but well-defined motifs and compared with the other methods for searching sequence motifs. The method was also applied to assign the bHLH motif in Epstein-Barr virus nuclear antigen 1 (EBNA-1). This application revealed one candidate for the basic/helix 1 region and two candidates for the helix 2 region in the bHLH motif, within the region from amino acid residues 460 to 600, which is in good agreement with our previous experimental studies on the DNA binding region of EBNA-1. The basic/helix-loop-helix-loop-helix structure thus assigned suggests a function of EBNA-1 which is associated with both replication and transcription.

Amino Acid Sequence

Comparison of conservation within and between the Ser/Thr and Tyr protein kinase family: proposed model for the catalytic domain of the epidermal growth factor receptor.

The protein kinase family can be subdivided into two main groups based on their ability to phosphorylate Ser/Thr or Tyr substrates. In order to understand the basis of this functional difference, we have carried out a comparative analysis of sequence conservation within and between the Ser/Thr and Tyr protein kinases. A multiple sequence alignment of 86 protein kinase sequences was generated. For each position in the alignment we have computed the conservation of residue type in the Ser/Thr, in the Tyr and in both of the kinase subfamilies. To understand the structural and/or functional basis for the conservation, we have mapped these conservation properties onto the backbone of the recently determined structure of the cAMP-dependent Ser/Thr kinase. The results show that the kinase structure can be roughly segregated, based upon conservation, into three zones. The inner zone contains residues highly conserved in all the kinase family and describes the hydrophobic core of the enzyme together with residues essential for substrate and ATP binding and catalysis. The outer zone contains residues highly variable in all kinases and represents the solvent-exposed surface of the protein. The third zone is comprised of residues conserved in either the Ser/Thr or Tyr kinases or in both, but which are not conserved between them. These are sandwiched between the hydrophobic core and the solvent-exposed surface. In addition to analyzing overall conservation in the kinase family, we have also looked at conservation of its substrate and ATP binding sites. The ATP site is highly conserved throughout the kinases, whereas the substrate binding site is more variable. The active site contains several positions which differ between the Ser/Thr and Tyr kinases and may be responsible for discriminating between hydroxyl bearing side chains. Using this information we propose a model for Tyr substrate binding to the catalytic domain of the epidermal growth factor receptor (EGFR).

Amino Acid Sequence

A 3D model of the delta opioid receptor and ligand-receptor complexes.

A model for the 3D structure of the transmembrane domain of the delta opioid receptor was predicted from the sequence divergence analysis of 42 sequences of G-protein coupled peptide hormone receptors belonging to the opioid, somatostatin and angiotensin receptor families. No template was used in the prediction steps, which include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, prediction of their secondary structure, optimization of the packing shape in a helix bundle, prediction of side chain conformations and structural refinement. The general shape of the model is similar to that of the low resolution rhodopsin structure in that the TM3 and TM7 helices are most buried in the bundle and the TM1 and TM4 helices are most exposed to the lipid phase. An initial assessment of this model was made by determining to what extent a binding site identified using four structurally disparate high affinity delta opioid ligands was consistent with known mutational studies. With the assumption that the protonated amine nitrogen, a feature common to all delta opioid ligands, interacts with the highly conserved Asp127 in TM3, a pocket was found that satisfied the criteria of complementarity to the requirements for receptor recognition for these four diverse ligands, two delta selective antagonists (the fused ring naltrindole and the peptide Tyr-Tic-Phe-Phe-NH2) and the two agonists lofentanil and BW373U86 deduced from previous studies of the ligands alone. These ligands could be accommodated in a similar region of the receptor. The receptor binding site identified in the optimized complexes contained many residues in positions known to affect ligand binding in G-protein coupled receptors. These results also allowed identification of key residues as candidates for point mutations for further assessment and refinement of this model as well as preliminary indications of the requirements for recognition of this receptor.

Amino Acid Sequence

Structural and evolutionary relationships among the immunophilins: two ubiquitous families of peptidyl-prolyl cis-trans isomerases.

The immunophilins, protein receptors for the immunosuppressing drugs cyclosporin A and FK506 and related proteins from plants, fungi, and bacteria, have been analyzed structurally and evolutionarily. The cyclosporin A binding proteins (cyclophilins) represent one ubiquitous family of homologous proteins, and the FK506- and rapamycin-binding proteins (FKBPs) constitute a second, unrelated family. Multiple sequence alignments of members of each of these two protein families define the highly conserved residues that are likely to play important structural and functional roles, and mutations in representative members of these two families that abolish or alter function have been evaluated. FKBPs have undergone greater evolutionary divergence than the cyclophilins. Evolutionary trees were constructed using two distinct programs, and these trees establish the structural relationships that allow division of each of these families into subgroups. The results lead to the suggestion that several genes encoding isozymic forms of the FKBPs and possibly also of the cyclophilins existed in prokaryotes before the emergence of eukaryotes on earth and that representatives of these genes were transmitted to both kingdoms to give rise to current subfamilies of these proteins. By contrast, compartmentalization of both classes of immunophilins appears to have arisen independently in prokaryotes and eukaryotes, late in evolutionary history.

Amino Acid Isomerases

Genomic divergence of an HIV-2 from a German AIDS patient probably infected in Mali.

The complete nucleotide sequence of an HIV-2 isolate derived from a German AIDS patient with predominantly neurological symptoms is reported. The HIV-2BEN sequence is highly divergent from those of previously described HIV-2 and SIV strains. Evolutionary tree analysis of eight HIV-2 sequences reveals the existence of three HIV-2 groups. HIV-2BEN belongs to a group with two isolates from Ghana and The Gambia. Based on a comparison of HIV-2BEN with six HIV-2 isolates, SIVsmm and SIVmac, the variability of the structural env and gag proteins is similar within the HIV-2/SIVsmm/mac and HIV-1 groups. In contrast, the regulatory HIV-1 proteins are more highly conserved than those from HIV-2 strains. Multiple sequence alignments reveal that some domains of the envelope and regulatory proteins are well conserved among HIV-1, HIV-2/SIVsmm/mac, SIVagm and SIVmnd. The identification of conserved domains within the external glycoprotein could help to develop broadly active vaccines.

Acquired Immunodeficiency Syndrome

Phylogenetic analysis of gag genes from 70 international HIV-1 isolates provides evidence for multiple genotypes.

OBJECTIVE: To determine the extent of genetic variation among internationally collected HIV-1 isolates, to analyse phylogenetic relationships and the geographic distribution of different variants. DESIGN: Phylogenetic comparison of 70 HIV-1 isolates collected in 15 countries on four continents. METHODS: To sequence the complete gag genome of HIV-1 isolates, build multiple sequence alignments and construct phylogenetic trees using distance matrix methods and maximum parsimony algorithms. RESULTS: Phylogenetic tree analysis identified seven distinct genotypes. The seven genotypes were evident by both distance matrix methods and maximum parsimony analysis, and were strongly supported by bootstrap resampling of the data. The intra-genotypic gag distances averaged 7%, whereas the inter-genotypic distances averaged 14%. The geographic distribution of variants was complex. Some genotypes have apparently migrated to several continents and many areas harbor a mixture of genotypes. Related variants may cluster in certain areas, particularly isolates from a single city collected over a short time. CONCLUSIONS: The genetic variation among HIV-1 isolates is more extensive than previously appreciated. At least seven distinct HIV-1 genotypes can be identified. Diversification, migration and establishment of local, temporal 'blooms' of particular variants may all occur concomitantly.

Africa

Site-directed mutagenesis of the lipoate acetyltransferase of Escherichia coli.

Remote but significant similarities between the primary and predicted secondary structures of the chloramphenicol acetyltransferases (CAT) and lipoate acyltransferase subunits (LAT, E2) of the 2-oxo acid dehydrogenase complexes, have suggested that both types of enzyme may use similar catalytic mechanisms. Multiple sequence alignments for CAT and LAT have highlighted two conserved motifs that contain the active-site histidine and serine residues of CAT. Site-directed replacement of Ser550 in the E2p subunit (LAT) of the pyruvate dehydrogenase complex of Escherichia coli, deemed to be equivalent to the active-site Ser148 of CAT, supported the CAT-based model of LAT catalysis. The effects of other substitutions were also consistent with the predicted similarity in catalytic mechanism although specific details of active-site geometry may not be conserved.

Acetyltransferases

Identification of a gag protein epitope conserved among all four groups of primate immunodeficiency viruses by using monoclonal antibodies.

Five monoclonal antibodies (MAbs) were raised against the gag proteins of simian immunodeficiency virus (SIV) from African green monkey (SIVagmTYO-7). Two MAbs reacted with the matrix protein p17 and the other three with the core protein p24. Studies on the cross-reactivity of the MAbs revealed that the anti-p24 MAbs detected an epitope shared by the viruses belonging to the human immunodeficiency virus type 2 (HIV-2)/SIVmac group and SIVagmTYO-7 and SIVagmTYO-5. The anti-p17 MAbs recognized an epitope present on all these viruses and on SIVagmTYO-1, HIV-1 and SIVmnd. This finding demonstrates for the first time that the matrix protein, p17 or p18, respectively, of all nine HIV and SIV isolates tested in this study expresses at least one conserved immunogenic epitope recognized serologically. By using synthetic peptides, this epitope was identified at the N terminus of p17. Furthermore, this epitope was analysed by multiple sequence alignments of the peptide with homologous sequences of HIV and SIV p17.

Amino Acid Sequence

Sequence and structural analysis of murine adenovirus type 1 hexon.

The genomic region encoding the major capsid protein (hexon) of murine adenovirus type 1 (MAV-1) has been isolated and sequenced. The sequence predicts a 908 residue MAV-1 hexon protein and is flanked by a portion of the upstream pVI gene and the downstream endoproteinase gene. The order of these genes and their location in the middle of the genome are the same as those found in other adenoviruses sequenced to date. Multiple sequence alignment with the other five known hexon protein sequences reveals an overall residue identity of 51% and residue conservation of 66%. In comparison with human adenovirus type 2 (Ad2), MAV-1 hexon has major deletions between residues 141 to 170, 270 to 284 and 446 to 455. Since these regions in the Ad2 hexon are partially exposed on the outer surface of the virion, they may represent type-specific antigenic determinants. The MAV-1 hexon sequence has been modelled using the known three-dimensional structure of the Ad2 hexon. The variable regions in which the mutations, deletions and insertions occur are located in the l1 and l2 loops of the molecule that form the protruding hexon towers on the external surface of the virion.

Amino Acid Sequence

Genome characterization and taxonomy of Plantago asiatica mosaic potexvirus.

The complete nucleotide sequence of Plantago asiatica mosaic virus (P1AMV) genomic RNA has been determined. The 6128 nucleotide sequence contains five open reading frames (ORFs) coding for proteins of M(r) 156K (ORF1), 25K (ORF2), 12K (ORF3), 13K (ORF4) and 22K (ORF5). The sequences of these P1AMV proteins exhibit strong homology to the proteins of the other potexviruses. Phylogenetic trees based on the multiple sequence alignments of three conserved domains in ORF1 product and capsid protein reveal a close relationship of P1AMV to papaya mosaic virus and clover yellow mosaic virus. The P1AMV genomic RNA and a major subgenomic RNA (sgRNA) of 0.9 kb have been detected in infected leaves by Northern blot hybridization. The latter sgRNA is the messenger for virus capsid protein and its 5' terminus has been located 23 nucleotides upstream of the initiator codon of the coat protein gene. The P1AMV virion RNA and RNA transcript resembling the 0.9 kb sgRNA have been translated in vitro giving rise to a single major 170K product and a major 22K product, respectively.

Amino Acid Sequence

Citrate synthase from the thermophilic archaebacterium Thermoplasma acidophilium. Cloning and sequencing of the gene.

The gene encoding the citric acid cycle enzyme, citrate synthase, has been cloned from the thermoacidophilic archaebacterium, Thermoplasma acidophilum. We report the sequencing of this gene and its flanking regions, and the derived amino acid sequence of the enzyme is compared by multiple-sequence alignment analysis with those of citrate synthases from eubacterial and eukaryotic organisms. The similarity is less than 30% between the archaebacterial and non-archaebacterial sequences, although the majority of residues implicated in the catalytic action of the enzyme have been conserved across all three kingdoms. The cloned archaebacterial gene has been expressed in Escherichia coli to produce catalytically active citrate synthase. This is the first reported sequence of citrate synthase from the archaebacteria.

Amino Acid Sequence

Characteristics of an exochitinase from Streptomyces olivaceoviridis, its corresponding gene, putative protein domains and relationship to other chitinases.

Streptomyces olivaceoviridis efficiently degrades chitin. Shotgun cloning of partially Sau3A-cleaved DNA using the multicopy vector pIJ702 and Streptomyces lividans 66 as host resulted in the identification of the plasmid pCHI O1 which harbours an insert of 4.6 kb. In the presence of chitin as sole carbon source, transformants of S. lividans 66 carrying pCHI O1 or its derivatives with smaller inserts overproduced an exochitinase which was purified to homogeneity. The chitin-inducible enzyme with an isoelectric point of 4.0 shows optimal activity at pH 7.3 and 55 degrees C, has an apparent molecular mass of 47 kDa and is competitively inhibited by the pseudosugar allosamidin. The enzyme was identified as an exochitinase since it generates exclusively chitobiose from chitotetraose, chitohexaose, and colloidal high-molecular mass chitin. Sequence analysis of a reading frame of 1794 base pairs and comparison of the deduced amino-acid sequence allowed the identification of the putative catalytic domain, one region with significant similarity to the type-III module of fibronectin and one domain of unknown function. Multiple sequence alignment and hydrophobic-cluster analysis of 25 chitinolytic enzymes from bacteria, fungi and plants allowed the identification of their characteristic domains. The exochitinase from S. olivaceoviridis shares highest similarity with the chitinase D from Bacillus circulans.

Acetylglucosamine

The purification, characterization and analysis of primary and secondary-structure of prolyl oligopeptidase from human lymphocytes. Evidence that the enzyme belongs to the alpha/beta hydrolase fold family.

Prolyl oligopeptidase was isolated and purified to homogeneity from human lymphocytes, yielding a specific activity of 7780 mU/mg. The molecular mass using size-exclusion chromatography matches the 76 kDa obtained by SDS/PAGE. This provides evidence that prolyl oligopeptidase is a monomer. The isoelectric point is 4.8 as judged by isoelectric focusing in free solution. Di-isopropyl fluorophosphate and phenylmethylsulphonyl fluoride completely abolish the activity, classifying the enzyme as a serine proteinase. The inhibition by p-chloromercuribenzoic acid indicates the importance of a free sulfhydryl group near the active-site. alpha 1-Casein and ornithine decarboxylase, two proteins containing a PEST sequence, inhibit prolyl oligopeptidase, but were not hydrolyzed. This demonstrates that prolyl oligopeptidase is not participating in the metabolism of proteins according to a PEST-dependent pathway. alpha 1-Antitrypsin partially inhibits the enzyme but in contrast, aprotinin does not. Its inability to cleave corticotropin-releasing factor, ubiquitin, albumin and aprotinin, together with the hydrolysis of bradykinin between Pro7-Arg8 confirms the affinity of prolyl oligopeptidase for small peptides. Multiple sequence alignment does not reveal any similarity with proteases of known tertiary structure. Secondary-structure prediction displays striking similarity with dipeptidyl peptidase IV and acylaminoacyl peptidase. Two characteristic features of the members of the prolyl oligopeptidase family of serine proteases are high-lighted: the linear arrangement of the catalytic triad is nucleophile-acid-base and the proteolytic cleavage releasing the catalytically active C-terminal region of around 500 amino acids from the N-terminal sequence. Secondary structure prediction and comparison of the active-site of serine proteinases with known three-dimensional coordinates prove that Asp641 is the third member of the catalytic triad. The secondary structural organization of the protease domain of prolyl oligopeptidase is in accordance with the alpha/beta hydrolase fold.

Amino Acid Sequence