Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Properties and partial protein sequence of plant annexins.

We have examined the characteristics of Ca(2+)-dependent phospholipid-binding proteins (annexins) in maize (Zea mays L.) coleoptiles and tip-growing pollen tubes of Lilium longiflorum. In maize, there are three such proteins, p35, p33, and p23. Partial sequence analysis reveals that peptides from p35 and p33 have identity to members of the annexin family of animal proteins and to annexins from tomato. Interestingly, multiple sequence alignments reveal that the domain responsible for Ca(2+) binding in animal annexins is not conserved in these plant peptide sequences. Although p33 and p35 share the annexin characteristic of binding to membrane lipid, unlike annexins II and VI they do not associate with detergent-insoluble cytoskeletal proteins or with F-actin from either plants or animals. Immunoblotting with antiserum raised to p33/p35 from maize reveals that cross-reactive polypeptides of 33 to 35 kilodaltons are also present in protein extracts from pollen tubes of L. longiflorum. Immunolocalization at the light microscope level suggests that these proteins are predominantly confined to the nongranular zone at the tube tip, a region rich in secretory vesicles. Our hypothesis that plant annexins mediate exocytotic events is supported by the finding that p23, p33, and p35 bind to these secretory vesicles in a Ca(2+)-dependent manner.

Journal Article↗

The SUPERFAMILY database in structural genomics.

The SUPERFAMILY hidden Markov model library representing all proteins of known structure predicts the domain architecture of protein sequences and classifies them at the SCOP superfamily level. This analysis has been carried out on all completely sequenced genomes. The ways in which the database can be useful to crystallographers is discussed, in particular with a view to high-throughput structure determination. The application of the SUPERFAMILY database to different target-selection strategies is suggested: novel folds, novel domain combinations and targeted attacks on genomes. Use of the database for more general inquiry in the context of structural studies is also explained. The database provides evolutionary relationships between target proteins and other proteins of known structure through the SCOP database, genome assignments and multiple sequence alignments.

Amino Acid Sequence↗

Structure of 2C-methyl-D-erythritol-2,4-cyclodiphosphate synthase from Shewanella oneidensis at 1.6 A: identification of farnesyl pyrophosphate trapped in a hydrophobic cavity.

Isopentenyl pyrophosphate (IPP) is a universal building block for the ubiquitous isoprenoids that are essential to all organisms. The enzymes of the non-mevalonate pathway for IPP synthesis, which is unique to many pathogenic bacteria, have recently been explored as targets for antibiotic development. Several crystal structures of 2C-methyl-D-erythritol-2,4-cyclophosphate (MECDP) synthase, the fifth of seven enzymes involved in the non-mevalonate pathway for synthesis of IPP, have been reported; however, the composition of metal ions in the active site and the presence of a hydrophobic cavity along the non-crystallographic threefold symmetry axis has varied between the reported structures. Here, the structure of MEDCP from Shewanella oneidensis MR1 (SO3437) was determined to 1.6 A resolution in the absence of substrate. The presence of a zinc ion in the active-site cleft, tetrahedrally coordinated by two histidine side chains, an aspartic acid side chain and an ambiguous fourth ligand, was confirmed by zinc anomalous diffraction. Based on analysis of anomalous diffraction data and typical metal-to-ligand bond lengths, it was concluded that an octahedral sodium ion was 3.94 A from the zinc ion. A hydrophobic cavity was observed along the threefold non-crystallographic symmetry axis, filled by a well defined non-protein electron density that could be modeled as farnesyl pyrophosphate (FPP), a downstream product of IPP, suggesting a possible feedback mechanism for enzyme regulation. The high-resolution data clarified the FPP-binding mode compared with previously reported structures. Multiple sequence alignment indicated that the residues critical to the formation of the hydrophobic cavity and for coordinating the pyrophosphate group of FPP are present in the majority of MEDCP synthase enzymes, supporting the idea of a specialized biological function related to FPP binding in a subfamily of MEDCP synthase homologs.

Aldose-Ketose Isomerases↗

The structure at 1.7 A resolution of the protein product of the At2g17340 gene from Arabidopsis thaliana.

The crystal structure of the At2g17340 protein from A. thaliana was determined by the multiple-wavelength anomalous diffraction method and was refined to an R factor of 16.9% (Rfree = 22.1%) at 1.7 A resolution. At2g17340 is a member of the Pfam01937.11 protein family and its structure provides the first insight into the structural organization of this family. A number of fully and highly conserved residues defined by multiple sequence alignment of members of the Pfam01937.11 family were mapped onto the structure of At2g17340. The fully conserved residues are involved in the coordination of a metal ion and in the stabilization of loops surrounding the metal site. Several additional highly conserved residues also map into the vicinity of the metal-binding site, while others are clearly involved in stabilizing the hydrophobic core of the protein. The structure of At2g17340 represents a new fold in protein conformational space.

Amino Acid Sequence↗

Automated selection of positions determining functional specificity of proteins by comparative analysis of orthologous groups in protein families.

The increasing volume of genomic data opens new possibilities for analysis of protein function. We introduce a method for automated selection of residues that determine the functional specificity of proteins with a common general function (the specificity-determining positions [SDP] prediction method). Such residues are assumed to be conserved within groups of orthologs (that may be assumed to have the same specificity) and to vary between paralogs. Thus, considering a multiple sequence alignment of a protein family divided into orthologous groups, one can select positions where the distribution of amino acids correlates with this division. Unlike previously published techniques, the introduced method directly takes into account nonuniformity of amino acid substitution frequencies. In addition, it does not require setting arbitrary thresholds. Instead, a formal procedure for threshold selection using the Bernoulli estimator is implemented. We tested the SDP prediction method on the LacI family of bacterial transcription factors and a sample of bacterial water and glycerol transporters belonging to the major intrinsic protein (MIP) family. In both cases, the comparison with available experimental and structural data strongly supported our predictions.

Automation↗

Rapid evolution in conformational space: a study of loop regions in a ubiquitous GTP binding domain.

The rapidly evolving subsets of a protein are often evident in multiple sequence alignments as poorly defined, gap-containing regions. We investigated the 3D context of these regions observed in 28 protein structures containing a GTP-binding domain assumed to be homologous to the transforming factor p21-RAS. The phylogenetic depth of this data set is such that it is possible to observe lineages sharing a common protein core that diverged early in the eukaryotic cell history. The sequence variability among these homolog proteins is directly linked to the structural variability of surface loops. We demonstrate that these regions are self-contained and thus mostly free of the evolutionary constraints imposed by the conserved core of the domain. These intraloop interactions have the property to create stem-like structures. Interestingly, these stem-like structures can be observed in loops of varying size, up to the size of small protein domains. We propose a model under which the diversity of protein topologies observed in these loops can be the product of a stochastic sampling of sequence and conformational space in a near-neutral fashion, while the proximity of the functional features of the domain core allows novel beneficial traits to be fixed. Our comparative observations, limited here to the proteins containing the RAS-like GTP-binding domain, suggest that a stochastic process of insertion/deletion analogous to "budding" of loops is a likely mechanism of structural innovation. Such a framework could be experimentally exploited to investigate the folding of increasingly complex model inserts.

Amino Acid Sequence↗

Are protein-protein interfaces more conserved in sequence than the rest of the protein surface?

Protein interfaces are thought to be distinguishable from the rest of the protein surface by their greater degree of residue conservation. We test the validity of this approach on an expanded set of 64 protein-protein interfaces using conservation scores derived from two multiple sequence alignment types, one of close homologs/orthologs and one of diverse homologs/paralogs. Overall, we find that the interface is slightly more conserved than the rest of the protein surface when using either alignment type, with alignments of diverse homologs showing marginally better discrimination. However, using a novel surface-patch definition, we find that the interface is rarely significantly more conserved than other surface patches when using either alignment type. When an interface is among the most conserved surface patches, it tends to be part of an enzyme active site. The most conserved surface patch overlaps with 39% (+/- 28%) and 36% (+/- 28%) of the actual interface for diverse and close homologs, respectively. Contrary to results obtained from smaller data sets, this work indicates that residue conservation is rarely sufficient for complete and accurate prediction of protein interfaces. Finally, we find that obligate interfaces differ from transient interfaces in that the former have significantly fewer alignment gaps at the interface than the rest of the protein surface, as well as having buried interface residues that are more conserved than partially buried interface residues.

Amino Acid Sequence↗

Soluble domains of telomerase reverse transcriptase identified by high-throughput screening.

Telomerase is a ribonucleoprotein complex responsible for extending the ends of eukaryotic chromosomes. Structural and biophysical studies of this enzyme have been limited by the inability to produce large amounts of recombinant protein. Here we perform a high-throughput screen to map regions of the Tetrahymena thermophila TERT (Telomerase Reverse Transcriptase) protein that are overexpressed in a soluble form in Escherichia coli using a GFP-fusion system. Many of the soluble protein domains identified do not coincide with domains inferred from multiple sequence alignment, so screening for fluorescent colonies provided information not otherwise readily obtained. The method revealed an essential, independently folded N-terminal domain that was expressed and purified with high yield and found to be suitable for structural analysis. These results provide a tool for future structural and biophysical studies of TERT.

Amino Acid Sequence↗

A Consensus Data Mining secondary structure prediction by combining GOR V and Fragment Database Mining.

The major aim of tertiary structure prediction is to obtain protein models with the highest possible accuracy. Fold recognition, homology modeling, and de novo prediction methods typically use predicted secondary structures as input, and all of these methods may significantly benefit from more accurate secondary structure predictions. Although there are many different secondary structure prediction methods available in the literature, their cross-validated prediction accuracy is generally <80%. In order to increase the prediction accuracy, we developed a novel hybrid algorithm called Consensus Data Mining (CDM) that combines our two previous successful methods: (1) Fragment Database Mining (FDM), which exploits the Protein Data Bank structures, and (2) GOR V, which is based on information theory, Bayesian statistics, and multiple sequence alignments (MSA). In CDM, the target sequence is dissected into smaller fragments that are compared with fragments obtained from related sequences in the PDB. For fragments with a sequence identity above a certain sequence identity threshold, the FDM method is applied for the prediction. The remainder of the fragments are predicted by GOR V. The results of the CDM are provided as a function of the upper sequence identities of aligned fragments and the sequence identity threshold. We observe that the value 50% is the optimum sequence identity threshold, and that the accuracy of the CDM method measured by Q(3) ranges from 67.5% to 93.2%, depending on the availability of known structural fragments with sufficiently high sequence identity. As the Protein Data Bank grows, it is anticipated that this consensus method will improve because it will rely more upon the structural fragments.

Algorithms↗

A novel clan of zinc metallopeptidases with possible intramembrane cleavage properties.

Computer-based database searching and protein multiple sequence alignment has identified a novel clan of zinc metallopeptidases, which, by phylogenetic analysis, has been shown to contain six subfamilies. The family is characterized by four common transmembrane segments and three conserved sequence motifs. The combination of topology analysis and motif identification has detected three potential Zn2+ coordinating residues. Only two of the sequences of this novel zinc metallopeptidase clan possess any functional annotation, one of which is able to cleave its substrate within a cytosol/transmembrane segment junction. A number of observations suggest that the remaining members of this novel clan may also cleave their substrates within transmembrane segments.

Animals↗

A secondary structural model of the 28S rRNA expansion segments D2 and D3 from rootworms and related leaf beetles (Coleoptera: Chrysomelidae; Galerucinae).

We analysed the secondary structure of two expansion segments (D2, D3) of the 28S rRNA gene from 229 leaf beetles (Coleoptera: Chrysomelidae), the majority of which are in the subfamily Galerucinae. The sequences were compared in a multiple sequence alignment, with secondary structure inferred primarily from the compensatory base changes in the conserved helices of the rRNA molecules. This comparative approach yielded thirty helices comprised of base pairs with positional covariation. Based on these leaf beetle sequences, we report an annotated secondary structural model for the D2 and D3 expansion segments that will prove useful in assigning positional nucleotide homology for phylogeny reconstruction in these and closely related beetle taxa. This predicted structure, consisting of seven major compound helices, is mostly consistent with previously proposed models for the D2 and D3 expansion segments in insects. Despite a lack of conservation in the primary structure of these regions of insect 28S rRNA, the evolution of the secondary structure of these seven major motifs may be informative above the nucleotide level for higher-order phylogeny reconstruction of major insect lineages.

Animals↗

VH gene usage in immunoglobulin E responses of seasonal rhinitis patients allergic to grass pollen is oligoclonal and antigen driven.

BACKGROUND: IgE is the pivotal-specific effector molecule of allergic reactions yet it remains unclear whether the elevated production of IgE in atopic individuals is due to superantigen activation of B cell populations, increased antibody class switching to IgE or oligoclonal allergen-driven IgE responses. OBJECTIVES: To increase our understanding of the mechanisms driving IgE responses in allergic disease we examined immunoglobulin variable regions of IgE heavy chain transcripts from three patients with seasonal rhinitis due to grass pollen allergy. METHODS: Variable domain of heavy chain-epsilon constant domain 1 cDNAs were amplified from peripheral blood using a two-step semi-nested PCR, cloned and sequenced. RESULTS: The VH gene family usage in subject A was broadly based, but there were two clusters of sequences using genes VH 3-9 and 3-11 with unusually low levels of somatic mutations, 0-3%. Subject B repeatedly used VH 1-69 and subject C repeatedly used VH 1-02, 1-46 and 5a genes. Most clones were highly mutated being only 86-95% homologous to their germline VH gene counterparts and somatic mutations were more abundant at the complementarity determining rather than framework regions. Multiple sequence alignment revealed both repeated use of particular VH genes as well as clonal relatedness among clusters of IgE transcripts. CONCLUSION: In contrast to previous studies we observed no preferred VH gene common to IgE transcripts of the three subjects allergic to grass pollen. Moreover, most of the VH gene characteristics of the IgE transcripts were consistent with oligoclonal antigen-driven IgE responses.

Adult↗

Comparative analysis of serine protease-related genes in the honey bee genome: possible involvement in embryonic development and innate immunity.

We have identified 44 serine protease (SP) and 13 serine protease homolog (SPH) genes in the genome of Apis mellifera. Most of these genes encode putative secreted proteins, but four SPs and three SPHs may associate with the plasma membrane via a transmembrane region. Clip domains represent the most abundant non-catalytic structural units in these SP-like proteins -12 SPs and six SPHs contain at least one clip domain. Some of the family members contain other modules for protein-protein interactions, including disulphide-stabilized structures (LDL(r)A, SRCR, frizzled, kringle, Sushi, Wonton and Pan/apple), carbohydrate-recognition domains (C-type lectin and chitin-binding), and other modules (such as zinc finger, CUB, coiled coil and Sina). Comparison of the sequences with those from Drosophila led to a proposed SP pathway for establishing the dorsoventral axis of honey bee embryos. Multiple sequence alignments revealed evolutionary relationships of honey bee SPs and SPHs with those in Drosophila melanogaster, Anopheles gambiae, and Manduca sexta. We identified homologs of D. melanogaster persephone, M. sexta HP14, PAP-1 and SPH-1. A. mellifera genome includes at least five genes for potential SP inhibitors (serpin-1 through -5) and three genes of SP putative substrates (prophenoloxidase, spätzle-1 and spätzle-2). Quantitative RT-PCR analyses showed an elevation in the mRNA levels of SP2, SP3, SP9, SP10, SPH41, SPH42, SP49, serpin-2, serpin-4, serpin-5, and spätzle-2 in adults after a microbial challenge. The SP41 and SP6 transcripts significantly increased after an injection of Paenibacillus larva, but there was no such increase after injection of saline or Escherichia coli. mRNA levels of most SPs and serpins significantly increased by 48 h after the pathogen infection in 1st instar larvae. On the contrary, SP1, SP3, SP19 and serpin-5 transcript levels reduced. These results, taken together, provide a framework for designing experimental studies of the roles of SPs and related proteins in embryonic development and immune responses of A. mellifera.

Amino Acid Sequence↗

Characteristics of the nuclear (18S, 5.8S, 28S and 5S) and mitochondrial (12S and 16S) rRNA genes of Apis mellifera (Insecta: Hymenoptera): structure, organization, and retrotransposable elements.

As an accompanying manuscript to the release of the honey bee genome, we report the entire sequence of the nuclear (18S, 5.8S, 28S and 5S) and mitochondrial (12S and 16S) ribosomal RNA (rRNA)-encoding gene sequences (rDNA) and related internally and externally transcribed spacer regions of Apis mellifera (Insecta: Hymenoptera: Apocrita). Additionally, we predict secondary structures for the mature rRNA molecules based on comparative sequence analyses with other arthropod taxa and reference to recently published crystal structures of the ribosome. In general, the structures of honey bee rRNAs are in agreement with previously predicted rRNA models from other arthropods in core regions of the rRNA, with little additional expansion in non-conserved regions. Our multiple sequence alignments are made available on several public databases and provide a preliminary establishment of a global structural model of all rRNAs from the insects. Additionally, we provide conserved stretches of sequences flanking the rDNA cistrons that comprise the externally transcribed spacer regions (ETS) and part of the intergenic spacer region (IGS), including several repetitive motifs. Finally, we report the occurrence of retrotransposition in the nuclear large subunit rDNA, as R2 elements are present in the usual insertion points found in other arthropods. Interestingly, functional R1 elements usually present in the genomes of insects were not detected in the honey bee rRNA genes. The reverse transcriptase products of the R2 elements are deduced from their putative open reading frames and structurally aligned with those from another hymenopteran insect, the jewel wasp Nasonia (Pteromalidae). Stretches of conserved amino acids shared between Apis and Nasonia are illustrated and serve as potential sites for primer design, as target amplicons within these R2 elements may serve as novel phylogenetic markers for Hymenoptera. Given the impending completion of the sequencing of the Nasonia genome, we expect our report eventually to shed light on the evolution of the hymenopteran genome within higher insects, particularly regarding the relative maintenance of conserved rDNA genes, related variable spacer regions and retrotransposable elements.

3' Untranslated Regions↗

Cloning and expression of a novel esterase gene cpoA from Burkholderia cepacia.

AIMS: To screen and clone a novel enzyme with specific activity for the resolution of (R)-beta-acetylmercaptoisobutyrate (RAM) from (R,S)-beta-acetylmercaptoisobutyrate [(R,S)-ester]. METHODS AND RESULTS: A micro-organism that produces a novel esterase was isolated and identified as the bacterium Burkholderia cepacia by using the analysis of cellular fatty acids, Biolog automated microbial identification/characterization system, and 16S rRNA gene sequence analysis. A novel esterase gene was cloned from the chromosomal DNA of B. cepacia and was designated as cpoA. The cpoA encodes a polypeptide of 273 amino acids which shows a strong sequence homology with many bacterial nonhaeme chloroperoxidases. In addition, a typical serine-hydrolase motif, Gly-X-Ser-X-Gly, and the highly conserved catalytic triad, Ser95, Asp224, and His253, were identified in the deduced amino acid sequence of cpoA by multiple sequence alignment. CONCLUSION: The cpoA cloned from B. cepacia encodes a novel esterase which is highly related to the nonhaeme chloroperoxidases. SIGNIFICANCE AND IMPACT OF THE STUDY: This is the first report that describes the isolation and cloning of a serine esterase gene from B. cepacia, which is useful in the chiral resolution of (R,S)-ester. The cloned gene will allow additional research on the bifunctionality of the enzyme with esterase and chloroperoxidase activity at the structural and functional levels.

Amino Acid Sequence↗

Citrate synthase from the thermophilic archaebacterium Thermoplasma acidophilium. Cloning and sequencing of the gene.

The gene encoding the citric acid cycle enzyme, citrate synthase, has been cloned from the thermoacidophilic archaebacterium, Thermoplasma acidophilum. We report the sequencing of this gene and its flanking regions, and the derived amino acid sequence of the enzyme is compared by multiple-sequence alignment analysis with those of citrate synthases from eubacterial and eukaryotic organisms. The similarity is less than 30% between the archaebacterial and non-archaebacterial sequences, although the majority of residues implicated in the catalytic action of the enzyme have been conserved across all three kingdoms. The cloned archaebacterial gene has been expressed in Escherichia coli to produce catalytically active citrate synthase. This is the first reported sequence of citrate synthase from the archaebacteria.

Amino Acid Sequence↗

Characteristics of an exochitinase from Streptomyces olivaceoviridis, its corresponding gene, putative protein domains and relationship to other chitinases.

Streptomyces olivaceoviridis efficiently degrades chitin. Shotgun cloning of partially Sau3A-cleaved DNA using the multicopy vector pIJ702 and Streptomyces lividans 66 as host resulted in the identification of the plasmid pCHI O1 which harbours an insert of 4.6 kb. In the presence of chitin as sole carbon source, transformants of S. lividans 66 carrying pCHI O1 or its derivatives with smaller inserts overproduced an exochitinase which was purified to homogeneity. The chitin-inducible enzyme with an isoelectric point of 4.0 shows optimal activity at pH 7.3 and 55 degrees C, has an apparent molecular mass of 47 kDa and is competitively inhibited by the pseudosugar allosamidin. The enzyme was identified as an exochitinase since it generates exclusively chitobiose from chitotetraose, chitohexaose, and colloidal high-molecular mass chitin. Sequence analysis of a reading frame of 1794 base pairs and comparison of the deduced amino-acid sequence allowed the identification of the putative catalytic domain, one region with significant similarity to the type-III module of fibronectin and one domain of unknown function. Multiple sequence alignment and hydrophobic-cluster analysis of 25 chitinolytic enzymes from bacteria, fungi and plants allowed the identification of their characteristic domains. The exochitinase from S. olivaceoviridis shares highest similarity with the chitinase D from Bacillus circulans.

Acetylglucosamine↗

The purification, characterization and analysis of primary and secondary-structure of prolyl oligopeptidase from human lymphocytes. Evidence that the enzyme belongs to the alpha/beta hydrolase fold family.

Prolyl oligopeptidase was isolated and purified to homogeneity from human lymphocytes, yielding a specific activity of 7780 mU/mg. The molecular mass using size-exclusion chromatography matches the 76 kDa obtained by SDS/PAGE. This provides evidence that prolyl oligopeptidase is a monomer. The isoelectric point is 4.8 as judged by isoelectric focusing in free solution. Di-isopropyl fluorophosphate and phenylmethylsulphonyl fluoride completely abolish the activity, classifying the enzyme as a serine proteinase. The inhibition by p-chloromercuribenzoic acid indicates the importance of a free sulfhydryl group near the active-site. alpha 1-Casein and ornithine decarboxylase, two proteins containing a PEST sequence, inhibit prolyl oligopeptidase, but were not hydrolyzed. This demonstrates that prolyl oligopeptidase is not participating in the metabolism of proteins according to a PEST-dependent pathway. alpha 1-Antitrypsin partially inhibits the enzyme but in contrast, aprotinin does not. Its inability to cleave corticotropin-releasing factor, ubiquitin, albumin and aprotinin, together with the hydrolysis of bradykinin between Pro7-Arg8 confirms the affinity of prolyl oligopeptidase for small peptides. Multiple sequence alignment does not reveal any similarity with proteases of known tertiary structure. Secondary-structure prediction displays striking similarity with dipeptidyl peptidase IV and acylaminoacyl peptidase. Two characteristic features of the members of the prolyl oligopeptidase family of serine proteases are high-lighted: the linear arrangement of the catalytic triad is nucleophile-acid-base and the proteolytic cleavage releasing the catalytically active C-terminal region of around 500 amino acids from the N-terminal sequence. Secondary structure prediction and comparison of the active-site of serine proteinases with known three-dimensional coordinates prove that Asp641 is the third member of the catalytic triad. The secondary structural organization of the protease domain of prolyl oligopeptidase is in accordance with the alpha/beta hydrolase fold.

Amino Acid Sequence↗