Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

Beyond consensus: statistical free energies reveal hidden interactions in the design of a TPR motif.

Consensus design methods have been used successfully to engineer proteins with a particular fold, and moreover to engineer thermostable exemplars of particular folds. Here, we consider how a statistical free energy approach can expand upon current methods of phylogenetic design. As an example, we have analyzed the tetratricopeptide repeat (TPR) motif, using multiple sequence alignment to identify the significance of each position in the TPR. The results provide information above and beyond that revealed by consensus design alone, especially at poorly conserved positions. A particularly striking finding is that certain residues, which TPR-peptide co-crystal structures show are in direct contact with the ligand, display a marked hypervariability. This suggests a novel means of identifying ligand-binding sites, and also implies that TPRs generally function as ligand-binding domains. Using perturbation analysis (or statistical coupling analysis), we examined site-site interactions within the TPR motif. Correlated occurrences of amino acid residues at poorly conserved positions explain how TPRs achieve their near-neutral surface charge distributions, and why a TPR designed from straight consensus has an unusually high net charge. Networks of interacting sites revealed that TPRs fall into two unrecognized families with distinct sets of interactions related to the identity of position 7 (Leu or Lys/Arg). Statistical free energy analysis provides a more complete description of "What makes a TPR a TPR?" than consensus alone, and it suggests general approaches to extend and improve the phylogenetic design of proteins.

Amino Acid Motifs↗

Sequence correlations between Cro recognition helices and cognate O(R) consensus half-sites suggest conserved rules of protein-DNA recognition.

The O(R) regions from several lambdoid bacteriophages contain the three regulatory sites O(R)1, O(R)2 and O(R)3, to which the Cro and CI proteins can bind. These sites show imperfect dyad symmetry, have similar sequences, and generally lie on the same face of the DNA double helix. We have developed a computational method, which analyzes the O(R) regions of additional phages and predicts the location of these three sites. After tuning the method to predict known O(R) sites accurately, we used it to predict unknown sites, and ultimately compiled a database of 32 known and predicted O(R) binding site sets. We then identified sequences of the recognition helices (RH) for the cognate Cro proteins through manual inspection of multiple sequence alignments. Comparison of Cro RH and consensus O(R) half-site sequences revealed strong one-to-one correlations between two amino acids at each of three RH positions and two bases at each of three half-site positions (H1-->2, H3-->5 and H6-->6). In each of these three cases, one of the two amino acid/base-pairings corresponds to a contact observed in the crystal structure of a lambda Cro/consensus operator complex. The alternate amino acid/base combinations were rationalized using structural models. We suggest that the pairs of amino acid residues act as binary switches that efficiently modulate specificity for different consensus half-site variants during evolution. The observation of structurally reasonable amino acid-to-base correlations suggests that Cro proteins share some common rules of recognition despite their functional and structural diversity.

Bacteriophage lambda↗

The structure of Sif2p, a WD repeat protein functioning in the SET3 corepressor complex.

In Saccharomyces cerevisiae, the SIF2 gene product is an integral component of the Set3 complex (SET3C), an assembly of proteins with some homology to the human SMRT and N-CoR corepressor complexes. SET3C has histone deacetylase activity that is responsible for repressing a set of meiotic genes. We have determined the X-ray crystal structure of a 46 kDa C-terminal domain of a SET3C core protein, Sif2p to 1.55 A resolution and a crystallographic R-factor of 19.0%. This domain contains an unusual eight-bladed beta-propeller structure, which differs from other transcriptional corepressor structures such as yeast Tup1p and human groucho (Gro)/TLE1, which have only seven. We have demonstrated intact Sif2p is a tetramer and the N-terminal LisH (Lis-homology)-containing domain mediates tetramerization and interaction with another component of SET3C, Snt1p. Multiple sequence alignments indicate that a surface on the "top" of the protein is conserved among species, suggesting that it may play a common role in binding partner proteins. Since Sif2p appears to be the yeast homolog of human TBL1 and TBLR1, which function in the N-CoR/SMRT complexes, its structural and oligomeric properties are likely to be very similar.

Amino Acid Sequence↗

High-resolution crystal structure of AKR11C1 from Bacillus halodurans: an NADPH-dependent 4-hydroxy-2,3-trans-nonenal reductase.

Aldo-keto reductase AKR11C1 from Bacillus halodurans, a new member of aldo-keto reductase (AKR) family 11, has been characterized structurally and biochemically. The structures of the apo and NADPH bound form of AKR11C1 have been solved to 1.25 A and 1.3 A resolution, respectively. AKR11C1 possesses a novel non-aromatic stacking interaction of an arginine residue with the cofactor, which may favor release of the oxidized cofactor. Our biochemical studies have revealed an NADPH-dependent activity of AKR11C1 with 4-hydroxy-2,3-trans-nonenal (HNE). HNE is a cytotoxic lipid peroxidation product, and detoxification in alkaliphilic bacteria, such as B.halodurans, plays a crucial role in survival. AKR11C1 could thus be part of the detoxification system, which ensures the well being of the microorganism. The very poor activity of AKR11C1 on standard, small substrates such as benzaldehyde or DL-glyeraldehyde is consistent with the observed, very open active site lacking a binding pocket for these substrates. In contrast, modeling of HNE with its aldehyde function suitably positioned in the active site suggests that its elongated hydrophobic tail occupies a groove defined by hydrophobic side-chains. Multiple sequence alignment of AKR11C1 with the highly homologous iolS and YqkF proteins shows a high level of conservation in this putative substrate-binding site. We suggest that AKR11C1 is the first structurally characterized member of a new class of AKRs with specificity for substrates with long aliphatic tails.

Alcohol Oxidoreductases↗

Crystal structure of papaya glutaminyl cyclase, an archetype for plant and bacterial glutaminyl cyclases.

Glutaminyl cyclases (QCs) (EC 2.3.2.5) catalyze the intramolecular cyclization of protein N-terminal glutamine residues into pyroglutamic acid with the concomitant liberation of ammonia. QCs may be classified in two groups containing, respectively, the mammalian enzymes, and the enzymes from plants, bacteria, and parasites. The crystal structure of the QC from the latex of Carica papaya (PQC) has been determined at 1.7A resolution. The structure was solved by the single wavelength anomalous diffraction technique using sulfur and zinc as anomalous scatterers. The enzyme folds into a five-bladed beta-propeller, with two additional alpha-helices and one beta hairpin. The propeller closure is achieved via an original molecular velcro, which links the last two blades into a large eight stranded beta-sheet. The zinc ion present in the PQC is bound via an octahedral coordination into an elongated cavity located along the pseudo 5-fold axis of the beta-propeller fold. This zinc ion presumably plays a structural role and may contribute to the exceptional stability of PQC, along with an extended hydrophobic packing, the absence of long loops, the three-joint molecular velcro and the overall folding itself. Multiple sequence alignments combined with structural analyses have allowed us to tentatively locate the active site, which is filled in the crystal structure either by a Tris molecule or an acetate ion. These analyses are further supported by the experimental evidence that Tris is a competitive inhibitor of PQC. The active site is located at the C-terminal entrance of the PQC central tunnel. W83, W110, W169, Q24, E69, N155, K225, F22 and F67 are highly conserved residues in the C-terminal entrance, and their putative role in catalysis is discussed. The PQC structure is representative of the plants, bacterial and parasite enzymes and contrasts with that of mammalian enzymes, that may possibly share a conserved scaffold of the bacterial aminopeptidase.

Amino Acid Sequence↗

Nanopore-protein interactions dramatically alter stability and yield of the native state in restricted spaces.

We have studied the stability and the yield of the folded WW domains in a spherical nanopore to provide insights into the changes in the folding characteristics due to interactions of the polypeptide (SP) with the walls of the pore. Using different models for the interactions between the nanopore and the polypeptide chain we have obtained results that are relevant to a broad range of experiments. (a) In the temperature and the strength of the SP-pore interaction plane (lambda), there are four "phases," namely, the unfolded state, the native state, the molten globule phase (MG), and the surface interaction-stabilized (SIS) state. The MG and SIS states are populated at moderate and large values of lambda, respectively. For a fixed pore size, the folding rates vary non-monotonically as lambda is varied with a maximum at lambda approximately 1 at which the SP-nanopore interaction is comparable to the stability of the native state. At large lambda values, the WW domain is kinetically trapped in the SIS states. Using multiple sequence alignment, we conclude that similar folding mechanism should be observed in other WW domains as well. (b) To mimic the changes in the nature of the allosterically driven SP-GroEL interactions we consider two models for the dynamic Anfinsen cage (DAC). In DAC1, the SP-cavity interaction cycles between hydrophobic (lambda>0) and hydrophilic (lambda=0) with a period tau. The yield of the native state is a maximum for an optimum value of tau=tau(OPT). At tau=tau(OPT), the largest yield of the native state is obtained when tau(H) approximately tau(P) where tau(H)(tau(P)) is the duration for which the cavity is hydrophobic (hydrophilic). Thus, in order to enhance the native state yield, the cycling rate, for a given loading rate of the GroEL nanomachine, should be maximized. In DAC2, the volume of the cavity is doubled (as happens when ATP and GroES bind to GroEL) and the SP-pore interaction simultaneously changes from hydrophobic to hydrophilic. In this case, we find greater increase in yield of the native state compared to DAC1 at all values of tau.

Mathematics↗

Role of structural and dynamical plasticity in Sin3: the free PAH2 domain is a folded module in mSin3B.

The co-repressor Sin3 is the essential scaffold protein of the Sin3/HDAC co-repressor complex, which is recruited to the DNA by a diverse group of transcriptional repressors, targeting genes involved in the regulation of the cell cycle, proliferation and differentiation. Sin3 contains four repeats commonly denoted as paired amphipathic helix (PAH1-4) domains that provide the principal interaction surface for various repressors. Here, we present the first structure of the free state of the PAH2 domain and discuss its implications for interaction with the repressors. The unbound conformation is very similar to the conformation observed when bound to either the Mad1 or HBP1 repressor, suggesting that the PAH2 domain serves as a template that guides proper folding of the unstructured repressor. The free PAH2 domain shows micro- to millisecond conformational exchange between the folded, major state and a partially unfolded, minor state. Upon complex formation, we observe a significant decrease in fast time-scale flexibility of local regions of the protein, correlated with the formation of intermolecular contacts, and an overall decrease in the slow time-scale conformational exchange. On the basis of our data and using a multiple sequence alignment of all PAH domains, we suggest that the PAH1, PAH2 and PAH3 domains form pre-folded binding modules in full-length Sin3 like beads-on-a-string, and act as folding templates for the interaction domains of their targets.

Amino Acid Sequence↗

Two-rung model of a left-handed beta-helix for prions explains species barrier and strain variation in transmissible spongiform encephalopathies.

In this study, a new beta-helical model is proposed that explains the species barrier and strain variation in transmissible spongiform encephalopathies. The left-handed beta-helix serves as a structural model that can explain the seeded growth characteristics of beta-sheet structure in PrP(Sc) fibrils. Molecular dynamics simulations demonstrate that the left-handed beta-helix is structurally more stable than the right-handed beta-helix, with a higher beta-sheet content during the simulation and a better distributed network of inter-strand backbone-backbone hydrogen bonds between parallel beta-strands of different rungs. Multiple sequence alignments and homology modelling of prion sequences with different rungs of left-handed beta-helices illustrate that the PrP region with the highest beta-helical propensity (residues 105-143) can fold in just two rungs of a left-handed beta-helix. Even if no other flanking sequence participates in the beta-helix, the two rungs of a beta-helix can give the growing fibril enough elevation to accommodate the rest of the PrP protein in a tight packing at the periphery of a trimeric beta-helix. The folding of beta-helices is driven by backbone-backbone hydrogen bonding and stacking of side-chains in adjacent rungs. The sequence and structure of the last rung at the fibril end with unprotected beta-sheet edges selects the sequence of a complementary rung and dictates the folding of the new rung with optimal backbone hydrogen bonding and side-chain stacking. An important side-chain stack that facilitates the beta-helical folding is between methionine residues 109 and 129, which explains their importance in the species barrier of prions. Because the PrP sequence is not evolutionarily optimised to fold in a beta-helix, and because the beta-helical fold shows very little sequence preference, alternative alignments are possible that result in a different rung able to select for an alternative complementary rung. A different top rung results in a new strain with different growth characteristics. Hence, in the present model, sequence variation and alternative alignments clarify the basis of the species barrier and strain specificity in PrP-based diseases.

Amino Acid Sequence↗

The solution structure of antigen MPT64 from Mycobacterium tuberculosis defines a new family of beta-grasp proteins.

The MPT64 protein and its homologs form a highly conserved family of secreted proteins with unknown function that are found within the pathogenic Mycobacteria genus. The founding member of this family from Mycobacterium tuberculosis (MPT64 or protein Rv1980c) is expressed only when Mycobacteria cells are actively dividing. By virtue of this relatively unique expression profile, Rv1980c is currently under phase III clinical trials to evaluate its potential to replace tuberculin, or purified protein derivative, as the rapid diagnostic of choice for detection of active tuberculosis infection. We describe here the NMR solution structure of Rv1980c. This structure reveals a previously undescribed fold that is based upon a variation of a beta-grasp motif most commonly found in protein-protein interaction domains. Examination of this structure in conjunction with multiple sequence alignments of MPT64 homologs identifies a candidate ligand-binding site, which may help guide future studies of Rv1980c function. The work presented here also suggests structure-based approaches for increasing the antigenic potency of a Rv1980c-based diagnostic.

Amino Acid Motifs↗

Molecular modelling of the GABAA ion channel protein.

The GABAA ion channel protein is central to the mechanism of action of general anaesthetics and thus to the phenomenon of human consciousness. A molecular model of the alpha1beta2gamma2 gamma-aminobutyric acid type-A (GABAA) ligand-gated ion channel protein has been constructed. The cryo-electron microscopy structure of the nicotinic acetylcholine receptor (nAChR) from Torpedo marmorata and the X-ray crystal structure of the acetylcholine binding protein (AChBP) from Lymnaea stagnalis were used as starting templates for comparative modelling. Features of the modelling approach used in the development of this GABAA model include: (1) multiple sequence alignment of members of the Cys-loop superfamily; (2) the design and implementation of a quasi-ab initio loop modelling algorithm; (3) expansion of the transmembrane domain (TMD) ion pore to model the open-state of the GABAA channel; (4) hydrophobicity analysis of the TMD to refine the structure in regions involved in general anaesthetic binding. The final model of the alpha1beta2gamma2 GABAA protein agrees with available experimental data concerning general anaesthetics.

Amino Acid Sequence↗

Identification and characterization of the novel gene GhDBP2 encoding a DRE-binding protein from cotton (Gossypium hirsutum).

A cDNA encoding one novel DRE-binding protein, GhDBP2, was isolated from cotton seedlings. It is classified into the A-6 group of DREB subfamily based on multiple sequence alignment and phylogenetic characterization. Using semi-quantitative RT-PCR, we found that the GhDBP2 transcripts were greatly induced by drought, NaCl, low temperature and ABA treatments in cotton cotyledons. The DNA-binding properties of GhDBP2 were analyzed by electrophoretic mobility shift assay (EMSA), showing that GhDBP2 successfully binds to the previously characterized DRE cis-element as well as the promoter region of the LEA D113 gene. Consistent with its role as a DNA-binding protein, GhDBP2 is preferentially localized to the nucleus of onion epidermal cells. In addition, when GhDBP2 is transiently expressed in tobacco cells, it activates reporter gene expression driven by the LEA D113 promoter. Taken together, our results indicate that GhDBP2 is a DRE-binding transcriptional activator involved in activation of down-stream genes such as LEA D113 expression through interaction with the DRE element, in response to environmental stresses as well as ABA treatment.

Amino Acid Sequence↗

Development of a rapid, sensitive and specific diagnostic assay for fish Aquareovirus based on RT-PCR.

A rapid, sensitive and highly specific detection method for Aquareovirus based on reverse-transcription polymerase chain reaction (RT-PCR) was developed. Based on multiple sequence alignment of the cloned sequences of a local isolates, the Threadfin reovirus (TFV) and Guppy reovirus (GPV) with Grass carp reovirus (GCRV), a pair of degenerate primers was selected carefully and synthesized. Using this primer combination, only one specific product, approximately 450 bp in length was obtained when RT-PCR was carried out using the genomic double-stranded RNA (dsRNA) of TFV, GPV and GCRV. Similar results were also obtained when Chum salmon reovirus (CSRV) and Striped bass reovirus (SBRV) dsRNA were used as templates. No products were observed when nucleic acids other than the dsRNA of the aquareoviruses described above were used as RT-PCR templates. This technique could detect not only TFV but also GPV and GCRV in low titer virus-infected cell cultured cells. Furthermore, this method has also been shown to be able to diagnose GPV-infected guppy (Poecilia reticulata) that exhibit clinical symptoms as well as GPV-carrier guppy. Collectively, these results showed that the RT-PCR amplification method using specific degenerate primers described below is very useful for rapid and accurate detection of a variety of aquareovirus strains isolated from different host species and origin.

Animals↗

A broadly reactive one-step real-time RT-PCR assay for rapid and sensitive detection of hepatitis E virus.

Hepatitis E virus (HEV) is transmitted by the fecal-oral route and causes sporadic and epidemic forms of acute hepatitis. Large waterborne HEV epidemics have been documented exclusively in developing countries. At least four major genotypes of HEV have been reported worldwide: genotype 1 (found primarily in Asian countries), genotype 2 (isolated from a single outbreak in Mexico), genotype 3 (identified in swine and humans in the United States and many other countries), and genotype 4 (identified in humans, swine and other animals in Asia). To better detect and quantitate different HEV strains that may be present in clinical and environmental samples, we developed a rapid and sensitive real-time RT-PCR assay for the detection of HEV RNA. Primers and probes for the real-time RT-PCR were selected based on the multiple sequence alignments of 27 sequences of the ORF3 region. Thirteen HEV isolates representing genotypes 1-4 were used to standardize the real-time RT-PCR assay. The TaqMan assay detected as few as four genome equivalent (GE) copies of HEV plasmid DNA and detected as low as 0.12 50% pig infectious dose (PID50) of swine HEV. Different concentrations of swine HEV (120-1.2PID50) spiked into a surface water concentrate were detected in the real-time RT-PCR assay. This is the first reporting of a broadly reactive TaqMan RT-PCR assay for the detection of HEV in clinical and environmental samples.

Animals↗

A multiplex RT-PCR method for screening of reassortant live influenza vaccine virus strains.

A system based on reverse transcription polymerase chain reaction (RT-PCR) of the RNA genome was established to identify genetic composition of influenza viruses generated by reassortment between an attenuated donor virus and virulent wild type virus. The primers were designed, by multiple sequence alignment of variable regions, specific for cold-adapted donor virus HTCA-A 101, as compared to other influenza A viruses. The specificity of each primer set was confirmed and the primers were combined to perform RT-PCR in multiplex manner. The multiplex PCR was adopted to distinguish the 6:2 reassortant viruses containing six internal genome segments of attenuated donor virus and two surface antigens of virulent strain from the wild type viruses. The method allowed us to optimize the reassorting process on a routine basis and to confirm the selection of reassortant clones efficiently. The method is suitable for analyzing the contribution of specific gene segments for growth and attenuating characteristics and for generation of live attenuated vaccine by annual reassortment.

Adaptation, Physiological↗

Computational prediction of the effects of non-synonymous single nucleotide polymorphisms in human DNA repair genes.

Non-synonymous single nucleotide polymorphisms (nsSNPs) represent common genetic variation that alters encoded amino acids in proteins. All nsSNPs may potentially affect the structure or function of expressed proteins and could therefore have an impact on complex diseases. In an effort to evaluate the phenotypic effect of all known nsSNPs in human DNA repair genes, we have characterized each polymorphism in terms of different functional properties. The properties are computed based on amino acid characteristics (e.g. residue volume change); position-specific phylogenetic information from multiple sequence alignments and from prediction programs such as SIFT (Sorting Intolerant From Tolerant) and PolyPhen (Polymorphism Phenotyping). We provide a comprehensive, updated list of all validated nsSNPs from dbSNP (public database of human single nucleotide polymorphisms at National Center for Biotechnology Information, USA) located in human DNA repair genes. The list includes repair enzymes, genes associated with response to DNA damage as well as genes implicated with genetic instability or sensitivity to DNA damaging agents. Out of a total of 152 genes involved in DNA repair, 95 had validated nsSNPs in them. The fraction of nsSNPs that had high probability of being functionally significant was predicted to be 29.6% and 30.9%, by SIFT and PolyPhen respectively. The resulting list of annotated nsSNPs is available online (http://dna.uio.no/repairSNP), and is an ongoing project that will continue assessing the function of coding SNPs in human DNA repair genes.

Computational Biology↗

cDNA sequences, MALDI-TOF analyses, and molecular modelling of barley PR-5 proteins.

Barley plants are known to produce various PR-5 proteins. Transcripts encoding eight different barley PR-5 proteins (TLPs 1-8, TLP for thaumatin-like protein) were identified and cloned - seven from infected leaves and one from developing grains. Here, we describe the cDNA sequences of four of these TLP isoforms. Moreover, the TLPs from the infected leaves (TLPs 1, 2, and TLPs 4-8) were subjected to MALDI-TOF mass spectrometric measurements that resulted in protein fragments consistent with their deduced peptide sequences. Multiple sequence alignment analysis revealed that the TLPs in barley fall into two groups: long-chain proteins (TLPs 5-8) having 16 cysteine residues and short-chain proteins (TLPs 1-4) with only 10 cysteine residues. Finally, modelling experiments highlighted the effects of sequence differences between the TLP isoforms in terms of their secondary structures and their molecular electrostatic potentials. We propose that these sequence differences have implications for the target preferences of the different isomers.

Amino Acid Sequence↗

Comparative genomics and functional roles of the ATP-dependent proteases Lon and Clp during cytosolic protein degradation.

The general pathway involving adenosine triphosphate (ATP)-dependent proteases and ATP-independent peptidases during cytosolic protein degradation is conserved, with differences in the enzymes utilized, in organisms from different kingdoms. Lon and caseinolytic protease (Clp) are key enzymes responsible for the ATP-dependent degradation of cytosolic proteins in Escherichia coli. Orthologs of E. coli Lon and Clp were searched for, followed by multiple sequence alignment of active site residues, in genomes from seventeen organisms, including representatives from eubacteria, archaea, and eukaryotes. Lon orthologs, unlike ClpP and ClpQ, are present in most organisms studied. The roles of these proteases as essential enzymes and in the virulence of some organisms are discussed.

Adenosine Triphosphate↗

Transmembrane domain prediction and consensus sequence identification of the oligopeptide transport family.

Few polytopic membrane proteins have had their topology determined experimentally. Often, researchers turn to an algorithm to predict where the transmembrane domains might lie. Here we use a consensus method, using six different transmembrane domain prediction algorithms on six members of the oligopeptide transport family, all of which have been experimentally characterized. PSI-BLAST results indicate that the six chosen oligopeptide transport family members are distributed throughout most branches of the phylogram, suggesting that these members represent a broad view of the oligopeptide transport family. We combined the prediction algorithms with a multiple sequence alignment, and consensus transmembrane domains were assigned not only based on algorithmic output, but also based on conserved familial motifs found by analysis of the PSI-BLAST results. The consensus method combined with the "charge-difference rule" yields a model topology for the family containing 12 transmembrane domains with the N- and C-termini facing extracellular.

Algorithms↗