Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Characterization and modeling of membrane proteins using sequence analysis.

The current libraries of amino acid sequences of membrane proteins are a valuable resource for the analysis of elements common to these proteins. Multiple-sequence alignment techniques and the identification of conserved features of transmembrane segments have improved the prediction of membrane protein topology. Molecular modeling in combination with structural studies or site-directed mutagenesis is proving to be a powerful link between theory and experiment. Unfortunately, the number of high-resolution structures of intrinsic membrane proteins, although increased recently, presents a restricted and perhaps biased view of membrane protein structure.

Amino Acid Sequence↗

Progress in protein structure prediction?

Prediction of protein secondary structure is an old problem and progress has been slow. Recently, spectacular success has been claimed in the blind prediction of the catalytic subunit of the cAMP-dependent protein kinase. When predictions in this and other test cases are assessed critically, some claims of prediction success turn out to be exaggerated, but a kernel of real progress remains: protein structure prediction can be improved substantially when a family of related sequences is available. Enough so that molecular biologists equipped with a new amino acid sequence and a multiple sequence alignment in hand may be tempted to test the new prediction methods.

Amino Acid Sequence↗

Sequence analysis of cytochrome bd oxidase suggests a revised topology for subunit I.

Numerous sequences of the cytochrome bd quinol oxidase (cytochrome bd) have recently become available for analysis. The analysis has revealed a small number of conserved residues, a new topology for subunit I and a phylogenetic tree involving extensive horizontal gene transfer. There are 20 conserved residues in subunit I and two in subunit II. Algorithms utilizing multiple sequence alignments predicted a revised topology for cytochrome bd, adding two transmembrane helices to subunit I to the seven that were previously indicated by the analysis of the sequence of the oxidase from E. coli. This revised topology has the effect of relocating the N-terminus and C-terminus to the periplasmic and cytoplasmic sides of the membrane, respectively. The new topology repositions I-H19, the putative ligand for heme b595, close to the periplasmic edge of the membrane, which suggests that the heme b595/heme d active site of the oxidase is located near the outer (periplasmic) surface of the membrane. The most highly conserved region of the sequence of subunit I contains the sequence GRQPW and is located in a predicted periplasmic loop connecting the eighth and ninth transmembrane helices. The potential importance of this region of the protein was previously unsuspected, and it may participate in the binding of either quinol or heme d. There are two very highly conserved glutamates in subunit I, E99 and E107, within the third transmembrane helix (E. coli cytochrome bd-I numbering). It is speculated that these glutamates may be part of a proton channel leading from the cytoplasmic side of the membrane to the heme d oxygen-reactive site, now placed near the periplasmic surface. The revised topology and newly revealed conserved residues provide a clear basis for further experimental tests of these hypotheses. Phylogenetic analysis of the new sequences of cytochrome bd reveals considerable deviation from the 16sRNA tree, suggesting that a large amount of horizontal gene transfer has occurred in the evolution of cytochrome bd.

Amino Acid Sequence↗

Evolutionary relationship between K(+) channels and symporters.

The hypothesis is presented that at least four families of putative K(+) symporter proteins, Trk and KtrAB from prokaryotes, Trk1,2 from fungi, and HKT1 from wheat, evolved from bacterial K(+) channel proteins. Details of this hypothesis are organized around the recently determined crystal structure of a bacterial K(+) channel: i. e., KcsA from Streptomyces lividans. Each of the four identical subunits of this channel has two fully transmembrane helices (designated M1 and M2), plus an intervening hairpin segment that determines the ion selectivity (designated P). The symporter sequences appear to contain four sequential M1-P-M2 motifs (MPM), which are likely to have arisen from gene duplication and fusion of the single MPM motif of a bacterial K(+) channel subunit. The homology of MPM motifs is supported by a statistical comparison of the numerical profiles derived from multiple sequence alignments formed for each protein family. Furthermore, these quantitative results indicate that the KtrAB family of symporters has remained closest to the single-MPM ancestor protein. Strong sequence evidence is also found for homology between the cytoplasmic C-terminus of numerous bacterial K(+) channels and the cytoplasm-resident TrkA and KtrA subunits of the Trk and KtrAB symporters, which in turn are homologous to known dinucleotide-binding domains of other proteins. The case for homology between bacterial K(+) channels and the four families of K(+) symporters is further supported by the accompanying manuscript, in which the patterns of residue conservation are demonstrated to be similar to each other and consistent with the known 3D structure of the KcsA K(+) channel.

Amino Acid Sequence↗

The molecular peculiarities of catalase-peroxidases.

In developing ideas of how protein structure modifies haem reactivity, the activity of Class I of the plant peroxidase superfamily (including cytochrome c peroxidase, ascorbate peroxidase and catalase-peroxidases (KatGs)) is an exciting field of research. Despite striking sequence homologies, there are dramatic differences in catalytic activity and substrate specificity with KatGs being the only member with substantial catalase activity. Based on multiple sequence alignment performed for Class I peroxidases, we present a hypothesis for the pronounced catalase activity of KatGs. In their catalytic domains KatGs are shown to possess three large insertions, two of them are typical for KatGs showing highly conserved sequence patterns. Besides an extra C-terminal copy of the ancestral hydroperoxidase gene resulting from gene duplication, these two large loops are likely to control the orientation of both the haem group and of essential residues in the active site. They seem to modulate the access of substrates to the prosthetic group at the distal side as well as the flexibility and character of the bond between the proximal histidine and the ferric iron. The hypothesis presented opens new possibilities in the rational engineering of peroxidases.

Amino Acid Sequence↗

A refined structure of human aquaporin-1.

A refined structure of the human water channel aquaporin-1 is presented. The model rests on the high resolution X-ray structure of the homologous bacterial glycerol transporter GlpF, electron crystallographic data at 3.8 A resolution and a multiple sequence alignment of the aquaporin superfamily. The crystallographic R and free R values (36.7% and 37.8%) for the refined structure are significantly lower than for previous models. Improved geometry and enhanced stability in molecular dynamics simulations demonstrate a significant improvement of the aquaporin-1 structure. Comparison with previous aquaporin-1 models shows significant differences, not only in the loop regions, but also in the core of the water channel.

Aquaporin 1↗

Evolution of dnaQ, the gene encoding the editing 3' to 5' exonuclease subunit of DNA polymerase III holoenzyme in Gram-negative bacteria.

The nucleotide sequences of the dnaQ genes from Salmonella typhimurium and Buchnera aphidicola, encoding the epsilon-subunit of the DNA polymerase III holoenzyme, have been determined. The Salmonella typhimurium dnaQ protein consists of 243 amino acid residues with a calculated molecular weight of 27224. The Buchnera aphidicola dnaQ protein contains 233 amino acid residues with a calculated molecular weight of 27170. A multiple sequence alignment of the amino acid sequences of the dnaQ proteins and those of DNA polymerase IIIs from Gram-positive bacteria produced six homologous segments. These homologous segments contain highly conserved amino acid sequence motifs involved in catalytically important metal ion bindings (ligands 1, 2 and 3). However, metal ligand 4 is found to be altered in the 3'-5' exonuclease domain of the family C DNA polymerases and dnaQ proteins in Gram-negative bacteria. From these results, we propose that the last common ancestor of the dnaQ gene of Gram-negative bacteria and the DNA polymerase III gene (pol C gene) of Gram-positive bacteria was a single gene containing both 3'-5' exonuclease and DNA polymerase domains and then the dnaQ gene separated from the polymerase gene in Gram-negative bacteria.

Amino Acid Sequence↗

The arrangement of the transmembrane helices in the secretin receptor family of G-protein-coupled receptors.

The members of the secretin receptor family of G-protein-coupled receptors share no significant sequence similarity to the more familiar rhodopsin-like family. However, multiple sequence alignment analysis reveals seven hydrophobic regions with significant alpha-helical periodicity. Residues that are likely to be buried on the interior of the helical bundle and others that are likely to contact the lipid bilayer are identified. A predicted arrangement of the helical bundle is described in which, by comparison with the arrangement in the rhodopsin family, helices 2 and 7 are more buried within the bundle while helix 3 is more exposed to the lipid bilayer.

Amino Acid Sequence↗

Evidence for a new class of scorpion toxins active against K+ channels.

cDNAs encoding novel long-chain scorpion toxins (64 amino acid residues, including only six cysteines) were isolated from cDNA libraries produced from the venom glands of the scorpions Androctonus australis from Old World and Tityus serrulatus from New World. The encoded peptides were very similar to a recently identified toxin from T. serrulatus, which is active against the voltage-sensitive 'delayed-rectifier' potassium channel, but they were completely different from the long-chain and short-chain scorpion toxins already characterised. However, there was some sequence similarity (42%) between these new toxins, Aa TX Kbeta and Ts TX Kbeta, and scorpion defensins purified from the hemolymph of Buthidae scorpions Leiurus quinquestriatus and A. australis. Thus, according to a multiple sequence alignment using CLUSTAL, these new toxins seem to be related to the scorpion defensins.

Amino Acid Sequence↗

Association between fulminant hepatic failure and a strain of GBV virus C.

BACKGROUND: The GB virus C (GBV-C) and the hepatitis G virus (HGV) have been detected in patients with acute indeterminant hepatitis and post-transfusion hepatitis. However, the role of the new hepatitis viruses in the aetiology of fulminant hepatitis is little understood. We investigated the presence of GBV-C/HGV in patients with fulminant hepatic failure. METHODS: Serum samples from 22 German patients with fulminant hepatic failure and 106 symptom-free blood donors (controls) were studied for presence of GBV-C RNA by seminested reverse transcriptase PCR. Primer sequences were derived from the published gene sequences of the conserved NS3 region of the GBV-C prototype and the published isolates. Nucleotide and amino acid sequences of GBV-C-positive isolates, the control RNA, and the published HGV and GBV-C prototype sequences were compared by multiple sequence alignment. We also compared the GBV-C sequences of virus-positive patients who had fulminant hepatic failure with those of 19 patients with chronic hepatitis from our centre. In addition, we searched databases and published papers for further GBV-C helicase sequences in patients with non-fulminant hepatitis. FINDINGS: GBV-C RNA was detected in 11 (50%) of the 22 patients with fulminant hepatic failure and in five (4.7%) of 106 control-group blood donors. Among the patients with fulminant hepatic failure, six of seven with fulminant hepatitis B and five of ten with fulminant non-A-E hepatitis were positive for GBV-C RNA. Analysis of nucleic acid sequences showed six mutations at defined positions in all 11 patients with fulminant hepatic failure who were positive for GBV-C. None of these mutations were found in the five GBV-C-positive control-group blood donors. Of the six nucleotide changes, four caused no amino acid changes, whereas two mutations at position 100 (G to T) and 102 (T to C) led to an alanine to serine change in the predicted translation product. However, comparison with GBV-C sequences of patients with non-fulminant hepatitis showed that this amino acid mutation was not specific for fulminant hepatic failure. The sequence-motif containing the six nucleotide mutations detected in all patients with fulminant hepatic failure was found in only two of 19 German patients with chronic hepatitis from our centre, and in only one of 88 GBV-C sequences from non-fulminant patients reported by others. INTERPRETATION: The frequency of GBV-C RNA is higher in fulminant hepatic failure than in any other group of patients with hepatitis, particularly in patients with fulminant hepatitis B or fulminant non-A-E hepatitis. A specific strain of GBV-C may occur in serum of German patients with fulminant hepatic failure.

Adult↗

Human complement factor I: its expression by insect cells and its biochemical and structural characterisation.

Factor I is a five-domain plasma serine protease which is essential for the regulation of the complement system. In order to express this, the factor I coding sequence was cloned into a recombinant baculovirus system, which was used to infect Trichoplusia ni cells. Using the native factor I leader sequence, recombinant factor I (rFI) was secreted into the culture medium. Purified rFI was recognised by polyclonal antisera and by the factor I-specific monoclonal antibody MRC-OX21. SDS PAGE showed that rFI was processed into two chains with molecular weights of 48,000 and 36,000. Amino acid sequence analysis showed that the N-terminal sequences of the rFI chains were the same as those of serum-derived factor I (sFI), confirming that processing was correct. Since both molecular weights were less than those observed for sFI, this is attributed to the replacement of complex-type oligosaccharides by high mannose ones in rFI. C3(NH,) cleavage assays showed that rFI had 55% the activity of sFI. Circular dichroism and Fourier transform infrared spectroscopy showed that the protein folding of rFI and sFI were very similar. Both had a secondary structure low in alpha-helix and high in beta-sheet, as expected from crystal structure and multiple sequence alignment analyses. It is inferred that the reduced activity of rFI is attributable to its changed glycosylation. The availability of rFI and structures for the domains in factor I makes possible new approaches to determine the molecular basis of its interactions with factor H and C3b.

Animals↗

The chicken genome contains no HMG1 retropseudogenes but a functional HMG1 gene with long introns.

We have cloned the genomic sequence coding for the high mobility group 1 (HMG1) protein in chickens. Multiple sequence alignment shows that the chicken HMG1 gene is highly homologous to the human and the mouse HMG1 genes. The gene structure of chicken HMG1 is similar to that of the mouse and the human HMG1 genes, with the same exon-intron boundaries. However, in contrast to other avian genes that have shorter introns, the chicken HMG1 gene has introns that are twice as long as their mammalian homologues. In addition to the functional, intron-containing HMG1 gene, all mammalian genomes contain more than 50 copies of HMG1 retropseudogenes each, while in the chicken genome there are no HMG1 retropseudogenes. This finding suggests that the HMG1 retropseudogenes arose in mammals after their divergence away from the birds.

Amino Acid Sequence↗

Nucleotide sequence of a cDNA clone coding for an intestinal-type fatty acid binding protein and its tissue-specific expression in zebrafish (Danio rerio).

We have cloned a cDNA from zebrafish (Danio rerio) that contains an open-reading frame of 132 amino acids coding for a fatty acid binding protein (FABP) of approximately 15 kDa. Multiple sequence alignment revealed extensive amino acid identity between this zebrafish FABP and intestinal-like FABPs (I-FABP) from other species. The zebrafish I-FABP cDNA hybridized to single restriction fragments of total zebrafish genomic DNA digested with the restriction endonucleases PstI Bg/II or EcoRI suggesting that a single copy of the I-FABP gene is present in the zebrafish genome. An oligonucleotide probe complementary to the zebrafish I-FABP mRNA hybridized to an mRNA of approximately 800 bases in Northern blot analysis. In situ hybridization revealed that the I-FABP mRNA was expressed exclusively in the intestine of the adult zebrafish.

Amino Acid Sequence↗

Primary structure and unusual carbohydrate moiety of functional unit 2-c of keyhole limpet hemocyanin (KLH).

The complete amino acid sequence of the Megathura crenulata hemocyanin functional unit KLH2-c was determined by direct sequencing and matrix-assisted laser desorption ionization mass spectrometry of the protein, and of peptides obtained by cleavage with EndoLysC proteinase, chymotrypsin and cyanogen bromide. This is the first complete primary structure of a functional unit c from a gastropod hemocyanin. KLH2-c consists of 420 amino acid residues. Circular dichroism spectra indicated approx. 31% beta-sheet and 29% alpha-helix contents. A multiple sequence alignment with other molluscan hemocyanin functional units revealed average identities between 41 and 49%, but 55% in case of Octopus hemocyanin functional unit c which is the structural equivalent to KLH2-c. KLH2-c has a molecular mass of approx. 48 kDa as calculated from its sequence and a measured mass of approx. 56 kDa; the mass difference is attributed to the sugar side chains usually decorating molluscan hemocyanin. However, inspection of the sequence of KLH2-c revealed no potential N-linked carbohydrate attachment sites, and this was supported by its inability to bind concanavalin A. Also KLH1-c was unreactive, whereas most, if not all, other functional units of KLH1 and KLH2 reacted positively to this lectin. On the other hand, peanut agglutinin specifically binds KLH2-c, indicating the presence of O-glycosidically linked carbohydrates in this functional unit. This contrasts to all other KLH functional units (including KLH1-c), which lack O-linked glycosides. The present results are discussed in view of the recent X-ray structure of the functional unit g from Octopus hemocyanin, and a published record of the Thomsen Friedenreich tumor antigenic epitope in KLH.

Amino Acid Sequence↗

PCR primers and functional probes for amplification and detection of bacterial genes for extracellular peptidases in single strains and in soil.

A set of primers and functional probes was developed for the detection of peptidase gene fragments of proteolytic bacteria. Based on DNA sequence data, degenerate PCR primers and internal DIG-labeled probes specific for genes encoding alkaline metallopeptidases (apr) (E.3.4.24), neutral metallopeptidases (npr) (E.3.4.24) and serine peptidases (sub) (E.3.4.21) were derived by multiple sequence alignments. Type strains with known peptidase genes and proteolytic bacteria from a grassland rhizosphere soil, a garden soil and an arable field were investigated for their genotypic proteolytic potential. For 52 out of 53 proteolytic bacterial isolates, at least one of the three peptidase classes could be identified by this approach. The amplified gene fragments were of the expected sizes with each of the three primer sets. The functional probes APR, NPR and SUB have been shown to hybridize specifically to the corresponding gene fragments. sub and npr genes were mainly found in Bacillus species. apr genes were only found in the Pseudomonas fluorescens biotypes and in two morphologically identical Flavobacterium-Cytophaga strains from two different sites. In most of the Bacillus spp., both sub and the npr and in the Flavobacterium-Cytophaga strains even all the three genes could be detected. PCR with DNA isolated from soil led to one main product of the expected size with each primer pair whose identity was additionally confirmed by Southern blot hybridization with the corresponding probes.

Bacteria↗

Evidence for the assignment of two strains of SPLV to the genus Potyvirus based on coat protein and 3' non-coding region sequence data.

The use of potyvirus-specific primers and subsequent application of the RACE procedure allowed the cloning of the 3' terminal 1088 nucleotides of the genomic RNA of the Taiwan isolate of sweetpotato latent virus (SPLV-T) and the 3' genomic 1085 nucleotides of a SPLV-like virus from China (SPLV-CH). The sequence of an internal part of the presumptive nuclear inclusion b gene was also determined for both isolates. Detailed sequence analyses revealed the presence of consensus motifs which indicated that SPLV-CH and SPLV-T should be regarded as members of the genus Potyvirus. Multiple sequence alignments and phylogenetic analyses were also performed and unambiguously assessed these isolates as strains of a distinct Potyvirus. SPLV was not related to other potyviruses infecting sweetpotato nor to any other sequenced virus. From the presence of the DAG box, SPLV-CH is expected to be a typical aphid transmitted Potyvirus whereas a conceivable explanation is proposed for the non-aphid transmission of SPLV-T.

Amino Acid Sequence↗

Use of a neural network secondary structure prediction to define targets for mutagenesis of herpes simplex virus glycoprotein B.

Herpes simplex virus glycoprotein B (HSV gB) is essential for penetration of virus into cells, for cell-to-cell spread of virus, and for cell-cell fusion. Every member of the family Herpesviridae has a gB homolog, underlining its importance. The antigenic structure of gB has been studied extensively, but little is known about which regions of the protein are important for its roles in virus entry and spread. In contrast to successes with other HSV glycoproteins, attempts to map functional domains of gB by insertion mutagenesis have been largely frustrated by the misfolding of most mutants. The present study shows that this problem can be overcome by targeting mutations to the loop regions that connect alpha-helices and beta-strands, avoiding the helices and strands themselves. The positions of loops in the primary sequence were predicted by the PHD neural network procedure, using a multiple sequence alignment of 19 alphaherpesvirus gB sequences as input. Comparison of the prediction with a panel of insertion mutants showed that all mutants with insertions in predicted alpha-helices or beta-strands failed to fold correctly and consequently had no activity in virus entry; in contrast, half the mutants with insertions in predicted loops were able to fold correctly. There are 27 predicted loops of four or more residues in gB; targeting of mutations to these regions will minimize the number of misfolded mutants and maximize the likelihood of identifying functional domains of the protein.

Amino Acid Sequence↗

Knowledge-based grouping of modeled HLA peptide complexes.

Human leukocyte antigens are the most polymorphic of human genes and multiple sequence alignment shows that such polymorphisms are clustered in the functional peptide binding domains. Because of such polymorphism among the peptide binding residues, the prediction of peptides that bind to specific HLA molecules is very difficult. In recent years two different types of computer based prediction methods have been developed and both the methods have their own advantages and disadvantages. The nonavailability of allele specific binding data restricts the use of knowledge-based prediction methods for a wide range of HLA alleles. Alternatively, the modeling scheme appears to be a promising predictive tool for the selection of peptides that bind to specific HLA molecules. The scoring of the modeled HLA-peptide complexes is a major concern. The use of knowledge based rules (van der Waals clashes and solvent exposed hydrophobic residues) to distinguish binders from nonbinders is applied in the present study. The rules based on (1) number of observed atomic clashes between the modeled peptide and the HLA structure, and (2) number of solvent exposed hydrophobic residues on the modeled peptide effectively discriminate experimentally known binders from poor/nonbinders. Solved crystal complexes show no vdW Clash (vdWC) in 95% cases and no solvent exposed hydrophobic peptide residues (SEHPR) were seen in 86% cases. In our attempt to compare experimental binding data with the predicted scores by this scoring scheme, 77% of the peptides are correctly grouped as good binders with a sensitivity of 71%.

Alleles↗