Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Evolution of dnaQ, the gene encoding the editing 3' to 5' exonuclease subunit of DNA polymerase III holoenzyme in Gram-negative bacteria.

The nucleotide sequences of the dnaQ genes from Salmonella typhimurium and Buchnera aphidicola, encoding the epsilon-subunit of the DNA polymerase III holoenzyme, have been determined. The Salmonella typhimurium dnaQ protein consists of 243 amino acid residues with a calculated molecular weight of 27224. The Buchnera aphidicola dnaQ protein contains 233 amino acid residues with a calculated molecular weight of 27170. A multiple sequence alignment of the amino acid sequences of the dnaQ proteins and those of DNA polymerase IIIs from Gram-positive bacteria produced six homologous segments. These homologous segments contain highly conserved amino acid sequence motifs involved in catalytically important metal ion bindings (ligands 1, 2 and 3). However, metal ligand 4 is found to be altered in the 3'-5' exonuclease domain of the family C DNA polymerases and dnaQ proteins in Gram-negative bacteria. From these results, we propose that the last common ancestor of the dnaQ gene of Gram-negative bacteria and the DNA polymerase III gene (pol C gene) of Gram-positive bacteria was a single gene containing both 3'-5' exonuclease and DNA polymerase domains and then the dnaQ gene separated from the polymerase gene in Gram-negative bacteria.

Amino Acid Sequence↗

The arrangement of the transmembrane helices in the secretin receptor family of G-protein-coupled receptors.

The members of the secretin receptor family of G-protein-coupled receptors share no significant sequence similarity to the more familiar rhodopsin-like family. However, multiple sequence alignment analysis reveals seven hydrophobic regions with significant alpha-helical periodicity. Residues that are likely to be buried on the interior of the helical bundle and others that are likely to contact the lipid bilayer are identified. A predicted arrangement of the helical bundle is described in which, by comparison with the arrangement in the rhodopsin family, helices 2 and 7 are more buried within the bundle while helix 3 is more exposed to the lipid bilayer.

Amino Acid Sequence↗

Evidence for a new class of scorpion toxins active against K+ channels.

cDNAs encoding novel long-chain scorpion toxins (64 amino acid residues, including only six cysteines) were isolated from cDNA libraries produced from the venom glands of the scorpions Androctonus australis from Old World and Tityus serrulatus from New World. The encoded peptides were very similar to a recently identified toxin from T. serrulatus, which is active against the voltage-sensitive 'delayed-rectifier' potassium channel, but they were completely different from the long-chain and short-chain scorpion toxins already characterised. However, there was some sequence similarity (42%) between these new toxins, Aa TX Kbeta and Ts TX Kbeta, and scorpion defensins purified from the hemolymph of Buthidae scorpions Leiurus quinquestriatus and A. australis. Thus, according to a multiple sequence alignment using CLUSTAL, these new toxins seem to be related to the scorpion defensins.

Amino Acid Sequence↗

Association between fulminant hepatic failure and a strain of GBV virus C.

BACKGROUND: The GB virus C (GBV-C) and the hepatitis G virus (HGV) have been detected in patients with acute indeterminant hepatitis and post-transfusion hepatitis. However, the role of the new hepatitis viruses in the aetiology of fulminant hepatitis is little understood. We investigated the presence of GBV-C/HGV in patients with fulminant hepatic failure. METHODS: Serum samples from 22 German patients with fulminant hepatic failure and 106 symptom-free blood donors (controls) were studied for presence of GBV-C RNA by seminested reverse transcriptase PCR. Primer sequences were derived from the published gene sequences of the conserved NS3 region of the GBV-C prototype and the published isolates. Nucleotide and amino acid sequences of GBV-C-positive isolates, the control RNA, and the published HGV and GBV-C prototype sequences were compared by multiple sequence alignment. We also compared the GBV-C sequences of virus-positive patients who had fulminant hepatic failure with those of 19 patients with chronic hepatitis from our centre. In addition, we searched databases and published papers for further GBV-C helicase sequences in patients with non-fulminant hepatitis. FINDINGS: GBV-C RNA was detected in 11 (50%) of the 22 patients with fulminant hepatic failure and in five (4.7%) of 106 control-group blood donors. Among the patients with fulminant hepatic failure, six of seven with fulminant hepatitis B and five of ten with fulminant non-A-E hepatitis were positive for GBV-C RNA. Analysis of nucleic acid sequences showed six mutations at defined positions in all 11 patients with fulminant hepatic failure who were positive for GBV-C. None of these mutations were found in the five GBV-C-positive control-group blood donors. Of the six nucleotide changes, four caused no amino acid changes, whereas two mutations at position 100 (G to T) and 102 (T to C) led to an alanine to serine change in the predicted translation product. However, comparison with GBV-C sequences of patients with non-fulminant hepatitis showed that this amino acid mutation was not specific for fulminant hepatic failure. The sequence-motif containing the six nucleotide mutations detected in all patients with fulminant hepatic failure was found in only two of 19 German patients with chronic hepatitis from our centre, and in only one of 88 GBV-C sequences from non-fulminant patients reported by others. INTERPRETATION: The frequency of GBV-C RNA is higher in fulminant hepatic failure than in any other group of patients with hepatitis, particularly in patients with fulminant hepatitis B or fulminant non-A-E hepatitis. A specific strain of GBV-C may occur in serum of German patients with fulminant hepatic failure.

Adult↗

Human complement factor I: its expression by insect cells and its biochemical and structural characterisation.

Factor I is a five-domain plasma serine protease which is essential for the regulation of the complement system. In order to express this, the factor I coding sequence was cloned into a recombinant baculovirus system, which was used to infect Trichoplusia ni cells. Using the native factor I leader sequence, recombinant factor I (rFI) was secreted into the culture medium. Purified rFI was recognised by polyclonal antisera and by the factor I-specific monoclonal antibody MRC-OX21. SDS PAGE showed that rFI was processed into two chains with molecular weights of 48,000 and 36,000. Amino acid sequence analysis showed that the N-terminal sequences of the rFI chains were the same as those of serum-derived factor I (sFI), confirming that processing was correct. Since both molecular weights were less than those observed for sFI, this is attributed to the replacement of complex-type oligosaccharides by high mannose ones in rFI. C3(NH,) cleavage assays showed that rFI had 55% the activity of sFI. Circular dichroism and Fourier transform infrared spectroscopy showed that the protein folding of rFI and sFI were very similar. Both had a secondary structure low in alpha-helix and high in beta-sheet, as expected from crystal structure and multiple sequence alignment analyses. It is inferred that the reduced activity of rFI is attributable to its changed glycosylation. The availability of rFI and structures for the domains in factor I makes possible new approaches to determine the molecular basis of its interactions with factor H and C3b.

Animals↗

The chicken genome contains no HMG1 retropseudogenes but a functional HMG1 gene with long introns.

We have cloned the genomic sequence coding for the high mobility group 1 (HMG1) protein in chickens. Multiple sequence alignment shows that the chicken HMG1 gene is highly homologous to the human and the mouse HMG1 genes. The gene structure of chicken HMG1 is similar to that of the mouse and the human HMG1 genes, with the same exon-intron boundaries. However, in contrast to other avian genes that have shorter introns, the chicken HMG1 gene has introns that are twice as long as their mammalian homologues. In addition to the functional, intron-containing HMG1 gene, all mammalian genomes contain more than 50 copies of HMG1 retropseudogenes each, while in the chicken genome there are no HMG1 retropseudogenes. This finding suggests that the HMG1 retropseudogenes arose in mammals after their divergence away from the birds.

Amino Acid Sequence↗

Nucleotide sequence of a cDNA clone coding for an intestinal-type fatty acid binding protein and its tissue-specific expression in zebrafish (Danio rerio).

We have cloned a cDNA from zebrafish (Danio rerio) that contains an open-reading frame of 132 amino acids coding for a fatty acid binding protein (FABP) of approximately 15 kDa. Multiple sequence alignment revealed extensive amino acid identity between this zebrafish FABP and intestinal-like FABPs (I-FABP) from other species. The zebrafish I-FABP cDNA hybridized to single restriction fragments of total zebrafish genomic DNA digested with the restriction endonucleases PstI Bg/II or EcoRI suggesting that a single copy of the I-FABP gene is present in the zebrafish genome. An oligonucleotide probe complementary to the zebrafish I-FABP mRNA hybridized to an mRNA of approximately 800 bases in Northern blot analysis. In situ hybridization revealed that the I-FABP mRNA was expressed exclusively in the intestine of the adult zebrafish.

Amino Acid Sequence↗

Primary structure and unusual carbohydrate moiety of functional unit 2-c of keyhole limpet hemocyanin (KLH).

The complete amino acid sequence of the Megathura crenulata hemocyanin functional unit KLH2-c was determined by direct sequencing and matrix-assisted laser desorption ionization mass spectrometry of the protein, and of peptides obtained by cleavage with EndoLysC proteinase, chymotrypsin and cyanogen bromide. This is the first complete primary structure of a functional unit c from a gastropod hemocyanin. KLH2-c consists of 420 amino acid residues. Circular dichroism spectra indicated approx. 31% beta-sheet and 29% alpha-helix contents. A multiple sequence alignment with other molluscan hemocyanin functional units revealed average identities between 41 and 49%, but 55% in case of Octopus hemocyanin functional unit c which is the structural equivalent to KLH2-c. KLH2-c has a molecular mass of approx. 48 kDa as calculated from its sequence and a measured mass of approx. 56 kDa; the mass difference is attributed to the sugar side chains usually decorating molluscan hemocyanin. However, inspection of the sequence of KLH2-c revealed no potential N-linked carbohydrate attachment sites, and this was supported by its inability to bind concanavalin A. Also KLH1-c was unreactive, whereas most, if not all, other functional units of KLH1 and KLH2 reacted positively to this lectin. On the other hand, peanut agglutinin specifically binds KLH2-c, indicating the presence of O-glycosidically linked carbohydrates in this functional unit. This contrasts to all other KLH functional units (including KLH1-c), which lack O-linked glycosides. The present results are discussed in view of the recent X-ray structure of the functional unit g from Octopus hemocyanin, and a published record of the Thomsen Friedenreich tumor antigenic epitope in KLH.

Amino Acid Sequence↗

PCR primers and functional probes for amplification and detection of bacterial genes for extracellular peptidases in single strains and in soil.

A set of primers and functional probes was developed for the detection of peptidase gene fragments of proteolytic bacteria. Based on DNA sequence data, degenerate PCR primers and internal DIG-labeled probes specific for genes encoding alkaline metallopeptidases (apr) (E.3.4.24), neutral metallopeptidases (npr) (E.3.4.24) and serine peptidases (sub) (E.3.4.21) were derived by multiple sequence alignments. Type strains with known peptidase genes and proteolytic bacteria from a grassland rhizosphere soil, a garden soil and an arable field were investigated for their genotypic proteolytic potential. For 52 out of 53 proteolytic bacterial isolates, at least one of the three peptidase classes could be identified by this approach. The amplified gene fragments were of the expected sizes with each of the three primer sets. The functional probes APR, NPR and SUB have been shown to hybridize specifically to the corresponding gene fragments. sub and npr genes were mainly found in Bacillus species. apr genes were only found in the Pseudomonas fluorescens biotypes and in two morphologically identical Flavobacterium-Cytophaga strains from two different sites. In most of the Bacillus spp., both sub and the npr and in the Flavobacterium-Cytophaga strains even all the three genes could be detected. PCR with DNA isolated from soil led to one main product of the expected size with each primer pair whose identity was additionally confirmed by Southern blot hybridization with the corresponding probes.

Bacteria↗

Evidence for the assignment of two strains of SPLV to the genus Potyvirus based on coat protein and 3' non-coding region sequence data.

The use of potyvirus-specific primers and subsequent application of the RACE procedure allowed the cloning of the 3' terminal 1088 nucleotides of the genomic RNA of the Taiwan isolate of sweetpotato latent virus (SPLV-T) and the 3' genomic 1085 nucleotides of a SPLV-like virus from China (SPLV-CH). The sequence of an internal part of the presumptive nuclear inclusion b gene was also determined for both isolates. Detailed sequence analyses revealed the presence of consensus motifs which indicated that SPLV-CH and SPLV-T should be regarded as members of the genus Potyvirus. Multiple sequence alignments and phylogenetic analyses were also performed and unambiguously assessed these isolates as strains of a distinct Potyvirus. SPLV was not related to other potyviruses infecting sweetpotato nor to any other sequenced virus. From the presence of the DAG box, SPLV-CH is expected to be a typical aphid transmitted Potyvirus whereas a conceivable explanation is proposed for the non-aphid transmission of SPLV-T.

Amino Acid Sequence↗

Use of a neural network secondary structure prediction to define targets for mutagenesis of herpes simplex virus glycoprotein B.

Herpes simplex virus glycoprotein B (HSV gB) is essential for penetration of virus into cells, for cell-to-cell spread of virus, and for cell-cell fusion. Every member of the family Herpesviridae has a gB homolog, underlining its importance. The antigenic structure of gB has been studied extensively, but little is known about which regions of the protein are important for its roles in virus entry and spread. In contrast to successes with other HSV glycoproteins, attempts to map functional domains of gB by insertion mutagenesis have been largely frustrated by the misfolding of most mutants. The present study shows that this problem can be overcome by targeting mutations to the loop regions that connect alpha-helices and beta-strands, avoiding the helices and strands themselves. The positions of loops in the primary sequence were predicted by the PHD neural network procedure, using a multiple sequence alignment of 19 alphaherpesvirus gB sequences as input. Comparison of the prediction with a panel of insertion mutants showed that all mutants with insertions in predicted alpha-helices or beta-strands failed to fold correctly and consequently had no activity in virus entry; in contrast, half the mutants with insertions in predicted loops were able to fold correctly. There are 27 predicted loops of four or more residues in gB; targeting of mutations to these regions will minimize the number of misfolded mutants and maximize the likelihood of identifying functional domains of the protein.

Amino Acid Sequence↗

Knowledge-based grouping of modeled HLA peptide complexes.

Human leukocyte antigens are the most polymorphic of human genes and multiple sequence alignment shows that such polymorphisms are clustered in the functional peptide binding domains. Because of such polymorphism among the peptide binding residues, the prediction of peptides that bind to specific HLA molecules is very difficult. In recent years two different types of computer based prediction methods have been developed and both the methods have their own advantages and disadvantages. The nonavailability of allele specific binding data restricts the use of knowledge-based prediction methods for a wide range of HLA alleles. Alternatively, the modeling scheme appears to be a promising predictive tool for the selection of peptides that bind to specific HLA molecules. The scoring of the modeled HLA-peptide complexes is a major concern. The use of knowledge based rules (van der Waals clashes and solvent exposed hydrophobic residues) to distinguish binders from nonbinders is applied in the present study. The rules based on (1) number of observed atomic clashes between the modeled peptide and the HLA structure, and (2) number of solvent exposed hydrophobic residues on the modeled peptide effectively discriminate experimentally known binders from poor/nonbinders. Solved crystal complexes show no vdW Clash (vdWC) in 95% cases and no solvent exposed hydrophobic peptide residues (SEHPR) were seen in 86% cases. In our attempt to compare experimental binding data with the predicted scores by this scoring scheme, 77% of the peptides are correctly grouped as good binders with a sensitivity of 71%.

Alleles↗

Molecular evolution of thyroid peroxidase.

Thyroid peroxidase is a member of a family of mammalian peroxidases that includes myeloperoxidase, lactoperoxidase, eosinophil peroxidase, and salivary peroxidase. Protein sequences showing a high degree of sequence similarity with mammalian peroxidases have recently been observed in several invertebrate species. A multiple sequence alignment prepared with five mammalian and six invertebrate peroxidases shows complete conservation of amino acid residues considered to be important in the formation of peroxidase compound 1. These include the distal and proximal histidines, a catalytic arginine residue, and an asparagine residue hydrogen bonded to the proximal histidine. TPO-2, an alternatively spliced form of TPO, lacks the essential asparagine (Asn 579). It is now possible to speak more broadly of the family of animal peroxidases, rather than mammalian peroxidases. The animal peroxidases comprise a group of homologous proteins that differ markedly from the plant/fungal/bacterial peroxidases in primary, secondary and tertiary structure, but which share with them a common function. Animal peroxidases probably arose independently of the plant/fungal/bacterial peroxidase superfamily and most likely belong to a different gene family. The relationship between animal and non-animal peroxidases probably represents an example of convergent evolution to a common enzymatic mechanism.

Amino Acid Sequence↗

Oxygen transport proteins: III. Structural studies of the scorpion (Buthus sindicus) hemocyanin, partial primary structure of its subunit Bsin1.

The hemocyanin (Hc) from Buthus sindicus, studied in the native state, demonstrated to be an aggregate of eight different types of subunits arranged in four cubic hexamers. Both, the 'top' and the 'side' views of the native molecule have been identified from the negatively stained specimens using transmission electron microscopy. Out of these, eight different polypeptide chains, the partial primary structure (68%) of a subunit Bsin1 (Mr = 72422.7 Da) was established using a combination of automated Edman degradation and mass spectrometry. A multiple sequence alignment with other closely related cheliceratan Hc subunits revealed average identities of ca. 60%. Most of the structurally important residues, i.e. copper and calcium-binding ligands, as well as the residues involved in the presumed oxygen entrance pathway, proved to be strictly conserved in Bsin1. Sequence variations have been observed around the functionally important chloride-binding site, not only for the B. sindicus subunit Bsin1, but also for the subunit Aaus-6 of the scorpion A. australis and the subunit Ecal-a from the spider Eurypelma californicum Hcs. Deviation in the primary structure related to the chloride-binding site suggest that the effect of chloride ions may vary in different hemocyanins. Furthermore, the secondary structural contents of the Hc subunit Bsin1 were determined by circular dichroism revealing ca. 33% alpha-helix, 18%, beta-sheet, 19% beta-turn, and 30% random coil composition. These values are in good agreement with the crystal structure of the closely related Hc subunit Lpol-II from horseshoe crab L. polyphemus. Electron microscopic studies of the purified Hc subunit under native conditions revealed that Bsin1 has self aggregation properties. Results of these studies are discussed.

Amino Acid Sequence↗

Structural similarities and evolutionary relationships in chloride-dependent alpha-amylases.

The alpha-amylase sequences contained in databanks were screened for the presence of amino acid residues Arg195, Asn298 and Arg/Lys337 forming the chloride-binding site of several specialized alpha-amylases allosterically activated by this anion. This search provides 38 alpha-amylases potentially binding a chloride ion. All belong to animals, including mammals, birds, insects, acari, nematodes, molluscs, crustaceans and are also found in three extremophilic Gram-negative bacteria. An evolutionary distance tree based on complete amino acid sequences was constructed, revealing four distinct clusters of species. On the basis of multiple sequence alignment and homology modeling, invariable structural elements were defined, corresponding to the active site, the substrate binding site, the accessory binding sites, the Ca(2+) and Cl(-) binding sites, a protease-like catalytic triad and disulfide bonds. The sequence variations within functional elements allowed engineering strategies to be proposed, aimed at identifying and modifying the specificity, activity and stability of chloride-dependent alpha-amylases.

Amino Acid Sequence↗

Structural/functional assignment of unknown bacteriophage T4 proteins by iterative database searches.

Among the total of 274 orfs within bacteriophage T4, only half have been reasonably well characterized, and the functions of the rest have remained obscure. In order to predict the molecular functions of the orfs, a position-specific iterated (PSI)-BLAST search of bacteriophage T4 against the sequence database of known 3D structures was carried out. PSI-BLAST is one of the most powerful iterative sequence search methods using multiple sequence alignment, with the ability to detect many more proteins with distant homology than standard pairwise methods. The 3D structures of proteins are considered to be better preserved than the sequences, and the detected distantly homologous proteins are likely to possess highly similar 3D structures. Thirteen orfs of phage T4, whose homologues were not detected by standard pairwise methods, were found to have significantly homologous counterparts by this method. The plausibility of the results was confirmed by checking whether important residues at substrate/ligand-binding sites were conserved. Among them, two orfs, vs.1 and e.1, which are similar to Escherichia coli lytic enzyme and MutT protein, respectively, had not been studied previously. Also, gp rIIA, a rapid lysis protein, whose gene structure had been intensively studied during the development of molecular biology in the 1950s and yet whose molecular function remains unknown, has an N-terminal domain that is significantly similar to the N-terminal region of the heat shock protein Hsp90.

Amino Acid Sequence↗

Cloning and expression of a nuclear encoded plastid specific 33 kDa ribonucleoprotein gene (33RNP) from pea that is light stimulated.

We report the cloning and sequencing of both cDNA and genomic DNA of a 33 kDa chloroplast ribonucleoprotein (33RNP) from pea. The analysis of the predicted amino acid sequence of the cDNA clone revealed that the encoded protein contains two RNA binding domains, including the conserved consensus ribonucleoprotein sequences CS-RNP1 and CS-RNP2, on the C-terminus half and the presence of a putative transit peptide sequence in the N-terminus region. The phylogenetic and multiple sequence alignment analysis of pea chloroplast RNP along with RNPs reported from the other plant sources revealed that the pea 33RNP is very closely related to Nicotiana sylvestris 31RNP and 28RNP and also to 31RNP and 28RNP of Arabidopsis and spinach, respectively. The pea 33RNP was expressed in Escherichia coli and purified to homogeneity. The in vitro import of precursor protein into chloroplasts confirmed that the N-terminus putative transit peptide is a bona fide transit peptide and 33RNP is localized in the chloroplast. The nucleic acid-binding properties of the recombinant protein, as revealed by South-Western analysis, showed that 33RNP has higher binding affinity for poly (U) and oligo dT than for ssDNA and dsDNA. The steady state transcript level was higher in leaves than in roots and the expression of this gene is light stimulated. Sequence analysis of the genomic clone revealed that the gene contains four exons and three introns. We have also isolated and analyzed the 5' flanking region of the pea 33RNP gene.

Amino Acid Sequence↗

Identification of a novel member of the TGF-beta superfamily highly expressed in human placenta.

While conducting a gene discovery effort targeted to transcripts of the prevalent and intermediate frequency classes in placenta throughout gestation, we identified a novel member of the TGF-beta superfamily that is expressed at high levels in human placenta. Hence, we named this factor 'Placental Transforming Growth Factor Beta' (PTGFB). The full-length sequence of the 1.2-kb PTGFB mRNA has the potential of encoding a putative pre-pro-PTGFB protein of 295 amino acids and a putative mature PTGFB protein of 112 amino acids. Multiple sequence alignments of PTGFB and representative members of all TGF-beta subfamilies evidenced a number of conserved residues, including the seven cysteines that are almost invariant in all members of the TGF-beta superfamily. The single-copy PTGFB gene was shown to be composed of only two exons of 309 bp and 891 bp, separated by a 2.9-kb intron. The gene was localized to chromosome 19p12-13.1 by fluorescence in-situ hybridization. Northern analyses revealed a complex tissue-specific pattern of expression and a second transcript of 1.9 kb that is predominant in adult skeletal muscle. Most importantly, the 1.2-kb PTGFB transcript was shown to be expressed in placenta at much higher levels than in any other human fetal or adult tissue surveyed.

Adult↗