Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Purification, sequencing and structural analysis of two acidic phospholipases A2 from the venom of Bothrops insularis (jararaca ilhoa).

Bothrops snake venoms contain a variety of phospholipases (PLA(2)), some of which are myotoxic. In this work, we used reverse-phase HPLC and mass spectrometry to purify and sequence two PLA(2) from the venom of Bothrops insularis. The two enzymes, designated here as BinTX-I and BinTx-II, were acidic (pI 5.05 and 4.49) Asp49 PLA(2), with molecular masses of 13,975 and 13,788, respectively. The amino acid sequence and molecular mass of BinTX-I were identical to those of a PLA(2) previously isolated from this venom (PA2_BOTIN, SwissProt accession number ) while those of BinTX-II indicated that this was a new enzyme. Multiple sequence alignments with other Bothrops PLA(2) showed that the amino acids His48, Asp49, Tyr52 and Asp99, which are important for enzymatic activity, were fully conserved, as were the 14 cysteine residues involved in disulfide bond formation, in addition to various other residues. A phylogenetic analysis showed that BinTX-I and BinTX-II grouped with other acidic Asp49 PLA(2) from Bothrops venoms, and computer modeling indicated that these enzymes had the characteristic structure of bothropic PLA(2) that consisted of three alpha-helices, a beta-wing, a short helix and a calcium-binding loop. BinTX-I (30 microg/paw) produced mouse hind paw edema that was maximal after 1h compared to after 3h with venom (10 and 100 microg/paw); in both cases, the edema decreased after 6h. BinTX-1 and venom (40 microg/ml each) produced time-dependent neuromuscular blockade in chick biventer cervicis preparations that reached 40% and 95%, respectively, after 120 min. BinTX-I also produced muscle fiber damage and an elevation in CK, as also seen with venom. These results indicate that BinTX-I contributes to the neuromuscular activity and tissue damage caused by B. insularis venom in vitro and in vivo.

Amino Acid Sequence↗

Characterization and cDNA cloning of hinnavin II, a cecropin family antibacterial peptide from the cabbage butterfly, Artogeia rapae.

Hinnavins, together with lysozymes, are the main types of antibacterial peptides/proteins previously isolated from the larval haemolymph of the cabbage butterfly, Artogeia rapae as part of the humoral immune response to a bacterial invasion. One of these antibacterial peptides, named hinnavin II, was purified and characterized after cDNA cloning. The purified hinnavin II was more active against Gram negative than against Gram positive bacteria. Hinnavin II also showed a powerful synergistic effect on the inhibition of bacterial growth with purified lysozyme. The cDNA has a total length of 186 bp with a 114 coding region. The deduced protein sequence contains 38 amino acids with a coding capacity of 4142.8 Da. The result of a multiple sequence alignment and phylogenetic analysis with Clustal W indicated that mature hinnavin II showed an approximately 78.9% amino acid sequence identity with cecropin A and originated from a group containing mostly lepidopteran cecropins.

Amino Acid Sequence↗

Domain boundary prediction based on profile domain linker propensity index.

Successful prediction of protein domain boundaries provides valuable information not only for the computational structure prediction of multi-domain proteins but also for the experimental structure determination. In this work, a novel index at the profile level is presented, namely, the profile domain linker propensity index (PDLI), which uses the evolutionary information of profiles for domain linker prediction. The frequency profiles are directly calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into binary profiles with a probability threshold. PDLI is then obtained by the frequencies of binary profiles in domain linkers as compared to those in domains. A smooth and normalized numeric profile is generated for any amino acid sequences from which the domain linkers can be predicted. Testing on the Structural Classification of Proteins (SCOP) database and CASP6 targets shows that PDLI outperforms other indexes at the amino acid level.

Computational Biology↗

Involvement of some large immunophilins and their ligands in the protection and regeneration of neurons: a hypothetical mode of action.

The powerful immunosuppressive drugs such as FK506 and its derivatives induce some regeneration and protection of neurons from ischaemic brain injury and some other neurological disorders. The drugs form complexes with diverse FKBPs but apparently the FKBP52/FK506 complex was shown to be involved in the protection and regeneration of neurons. We used several different sequence attributes in searching diverse genomic databases for similar motifs as those present in the FKBPs. A Fortran library of algorithms (Par_Seq) has been designed and used in searching for the similarity of sequence motifs extracted from the multiple sequence alignments of diverse groups of proteins (query motifs) and the target motifs which are encoded in various genomes. The following sequence attributes were used in the establishment of the degree of convergence between: (A) amino acid (AA) sequence similarity (ID) of the query/target motifs and (B) their: (1) AA composition (AAC); (2) hydrophobicity (HI); (3) Jensen-Shannon entropy; and (4) AA propensity to form a particular secondary structure. The sequence hallmark of two different groups of peptidylprolyl cis/trans isomerases (PPIases), namely tetratricopetide repeat (TPR) motifs, which are present in the heat-shock cyclophilins and in the large FK506-binding proteins (FKBPs) were used to search various genomic databases. The Par_Seq algorithm has revealed that the TPR motifs have similar sequence attributes as a number of hydrophobic sequence segments of functionally unrelated membrane proteins, including some of the TMs from diverse G protein-coupled receptors (GPCRs). It is proposed that binding of the FKBP52/FK506 complex to the membranes via the TPR motifs and its interaction with some membrane proteins could be in part responsible for some neuro-regeneration and neuro-protection of the brain during some ischaemia-induced stresses.

Algorithms↗

Evolution of tissue-specific keratins as deduced from novel cDNA sequences of the lungfish Protopterus aethiopicus.

Lungfishes are possibly the closest extant relatives of the land vertebrates (tetrapods). We report here the cDNA and predicted amino acid sequences of 13 different keratins (ten type I and three type II) of the lungfish Protopterus aethiopicus. These keratins include the orthologs of human K8 and K18. The lungfish keratins were also identified in tissue extracts using two-dimensional polyacrylamide gel electrophoresis, keratin blot binding assays and immunoblotting. The identified keratin spots were analyzed by peptide mass fingerprinting which assigned seven sequences (inclusively Protopterus K8 and K18) to their respective protein spot. The peptide mass fingerprints also revealed the fact that the major epidermal type I and type II keratins of this lungfish have not yet been sequenced. Nevertheless, phylogenetic trees constructed from multiple sequence alignments of keratins from lungfish and distantly related vertebrates such as lamprey, shark, trout, frog, and human reveal new insights into the evolution of K8 and K18, and unravel a variety of independent keratin radiation events.

Amino Acid Sequence↗

Predicting protein interaction sites from residue spatial sequence profile and evolution rate.

This paper proposes a novel method that can predict protein interaction sites in heterocomplexes using residue spatial sequence profile and evolution rate approaches. The former represents the information of multiple sequence alignments while the latter corresponds to a residue's evolutionary conservation score based on a phylogenetic tree. Three predictors using a support vector machines algorithm are constructed to predict whether a surface residue is a part of a protein-protein interface. The efficiency and the effectiveness of our proposed approach is verified by its better prediction performance compared with other models. The study is based on a non-redundant data set of heterodimers consisting of 69 protein chains.

Algorithms↗

The first potassium channel toxin from the venom of the Iranian scorpion Odonthobuthus doriae.

The very first member of K(+) channels toxins from the venom of the Iranian scorpion Odonthobuthus doriae (OdK1) was purified, sequenced and characterized physiologically. OdK1 has 29 amino acids, six conserved cysteines and a pI value of 4.95. Based on multiple sequence alignments, OdK1 was classified as alpha-KTx 8.5. The pharmacological effects of OdK1 were studied on six different cloned K(+) channels (vertebrate Kv1.1-Kv1.5 and Shaker IR) expressed in Xenopus laevis oocytes. Interestingly, OdK1 selectively inhibited the currents through Kv1.2 channels with an IC50 value of 183+/-3 nM but did not affect any of the other channels.

Animals↗

Divergent evolutionary lines of fungal cytochrome c peroxidases belonging to the superfamily of bacterial, fungal and plant heme peroxidases.

Novel open reading frames coding for cytochrome c peroxidase (CcP) belonging to the superfamily of bacterial, fungal, and plant heme peroxidases were analyzed in the available fungal genomes. Multiple sequence alignment of 71 selected peroxidase genes revealed the presence of three conserved regions essential for their function: one on the distal and two on the proximal side of the prosthetic heme group. Conserved sequence motifs on the proximal heme side are peculiar for CcPs and are responsible for their reactivity. Phylogenetic analysis performed with the distance method as well as with the maximum likelihood method revealed the existence of three distinct subfamilies of fungal CcP and their relationship to other members of the peroxidase superfamily. These divergent CcP evolutionary lines apparently evolved from a single primordial heme peroxidase gene in parallel with the evolution of ascorbate peroxidase genes. Analyzed CcPs differ significantly in their N-terminal sequences. Only subfamily I did not exhibit a presence of any signal sequence. Subfamily II members possess a well defined signal sequence allowing processing and release into mitochondrion and also in subfamily III a signal sequence was detected. Several here analyzed peroxidase genes mainly from Candida albicans and from Rhizopus oryzae can be considered interesting for the investigation of the structure-function relationship of novel CcPs revealing differences to the well documented properties of cytochrome c peroxidase from Saccharomyces cerevisiae.

Amino Acid Motifs↗

Glutathione transferase-like proteins encoded in genomes of yeasts and fungi: insights into evolution of a multifunctional protein superfamily.

Most fungal glutathione transferases (GSTs) do not fit easily into any of the previously characterised classes by immunological, sequence or catalytic criteria. In contrast to the paucity of studies on GSTs cloned or isolated from fungal sources, a screen of databases revealed 67 GST-like sequences from 21 fungal species. Comparison by multiple sequence alignment generated a dendrogram revealing five clusters of GST-like proteins designated clusters 1, 2, EFIBgamma, Ure2p and MAK16, the last three of which have previously been related to the GST superfamily. Surprisingly, a relatively small number of fungal GSTs belong to mainstream classes and the previously-described fungal Gamma class is not widespread in the 21 species studied. Representative crystal structures are available for the EFIBgamma and Ure2p classes and the domain structures of representative sequences are compared with these. In addition, there are some "orphan" sequences that do not fit into any previously-described class, but show similarity to genes implicated in fungal biosynthetic gene clusters. We suggest that GST-like sequences are widespread in fungi, participating in a wide range of functions. They probably evolved by a process similar to domain "shuffling".

Amino Acid Sequence↗

Characterization of the porcine alpha interferon multigene family.

The availability of data on the pig genome sequence prompted us to characterize the porcine IFN-alpha (PoIFN-alpha) multigene family. Fourteen functional PoIFN-alpha genes and two PoIFN-alpha pseudogenes were detected in the porcine genome. Multiple sequence alignment revealed a C-terminal deletion of eight residues in six subtypes. A phylogenetic tree of the porcine IFN-alpha gene family defined the evolutionary relationship of the various subtypes. In addition, analysis of the evolutionary rate and the effect of positive selection suggested that the C-terminal deletion is a strategy for preservation in the genome. Eight PoIFN-alpha subtypes were isolated from the porcine liver genome and expressed in BHK-21 cells line. We detected the level of transcription by real-time quantitative RT-PCR analysis. The antiviral activities of the products were determined by WISH cells/Vesicular Stomatitis Virus (VSV) and PK 15 cells/Pseudorabies Virus (PRV) respectively. We found the antiviral activities of intact PoIFN-alpha genes are approximately 2-50 times higher than those of the subtypes with C-terminal deletions in WISH cells and 15-55 times higher in PK 15 cells. There was no obvious difference between the subtypes with and without C-terminal deletion on acid susceptibility.

Amino Acid Sequence↗

cDNA cloning, gene organization and variant specific expression of HIF-1 alpha in high altitude yak (Bos grunniens).

Hypoxia-inducible factor 1 (HIF-1) is a heterodimeric basic-helix-loop-helix-PER-ARNT-SIM (bHLH-PAS) transcription factor consisting of HIF-1alpha and HIF-1beta subunits. HIF-1alpha is the oxygen-regulated subunit of HIF-1, which regulates the transcription of genes involved in oxygen homeostasis in response to hypoxia. Yak (Bos grunniens), a mammal native to high altitude (HA) region ( approximately 3500-5500 m), has successfully adapted over many generations to the chronic hypoxia of HA. In the present work, cDNA encoding HIF-1alpha has been cloned from the blood of yak. Tissue specific expression of the mRNA was analyzed in blood, heart, lung, liver and kidney by RT-PCR with primers from three different regions of cDNA. The HIF-1alpha expression was liver and blood specific. The HIF-1alpha mRNA contains 823 bp long 3'UTR that is AU-rich and contains ten AUUUA pentamers and two overlapping copies of the nonamer UUAUUUAUUUAUU. Three potential microRNAs, hsa-miR-107/mmu-miR-107/rno-miR-107, hsa-miR-18b and hsa-miR-135a/mmu-miR-135a/rno-miR-135a, targeting 3'UTR of yak HIF-1alpha, were identified by using target prediction software. The CDS encodes for 823 residues of amino acids and showed 99%, 95%, 92%, 90% and 90% similarity to domestic cattle, human, plateau pika, mouse and rat HIF-1alpha, respectively. HIF-1alpha cDNA, cloned and sequenced in the present work has revealed the evolutionary conservation through multiple sequence alignment. Liver and blood specific stability of HIF-1alpha mRNA appears miR-107 regulated.

Altitude↗

Molecular identification of a bevy of serine proteinases in Manduca sexta hemolymph.

Extracellular serine proteinase pathways control immune and homeostatic processes in insects. Our current knowledge of their components is limited-prophenoloxidase-activating proteinases (PAPs) are among the few hemolymph proteinases (HPs) with known functions. To identify components of proteinase systems in the hemolymph of Manduca sexta, we amplified cDNAs from larval fat body or hemocytes using degenerate primers coding for two conserved regions in S1 family serine proteinases. PCR yielded fragments encoding seven known (HP1-HP4, PAP-1, PAP-2 and PAP-3) and 18 unknown (HP5-HP22) serine proteinases. We screened cDNA libraries and isolated clones for 17 of the newly discovered HPs (HP5-HP22 except for HP11) and prepared antibodies to 14 recombinant proteins (HP6, HP8-HP10, HP12, HP14-HP19, HP21 and HP22). Fourteen of the HPs contain regulatory clip domain(s) at their amino-terminus--HP1, HP2, HP6, HP8, HP13, HP17, HP18, HP21, HP22 and PAP-1 have one, whereas HP12, HP15, PAP-2 and PAP-3 have two clip domains. Multiple sequence alignment of catalytic domains in these and other arthropod serine proteinases provided useful clues for future functional analysis. Northern blot and reverse transcription PCR (RT-PCR) analyses showed increases in HP2, HP7, HP9, HP10, HP12-HP22 mRNA levels at 24h after a bacterial challenge, and immunoblot analysis confirmed elevated concentrations of HP12, HP14-HP19, HP21 and HP22 proteins in plasma in response to injected bacteria. Hemocytes express HP13 and HP18; fat body produces HP12, HP20-HP22; both tissues synthesize the other HPs. These results collectively indicate the existence of a complex serine proteinase network in M. sexta hemolymph, predicted to mediate rapid defense responses upon wounding and/or microbial infection.

Amino Acid Sequence↗

Identifying protein-protein interfacial residues in heterocomplexes using residue conservation scores.

Identifying protein-protein interfaces is crucial for structural biology. Because of the constraints in wet experiments, many computational methods have been proposed. Without knowing any information about the partner chains, a new method of predicting protein-protein interaction interface residues purely based on evolutionary information in heterocomplexes is proposed here. Unlike traditional approaches using multiple sequence alignment profiles to represent the conservation level for each residue, we make predictions based on the concept of residue conservation scores so that the dimension of the feature vector for each residue can be drastically reduced, at least 20 times less than conventional methods. Based on the representation approach, a simple linear discriminant function is used to make predictions, so the computational complexity of the whole prediction procedure can also be greatly decreased. By testing our approach on 69 heterocomplex chains, experimental results demonstrate the performance of our approach is indeed superior to current existing methods.

Computational Biology↗

Proteolytic activity in Encephalitozoon cuniculi sporogonial stages: predominance of metallopeptidases including an aminopeptidase-P-like enzyme.

A fraction enriched in spore precursor cells (sporoblasts) of the microsporidian Encephalitozoon cuniculi, an intracellular parasite of mammals, was obtained by Percoll gradient centrifugation. Soluble extracts of these cells exhibited proteolytic activity towards azocasein, with an alkaline optimum pH range (9-10). Prevalence of some metallopeptidases was supported by the stimulating effect of Ca2+, Mg2+, Mn2+ and Zn2+ ions, and inhibition by two chelating agents (EDTA and 1,10-phenanthroline), a thiol reductant (dithiothreitol) and two aminopeptidase inhibitors (bestatin and apstatin). Zymographic analysis revealed four caseinolytic bands at about 76, 70, 55 and 50 kDa. Mass spectrometry of tryptic peptides from one-dimensional gel slices identified a cytosol (leucine) aminopeptidase homologue (M17 family) in 50-kDa band and an enzyme similar to aminopeptidase P (AP-P) of cytosolic type (M24B subfamily) in 70-kDa band. Multiple sequence alignments showed conservation of critical residues for catalysis and metal binding. A long insertion in a common position was found in AP-P sequences from E. cuniculi and Nosema locustae, an insect-infecting microsporidian. The expression of cytosolic AP-P in sporogonial stages of microsporidia may suggest a key role in the attack of proline-containing peptides as a prerequisite to long-duration biosynthesis of structural proteins destined to the sporal polar tube.

Amino Acid Sequence↗

Are residues in a protein folding nucleus evolutionarily conserved?

Protein is the working molecule of the cell, and evolution is the hallmark of life. It is important to understand how protein folding and evolution influence each other. Several studies correlating experimental measurement of residue participation in folding nucleus and sequence conservation have reached different conclusions. These studies are based on assessment of sequence conservation at folding nucleus sites using entropy or relative entropy measurement derived from multiple sequence alignment. Here we report analysis of conservation of folding nucleus using an evolutionary model alternative to entropy-based approaches. We employ a continuous time Markov model of codon substitution to distinguish mutation fixed by evolution and mutation fixed by chance. This model takes into account bias in codon frequency, bias-favoring transition over transversion, as well as explicit phylogenetic information. We measure selection pressure using the ratio omega of synonymous versus non-synonymous substitution at individual residue site. The omega-values are estimated using the PAML method, a maximum-likelihood estimator. Our results show that there is little correlation between the extent of kinetic participation in protein folding nucleus as measured by experimental phi-value and selection pressure as measured by omega-value. In addition, two randomization tests failed to show that folding nucleus residues are significantly more conserved than the whole protein, or the median omega value of all residues in the protein. These results suggest that at the level of codon substitution, there is no indication that folding nucleus residues are significantly more conserved than other residues. We further reconstruct candidate ancestral residues of the folding nucleus and suggest possible test tube mutation studies for testing folding behavior of ancient folding nucleus.

Amino Acid Sequence↗

A family of evolution-entropy hybrid methods for ranking protein residues by importance.

In order to identify the amino acids that determine protein structure and function it is useful to rank them by their relative importance. Previous approaches belong to two groups; those that rely on statistical inference, and those that focus on phylogenetic analysis. Here, we introduce a class of hybrid methods that combine evolutionary and entropic information from multiple sequence alignments. A detailed analysis in insulin receptor kinase domain and tests on proteins that are well-characterized experimentally show the hybrids' greater robustness with respect to the input choice of sequences, as well as improved sensitivity and specificity of prediction. This is a further step toward proteome scale analysis of protein structure and function.

Amino Acid Sequence↗

phi29 DNA polymerase-terminal protein interaction. Involvement of residues specifically conserved among protein-primed DNA polymerases.

By multiple sequence alignments of DNA polymerases from the eukaryotic-type (family B) subgroup of protein-primed DNA polymerases we have identified five positively charged amino acids, specifically conserved, located N-terminally to the (S/T)Lx(2)h motif. Here, we have studied, by site-directed mutagenesis, the functional role of phi29 DNA polymerase residues Arg96, Lys110, Lys112, Arg113 and Lys114 in specific reactions dependent on a protein-priming event. Mutations introduced at residues Arg96, Arg113 and Lys114 and to a lower extent Lys110 and Lys112, showed a defective protein-primed initiation step. Analysis of the interaction with double-stranded DNA and terminal protein (TP) displayed by mutant derivatives R96A, K110A, K112A, R113A and K114A allows us to conclude that phi29 DNA polymerase residue Arg96 is an important DNA/TP-ligand residue, essential to form stable DNA polymerase/DNA(TP) complexes, while residues Lys110, Lys112 and Arg113 could be playing a role in establishing contacts with the TP-DNA template during the first step of DNA replication. The importance of residue Lys114 to make a functionally active DNA polymerase/TP complex is also discussed. These results, together with the high degree of conservation of those residues among protein-primed DNA polymerases, strongly suggest a functional role of those amino acids in establishing the appropriate interactions with DNA polymerase substrates, DNA and TP, to successfully accomplish the first steps of TP-DNA replication.

Amino Acid Sequence↗

A family-based approach reveals the function of residues in the nuclear receptor ligand-binding domain.

Literature studies, 3D structure data, and a series of sequence analysis techniques were combined to reveal important residues in the structure and function of the ligand-binding domain of nuclear hormone receptors. A structure-based multiple sequence alignment allowed for the seamless combination of data from many different studies on different receptors into one single functional model. It was recently shown that a combined analysis of sequence entropy and variability can divide residues in five classes; (1) the main function or active site, (2) support for the main function, (3) signal transduction, (4) modulator or ligand binding and (5) the rest. Mutation data extracted from the literature and intermolecular contacts observed in nuclear receptor structures were analyzed in view of this classification and showed that the main function or active site residues of the nuclear receptor ligand-binding domain are involved in cofactor recruitment. Furthermore, the sequence entropy-variability analysis identified the presence of signal transduction residues that are located between the ligand, cofactor and dimer sites, suggesting communication between these regulatory binding sites. Experimental and computational results agreed well for most residues for which mutation data and intermolecular contact data were available. This allows us to predict the role of the residues for which no functional data is available yet. This study illustrates the power of family-based approaches towards the analysis of protein function, and it points out the problems and possibilities presented by the massive amounts of data that are becoming available in the "omics era". The results shed light on the nuclear receptor family that is involved in processes ranging from cancer to infertility, and that is one of the more important targets in the pharmaceutical industry.

Amino Acids↗