Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Molecular cloning of giant panda pituitary prolactin cDNA and its expression in Escherichia coli.

cDNA encoding pituitary (PRL) of giant panda was obtained using RT-PCR and expressed in E. coli. The results revealed that panda PRL cDNA encodes a precursor protein of 229 amino acids including a putative signal peptide of 30 amino acids and a mature protein of 199 residues with one potential N-glycosylation site. Sequence comparison indicated that panda PRL shares a high degree of identity to other known PRL sequences ranging from 98% with mink PRL to about 50% with rodent PRL. Six cysteine residues and 29 conserved residues distributed in four domains (PD1, PD2, PD3, and PD4) of PRL were observed. through multiple sequence alignment. Fourteen key residues of binding sites 1 and 2 involved in receptor binding are conserved in panda PRL. GST fused recombinant panda PRL protein was efficiently expressed with the form of insoluble inclusion bodies in E. coli BL21 transformed with a pGEX-4T-1 expression vector containing the DNA sequence encoding mature panda PRL. Western blot analysis indicated that GST-panda PRL recombinant protein could be recognized by antibody against human PRL. Our results would contribute to further elucidating the structural and functional characteristics of pituitary PRL and provide a basis for the production of recombinant panda prolactin for future use in the breeding of giant panda.

Amino Acid Sequence↗

A novel method for GPCR recognition and family classification from sequence alone using signatures derived from profile hidden Markov models.

G-protein coupled receptors (GPCRs) constitute a broad class of cell-surface receptors, including several functionally distinct families, that play a key role in cellular signalling and regulation of basic physiological processes. GPCRs are the focus of a significant amount of current pharmaceutical research since they interact with more than 50% of prescription drugs, whereas they still comprise the best potential targets for drug design. Taking into account the excess of data derived by genome sequencing projects, the use of computational tools for automated characterization of novel GPCRs is imperative. Typical computational strategies for identifying and classifying GPCRs involve sequence similarity searches (e.g. BLAST) coupled with pattern database analysis (e.g. PROSITE, BLOCKS). The diagnostic method presented here is based on a probabilistic approach that exploits highly discriminative profile Hidden Markov Models, excised from low entropy regions of multiple sequence alignments, to derive potent family signatures. For a given query, a P-value is obtained, combining individual hits derived from the same family. Hence a best-guess family membership is depicted, allowing GPCRs' classification at a family level, solely using primary structure information. A web-based version of the application is freely available at URL: http:/bioinformatics.biol.uoa.gr/PRED-GPCR.

Databases, Factual↗

Relationship between the DNA binding domains of SMAD and NFI/CTF transcription factors defines a new superfamily of genes.

Transcription factors of the SMAD family relay signals from cell surface receptors to the nucleus in response to TGF-beta related soluble factors. Members of the nuclear factor I/CAAT box binding family (NFI/CTF) have been implicated as regulators of diverse biological processes such as adenovirus replication and transcription of TGF-responsive genes. There are highly conserved DNA binding domains in SMAD and NFI/CTF transcription factors that allow sequence specific DNA binding for members of each family. However, no homology relationship has been established for the DNA binding domains present in these families. For a better understanding of the structure and evolution of SMAD genes, we carried out a sensitive PSI-BLAST database search. This revealed significant similarities between the DNA binding domains of SMADs and NFI/CTF transcription factors. Enhanced graphic matrix analysis and multiple sequence alignment of the amino acid sequences of the SMAD and NFI/CTF DNA binding domains also show that these two classes of domains share considerable structural similarity. These results strongly suggest that these two classes of factors share a homologous DNA binding domain presumably resulting from a common ancestry. In contrast, the C-terminal transcription modulation domains of both SMAD and NFI/CTF families do not show any sequence similarity. Based on the structural relationship of their DNA binding domains, we propose that the SMAD and NFI/CTF transcription factors belong to new superfamily of genes.

Amino Acid Sequence↗

A new family of plasma membrane polypeptides differentially regulated during plant development.

Two cDNAs encoding polypeptides identified in a tobacco leaf plasma membrane fraction prepared by phase partitioning were cloned. The deduced polypeptides, P16 and P17, exhibit a striking primary structure, similar to that of P19, a previously cloned plasma membrane polypeptide. Antibodies raised to the recombinant proteins were used to probe the cellular location of P16, P17 and P19 by means of western blotting of sucrose density-gradient fractions; all three polypeptides were found to be located solely at the plasma membrane. Furthermore, P19 antigen accumulated transiently at the time of floral induction while P16 and P17 antigens accumulated towards the end of the life-cycle. These results together with sequence database searches and multiple-sequence alignments suggest that we have identified a new family of plasma membrane polypeptides that (i) are putatively plant specific and (ii) are differentially regulated during plant development. These polypeptides are termed DREPPs for developmentally regulated plasma membrane polypeptides.

Amino Acid Sequence↗

Character and evolution of protein-protein interfaces.

Protein-protein interactions create the macromolecular assemblies and sequential signaling pathways essential for cell function. Their number far exceeds the number of proteins themselves and their experimental characterization, while improving, remains relatively slow. For these reasons, novel computational methods have important roles to play in understanding the physical basis of protein interactions, and in constraining the molecular basis of their specificity. This paper discusses methods based on multiple sequence alignments of protein homologues and phylogenetic trees.

Artificial Intelligence↗

HIV type 1 V3 domain serotyping and genotyping in Gauteng, Mpumalanga, KwaZulu-Natal, and Western Cape Provinces of South Africa.

More than 20.8 million people are living with HIV/AIDS in sub-Saharan Africa, with southern Africa the worst affected area and accounting for one of the fastest growing AIDS epidemics worldwide. Samples from 81 patients, including 25 from KwaZulu-Natal, 26 from Gauteng, 5 from Mpumalanga, and 25 from Western Cape Province, were serotyped using a competitive V3 peptide enzyme immunoassay (cPEIA). Viral RNA was also isolated from serum and the V3 region amplified by reverse transcriptase polymerase chain reaction (RT-PCR) to obtain a 240-bp product for direct sequencing of 29 samples. CLUSTAL W was used to make multiple sequence alignments. Distance calculation, tree construction methods, and bootstrap analysis were done using TREECON. Subtype C-like V3 loop sequences predominate in all provinces tested in South Africa. Discordant sero- and genotype results were observed in one patient only. The correlation between sero- and genotyping was 96% (24 of 25) in KwaZulu-Natal and 100% in Gauteng and Mpumalanga. In Western Cape Province 18% of patients were identified as sero/genotype B and 82% as sero/genotype C. Our data show that results of the second-generation V3 cPEIA correlated well with V3 sequencing and would be a rapid and affordable screening test to monitor the explosive southern African HIV-1 epidemic.

Adolescent↗

Representing and reasoning about protein families using generative and discriminative methods.

This work addresses the issues of data representation and incorporation of domain knowledge into the design of learning systems for reasoning about protein families. Given the limited expressive capacity of a particular method, a mixture of protein annotation and fold recognition experts, each implementing a different underlying representation, should provide a robust method for assigning sequences to families. These ideas are illustrated using two data-driven learning methods that make use of different prior information and employ independent, yet complementary, projections of a family: hidden Markov models (HMMs) based on a multiple sequence alignment and neural networks (NNs) based on global sequence descriptors of proteins. Examination of seven protein families indicates that combining a generative (HMM) and a discriminative (NN) method is better than either method on its own. Biologically, human 4-hydroxyphenylpyruvic acid dioxygenase, involved in tyrosinemia type 3, is predicted to be structurally and functionally related to the glyoxalase I family.

Amino Acid Sequence↗

Systematic and fully automated identification of protein sequence patterns.

We present an efficient algorithm to systematically and automatically identify patterns in protein sequence families. The procedure is based on the Splash deterministic pattern discovery algorithm and on a framework to assess the statistical significance of patterns. We demonstrate its application to the fully automated discovery of patterns in 974 PROSITE families (the complete subset of PROSITE families which are defined by patterns and contain DR records). Splash generates patterns with better specificity and undiminished sensitivity, or vice versa, in 28% of the families; identical statistics were obtained in 48% of the families, worse statistics in 15%, and mixed behavior in the remaining 9%. In about 75% of the cases, Splash patterns identify sequence sites that overlap more than 50% with the corresponding PROSITE pattern. The procedure is sufficiently rapid to enable its use for daily curation of existing motif and profile databases. Third, our results show that the statistical significance of discovered patterns correlates well with their biological significance. The trypsin subfamily of serine proteases is used to illustrate this method's ability to exhaustively discover all motifs in a family that are statistically and biologically significant. Finally, we discuss applications of sequence patterns to multiple sequence alignment and the training of more sensitive score-based motif models, akin to the procedure used by PSI-BLAST. All results are available at httpl//www.research.ibm.com/spat/.

Algorithms↗

CASTOR: clustering algorithm for sequence taxonomical organization and relationships.

Given a set of related proteins, two important problems in biology are the inference of protein subsets such that members of one subset share a common function and the identification of protein regions that possess functional significance. The former is typically approached by hierarchical bottom-up clustering based on pairwise sequence similarity and various linkage rules. The latter is typically approached in a supervised manner, based on global multiple sequence alignment. However, the two problems are inextricably linked, since functional subsets are usually characterized by distinctive functional regions. This paper introduces CASTOR, an automatic and unsupervised system that addresses both problems simultaneously and efficiently. It identifies protein regions that are likely to have functional significance by discovering and refining statistically significant motifs. It infers likely functional protein subsets and their relationships based on the presence of the discovered motifs in a top-down and recursive manner, allowing the identification of both hierarchical and nonhierarchical subset relationships. This is, to our knowledge, the first system that approaches both problems simultaneously in a top-down, systematic manner. CASTOR's performance is evaluated against the G-protein coupled receptor superfamily. The identified protein regions lead to a taxonomical organization of this superfamily that is in remarkable agreement with a biologically motivated one and which outperforms those produced by bottom-up clustering methods. We also find that conventional hierarchical representations may fail to accurately describe the complexity of evolutionary development responsible for the final organization of a complex protein family. In particular, many functional relationships governing distant subfamilies of such a protein family may not be represented hierarchically.

Algorithms↗

Cloning of a cDNA encoding bovine interleukin-18 and analysis of IL-18 expression in macrophages and its IFN-gamma-inducing activity.

Interleukin-18 (IL-18) is a recently described cytokine that enhances interferon-gamma (IFN-gamma) production, either independently or synergistically with IL-12. These properties identify IL-18 as an immunoregulatory cytokine that may be pivotal in host defense against intracellular pathogens. We have isolated and sequenced a cDNA encoding bovine IL-18. The open reading frame (ORF) is 582 bp in length, encoding a predicted 192 amino acid (aa) precursor protein. Multiple sequence alignment demonstrated that bovine IL-18 has 65% and 78% identity with the predicted amino acid sequences of murine and human IL-18, respectively. IL-18 mRNA was constitutively present in bovine peripheral blood monocyte-derived macrophages (MDM), with no upregulation on stimulation with lipopolysaccharide (LPS). IL-18 transcripts were weakly detected in B lymphocytes but inducible in the B cell line BL-3. Human recombinant IL-18 (rHuIL-18) induced IFN-gamma production by PHA-stimulated peripheral blood mononuclear cells (PBMC), which was potentiated by rHuIL-12. Further, rHuIL-12 and rHuIL-18 enhanced proliferation of untreated PBMC. Antigen-specific T cell lines demonstrated IL-18-dependent enhancement of IFN-gamma production, indicating that bovine T cells are one of the leukocyte subsets that respond to IL-18. Analysis of IL-18 expression and its ability to induce IFN-gamma production by bovine lymphocytes are important considerations for understanding mechanisms of protective immunity and designing vaccines for intracellular pathogens.

Amino Acid Sequence↗

Statistical modeling, phylogenetic analysis and structure prediction of a protein splicing domain common to inteins and hedgehog proteins.

Inteins, introns spliced at the protein level, and the hedgehog family of proteins involved in eucaryotic development both undergo autocatalytic proteolysis. Here, a specific and sensitive hidden Markov model (HMM) of protein splicing domain shared by inteins and the hedgehog proteins has been trained and employed for further analysis. The HMM characterizes the common features of this domain including the position where a site-specific DNA endonuclease domain is inserted in the majority of the inteins. The HMM was used to identify several new putative inteins, such as that in the Methanococcus jannaschii klbA protein, and to generate a multiple sequence alignment of sequences possessing this domain. Phylogenetic analysis suggests that hedgehog proteins evolved from inteins. Secondary and tertiary structure predictions suggest that the domain has a structure similar to a beta-sandwich. Similarities between the serine protease cleavage mechanism and the protein splicing reaction mechanism are discussed. Examination of the locations of inteins indicates that they are not inserted randomly in an extein, but are often inserted at functionally important positions in the host proteins. A specific and sensitive HMM for a domain present in klbA proteins identified several additional bacterial and archaeal family members, and analysis of the site of insertion of the intein suggests residues that may be functionally important. This domain may play a role in formation of surface-associated protein complexes.

Algorithms↗

PHD--an automatic mail server for protein secondary structure prediction.

By the middle of 1993, > 30,000 protein sequences has been listed. For 1000 of these, the three-dimensional (tertiary) structure has been experimentally solved. Another 7000 can be modelled by homology. For the remaining 21,000 sequences, secondary structure prediction provides a rough estimate of structural features. Predictions in three states range between 35% (random) and 88% (homology modelling) overall accuracy. Using information about evolutionary conservation as contained in multiple sequence alignments, the secondary structure of 4700 protein sequences was predicted by the automatic e-mail server PHD. For proteins with at least one known homologue, the method has an expected overall three-state accuracy of 71.4% for proteins with at least one known homologue (evaluated on 126 unique protein chains).

Algorithms↗

SEQSEE: a comprehensive program suite for protein sequence analysis.

SEQSEE (SEQuence SEEker) is a multi-purpose, menu-driven suite of programs designed to provide a fully integrated, state-of-the-art package for the analysis and display of protein sequences and protein databases. It is currently configured to run on most UNIX-based machines including Sun, SGI and NeXT workstations with conversion to other architectures (e.g. Vax or Cray) being a relatively simple task. SEQSEE is capable of performing nearly all of the analytical and comparative tasks found in most comprehensive commercially available software packages. These include sequence/database searching, sequence retrieval, sequence entry and editing, statistical sequence analysis, multiple sequence alignment, flexible pattern matching, and secondary structure prediction. SEQSEE also integrates a number of unique databases which allow it to perform many additional functions such as structure-based sequence alignments and homology-based secondary structure prediction. Additional enhancements to many previously published algorithms have substantially improved the performance of SEQSEE over that found for most other commercial products. The source code, the documentation and all of the required databases for SEQSEE are freely available and may be obtained by anonymous ftp.

Algorithms↗

Comparison of side chain interactions performed by structurally equivalent residues in homologous protein structures.

The present work describes the computer program Hom-Bond, which allows to identify and compare intra-molecular interactions performed by side chain polar atoms as observed in a family of homologous protein structures with known and conserved 3-D conformation. For this purpose, the side chain to side chain and the side chain to main chain hydrogen bonds, the disulfide and the salt bridges are identified in each considered protein structure. Subsequently, the side chain interactions are displayed according to the multiple sequence alignment. The presented approach allows to easily identify bonds which are conserved in homologous proteins and to analyse rearrangements of the network of side chain interactions that characterize each protein structure.

Amino Acid Sequence↗

Sisyphus and prediction of protein structure.

The problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and 1D prediction have reopened the field. Mean-force-potentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue contacts (2D) can be detected by analysis of correlated mutations, albeit with low accuracy. Secondary structure, solvent accessibility and transmembrane helices (1D) can be predicted with significantly improved accuracy using multiple sequence alignments. Some of these new prediction methods have proven accurate and reliable enough to be useful in genome analysis, and in experimental structure determination. Moreover, the new generation of theoretical methods is increasingly influencing experiments in molecular biology.

Computers↗

Efficient discovery of conserved patterns using a pattern graph.

MOTIVATION: We have previously reported an algorithm for discovering patterns conserved in sets of related unaligned protein sequences. The algorithm was implemented in a program called Pratt. Pratt allows the user to define a class of patterns (e.g. the degree of ambiguity allowed and the length and number of gaps), and is then guaranteed to find the conserved patterns in this class scoring highest according to a defined fitness measure. In many cases, this version of Pratt was very efficient, but in other cases it was too time consuming to be applied. Hence, a more efficient algorithm was needed. RESULTS: In this paper, we describe a new and improved searching strategy that has two main advantages over the old strategy. First, it allows for easier integration with programs for multiple sequence alignment and data base search. Secondly, it makes it possible to use branch-and-bound search, and heuristics, to speed up the search. The new search strategy has been implemented in a new version of the Pratt program.

Algorithms↗

VHMPT: a graphical viewer and editor for helical membrane protein topologies.

MOTIVATION: Lacking structures resolved at atomic resolution, the great majority of membrane proteins have typically been depicted in a schematic two-dimensional (2D) topology consisting of putative transmembrane domains predicted from hydropathy plots. As more and more sequences of membrane proteins become available from genome projects, there is a need to automate the process of generating the schematic topology while allowing important information, such as the individual amino acid and the extent to which it is conserved in evolution, to be conveniently inspected. We addressed this need by developing a program called VHMPT. RESULTS: VHMPT (a graphical V iewer and editor for H elical line M embrane P rotein T opologies) can automatically generate a schematic 2D topology for a protein with transmembrane helices. Through an interactive graphical interface, VHMPT allows users to modify the layout of the generated topology, label specific amino acid or amino acid groups, and annotate with arrows and texts. Given a multiple sequence alignment file, VHMPT can also color code a normalized conservation score for each amino acid on the generated topology, allowing ready visual recognition of highly conserved (or variable) topological regions. VHMPT is written in Tcl/Tk and can run on platforms that have installed the Tcl/Tk interpreter. AVAILABILITY: The source code and a user manual for VHMPT are available for download at http://www. ibms.sinica.edu.tw/mjhwang/vhmpt. CONTACT: mjhwang@mail.ibms.sinica.edu.tw

Computational Biology↗

TOPAL: recombination detection in DNA and protein sequences.

UNLABELLED: TOPAL scans a multiple sequence alignment for evidence of recombinant sequences, prior to phylogenetic analysis. AVAILABILITY: The TOPAL package may be accessed at http://www.bioss.sari.ac.uk/grainne, and by anonymous ftp at ftp.bioss. sari.ac.uk in the directory pub/phylogeny/topal. CONTACT: grainne@bioss.sari.ac.uk

Computational Biology↗