Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Fowlpox virus encodes a protein related to human deoxycytidine kinase: further evidence for independent acquisition of genes for enzymes of nucleotide metabolism by different viruses.

It is demonstrated that fowlpox virus (FPV) protein FP26 located in the HindIII D fragment of the genome is related to the human deoxycytidine kinase (dCK) and probably possesses the same enzymatic activity. A homologous protein is not encoded by vaccinia virus. A multiple alignment of the amino acid sequences of the human and FPV dCKs, the thymidine kinases (TK) of herpesviruses, and cellular and vaccinia virus thymidylate kinases (ThyK) was generated and the conserved motifs, at least two of which are implicated in ATP binding, were characterized. An apparent duplication of ATP-binding motif B in the dCKs was revealed, leading to the reassignment of one of the catalytic residues. Phylogenetic analysis based on the multiple alignment suggested that the putative dCK of FPV probably has diverged from the common ancestor with the human dCK at a later stage of evolution than the herpesvirus TKs, with the ThyKs being peripheral members of the family. These results are compatible with hypothesis that genes for enzymes of nucleotide metabolism could be acquired independently by different DNA viruses (Koonin, E.V. and Senkevich, T.G., Virus Genes 6:187-196, 1992).

Amino Acid Sequence↗

Comparative modeling of CASP4 target proteins: combining results of sequence search with three-dimensional structure assessment.

Comparative modeling aims at constructing molecular models for proteins of unknown structure, by using known structures of related proteins as templates. To test the comparative modeling approach reported here, predictions for 13 target proteins were submitted during the fourth round of "blind" protein structure prediction experiment (CASP4; http://PredictionCenter.llnl.gov/casp4). Sequence identity between these target proteins and the closest known structures ranged from 13 to 58%, indicating a broad spectrum of prediction difficulty. Although this broad difficulty range required addressing a variety of issues, the most important proved to be sequence-structure alignment for distant homology targets. The alignment step was based on structure-based evaluation of alignment variants produced mainly with PSI-BLAST intermediate sequence search procedure (PSI-BLAST-ISS). Although a fraction of correctly aligned residues in resulting models was markedly better than the average in all cases, for distant homology targets it was still considerably below the estimated achievable level. Results with CASP4 targets show that, along with the correctness of sequence-structure alignments, effective use of multiple template structures may significantly increase accuracy of the model structure. Improvement in this area should also result in more accurate loop modeling and side-chain prediction.

Bacterial Proteins↗

A cSNP map and database for human chromosome 21.

Single nucleotide polymorphisms (SNPs) are likely to contribute to the study of complex genetic diseases. The genomic sequence of human chromosome 21q was recently completed with 225 annotated genes, thus permitting efficient identification and precise mapping of potential cSNPs by bioinformatics approaches. Here we present a human chromosome 21 (HC21) cSNP database and the first chromosome-specific cSNP map. Potential cSNPs were generated using three approaches: (1) Alignment of the complete HC21 genomic sequence to cognate ESTs and mRNAs. Candidate cSNPs were automatically extracted using a novel program for context-dependent SNP identification that efficiently discriminates between true variation, poor quality sequencing, and paralogous gene alignments. (2) Multiple alignment of all known HC21 genes to all other human database entries. (3) Gene-targeted cSNP discovery. To date we have identified 377 cSNPs averaging ~1 SNP per 1.5 kb of transcribed sequence, covering 65% of known genes in the chromosome. Validation of our bioinformatics approach was demonstrated by a confirmation rate of 78% for the predicted cSNPs, and in total 32% of the cSNPs in our database have been confirmed. The database is publicly available at http://csnp.unige.ch or http://csnp.isb-sib.ch. These SNPs provide a tool to study the contribution of HC21 loci to complex diseases such as bipolar affective disorder and allele-specific contributions to Down syndrome phenotypes.

Base Composition↗

Measuring the fit of sequence data to phylogenetic model: allowing for missing data.

It is fundamentally important to assess the fit of data to model in phylogenetic and evolutionary studies. Phylogenetic methods using molecular sequences typically start with a multiple alignment. It is possible to measure the fit of data to model expectations of data, for example, via the likelihood-ratio (G) test or the X(2) test, if all sites in all sequences have an unambiguous residue. However, nearly all alignments of interest contain sites (columns of the alignment) with missing data, that is, ambiguous nucleotides, gaps, or unsequenced regions, which must presently be removed before using the above tests. Unfortunately, this is often either undesirable or impractical, as it will discard much of the data. Here, we show how iterative ML estimators may directly estimate the site-pattern probabilities for columns with missing data, given only standard i.i.d. assumptions. The optimization may use an EM or Newton algorithm, or any other hill-climbing approach. The resulting optimal likelihood under the unconstrained or multinomial model may be compared directly with the likelihood of the data coming from the model (a G statistic). Alternatively the modified observed and the expected frequencies of site patterns may be compared using a X(2) test. The distribution of such statistics is best assessed using appropriate simulations. The new method is applicable to models using codons or paired sites. The methods are also useful with Hadamard conjugations (spectral analysis) and are illustrated with these and with ML evolutionary models that allow site-rate variability.

Amino Acid Sequence↗

The genome of model malaria parasites, and comparative genomics.

The field of comparative genomics of malaria parasites has recently come of age with the completion of the whole genome sequences of the human malaria parasite Plasmodium falciparum and a rodent malaria model, Plasmodium yoelii yoelii. With several other genome sequencing projects of different model and human malaria parasite species underway, comparing genomes from multiple species has necessitated the development of improved informatics tools and analyses. Results from initial comparative analyses reveal striking conservation of gene synteny between malaria species within conserved chromosome cores, in contrast to reduced homology within subtelomeric regions, in line with previous findings on a smaller scale. Genes that elicit a host immune response are frequently found to be species-specific, although a large variant multigene family is common to many rodent malaria species and Plasmodium vivax. Sequence alignment of syntenic regions from multiple species has revealed the similarity between species in coding regions to be high relative to non-coding regions, and phylogenetic footprinting studies promise to reveal conserved motifs in the latter. Comparison of non-synonymous substitution rates between orthologous genes is proving a powerful technique for identifying genes under selection pressure, and may be useful for vaccine design. This is a stimulating time for comparative genomics of model and human malaria parasites, which promises to produce useful results for the development of antimalarial drugs and vaccines.

Animals↗

MATCH-BOX: a fundamentally new algorithm for the simultaneous alignment of several protein sequences.

Original algorithms for simultaneous alignment of protein sequences are presented, including sequence clustering and within- or between-groups multiple alignment. The way of matching similar regions is fundamentally new. Complete matches are formed by segments more similar than expected by random, according to a given probability limit. Any classic or user-defined score matrix can be used to express the similarity between the residues. The algorithm seeks for complete matches common to all the sequences without performing pairwise alignment and regardless of gap weighting. An automatic screening delineates all the similar regions (boxes) that may be defined for a given maximal shift between the sequences. The shift can be large enough to allow the matching of any region of a sequence with any region of another one. It can also be short and used to refine the alignment around anchor points. The algorithm provides the most likely optimal alignment and a comprehensive list of the alignment dilemma. Duality between automatism and interactivity is provided. Depending on the problem complexity, a final alignment is obtained fully automatically or requires some interactive handling to discriminate alternative pathways.

Algorithms↗

A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.

The expressed human genome is being sequenced and analyzed by disparate groups producing disparate data. The majority of the identified coding portion is in the form of expressed sequence tags (ESTs). The need to discover exonic representation and expression forms of full-length cDNAs for each human gene is frustrated by the partial and variable quality nature of this data delivery. A highly redundant human EST data set has been processed into integrated and unified expressed transcript indices that consist of hierarchically organized human transcript consensi reflecting gene expression forms and genetic polymorphism within an index class. The expression index and its intermediate outputs include cleaned transcript sequence, expression, and alignment information and a higher fidelity subset, SANIGENE. The STACK_PACK clustering system has been applied to dbEST release 121598 (GenBank version 110). Sixty-four percent of 1,313, 103 Homo sapiens ESTs are condensed into 143,885 tissue level multiple sequence clusters; linking through clone-ID annotations produces 68,701 total assemblies, such that 81% of the original input set is captured in a STACK multiple sequence or linked cluster. Indexing of alignments by substituent EST accession allows browsing of the data structure and its cross-links to UniGene. STACK metaclusters consolidate a greater number of ESTs by a factor of 1. 86 with respect to the corresponding UniGene build. Fidelity comparison with genome reference sequence AC004106 demonstrates consensus expression clusters that reflect significantly lower spurious repeat sequence content and capture alternate splicing within a whole body index cluster and three STACK v.2.3 tissue-level clusters. Statistics of a staggered release whole body index build of STACK v.2.0 are presented.

Algorithms↗

SABmark--a benchmark for sequence alignment that covers the entire known fold space.

The Sequence Alignment Benchmark (SABmark) provides sets of multiple alignment problems derived from the SCOP classification. These sets, Twilight Zone and Superfamilies, both cover the entire known fold space using sequences with very low to low, and low to intermediate similarity, respectively. In addition, each set has an alternate version in which unalignable but apparently similar sequences are added to each problem.

Algorithms↗

A strategy for the identification of T-cell epitopes on Leishmania cysteine proteinases.

In this study computational analysis was used to compile sequence alignments, construct a dendrogram and calculate physical data in order to predict potential T-cell epitopes of the Leishmania cysteine proteinase. Using multiple alignment of human and Leishmania proteinase sequences deposited on data bank sequences, it was possible to predict that the extreme C-terminus of cysteine proteinase (Cyspep, 355-444) contained three peptides (pI 361-370, pII 415-422 and pIII 431-444) with charge score, hydrophobicity and isoelectric points compatible for human leucocyte-associated antigen (HLA) class II binding. The prediction was confirmed in vitro through the ability of synthetic peptides corresponding to the predicted regions to stimulate peripheral blood mononuclear cells of patients with leishmaniasis.

Algorithms↗

Bipartite pattern discovery by entropy minimization-based multiple local alignment.

Many multimeric transcription factors recognize DNA sequence patterns by cooperatively binding to bipartite elements composed of half sites separated by a flexible spacer. We developed a novel bipartite algorithm, bipartite pattern discovery (Bipad), which produces a mathematical model based on information maximization or Shannon's entropy minimization principle, for discovery of bipartite sequence patterns. Bipad is a C++ program that applies greedy methods to search the bipartite alignment space and examines the upstream or downstream regions of co-regulated genes, looking for cis-regulatory bipartite patterns. An input sequence file with zero or one site per locus is required, and the left and right motif widths and a range of possible gap lengths must be specified. Bipad can run in either single-block or bipartite pattern search modes, and it is capable of comprehensively searching all four orientations of half-site patterns. Simulation studies showed that the accuracy of this motif discovery algorithm depends on sample size and motif conservation level, but results were independent of background composition. Bipad performed equivalent with or better than other pattern search algorithms in correctly identifying Escherichia coli cyclic AMP receptor protein and Bacillus subtilis sigma factor binding site sequences based on experimentally defined benchmarks. Finally, a new bipartite information weight matrix for vitamin D3 receptor/retinoid X receptor alpha (VDR/RXRalpha) binding sites was derived that comprehensively models the natural variability inherent in these sequence elements.

Algorithms↗

Comparative analysis of the catalytic domain of hemorrhagic and non-hemorrhagic snake venom metallopeptidases using bioinformatic tools.

Snake venom metalloproteases (SVMPs) are a set of interesting enzymes that are one of the major components of venom affecting hemostasis. A great challenge since their discovery has been to find molecular features responsible for their hemorrhagic potency and many attempts have been made without any consistent result. Here we describe a series of comparisons between the catalytic domains of hemorrhagic and non-hemorrhagic SVMPs made with the help of bioinformatics. These involved sequence and structure-based multiple alignments, phylogenetic reconstruction, predicted physical and chemical properties, motif scanning and structural analyses. Although hemorrhagic activity seems to be complex, involving multiple factors, we found some molecular characteristics that may influence the toxic effects. Among these findings, it was possible to use a molecular surface feature to subdivide the P-I class in hemorrhagic and non-hemorrhagic SVMPs. It was also possible to suggest a role for the conserved Asp148 and Ser176 residues in the stabilization of the active site.

Animals↗

BCM Search Launcher--an integrated interface to molecular biology data base search and analysis services available on the World Wide Web.

The BCM Search Launcher is an integrated set of World Wide Web (WWW) pages that organize molecular biology-related search and analysis services available on the WWW by function, and provide a single point of entry for related searches. The Protein Sequence Search Page, for example, provides a single sequence entry form for submitting sequences to WWW servers that offer remote access to a variety of different protein sequence search tools, including BLAST, FASTA, Smith-Waterman, BEAUTY, PROSITE, and BLOCKS searches. Other Launch pages provide access to (1) nucleic acid sequence searches, (2) multiple and pair-wise sequence alignments, (3) gene feature searches, (4) protein secondary structure prediction, and (5) miscellaneous sequence utilities (e.g., six-frame translation). The BCM Search Launcher also provides a mechanism to extend the utility of other WWW services by adding supplementary hypertext links to results returned by remote servers. For example, links to the NCBI's Entrez data base and to the Sequence Retrieval System (SRS) are added to search results returned by the NCBI's WWW BLAST server. These links provide easy access to auxiliary information, such as Medline abstracts, that can be extremely helpful when analyzing BLAST data base hits. For new or infrequent users of sequence data base search tools, we have preset the default search parameters to provide the most informative first-pass sequence analysis possible. We have also developed a batch client interface for Unix and Macintosh computers that allows multiple input sequences to be searched automatically as a background task, with the results returned as individual HTML documents directly to the user's system. The BCM Search Launcher and batch client are available on the WWW at URL http:@gc.bcm.tmc.edu:8088/search-launcher.html.

Animals↗

Secondary structural predictions for the clostridial neurotoxins.

The primary structures of a family of ten clostridial neurotoxins have recently been deduced yet little information is presently available concerning their secondary or tertiary structures. Because the overall similarity percentage of multiply aligned sequences is high, the secondary structures of these metalloendopeptidases are also expected to be conserved. The neural net program, PHD (Rost and Sander, Proc. Natl. Acad. Sci. USA 90:7558-7562, 1993), predicted that the secondary structures of the neurotoxins were indeed conserved in both single and multiple sequence modes of analysis. Predictions for the amounts of helical, extended, and loop states from the single sequence analyses were consistent with previously published data from circular dichroism studies on some of these neurotoxins. In the single analysis mode, only the aligned regions were predicted to show conservation of the three-state structure. In contrast, the multiple sequence analysis predicted that a conserved state (variable loops) also exists in non-aligned regions. Alignments with the primary structure of the prototypic metalloendopeptidase thermolysin showed that about 25% of the residues within this enzyme are similar to those in the neurotoxins. A comparison of thermolysin's known secondary structure with the predictions from this study showed that about 80% of thermolysin's residues could be structurally aligned with those in the neurotoxins. These predictions provide the necessary framework to build a homologous low-resolution tertiary structure of the neurotoxin active site that will be essential in the development of synthetic inhibitors.

Amino Acid Sequence↗

Genetic organization of the streptokinase region of the Streptococcus equisimilis H46A chromosome.

The complete nucleotide sequences of four genes and one open reading frame (ORF1) adjacent to the streptokinase gene, skc, from Streptococcus equisimilis H46A were determined. These genes are encoded on the opposite DNA strand to skc and are arranged as follows: dexB-abc-lrp-skc-ORF1-rel. The dexB gene, coding for an alpha-glucosidase (M(r) 61,733), and abc, encoding an ABC transporter (M(r) 42,080), are similar to the dexB and msmK genes, respectively, from the multiple sugar metabolism operon of S. mutans. The lrp gene specifies a leucine-rich protein (M(r) 32,302) that has a leucine-zipper motif at its C-terminus. The function of the Lrp protein is not known but appeared to be detrimental when overexpressed in Escherichia coli. Although lrp appears not to be an essential gene, as judged by plasmid insertion mutagenesis, it is conserved in all streptococcal strains carrying a streptokinase gene. The rel gene showed significant homology to the E. coli relA and spoT genes involved in the stringent response to amino acid deprivation. Multiple alignment of the amino acid sequences of Rel (M(r) 83,913), RelA and SpoT revealed 59.4% homology of the primary structures. Northern hybridization analyses of the genes in the skc region showed skc to be transcribed most abundantly. In addition to transcripts for skc, monocistronic mRNAs were detected for all three genes divergently transcribed from skc. Although there was also some read-through transcription from lrp into abc, and from abc into dexB, the transcription pattern suggests a high degree of transcriptional and functional independence not only of skc but also abc and dexB. Prominent structural features in intergenic regions included a static DNA bending locus located upstream and a putative bidirectional transcription terminator downstream of skc.

Amino Acid Sequence↗

Polyester synthases: natural catalysts for plastics.

Polyhydroxyalkanoates (PHAs) are biopolyesters composed of hydroxy fatty acids, which represent a complex class of storage polyesters. They are synthesized by a wide range of different Gram-positive and Gram-negative bacteria, as well as by some Archaea, and are deposited as insoluble cytoplasmic inclusions. Polyester synthases are the key enzymes of polyester biosynthesis and catalyse the conversion of (R)-hydroxyacyl-CoA thioesters to polyesters with the concomitant release of CoA. These soluble enzymes turn into amphipathic enzymes upon covalent catalysis of polyester-chain formation. A self-assembly process is initiated resulting in the formation of insoluble cytoplasmic inclusions with a phospholipid monolayer and covalently attached polyester synthases at the surface. Surface-attached polyester synthases show a marked increase in enzyme activity. These polyester synthases have only recently been biochemically characterized. An overview of these recent findings is provided. At present, 59 polyester synthase structural genes from 45 different bacteria have been cloned and the nucleotide sequences have been obtained. The multiple alignment of the primary structures of these polyester synthases show an overall identity of 8-96% with only eight strictly conserved amino acid residues. Polyester synthases can been assigned to four classes based on their substrate specificity and subunit composition. The current knowledge on the organization of the polyester synthase genes, and other genes encoding proteins related to PHA metabolism, is compiled. In addition, the primary structures of the 59 PHA synthases are aligned and analysed with respect to highly conserved amino acids, and biochemical features of polyester synthases are described. The proposed catalytic mechanism based on similarities to alpha/beta-hydrolases and mutational analysis is discussed. Different threading algorithms suggest that polyester synthases belong to the alpha/beta-hydrolase superfamily, with a conserved cysteine residue as catalytic nucleophile. This review provides a survey of the known biochemical features of these unique enzymes and their proposed catalytic mechanism.

Acyltransferases↗

The major opsin in bees (Insecta: Hymenoptera): A promising nuclear gene for higher level phylogenetics.

We report the phylogenetic utility of the nuclear gene encoding the long-wavelength opsin (LW Rh) for tribes of bees. Aligned nucleotide sequences were examined in multiple taxa from the four tribes comprising the corbiculate bees within the subfamily Apinae. Phylogenetic analyses of sequence variation in a 502-bp fragment (approx 40% of the coding region) strongly supported the monophyly of each of the four tribes, which are well established from previous studies of morphology and DNA. Trees estimated from parsimony and maximum likelihood analyses of LW Rh sequences show a strongly supported relationship between the tribes Meliponini and Bombini, a relationship that has been found uniformly in studies of other genes (28S, 16S, and cytochrome b). All of the tribal clades as well as relationships among the tribes are supported by high bootstrap values, suggesting the utility of LW Rh in estimating tribal and subfamily rank for these bees. The sequences exhibit minimal base composition bias. Both 1st + 2nd and 3rd position sites provide information for estimating a reliable tree topology. These results suggest that LW Rh, which has not been reported previously in studies of organismal phylogenetics, could provide important new data from the nuclear genome for phylogeny reconstruction.

Animals↗

Adaptations of the helix-grip fold for ligand binding and catalysis in the START domain superfamily.

With a protein structure comparison, an iterative database search with sequence profiles, and a multiple-alignment analysis, we show that two domains with the helix-grip fold, the star-related lipid-transfer (START) domain of the MLN64 protein and the birch allergen, are homologous. They define a large, previously underappreciated superfamily that we call the START superfamily. In addition to the classical START domains that are primarily involved in eukaryotic signaling mediated by lipid binding and the birch antigen family that consists of plant proteins implicated in stress/pathogen response, the START superfamily includes bacterial polyketide cyclases/aromatases (e.g., TcmN and WhiE VI) and two families of previously uncharacterized proteins. The identification of this domain provides a structural prediction of an important class of enzymes involved in polyketide antibiotic synthesis and allows the prediction of their active site. It is predicted that all START domains contain a similar ligand-binding pocket. Modifications of this pocket determine the ligand-binding specificity and may also be the basis for at least two distinct enzymatic activities, those of a cyclase/aromatase and an RNase. Thus, the START domain superfamily is a rare case of the adaptation of a protein fold with a conserved ligand-binding mode for both a broad variety of catalytic activities and noncatalytic regulatory functions. Proteins 2001;43:134-144.

Allergens↗