Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Phagotopes derived by antibody screening of phage-displayed random peptide libraries vary in immunoreactivity: studies using an exemplary monoclonal antibody, CII-C1, to type II collagen.

Antibody screening of phage-displayed random peptide libraries to identify mimotopes of conformational epitopes is promising. However, because interpretations can be difficult, an exemplary system has been used in the present study to investigate whether variation in the peptide sequences of selected phagotopes corresponded with variation in immunoreactivity. The phagotopes, derived using a well-characterized monoclonal antibody, CII-C1, to a known conformational epitope on type II collagen, C1, were tested by direct and inhibition ELISA for reactivity with CII-C1. A multiple sequence alignment algorithm, PILEUP, was used to sort the peptides expressed by the phagotopes into clusters. A model was prepared of the C1 epitope on type II collagen. The 12 selected phagotopes reacted with CII-C1 by both direct ELISA (titres from < 100-11 200) and inhibition ELISA (20-100% inhibition); the reactivity varied according to the peptide sequence and assay format. The differences in reactivity between the phagotopes were mostly in accord with the alignment, by PILEUP, of the peptide sequences. The finding that the phagotopes functionally mimicked the C1 epitope on collagen was validated in that amino acids RRL at the amino terminal of many of the peptides were topographically demonstrable on the model of the C1 epitope. Notably, one phagotope that expressed the widely divergent peptide C-IAPKRHNSA-C also mimicked the C1 epitope, as judged by reactivity in each of the assays used: these included cross-inhibition of CII-C1 reactivity with each of the other phagotopes and inhibition by a synthetic peptide corresponding to that expressed by the most frequently selected phagotope, RRLPFGSQM. Thus, it has been demonstrated that multiple phage-displayed peptides can mimic the same epitope and that observed immunoreactivity of selected phagotopes with the selecting mAb can depend on the primary sequence of the expressed peptide and also on the assay format used.

Amino Acid Sequence↗

CRYAB promoter polymorphisms: influence on multiple sclerosis susceptibility and clinical presentation.

BACKGROUND: alphaB-crystallin is a molecular chaperone and potential myelin antigen, up-regulated in the earlier stages of multiple sclerosis (MS) lesions. In the alphaB-crystallin gene (CRYAB), single nucleotide polymorphisms (SNPs) have been associated with MS susceptibility (g.CRYAB-652A>G) and a rapidly progressive clinical course (g.CRYAB-650C>G). METHOD: CRYAB was screened for mutations in 233 MS patients and 96 controls. Genomic DNA was extracted and the coding and 3' and 5' untranslated regions were amplified by PCR. Subsequently, the products were analysed by Single Strand Conformation Polymorphism technique followed by DNA sequencing of aberrant conformers. RESULTS: In CRYAB (Genbank ) no mutations were found but SNPs were identified in the promoter region (g.CRYAB-249C>G, g.CRYAB-650C>G and g.CRYAB-652A>G), and intronic region (g.CRYAB.2398T>G). The g.CRYAB-249C>G genotype distribution was significantly different between groups (chi(2), p=0.01), caused by differences between Relapsing Remitting MS (RRMS) and controls (chi(2), p=0.025) and Secondary Progressive MS (SPMS) and controls (chi(2), p=0.05). In addition, a significant difference was observed in the g.CRYAB-249C>G allele distribution (chi(2), p=0.04), caused by a difference between SPMS and controls (chi(2), p=0.01). In RRMS and SPMS a tendency of the g.CRYAB-249GG genotype being associated with an earlier age of onset (p=0.05) and a slowly progressive cause (p=0.07) was found. Multiple sequence alignment showed conservation of the g.CRYAB-249*C between mammalian CRAYB genes and within the small heat shock protein gene family. CONCLUSION: CRYAB polymorphisms may be involved in the pathogenesis of MS by mechanisms that could involve increased expression of the superantigen alphaB-crystallin and modulation of the immune response. CRYAB polymorphisms should be included in future multivariate biomaker studies in MS.

Adult↗

3MATRIX and 3MOTIF: a protein structure visualization system for conserved sequence motifs.

Computational methods such as sequence alignment and motif construction are useful in grouping related proteins into families, as well as helping to annotate new proteins of unknown function. These methods identify conserved amino acids in protein sequences, but cannot determine the specific functional or structural roles of conserved amino acids without additional study. In this work, we present 3MATRIX (http://3matrix.stanford.edu) and 3MOTIF (http://3motif.stanford.edu), a web-based sequence motif visualization system that displays sequence motif information in its appropriate three-dimensional (3D) context. This system is flexible in that users can enter sequences, keywords, structures or sequence motifs to generate visualizations. In 3MOTIF, users can search using discrete sequence motifs such as PROSITE patterns, eMOTIFs, or any other regular expression-like motif. Similarly, 3MATRIX accepts an eMATRIX position-specific scoring matrix, or will convert a multiple sequence alignment block into an eMATRIX for visualization. Each query motif is used to search the protein structure database for matches, in which the motif is then visually highlighted in three dimensions. Important properties of motifs such as sequence conservation and solvent accessible surface area are also displayed in the visualizations, using carefully chosen color shading schemes.

Amino Acid Motifs↗

Consensus-degenerate hybrid oligonucleotide primers for amplification of distantly related sequences.

We describe a new primer design strategy for PCR amplification of unknown targets that are related to multiply-aligned protein sequences. Each primer consists of a short 3' degenerate core region and a longer 5' consensus clamp region. Only 3-4 highly conserved amino acid residues are necessary for design of the core, which is stabilized by the clamp during annealing to template molecules. During later rounds of amplification, the non-degenerate clamp permits stable annealing to product molecules. We demonstrate the practical utility of this hybrid primer method by detection of diverse reverse transcriptase-like genes in a human genome, and by detection of C5DNA methyltransferase homologs in various plant DNAs. In each case, amplified products were sufficiently pure to be cloned without gel fractionation. This COnsensus-DEgenerate Hybrid Oligonucleotide Primer (CODEHOP) strategy has been implemented as a computer program that is accessible over the World Wide Web (http://blocks.fhcrc.org/codehop.html) and is directly linked from the BlockMaker multiple sequence alignment site for hybrid primer prediction beginning with a set of related protein sequences.

Amino Acid Sequence↗

Moderate mutation rate in the SARS coronavirus genome and its implications.

BACKGROUND: The outbreak of severe acute respiratory syndrome (SARS) caused a severe global epidemic in 2003 which led to hundreds of deaths and many thousands of hospitalizations. The virus causing SARS was identified as a novel coronavirus (SARS-CoV) and multiple genomic sequences have been revealed since mid-April, 2003. After a quiet summer and fall in 2003, the newly emerged SARS cases in Asia, particularly the latest cases in China, are reinforcing a wide-spread belief that the SARS epidemic would strike back. With the understanding that SARS-CoV might be with humans for years to come, knowledge of the evolutionary mechanism of the SARS-CoV, including its mutation rate and emergence time, is fundamental to battle this deadly pathogen. To date, the speed at which the deadly virus evolved in nature and the elapsed time before it was transmitted to humans remains poorly understood. RESULTS: Sixteen complete genomic sequences with available clinical histories during the SARS outbreak were analyzed. After careful examination of multiple-sequence alignment, 114 single nucleotide variations were identified. To minimize the effects of sequencing errors and additional mutations during the cell culture, three strategies were applied to estimate the mutation rate by 1) using the closely related sequences as background controls; 2) adjusting the divergence time for cell culture; or 3) using the common variants only. The mutation rate in the SARS-CoV genome was estimated to be 0.80 - 2.38 x 10-3 nucleotide substitution per site per year which is in the same order of magnitude as other RNA viruses. The non-synonymous and synonymous substitution rates were estimated to be 1.16 - 3.30 x 10-3 and 1.67 - 4.67 x 10-3 per site per year, respectively. The most recent common ancestor of the 16 sequences was inferred to be present as early as the spring of 2002. CONCLUSIONS: The estimated mutation rates in the SARS-CoV using multiple strategies were not unusual among coronaviruses and moderate compared to those in other RNA viruses. All estimates of mutation rates led to the inference that the SARS-CoV could have been with humans in the spring of 2002 without causing a severe epidemic.

Coronavirus↗

Linking enzyme sequence to function using Conserved Property Difference Locator to identify and annotate positions likely to control specific functionality.

BACKGROUND: Families of homologous enzymes evolved from common progenitors. The availability of multiple sequences representing each activity presents an opportunity for extracting information specifying the functionality of individual homologs. We present a straightforward method for the identification of residues likely to determine class specific functionality in which multiple sequence alignments are converted to an annotated graphical form by the Conserved Property Difference Locator (CPDL) program. RESULTS: Three test cases, each comprised of two groups of functionally-distinct homologs, are presented. Of the test cases, one is a membrane and two are soluble enzyme families. The desaturase/hydroxylase data was used to design and test the CPDL algorithm because a comparative sequence approach had been successfully applied to manipulate the specificity of these enzymes. The other two cases, ATP/GTP cyclases, and MurD/MurE synthases were chosen because they are well characterized structurally and biochemically. For the desaturase/hydroxylase enzymes, the ATP/GTP cyclases and the MurD/MurE synthases, groups of 8 (of approximately 400), 4 (of approximately 150) and 10 (of >400) residues, respectively, of interest were identified that contain empirically defined specificity determining positions. CONCLUSION: CPDL consistently identifies positions near enzyme active sites that include those predicted from structural and/or biochemical studies to be important for specificity and/or function. This suggests that CPDL will have broad utility for the identification of potential class determining residues based on multiple sequence analysis of groups of homologous proteins. Because the method is sequence, rather than structure, based it is equally well suited for designing structure-function experiments to investigate membrane and soluble proteins.

Algorithms↗

Phylogenetic supermatrix analysis of GenBank sequences from 2228 papilionoid legumes.

A comprehensive phylogeny of papilionoid legumes was inferred from sequences of 2228 taxa in GenBank release 147. A semiautomated analysis pipeline was constructed to download, parse, assemble, align, combine, and build trees from a pool of 11,881 sequences. Initial steps included all-against-all BLAST similarity searches coupled with assembly, using a novel strategy for building length-homogeneous primary sequence clusters. This was followed by a combination of global and local alignment protocols to build larger secondary clusters of locally aligned sequences, thus taking into account the dramatic differences in length of the heterogeneous coding and noncoding sequence data present in GenBank. Next, clusters were checked for the presence of duplicate genes and other potentially misleading sequences and examined for combinability with other clusters on the basis of taxon overlap. Finally, two supermatrices were constructed: a "sparse" matrix based on the primary clusters alone (1794 taxa x 53,977 characters), and a somewhat more "dense" matrix based on the secondary clusters (2228 taxa x 33,168 characters). Both matrices were very sparse, with 95% of their cells containing gaps or question marks. These were subjected to extensive heuristic parsimony analyses using deterministic and stochastic heuristics, including bootstrap analyses. A "reduced consensus" bootstrap analysis was also performed to detect cryptic signal in a subtree of the data set corresponding to a "backbone" phylogeny proposed in previous studies. Overall, the dense supermatrix appeared to provide much more satisfying results, indicated by better resolution of the bootstrap tree, excellent agreement with the backbone papilionoid tree in the reduced bootstrap consensus analysis, few problematic large polytomies in the strict consensus, and less fragmentation of conventionally recognized genera. Nevertheless, at lower taxonomic levels several problems were identified and diagnosed. A large number of methodological issues in supermatrix construction at this scale are discussed, including detection of annotation errors in GenBank sequences; the shortage of effective algorithms and software for local multiple sequence alignment; the difficulty of overcoming effects of fragmentation of data into nearly disjoint blocks in sparse supermatrices; and the lack of informative tools to assess confidence limits in very large trees.

Algorithms↗

Tree pattern matching in phylogenetic trees: automatic search for orthologs or paralogs in homologous gene sequence databases.

MOTIVATION: Comparative sequence analysis is widely used to study genome function and evolution. This approach first requires the identification of homologous genes and then the interpretation of their homology relationships (orthology or paralogy). To provide help in this complex task, we developed three databases of homologous genes containing sequences, multiple alignments and phylogenetic trees: HOBACGEN, HOVERGEN and HOGENOM. In this paper, we present two new tools for automating the search for orthologs or paralogs in these databases. RESULTS: First, we have developed and implemented an algorithm to infer speciation and duplication events by comparison of gene and species trees (tree reconciliation). Second, we have developed a general method to search in our databases the gene families for which the tree topology matches a peculiar tree pattern. This algorithm of unordered tree pattern matching has been implemented in the FamFetch graphical interface. With the help of a graphical editor, the user can specify the topology of the tree pattern, and set constraints on its nodes and leaves. Then, this pattern is compared with all the phylogenetic trees of the database, to retrieve the families in which one or several occurrences of this pattern are found. By specifying ad hoc patterns, it is therefore possible to identify orthologs in our databases.

Algorithms↗

Membrane topology of the multidrug resistance protein (MRP). A study of glycosylation-site mutants reveals an extracytosolic NH2 terminus.

Multidrug resistance protein, MRP, is a 190-kDa integral membrane phosphoglycoprotein that belongs to the ATP-binding cassette superfamily of transport proteins and is capable of conferring resistance to multiple chemotherapeutic agents. Previous studies have indicated that MRP consists of two membrane spanning domains (MSD) each followed by a nucleotide binding domain, plus an additional extremely hydrophobic NH2-terminal MSD. Computer-assisted hydropathy analyses and multiple sequence alignments suggest several topological models for MRP. To aid in determining the topology most likely to be correct, we have identified which of the 14 N-glycosylation sequons in this protein are utilized. Limited proteolysis of MRP-enriched membranes and deglycosylation of intact MRP and its tryptic fragments with PNGase F was carried out followed by immunoblotting with antibodies known to react with specific regions of MRP. The results obtained indicated that the sequon at Asn354 in the middle MSD is not utilized and suggested approximate sites of N-glycosylation. Subsequent site-directed mutagenesis studies established that Asn19 and Asn23 in the NH2-terminal MSD and Asn1006 in the COOH-terminal MSD are the only sites in MRP that are modified with N-linked oligosaccharides. N-Glycosylation of Asn19 and Asn23 provides the first direct experimental evidence that MRP has an extracytosolic NH2 terminus. This finding, together with those of previous studies, strongly suggests that the NH2-terminal MSD of MRP contains an odd number of transmembrane helices. These results may have important implications for the further understanding of the interaction of drugs with MRP.

ATP-Binding Cassette Transporters↗

CLAGen: a tool for clustering and annotating gene sequences using a suffix tree algorithm.

Most multiple gene sequence alignment methods rely on conventions regarding the score of a multiple alignment in pairwise fashion. Therefore, as the number of sequences increases, the runtime of sequencing expands exponentially. In order to solve the problem, this paper presents a multiple sequence alignment method using a linear-time suffix tree algorithm to cluster similar sequences at one time without pairwise alignment. After searching for common subsequences, cross-matching common subsequences were generated, and sometimes inexact matching was found. So, a procedure aimed at masking the inexact cross-matching pairs was suggested here. In addition, BLAST was combined with a clustering tool in order to annotate the clusters generated by suffix tree clustering. The proposed method for clustering and annotating genes consists of the following steps: (1) construction of a suffix tree; (2) searching and overlapping common subsequences; (3) grouping subsequence pairs; (4) masking cross-matching pairs; (5) clustering gene sequences; (6) annotating gene clusters by the BLAST search. The performance of the proposed system, CLAGen, was successfully evaluated with 42 gene sequences in a TCA cycle (a citrate cycle) of bacteria. The system generated 11 clusters and found the longest subsequences of each cluster, which are biologically significant.

Algorithms↗

Detection rates of TT virus among children who visited a general hospital in Japan.

Recently, genomic DNA of the novel TT virus (TTV) was isolated from patients suffering from posttransfusion hepatitis of unknown etiology. We examined sera from 197 children who visited the Department of Pediatrics at Toyohashi National Hospital. Sera were tested for TTV DNA by seminested polymerase chain reaction (PCR) using a set of primers synthesized according to the published TTV sequence. Ten children were found to be positive for TTV (5.1%). All positive PCR products were directly sequenced in both directions using a fluorescent dye terminator cycle sequencing system. The sequences were compared by a multiple sequence alignment and a phylogenetic tree was constructed. The phylogenetic tree showed that two of the TTV isolates found in the present experiment did not belong to any of the phylogenetic groups previously reported.

Base Sequence↗

Identification and characterization of the KlCMD1 gene encoding Kluyveromyces lactis calmodulin.

The KlCMD1 gene was isolated from a Kluyveromyces lactis genomic library as a suppressor of the Saccharomyces cerevisiae temperature-sensitive mutant spc110-124, an allele previously shown to be suppressed by elevated copy number of the S. cerevisiae calmodulin gene CMD1. The KlCMD1 gene encodes a polypeptide which is 95% identical to S. cerevisiae calmodulin and 55% identical to calmodulin from Schizosaccharomyces pombe. Complementation of a S. cerevisiae cdm1 deletion mutant by KlCMD1 demonstrates that this gene encodes a functional calmodulin homologue. Multiple sequence alignment of calmodulins from yeast and multicellular eukaryotes shows that the K. lactis and S. cerevisiae calmodulins are considerably more closely related to each other than to other calmodulins, most of which have four functional Ca2+-binding EF hand domains. Thus like its S. cerevisiae counterpart Cmd1p, the KlCMD1 product is predicted to form only three Ca2+-binding motifs.

Amino Acid Sequence↗

Refinement of 3D models of horseradish peroxidase isoenzyme C: predictions of 2D NMR assignments and substrate binding sites.

In this study, two alternative three-dimensional (3D) models of horseradish peroxidase (HRP-C)-differing mainly in the structure of a long untemplated insertion-were refined, systematically assessed, and used to make predictions that can both guide and be tested by future experimental studies. A key first step in the model-building process was a procedure for multiple sequence alignment based on structurally conserved regions and key conserved residues, including those side chains providing ligands to the two Ca2+ binding sites. The model refinements reported here include (1) optimization of side-chain conformations; (3) addition of structural waters using a template-independent procedure; (2) structural refinement of the untemplated 34 amino acid insertion located between the F and G helices, using both energy criteria and NMR data; (4) unconstrained energy optimization of the refined models. Using these procedures, two refined structures of HRP-C were obtained, differing mainly in the conformation of this long insertion. The presence of residues in this insertion that could potentially interact with bound substrates suggests a functional role that may be related to the general ability of class III peroxidases to form stable 1:1 complexes with a variety of substrates. The structural validity of the models was systematically assessed by a variety of criteria. Most notably, the ProsaII z scores and Profiles 3D scores of the two HRP-C models indicated that they are significantly better than would be obtained by simple amino acid replacement, using any of the known structures as a template. These two 3D HRP-C models, were then used to predict candidate residues for the assignment of NOESY cross-peaks previously noted in 2D-NMR studies. Specifically, the residues known as Ile X, Phe A, Phe B, aliphatic residue Q, and Ile T. Candidate substrate binding sites were also identified and compared with experimentally based predictions. This work is timely because new X-ray structures are anticipated that will facilitate the validation of these procedures.

Amino Acid Sequence↗

Better 1D predictions by experts with machines.

Accuracy of predicting protein secondary structure and solvent accessibility has been improved significantly by using evolutionary information contained in multiple sequence alignments. For the second Asilomar meeting, predictions were made automatically for all targets using the publicly available prediction service PredictProtein. Additionally, a semiautomatic procedure for generating more informative alignments was used in combination with the PHD prediction methods. Results confirmed the estimates for prediction accuracy. Furthermore, the more informative alignments yielded better predictions. The fairly accurate predictions of 1D structure were successfully used by various groups for the Asilomar meeting as first step toward predicting higher dimensions of protein structure.

Expert Systems↗

Improvement of protein secondary structure prediction using binary word encoding.

We propose a binary word encoding to improve the protein secondary structure prediction. A binary word encoding encodes a local amino acid sequence to a binary word, which consists of 0 or 1. We use an encoding function to map an amino acid to 0 or 1. Using the binary word encoding, we can statistically extract the multiresidue information, which depends on more than one residue. We combine the binary word encoding with the GOR method, its modified version, which shows better accuracy, and the neural network method. The binary word encoding improves the accuracy of GOR by 2.8%. We obtain similar improvement when we combine this with the modified GOR method and the neural network method. When we use multiple sequence alignment data, the binary word encoding similarly improves the accuracy. The accuracy of our best combined method is 68.2%. In this paper, we only show improvement of the GOR and neural network method, we cannot say that the encoding improves the other methods. But the improvement by the encoding suggests that the multiresidue interaction affects the formation of secondary structure. In addition, we find that the optimal encoding function obtained by the simulated annealing method relates to nonpolarity. This means that nonpolarity is important to the multiresidue interaction.

Amino Acid Sequence↗

Evaluation and improvement of multiple sequence methods for protein secondary structure prediction.

A new dataset of 396 protein domains is developed and used to evaluate the performance of the protein secondary structure prediction algorithms DSC, PHD, NNSSP, and PREDATOR. The maximum theoretical Q3 accuracy for combination of these methods is shown to be 78%. A simple consensus prediction on the 396 domains, with automatically generated multiple sequence alignments gives an average Q3 prediction accuracy of 72.9%. This is a 1% improvement over PHD, which was the best single method evaluated. Segment Overlap Accuracy (SOV) is 75.4% for the consensus method on the 396-protein set. The secondary structure definition method DSSP defines 8 states, but these are reduced by most authors to 3 for prediction. Application of the different published 8- to 3-state reduction methods shows variation of over 3% on apparent prediction accuracy. This suggests that care should be taken to compare methods by the same reduction method. Two new sequence datasets (CB513 and CB251) are derived which are suitable for cross-validation of secondary structure prediction methods without artifacts due to internal homology. A fully automatic World Wide Web service that predicts protein secondary structure by a combination of methods is available via http://barton.ebi.ac.uk/.

Algorithms↗

Structure modelling and site-directed mutagenesis of the rat aromatic L-amino acid pyridoxal 5'-phosphate-dependent decarboxylase: a functional study.

The pyridoxal-5'-phosphate-dependent enzymes (B6 enzymes) are grouped into three main families named alpha, beta, and gamma. Proteins in the alpha and gamma families share the same fold and might be distantly related, while those in the beta family exhibit specific structural features. The rat aromatic L-amino acid decarboxylase (AADC; EC(4.1.1.28)) catalyzes the synthesis of two important neurotransmitters: dopamine and serotonin. It binds the cofactor pyridoxal-5'-phosphate and belongs to the alpha family. Despite the low level of sequence identity (approximately 10%) shared by the rat AADC and the sequences of the enzymes belonging to the B6 enzymes family, including the known three-dimensional structures, a multiple sequence alignment was deduced. A model was built using segments belonging to seven of the eleven known structures. By homology, and based on knowledge of the biochemistry of the aspartate aminotransferase, structurally and functionally important residues were identified in the rat AADC. Site-directed mutagenesis of the conserved residues D271, T246, and C311 was carried out in order to confirm our predictions and highlight their functional role. Mutation of D271A and D271N resulted in complete loss of enzyme activity, while the D271E mutant exhibited 2% of the wild-type activity. Substitution of T246A resulted in 5% of the wild-type activity while the C311A mutant conserved 42% of the wild-type activity. A functional model of the AADC is discussed in view of the structural model and the complementary mutagenesis and labelling studies.

Amino Acid Sequence↗

Molecular dynamics simulations of isolated transmembrane helices of potassium channels.

In the middle of the S6 helix in voltage-gated potassium channels there is a highly conserved Pro-Val-Pro motif, while the equivalent M2 helix of inward rectifier potassium channels contains a conserved glycine residue in a comparable position. The structural implications of these conserved motifs are of interest given the evidence that S6 and M2 are components of the lining of their respective pores. Multiple sequence alignment and TM helix prediction methods were used to define consensus regions for S6 and M2. Ensembles of 50 structures for each helix were generated by simulated annealing and restrained molecular dynamics. Time-dependent fluctuations of S6 and M2 were investigated by long time scale molecular dynamics simulations on representative members of each ensemble carried out in vacuo in the presence and absence of a hydrophobic potential that mimics a lipid bilayer. The results are discussed in terms of the structural basis of the kink in S6 and M2 and of a putative functional role for flexible helices as "molecular swivels."

Amino Acid Sequence↗