Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Models of the primary and secondary structure for the 12S rRNA of birds: a guideline for sequence alignment.

Models of the primary and secondary structure for the 12S ribosomal RNA (rRNA) gene of birds is presented based on a comparison of 100 species. Preliminary higher-order structures were delimited following the model for vertebrates. Paired regions were refined following a complementary base-pairing criterion and compensatory mutations were considered as a further confirmation of their existence. The model shows 40 stems, 20 internal loops and 17 external loops, arranged in the typical four domains suggested for the small subunit rRNA. The higher-order structures recovered in the model were used to build a multiple sequence alignment appropriate for phylogenetic analysis. The phylogeny recovered from this alignment was compared with trees inferred from alignments assembled using different alignment parameters in the program ClustalW. The alignment based on the secondary structure is sensitive to positional covariation of stems. Nonetheless, the phylogeny recovered with this method resulted in relationships that are more congruent with non-molecular data than those inferred from alternative alignments.

Amino Acid Sequence↗

Sustained performance of knowledge-based potentials in fold recognition.

We describe the results obtained using fold recognition techniques in our third participation in the CASP experiment. The approach relies on knowledge-based potentials for alignment production and fold identification. As indicated by the increase in alignment quality and fold identification reliability, the predictions improved from CASP1 to CASP3. In particular, we identified structural relationships in which no known evolutionary link exists. Our predictions are based on single sequences rather than multiple sequence alignments. Additionally, we voluntarily submitted only a single model for each target because, in our view, submission of a single model is the most stringent test. We describe the methods used, the strategy adopted in the predictions, and the prediction results and discuss future work.

Algorithms↗

Integrated tools for structural and sequence alignment and analysis.

We have developed new computational methods for displaying and analyzing members of protein superfamilies. These methods (MinRMS, AlignPlot and MSFviewer) integrate sequence and structural information and are implemented as separate but cooperating programs to our Chimera molecular modeling system. Integration of multiple sequence alignment information and three-dimensional structural representations enable researchers to generate hypotheses about the sequence-structure relationship. Structural superpositions can be generated and easily tuned to identify similarities around important characteristics such as active sites or ligand binding sites. Information related to the release of Chimera, MinRMS, AlignPlot and MSFviewer can be obtained at http:¿www.cgl.ucsf.edu/chimera.

Amino Acid Sequence↗

Homology model of human corticosteroid binding globulin: a study of its steroid binding ability and a plausible mechanism of steroid hormone release at the site of inflammation.

Corticosteroid binding globulin (CBG) and thyroxin binding globulin (TBG) both belong to the same SERPIN superfamily of serine-proteinase inhibitors but in the course of evolution CBG has adapted to its new role as a transport agent of insoluble hormones. CBG binds corticosteroids in plasma, delivering them to sites of inflammation to modify the inflammatory response. CBG is an effective drug carrier for genetic manipulation, and hence there is immense biological interest in the location of the hormone binding site. The crystal structure of human CBG (hCBG) has not been determined, but sequence alignment with other SERPINs suggests that it conforms as a whole to the tertiary structure shared by the superfamily. Human CBG shares 52.15% and 55.50% sequence similarity with alpha1-antitrypsin and alpha1-antichymotrypsin, respectively. Multiple sequence alignment among the three sequences shows 73 conserved regions. The molecular structures of alpha1-antitrypsin and alpha1-antichymotrypsin, the archetype of the SERPIN superfamily, obtained by X-ray diffraction methods are used to develop a homology model of hCBG. Energy minimization was applied to the model to refine the structure further. The homology model of hCBG contains 371 residues (His13 to Val383 ). The secondary structure comprises 11 helices, 15 turns and 11 sheets. The putative corticosteroid binding region is found to exist in a pocket between beta-sheets S4, S10, S11 and alpha helix H10. Both cortisol and aldosterone are docked to the elongated hydrophobic ligand binding pocket with the polar residues at the two extremities. A difference accessible surface area (DASA) study revealed that cortisol binds with the native hCBG more tightly than aldosterone. Cleavage at the Val379-Met380 peptide bond causes a deformation of hCBG (also revealed through a DASA study). This deformation could probably trigger the release of the bound hormone. Figure Stereoscopic view of the ribbon diagram of hCBG complexed with cortisol. The bound cortisol is shown in space filling model in blue. Helices and sheets are shown in red and magenta respectively. Turns are shown in yellow.

Adrenal Cortex Hormones↗

An ATPase domain common to prokaryotic cell cycle proteins, sugar kinases, actin, and hsp70 heat shock proteins.

The functionally diverse actin, hexokinase, and hsp70 protein families have in common an ATPase domain of known three-dimensional structure. Optimal superposition of the three structures and alignment of many sequences in each of the three families has revealed a set of common conserved residues, distributed in five sequence motifs, which are involved in ATP binding and in a putative interdomain hinge. From the multiple sequence alignment in these motifs a pattern of amino acid properties required at each position is defined. The discriminatory power of the pattern is in part due to the use of several known three-dimensional structures and many sequences and in part to the "property" method of generalizing from observed amino acid frequencies to amino acid fitness at each sequence position. A sequence data base search with the pattern significantly matches sugar kinases, such as fuco-, glucono-, xylulo-, ribulo-, and glycerokinase, as well as the prokaryotic cell cycle proteins MreB, FtsA, and StbA. These are predicted to have subdomains with the same tertiary structure as the ATPase subdomains Ia and IIa of hexokinase, actin, and Hsc70, a very similar ATP binding pocket, and the capacity for interdomain hinge motion accompanying functional state changes. A common evolutionary origin for all of the proteins in this class is proposed.

Actins↗

Consensus shapes: an alternative to the Sankoff algorithm for RNA consensus structure prediction.

MOTIVATION: The well-known Sankoff algorithm for simultaneous RNA sequence alignment and folding is currently considered an ideal, but computationally over-expensive method. Available tools implement this algorithm under various pragmatic restrictions. They are still expensive to use, and it is difficult to judge if the moderate quality of results is because of the underlying model or to its imperfect implementation. RESULTS: We propose to redefine the consensus structure prediction problem in a way that does not imply a multiple sequence alignment step. For a family of RNA sequences, our method explicitly and independently enumerates the near-optimal abstract shape space, and predicts as the consensus an abstract shape common to all sequences. For each sequence, it delivers the thermodynamically best structure which has this common shape. Since the shape space is much smaller than the structure space, and identification of common shapes can be done in linear time (in the number of shapes considered), the method is essentially linear in the number of sequences. Our evaluation shows that the new method compares favorably with available alternatives. AVAILABILITY: The new method has been implemented in the program RNAcast and is available on the Bielefeld Bioinformatics Server. CONTACT: jreeder@TechFak.Uni-Bielefeld.DE, robert@TechFak.Uni-Bielefeld.DE SUPPLEMENTARY INFORMATION: Available at http://bibiserv.techfak.uni-bielefeld.de/rnacast/supplementary.html

Algorithms↗

A space-efficient algorithm for the constrained pairwise sequence alignment problem.

The constrained pairwise sequence alignment (CPSA) problem aims to align two given sequences by aligning their similar subsequences in the same region under the guidance of a given pattern (constraint). Let the lengths of the sequences be m, and n where n <or= m, and let r <or= n be the length of the given pattern. The optimum constrained pairwise alignment score can be computed using O(rn) space by a naive dynamic programming solution. If an optimal alignment path is desired then the space requirement of the naive dynamic programming algorithm is O(rnm). There is a divide-and-conquer algorithm that reduces the memory requirement of finding an optimal alignment for the CPSA problem to O(rn). In this paper, we present a space-efficient CPSA algorithm that returns an optimal alignment. Our analysis on real protein sequences suggests that our algorithm requires only O(n) space in practice. This algorithm is not only space efficient but also very fast. A generalization of the CPSA problem for multiple sequences is called the constrained multiple sequence alignment (CMSA) problem. Our CPSA algorithm also improves the space requirement of progressive CMSA algorithms that use solutions of CPSA problems.

Algorithms↗

Exploring substrate binding and discrimination in fructose1, 6-bisphosphate and tagatose 1,6-bisphosphate aldolases.

Fructose 1,6-bisphosphate aldolase catalyses the reversible condensation of glycerone-P and glyceraldehyde 3-phosphate into fructose 1,6-bisphosphate. A recent structure of the Escherichia coli Class II fructose 1,6-bisphosphate aldolase [Hall, D.R., Leonard, G.A., Reed, C.D., Watt, C.I., Berry, A. & Hunter, W.N. (1999) J. Mol. Biol. 287, 383-394] in the presence of the transition state analogue phosphoglycolohydroxamate delineated the roles of individual amino acids in binding glycerone-P and in the initial proton abstraction steps of the mechanism. The X-ray structure has now been used, together with sequence alignments, site-directed mutagenesis and steady-state enzyme kinetics to extend these studies to map important residues in the binding of glyceraldehyde 3-phosphate. From these studies three residues (Asn35, Ser61 and Lys325) have been identified as important in catalysis. We show that mutation of Ser61 to alanine increases the Km value for fructose 1, 6-bisphosphate 16-fold and product inhibition studies indicate that this effect is manifested most strongly in the glyceraldehyde 3-phosphate binding pocket of the active site, demonstrating that Ser61 is involved in binding glyceraldehyde 3-phosphate. In contrast a S61T mutant had no effect on catalysis emphasizing the importance of an hydroxyl group for this role. Mutation of Asn35 (N35A) resulted in an enzyme with only 1.5% of the activity of the wild-type enzyme and different partial reactions indicate that this residue effects the binding of both triose substrates. Finally, mutation of Lys325 has a greater effect on catalysis than on binding, however, given the magnitude of the effects it is likely that it plays an indirect role in maintaining other critical residues in a catalytically competent conformation. Interestingly, despite its proximity to the active site and high sequence conservation, replacement of a fourth residue, Gln59 (Q59A) had no significant effect on the function of the enzyme. In a separate study to characterize the molecular basis of aldolase specificity, the agaY-encoded tagatose 1,6-bisphosphate aldolase of E. coli was cloned, expressed and kinetically characterized. Our studies showed that the two aldolases are highly discriminating between the diastereoisomers fructose bisphosphate and tagatose bisphosphate, each enzyme preferring its cognate substrate by a factor of 300-1500-fold. This produces an overall discrimination factor of almost 5 x 105 between the two enzymes. Using the X-ray structure of the fructose 1,6-bisphosphate aldolase and multiple sequence alignments, several residues were identified, which are highly conserved and are in the vicinity of the active site. These residues might potentially be important in substrate recognition. As a consequence, nine mutations were made in attempts to switch the specificity of the fructose 1,6-bisphosphate aldolase to that of the tagatose 1,6-bisphosphate aldolase and the effect on substrate discrimination was evaluated. Surprisingly, despite making multiple changes in the active site, many of which abolished fructose 1, 6-bisphosphate aldolase activity, no switch in specificity was observed. This highlights the complexity of enzyme catalysis in this family of enzymes, and points to the need for further structural studies before we fully understand the subtleties of the shaping of the active site for complementarity to the cognate substrate.

Aldehyde-Lyases↗

The GPI1 homologue from Plasmodium falciparum complements a Saccharomyces cerevisiae GPI1 anchoring mutant.

Glycosylphosphatidylinositol (GPI) represents an important anchoring molecule for cell surface proteins. The first step in its synthesis is the transfer of N-acetylglucosamine (GlcNAc) from UDP-N-acetylglucosamine to phosphatidylinositol (PI). This chemically simple step is genetically complex because three or four genes are required in both yeast (GPI1, GPI2 and GPI3) and mammals (GPI1, PIG A, PIG H and PIG C), respectively. Here, we report cloning of a Plasmodium falciparum (P. falciparum) homologue of GPI1 (PfGPI1). Analysis showed that P. falciparum Gpi1p is somewhat more similar to the yeast proteins than human Gpi1p, showing 26 and 20% amino acid sequence identity with the Saccharomyces cerevisiae and Homo sapiens proteins, respectively. Multiple sequence alignment demonstrates also that the C-terminal half GPI1 proteins is much better conserved than the N-terminal half. The P. falciparum Gpi1p has a calculated molecular weight of 65 kDa and a predicted potential tyrosine phosphorylation site. The potential tyrosine phosphorylation site seems to occur in all other known Gpi1 proteins. Like the other GPI1 proteins, the predictive software revealed the absence of targeting signals such as organelle transit peptides, DNA binding sites, or N-terminal secretory signals. Hydrophobicity plots revealed multiple hydrophobic regions that could function as transmembrane segments. The cloned P. falciparum GPI1 gene complemented a gpi1 yeast mutant.

Amino Acid Sequence↗

S4: structure-based sequence alignments of SCOP superfamilies.

S4 is an automatically generated database of multiple structure-based sequence alignments of protein superfamilies in the SCOP database. All structural domains that do not share more than 40% sequence identity as defined by the ASTRAL compendium of protein structures are included. The alignments are constructed using pairwise structural alignments to generate residue equivalences that are then integrated into multiple alignments using sequence alignment tools. We describe the database and give examples showing how the automatically generated S4 alignments compare favourably to hand-crafted alignments. Available at: http://compbio.mds.qmw.ac.uk/S4.html.

Algorithms↗

An interactive bovine in silico SNP database (IBISS).

An interactive bovine in silico SNP (IBISS) database has been created through the clustering and aligning of bovine EST and mRNA sequences. Approximately 324,000 EST and mRNA sequences were clustered to produce 29,965 clusters (producing 48,679 consensus sequences) and 48,565 singletons. A SNP screening regime was placed on variations detected in the multiple sequence alignment files to determine which SNPs are more likely to be real rather than sequencing errors. A small subset of predicted SNPs was validated on a diverse set of bovine DNA samples using PCR amplification and sequencing. Fifty percent of the predicted SNPs in the "putative >1" category were polymorphic in the population sampled. The IBISS database represents more than just a SNP database; it is also a genomic database containing uniformly annotated predicted gene mRNA and protein sequences, gene structure, and genomic organization information.

Animals↗

Protein fold recognition by mapping predicted secondary structures.

A strategy is presented for protein fold recognition from secondary structure assignments (alpha-helix and beta-strand). The method can detect similarities between protein folds in the absence of sequence similarity. Secondary structure mapping first identifies all possible matches (maps) between a query string of secondary structures and the secondary structures of protein domains of known three-dimensional structure. The maps are then passed through a series of structural filters to remove those that do not obey simple rules of protein structure. The surviving maps are ranked by scores from the alignment of predicted and experimental accessibilities. Searches made with secondary structure assignments for a test set of 11 fold-families put the correct sequence-dissimilar fold in the first rank 8/11 times. With cross-validated predictions of secondary structure this drops to 4/11 which compares favourably with the widely used THREADER program (1/11). The structural class is correctly predicted 10/11 times by the method in contrast to 5/11 for THREADER. The new technique obtains comparable accuracy in the alignment of amino acid residues and secondary structure elements. Searches are also performed with published secondary structure predictions for the von-Willebrand factor type A domain, the proteasome 20 S alpha subunit and the phosphotyrosine interaction domain. These searches demonstrate how the method can find the correct fold for a protein from a carefully constructed secondary structure prediction, multiple sequence alignment and distant restraints. Scans with experimentally determined secondary structures and accessibility, recognise the correct fold with high alignment accuracy (86% on secondary structures). This suggests that the accuracy of mapping will improve alongside any improvements in the prediction of secondary structure or accessibility. Application to NMR structure determination is also discussed.

Algorithms↗

Molecular cloning, nucleotide sequencing, and expression of genes encoding alcohol dehydrogenases from the thermophile Thermoanaerobacter brockii and the mesophile Clostridium beijerinckii.

Proteins play a pivotal role in thermophily. Comparing the molecular properties of homologous proteins from thermophilic and mesophilic bacteria is important for understanding the mechanisms of microbial adaptation to extreme environments. The thermophile Thermoanaerobacter (Thermoanaerobium) brockii and the mesophile Clostridium beijerinckii contain an NADP(H)-linked, zinc-containing secondary alcohol dehydrogenase (TBADH and CBADH) showing a similarly broad substrate range. The structural genes encoding the TBADH and the CBADH were cloned, sequenced, and highly expressed in Escherichia coli. The coding sequences of the TB adh and the CB adh genes are, respectively, 1056 and 1053 nucleotides long. The TB adh gene encoded an amino acid sequence identical to that of the purified TBADH. Alignment of the deduced amino acid sequences of the TB and CB adh genes showed a 76% identity and a 86% similarity, and the two genes had a similar preference for codons with A or T in the third position. Multiple sequence alignment of ADHs from different sources revealed that two (Cys-46 and His-67) of the three ligands for the catalytic Zn atom of the horse-liver ADH are preserved in TBADH and CBADH. Both the TBADH and CBADH were homotetramers. The substrate specificities and thermostabilities of the TBADH and CBADH expressed inE. coli were identical to those of the enzymes isolated from T. brockii and C. beijerinckii, respectively. A comparison of the amino acid composition of the two ADHs suggests that the presence of eight additional proline residues in TBADH than in CBADH and the exchange of hydrophilic and large hydrophobic residues in CBADH for the small hydrophobic amino acids Pro, Ala, and Val in TBADH might contribute to the higher thermostability of the T. brockii enzyme.

Journal Article↗

Modeling Alternative Conformational States in CASP16.

The CASP16 Ensemble Prediction experiment assessed advances in methods for modeling proteins, nucleic acids, and their complexes in multiple conformational states. Targets included systems with experimental structures determined in two or three states, evaluated by direct comparison to experimental coordinates, as well as domain-linker-domain (D-L-D) targets assessed against statistical models from NMR and SAXS data. This paper focuses on the former class of multi-state targets. Ten ensembles were released as community challenges, including ligand-induced conformational changes, protein-DNA complexes, a trimeric protein, a stem-loop RNA, and multiple oligomeric states of a single RNA. For five targets, some groups produced reasonably accurate models of both reference states (best TM-score >0.75). However, with the exception of one protein-ligand complex (T1214), where an apo structure was available as a template, predictors generally failed to capture key structural details distinguishing the states. Overall, accuracy was significantly lower than for single-state targets in other CASP experiments. The most successful approaches generated multiple AlphaFold2 models using enhanced multiple sequence alignments and sampling protocols, followed by model quality based selection. While the AlphaFold3 server performed well on several targets, individual groups outperformed it in specific cases. By contrast, predictions for one protein-DNA complex, three RNA targets, and multiple oligomeric RNA states consistently fell short (TM-score <0.75). These results highlight both progress and persistent challenges in multi-state prediction. Despite recent advances, accurate modeling of conformational ensembles, particularly RNA and large multimeric assemblies, remains a critical frontier for structural biology.

AlphaFold2↗

Software for optimization of SNP and PCR-RFLP genotyping to discriminate many genomes with the fewest assays.

BACKGROUND: Microbial forensics is important in tracking the source of a pathogen, whether the disease is a naturally occurring outbreak or part of a criminal investigation. RESULTS: A method and SPR Opt (SNP and PCR-RFLP Optimization) software to perform a comprehensive, whole-genome analysis to forensically discriminate multiple sequences is presented. Tools for the optimization of forensic typing using Single Nucleotide Polymorphism (SNP) and PCR-Restriction Fragment Length Polymorphism (PCR-RFLP) analyses across multiple isolate sequences of a species are described. The PCR-RFLP analysis includes prediction and selection of optimal primers and restriction enzymes to enable maximum isolate discrimination based on sequence information. SPR Opt calculates all SNP or PCR-RFLP variations present in the sequences, groups them into haplotypes according to their co-segregation across those sequences, and performs combinatoric analyses to determine which sets of haplotypes provide maximal discrimination among all the input sequences. Those set combinations requiring that membership in the fewest haplotypes be queried (i.e. the fewest assays be performed) are found. These analyses highlight variable regions based on existing sequence data. These markers may be heterogeneous among unsequenced isolates as well, and thus may be useful for characterizing the relationships among unsequenced as well as sequenced isolates. The predictions are multi-locus. Analyses of mumps and SARS viruses are summarized. Phylogenetic trees created based on SNPs, PCR-RFLPs, and full genomes are compared for SARS virus, illustrating that purported phylogenies based only on SNP or PCR-RFLP variations do not match those based on multiple sequence alignment of the full genomes. CONCLUSION: This is the first software to optimize the selection of forensic markers to maximize information gained from the fewest assays, accepting whole or partial genome sequence data as input. As more sequence data becomes available for multiple strains and isolates of a species, automated, computational approaches such as those described here will be essential to make sense of large amounts of information, and to guide and optimize efforts in the laboratory. The software and source code for SPR Opt is publicly available and free for non-profit use at http://www.llnl.gov/IPandC/technology/software/softwaretitles/spropt.php.

Cluster Analysis↗

Prediction of protein structure by evaluation of sequence-structure fitness. Aligning sequences to contact profiles derived from three-dimensional structures.

The problem of protein structure prediction is formulated here as that of evaluating how well an amino acid sequence fits a hypothetical structure. The simplest and most complicated approaches, secondary structure prediction and all-atom free energy calculations, can be viewed as sequence-structure fitness problems. Here, an approach of intermediate complexity is described, which involves; (1) description of a protein structure in terms of contact interface vectors, with both intra-protein and protein-solvent contacts counted, (2) derivation of sequence preferences for 2 up to 29 contact interface types, (3) generation of numerous hypothetical model structures by placing the input sequence into a large set of known three-dimensional structures in all possible alignments, (4) evaluation of these models by summing the sequence preferences over all structural positions and (5) choice of predicted three-dimensional structure as that with the best sequence-structure fitness. Evolutionary information is incorporated by using position-dependent core weights derived from multiple sequence alignments. A number of tests of the method are performed: (1) evaluation of cyclic shifts of a sequence in its native structure; (2) alignment of a sequence in its native structure, allowing gaps; (3) alignment search with a sequence or sequence fragment in a database of structures; and (4) alignment search with a structure in a database of sequences. The main results are: (1) a native sequence can very well find its native structure among a large number of alternatives, in correct alignment; (2) substructures, such as (beta alpha)n units, can be detected in spite of very low sequence similarity; (3) remote homologous can be detected, with some dependence on the set of parameters used; (4) contact interface parameters are clearly superior to classical secondary structure parameters; (5) a simple interface description in terms of just two states, protein-protein and protein-water contacts, performs surprisingly well; (6) the use of core weights considerably improves accuracy in detection of remote homologues; (7) based on a sequence database search with a myoglobin contact profile, the C-terminal domain of a viral origin of replication binding protein is predicted to have an all-helical fold. The sequence-structure fitness concept is sufficiently general to accommodate a large variety of protein structure prediction methods, including new models of intermediate complexity currently being developed.

Amino Acid Sequence↗

Analysis and assessment of ab initio three-dimensional prediction, secondary structure, and contacts prediction.

CASP3 saw a substantial increase in the volume of ab initio 3D prediction data, with 507 datasets for fifteen selected targets and sixty-one groups participating. As with CASP2, methods ranged from computationally intensive strategies that attempt to recreate the physical and chemical forces involved in protein folding to the more recent knowledge-based approaches. These exploit information from the structure databases, extracting potentially similar fragments and/or distance constraints derived from multiple sequence alignments. The knowledge-based approaches generally gave more consistently successful predictions across the range of targets, particularly that of the Baker group (Bystroff and Baker, J Mol Biol 1998;281:565-577; Simons et al. Proteins Suppl 1999;3:171-176), which used a fragment library. In the secondary structure prediction category, the most successful approaches built on the concepts used in PHD (Rost et al. Comput Appl Biosci 1994;10:53-60), an accepted standard in this field. Like PHD, they exploit neural networks but have different strategies for incorporating multiple sequence data or position-dependent weight matrices for training the networks. Analysis of the contact data, for which only six groups participated, suggested that as yet this data provides a rather weak signal. However, in combination with other types of prediction data it can sometimes be a useful constraint for identifying the correct structure.

Animals↗

Relation between weight matrix and substitution matrix: motif search by similarity.

MOTIVATION: The discovery of patterns shared by several sequences that differ greatly is a basic task in sequence analysis, and still a challenge. Several methods have been developed for detecting patterns. Methods commonly used for motif search include the Gibbs sampler, Expectation-Maximization (EM) algorithm and some intuitive greedy approaches. One cannot guarantee the optimality of the result produced by the Gibbs sampler in a single run. The deterministic EM methods tend to get trapped by local optima. Solutions found by greedy approaches are rarely sufficiently good. RESULTS: A simple model describing a motif or a portion of local multiple sequence alignment is the weight matrix model, in which a motif is characterized with position-specific probabilities. Two substitution matrices are proposed to relate the sequence similarity with the weight matrix. Combining the substitution matrix and weight matrix, we examine three typical sets of protein sequences with increasing complexity. At a low score threshold for pair similarity, sliding windows are compared with a seed window to find the score sum, which provides a measure of statistical significance for multiple sequence comparison. Such a similarity analysis reveals many aspects of motifs. Blocks determined by similarity can be used to deduce a primary weight matrix or an improved substitution matrix. The algorithm successfully obtains the optimal solution for the test sets by just greedy iteration.

Algorithms↗