Search PubMedSearch

PubMed · 8743695

Using CLUSTAL for multiple sequence alignments.

Abstract

We have tested CLUSTAL W in a wide variety of situations, and it is capable of handling some very difficult protein alignment problems. If the data set consists of enough closely related sequences so that the first alignments are accurate, then CLUSTAL W will usually find an alignment that is very close to ideal. Problems can still occur if the data set includes sequences of greatly different lengths or if some sequences include long regions that are impossible to align with the rest of the data set. Trying to balance the need for long insertions and deletions in some alignments with the need to avoid them in others is still a problem. The default values for our parameters were tested empirically using test cases of sets of globular proteins where some information as to the correct alignment was available. The parameter values may not be very appropriate with nonglobular proteins. We have argued that using one weight matrix and two gap penalties is too simplistic to be of general use in the most difficult cases. We have replaced these parameters with a large number of new parameters designed primarily to help encourage gaps in loop regions. Although these new parameters are largely heuristic in nature, they perform surprisingly well and are simple to implement. The underlying speed of the progressive alignment approach is not adversely affected. The disadvantage is that the parameter space is now huge; the number of possible combinations of parameters is more than can easily be examined by hand. We justify this by asking the user to treat CLUSTAL W as a data exploration tool rather than as a definitive analysis method. It is not sensible to automatically derive multiple alignments and to trust particular algorithms as being capable of always getting the correct answer. One must examine the alignments closely, especially in conjunction with the underlying phylogenetic tree (or estimate of it) and try varying some of the parameters. Outliers (sequences that have no close relatives) should be aligned carefully, as should fragments of sequences. The program will automatically delay the alignment of any sequences that are less than 40% identical to any others until all other sequences are aligned, but this can be set from a menu by the user. It may be useful to build up an alignment of closely related sequences first and to then add in the more distant relatives one at a time or in batches, using the profile alignments and weighting scheme described earlier and perhaps using a variety of parameter settings. We give one example using SH2 domains. SH2 domains are widespread in eukaryotic signalling proteins where they function in the recognition of phosphotyrosine-containing peptides. In the chapter by Bork and Gibson ([11], this volume), Blast and pattern/profile searches were used to extract the set of known SH2 domains and to search for new members. (Profiles used in database searches are conceptually very similar to the profiles used in CLUSTAL W: see the chapters [11] and [13] for profile search methods.) The profile searches detected SH2 domains in the JAK family of protein tyrosine kinases, which were thought not to contain SH2 domains. Although the JAK family SH2 domains are rather divergent, they have the necessary core structural residues as well as the critical positively charged residue that binds phosphotyrosine, leaving no doubt that they are bona fide SH2 domains. The five new JAK family SH2 domains were added sequentially to the existing alignment of 65 SH2 domains using the CLUSTAL W profile alignment option. Figure 6 shows part of the resulting alignment. Despite their divergent sequences, the new SH2 domains have been aligned nearly perfectly with the old set. No insertions were placed in the original SH2 domains. In this example, the profile alignment procedure has produced better results than a one-step full alignment of all 70 SH2 domains, and in considerably less time. (ABSTRACT TRUNCATED)

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

D G Higgins, J D Thompson, T J Gibson. 1996. Using CLUSTAL for multiple sequence alignments.. https://doi.org/10.1016/s0076-6879(96)66024-8

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A new point mutation in the HC-Pro of potato virus Y is involved in tobacco vein necrosis.

Tobacco vein necrosis (TVN) is a complex phenomenon regulated by different genetic determinants mapped in the HC-Pro protein (amino acids N330, K391 and E410) and in two regions of potato virus Y (PVY) genome, corresponding to the cytoplasmic inclusion (CI) protein and the nuclear inclusion protein a-protease (NIa-Pro), respectively. A new determinant of TVN was discovered in the MK isolate of PVY which, although carried the HC-Pro determinants associated to TVN, did not induce TVN. The HC-Pro open reading frame (ORF) of the necrotic infectious clone PVY N605 was replaced with that of the non-necrotic MK isolate, which differed only by one amino acid at position 392 (T392 instead of I392). The cDNA clone N605_MKHCPro inoculated in tobacco induced only weak mosaics at the systemic level, demostrating that the amino acid at position 392 is a new determinant for TVN. No significant difference in accumulation in tobacco was observed between N605 and N605_MKHCPro. Since phylogenetic analyses showed that the loss of necrosis in tobacco has occurred several times independently during PVY evolution, these repeated evolutions strongly suggest that tobacco necrosis is a costly trait in PVY.

Amino Acid Sequence

Residues on Adeno-associated Virus Capsid Lumen Dictate Interactions and Compatibility with the Assembly-Activating Protein.

The adeno-associated virus (AAV) serves as a broadly used vector system for in vivo gene delivery. The process of AAV capsid assembly remains poorly understood. The viral cofactor assembly-activating protein (AAP) is required for maximum AAV production and has multiple roles in capsid assembly, namely, trafficking of the structural proteins (VP) to the nuclear site of assembly, promoting the stability of VP against multiple degradation pathways, and facilitating stable interactions between VP monomers. The N-terminal 60 amino acids of AAP (AAPN) are essential for these functions. Presumably, AAP must physically interact with VP to execute its multiple functions, but the molecular nature of the AAP-VP interaction is not well understood. Here, we query how structurally related AAVs functionally engage AAP from AAV serotype 2 (AAP2) toward virion assembly. These studies led to the identification of key residues on the lumenal capsid surface that are important for AAP-VP and for VP-VP interactions. Replacing a cluster of glutamic acid residues with a glutamine-rich motif on the conserved VP beta-barrel structure of variants incompatible with AAP2 creates a gain-of-function mutant compatible with AAP2. Conversely, mutating positively charged residues within the hydrophobic region of AAP2 and conserved core domains within AAPN creates a gain-of-function AAP2 mutant that rescues assembly of the incompatible variant. Our results suggest a model for capsid assembly where surface charge/neutrality dictates an interaction between AAPN and the lumenal VP surface to nucleate capsid assembly.IMPORTANCE Efforts to engineer the AAV capsid to gain desirable properties for gene therapy (e.g., tropism, reduced immunogenicity, and higher potency) require that capsid modifications do not affect particle assembly. The relationship between VP and the cofactor that facilitates its assembly, AAP, is central to both assembly preservation and vector production. Understanding the requirements for this compatibility can inform manufacturing strategies to maximize production and reduce costs. Additionally, library-based approaches that simultaneously examine a large number of capsid variants would benefit from a universally functional AAP, which could hedge against overlooking variants with potentially valuable phenotypes that were lost during vector library production due to incompatibility with the cognate AAP. Studying interactions between the structural and nonstructural components of AAV enhances our fundamental knowledge of capsid assembly mechanisms and the protein-protein interactions required for productive assembly of the icosahedral capsid.

Amino Acid Sequence

Characterization of Class III Peroxidases from Switchgrass.

Class III peroxidases (CIIIPRX) catalyze the oxidation of monolignols, generate radicals, and ultimately lead to the formation of lignin. In general, CIIIPRX genes encode a large number of isozymes with ranges of in vitro substrate specificities. In order to elucidate the mode of substrate specificity of these enzymes, we characterized one of the CIIIPRXs (PviPRX9) from switchgrass (Panicum virgatum), a strategic plant for second-generation biofuels. The crystal structure, kinetic experiments, molecular docking, as well as expression patterns of PviPRX9 across multiple tissues and treatments, along with its levels of coexpression with the majority of genes in the monolignol biosynthesis pathway, revealed the function of PviPRX9 in lignification. Significantly, our study suggested that PviPRX9 has the ability to oxidize a broad range of phenylpropanoids with rather similar efficiencies, which reflects its role in the fortification of cell walls during normal growth and root development and in response to insect feeding. Based on the observed interactions of phenylpropanoids in the active site and analysis of kinetics, a catalytic mechanism involving two water molecules and residues histidine-42, arginine-38, and serine-71 was proposed. In addition, proline-138 and gluntamine-140 at the 137P-X-P-X140 motif, leucine-66, proline-67, and asparagine-176 may account for the broad substrate specificity of PviPRX9. Taken together, these observations shed new light on the function and catalysis of PviPRX9 and potentially benefit efforts to improve biomass conservation properties in bioenergy and forage crops.

Amino Acid Sequence