Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Prediction of domain organisation and secondary structure of thyroid peroxidase, a human autoantigen involved in destructive thyroiditis.

Organ specific autoimmune diseases are relatively common immunological disorders in man which include thyroid autoimmune disease, insulin-dependent diabetes mellitus and myasthenia gravis. The target autoantigens in some of these diseases have recently been characterised. In thyroid autoimmune disease this includes the key enzyme, thyroid peroxidase (TPO), which is involved in the generation of thyroid hormone. Structural knowledge about autoantigens such as thyroid peroxidase will allow a greater understanding of the interaction between autoantigens and the aberrant immune response, and facilitate the development of strategies for antigen-specific therapeutic manipulation. We report here a prediction of the secondary structure of thyroid peroxidase, together with the results of circular dichroic spectroscopy of a homologous purified enzyme. A combination of 3 secondary structure prediction programs has been used, following multiple sequence alignment, and TPO has been found to consist mainly of alpha-helical conformation, with little beta-sheet present. This structure prediction, together with knowledge of the exon-intron boundaries allows a model for the domain organisation of the TPO molecule to be proposed.

Amino Acid Sequence↗

A structure-derived sequence pattern for the detection of type I copper binding domains in distantly related proteins.

A structure-based approach to the definition of sequence patterns characteristic of protein domains is presented by example. The approach requires a multiple sequence alignment of a family (or set of related families) as well as at least one three-dimensional structure. The pattern derived does not merely summarize the information in the known sequences but attempts to generalize the pattern specifications based on structural insight. In this example, the pattern-driven database search identified correctly most of the known type I copper-binding domains and detected the presence of a homologous domain in a previously unknown case (CopA protein). The significance of these results is discussed.

Amino Acid Sequence↗

Conservation analysis and structure prediction of the SH2 family of phosphotyrosine binding domains.

Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.

Amino Acid Sequence↗

Pattern recognition and self-correcting distance geometry calculations applied to myohemerythrin.

A topological list, consisting of segments of regular secondary structures and a list of buried and solvent accessible residues, is automatically predicted from multiple aligned sequences in a protein family. This topological list is translated into geometric constraints for distance geometry calculation in torsion angle space. A new self-correcting distance geometry method detects and eliminates false distance constraints. In an application to the four-helix bundle protein, myohem-erythrin, the right-handed global fold was correctly reproduced with a root-mean-square deviation of 2.6 A, when the topological list was derived from the X-ray structure. A predicted topological list, coupled with constraints from the residues in the active site of myohemerythrin, predicted the correct fold with a root-mean-square deviation of 4 A for backbone atoms.

Amino Acid Sequence↗

Specificity of the cytochrome P-450 interaction with cytochrome b5.

The specificity of the interaction of cytochrome b5 with different forms of cytochrome P-450 was examined. Immunopurification of cytochromes P-450 1A1, 2B1 and 2E1 from rat liver microsomes resulted in co-purification of cytochrome b5 with cytochrome P-450 forms 2B1 and 2E1 but not 1A1. This specificity was evaluated in conjunction with multiple sequence alignment of the three cytochrome P-450s and a molecular model of the cytochrome P-450-cytochrome b5 complex [(1989) Biochemistry 28, 8201-8205]. These analyses suggest two basic residues in the arginine cluster region of P-450, which are present in P-450s 2B1 and 2E1 but are absent in P-450 1A1, as potential binding sites for cytochrome b5.

Amino Acid Sequence↗

Hydrophobic cluster analysis and secondary structure predictions revealed that major and minor structural subunits of K88-related adhesins of Escherichia coli share a common overall fold and differ structurally from other fimbrial subunits.

The structural relatedness of K88-related major and minor subunits was deduced from their sequences by hydrophobic cluster analysis (HCA) and secondary structure predictions produced by the profile neural network prediction program (PHD) on multiple sequence alignments. Although the weak residue identity between major and minor subunits is evidence of a high evolutionary distance, an overall structural similarity was observed In addition, clear amphipathic conformations were conserved in predicted secondary structure. On the basis of this predicted structural similarity, a schematic 2D model of ClpG subunit was developed.

Amino Acid Sequence↗

Identification of additional homologues of subunits VII and VIII of the ubiquinol-cytochrome c oxidoreductase enables definition of consensus sequences.

The Candida utilis QCR7 gene encoding subunit VII of the ubiquinol-cytochrome c oxidoreductase was isolated by functional complementation of the Saccharomyces cerevisiae subunit VII-null mutant. Several other subunit VII homologues as well as homologues for subunit VIII were identified by screening the GenBank database. Some of these homologues for subunit VII could only be identified as such using a consensus sequence that was derived from the multiple sequence alignment. Definition of the consensus should facilitate further analysis of structure/function relationships in this protein.

Amino Acid Sequence↗

Ribosomal protein L22 from Thermus thermophilus: sequencing, overexpression and crystallisation.

The gene for the ribosomal protein L22 from Thermus thermophilus has been sequenced and overexpressed in Escherichia coli. A multiple sequence alignment was carried out for all proteins of the L22 family reported so far. The recombinant protein was purified and crystallized. The crystals belong to the space group P2(1)2(1)2(1), with cell parameters of a = 32.6 A, b = 66.0 A, c = 67.8 A.

Amino Acid Sequence↗

Spectroscopic study of an HIV-1 capsid protein major homology region peptide analog.

The capsid (CA) domain of retroviral Gag proteins possesses one subdomain, the major homology region (MHR), which is conserved among nearly all avian and mammalian retroviruses. While it is known that the mutagenesis of residues in the MHR will impair virus infectivity, the precise structure and function of the MHR is not known. In order to obtain further information on the MHR, we have examined the structure of a synthetic peptide encompassing the MHR of human immunodeficiency virus type I (HIV-1) CA protein. Multiple sequence alignment and secondary structure prediction indicate that the peptide could form 50% alpha-helix and 10% beta-sheet. In addition, circular dichroism studies indicate that, in the presence of 50% trifluoroethanol (TFE), the peptide adopts an alpha-helical structure over half of its length. Further analysis by proton nuclear magnetic resonance spectroscopy suggests that the C-terminal portion of the MHR forms a helix in aqueous solution. Upon the addition of TFE, the position of the helix remains nearly constant, but the magnitude of the changes in H alpha chemical shifts of the residues indicate a more stable helix. These results suggest that a helical C-terminus of retroviral MHRs may be integral to the function of this region.

Amino Acid Sequence↗

Cloning and sequencing of cDNA clones encoding chicken lamins A and B1 and comparison of the primary structures of vertebrate A- and B-type lamins.

Nuclear lamins are intermediate-filament-type proteins forming a fibrillar meshwork underlying the inner nuclear membrane. The existence of multiple isoforms of lamin proteins in vertebrates is believed to reflect functional specializations during cell division and differentiation. Although biochemical criteria may be used to classify many lamin isoforms into A- and B-type subfamilies, the structural features distinguishing the members of these subfamilies remain to be characterized fully. Here, we report the complete primary structures of chicken lamins A and B1, as they are deduced from cloned cDNAs; in the accompanying paper we present the complete sequence of lamin B2, a second avian B-type lamin. Comparisons of the chicken lamin sequences with each other and with those of other lamins allow us to establish structural features that are common to members of both subfamilies. Conversely, multiple sequence alignments make it possible to identify a number of structural motifs that clearly differentiate B-type lamins from A-type lamins. With this information at hand, we attempt to correlate different biochemical properties of A- and B-type lamins with the presence or absence of specific sequence motifs.

Amino Acid Sequence↗

Analysis of insertions/deletions in protein structures.

An analysis of insertions and deletions (indels) occurring in a databank of multiple sequence alignments based on protein tertiary structure is reported. Indels prefer to be short (1 to 5 residues). The average intervening sequence length between them versus the percentage of residue identity in pairwise alignments shows an exponential behaviour, suggesting a stochastic process such that nearly every loop in an ancestral structure is a possible target for indels during evolution. The results also suggest a limit to the average size of indels accommodated by protein structures. The preferred indel conformations are reverse turn and coil as are the preferred conformations at the indel edges (N- and C-terminal sides). Interruptions in helices and strands were observed as very rare events.

Amino Acid Sequence↗

Position-based sequence weights.

Sequence weighting methods have been used to reduce redundancy and emphasize diversity in multiple sequence alignment and searching applications. Each of these methods is based on a notion of distance between a sequence and an ancestral or generalized sequence. We describe a different approach, which bases weights on the diversity observed at each position in the alignment, rather than on a sequence distance measure. These position-based weights make minimal assumptions, are simple to compute, and perform well in comprehensive evaluations.

Amino Acid Sequence↗

Sequence analysis of steroid- and prostaglandin-metabolizing enzymes: application to understanding catalysis.

Amino acid sequence comparisons have revealed that mammalian 11 beta-hydroxysteroid and 17 beta-hydroxysteroid dehydrogenases and bacterial 3 alpha, 20 beta- and 3 beta-hydroxysteroid dehydrogenases are homologs; that is, these enzymes are descended from a common ancestor. These steroid dehydrogenases are also homologous to human 15-hydroxyprostaglandin dehydrogenase and to proteins found in Rhizobia, bacteria that form nitrogen-fixing nodules in the roots of legumes. We constructed a multiple sequence alignment of these proteins, which, when combined with the recently determined tertiary structure of Streptomyces hydrogenans 3 alpha, 20 beta-hydroxysteroid dehydrogenase and a homologous enzyme, rat dihydropteridine reductase, identifies segments and residues that are likely to be structurally important in the functioning of these enzymes especially regarding specificity for NADPH and NADH.

Amino Acid Sequence↗

Sequence conservation and correlation measures in protein structure prediction.

The rapid elucidation of protein sequences has allowed multiple sequence alignments to be calculated for a wide variety of proteins. Such alignments reveal positions that exhibit amino acid conservation--either of specific chemical groups in active and binding sites or of the more chemically inert hydrophobic residues that contribute to the protein core. The latter can provide constraints on the position of the protein chain and any local periodicity can suggest the type of secondary structure. Conservation measures, however, cannot provide specific pairwise packing information (each conserved hydrophobic position might pack against any other). However, if correlated changes between positions were observed then specific pairs of residue could be identified as interacting and therefore probably spatially adjacent. Most 'observations' of correlated changes have been anecdotal and of the few systematic studies that have been made, most have mistakenly incorporated a strong bias towards selecting conserved positions. When the conservation effect is separated (as best as possible) then little correlation signal remains to help identify adjacent positions.

Animals↗

Toward the unification of sequence and structural data for identification of structural and functional constraints.

The identification and characterization of local residue patterns or conserved segments shared by a set of biopolymers has provided a number of insights in molecular biology. Biopolymer sequences are observations from macro molecules that share common structural or function features. The approach taken here rests on the notion that information may be most efficiently extracted from these observations through the use of a model that faithfully represents macro-molecular characteristics. Accordingly, our efforts are focused on statistical models which attempt to capture central features of protein structure, function, and change. Here the assumptions that underlie two new methods for the analysis of protein sequence data are explicitly delineated. (1) Threading of a sequence through structural motifs seeks to determine if a protein sequence fits a known protein structure. The assumptions delineated here also generally apply to other contact based threading methods that have been recently described. (2) Multiple sequence alignment via the Gibbs sampling algorithm seeks to identify position specific empirical free energy models for residue sites in common motifs and simultaneously the align sequence observations form these motifs.

Algorithms↗

Structural analysis of homologous repeated domains in alpha-actinin and spectrin.

The amino acid sequences of chick and slime mould alpha-actinin each contain four repeats of approximately 122 residues. These repeats are homologous to the 18-22 repeats, each of approximately 106 residues, found in the alpha and beta subunits of spectrin and fodrin, and to the multiple repeats of approximately 110 residues found in the Duchenne muscular dystrophy protein (dystrophin). The repeats correspond to the elongated rod-like portion of these molecules. We present a multiple sequence alignment of 21 repeats from this superfamily (8 alpha-actinin and 13 spectrin/fodrin), based on optimal pairwise alignments, from which a characteristic consensus pattern of amino acid types is deduced. Trp 46 is invariant in all but one repeat, and physicochemical classes of amino acids are conserved at 25 other positions. Secondary structure prediction on both the alpha-actinin and spectrin repeats taken together with the distribution of proline residues in the sequences, strongly suggest that each repeated domain consists of a four-helix structure. Our predictions differ significantly from previous three-helix models based on analyses of fewer sequences. To determine possible interdomain regions, sites of limited proteolysis of the native chick alpha-actinin dimer were determined and located in the amino acid sequence. The majority of these sites were in corresponding positions in different repeats within a segment predicted as a long helix. We propose a model, consistent with the overall dimensions of the rod-like portions of the molecules, in which these long, probably interrupted helices, link adjacent domains.

Actinin↗

Comparative molecular modelling of the Fas-ligand and other members of the TNF family.

A number of proteins with significant similarity to the tumour necrosis factor (TNF) have been identified over the last years. Upon interaction with their cognate receptor (members of the TNF-receptor family), all members of this protein family induce either cell death or proliferation/differentiation of the receptor-bearing cells. One of the last identified members of the TNF family is the apoptosis-inducing ligand of the Fas-receptor, termed Fas-ligand (FasL). Here we report the cloning and sequencing of the mouse cDNA for the FasL. Using knowledge-based protein modelling, we demonstrate that all members of the TNF family form trimeric complexes, and define the residues located at the subunit interfaces. The resulting structurally corrected multiple sequence alignment allows the identification of residues potentially involved in receptor recognition, and should help design mutagenesis experiments for structure-function relationship studies.

Amino Acid Sequence↗

Identification of Zucchini yellow mosaic potyvirus by RT-PCR and analysis of sequence variability.

A reverse transcription-polymerase chain reaction (RT-PCR) method was used to identify Zucchini yellow mosaic virus (ZYMV) in leaves of infected cucurbits. Oligonucleotide primers which annealed to regions in the nuclear inclusion body (NIb) and the coat protein (CP) genes, generated a 300-bp product from ZYMV and also from the closely related watermelon mosaic virus type 2 (WMV-2). However, no product was obtained from papaya ringspot potyvirus which also infects cucurbits. ZYMV and WMV-2 were differentiated using a third primer which was complementary to a sequence in the 3'-untranslated region; a 1186-bp amplified product was obtained for ZYMV only. Nucleotide sequence analysis of the 300-bp fragments of Australian ZYMV and WMV-2 strains revealed 93.7-100% sequence identity between ZYMV strains. Multiple sequence alignments indicated that the nucleotide sequence which codes for the N-terminus of the CP was 74-100% identical for different isolates of ZYMV. The Australian isolate of WMV-2 was 43-46% identical to all isolates of ZYMV and was 84.6% identical to a Florida isolate of WMV-2.

Amino Acid Sequence↗