Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The evolution of rhodopsins and neurotransmitter receptors.

Rhodopsins share a limited number of amino acid identities with a variety of other integral membrane proteins. Most of these proteins have seven putative transmembrane segments and are likely to play a role in transmembrane signaling. We have undertaken a systematic series of comparisons of primary and secondary structure in order to clarify the functional and evolutionary significance of these sequence similarities. On the basis of consistently high similarity scores, we find that the most internally consistent definition of the rhodopsin gene family would include vertebrate rhodopsins, alpha- and beta-adrenergic receptors, M1 and M2 muscarinic acetylcholine receptors, substance K receptors, and insect rhodopsins, while excluding bacteriorhodopsin, the mass human oncogene, vertebrate and insect nicotinic acetylcholine receptors, and the yeast STE2 and STE3 peptide receptors. The rhodopsin gene family is highly diverged at the primary sequence level but has maintained a conserved secondary structure, including a previously unidentified hierarchy of transmembrane segment hydrophobicity. We have developed new computer algorithms for progressive multiple sequence alignment and the analysis of local conservation of protein domains, and we have used these algorithms to examine the phylogeny of the rhodopsin gene family and the changing domains of sequence conservation. The results show striking differences and similarities in the conserved domains in each of the three main branches of the rhodopsin gene family, and indicate that color vision arose independently in the lines of descent leading to modern humans and fruit flies.

Algorithms

Structural and functional relationships of human DNA polymerases.

A continuing theme of our laboratory has been the understanding of human DNA polymerases at the structural level. We have purified DNA polymerases delta, epsilon and alpha from human placenta. Monoclonal antibodies to these polymerases were isolated and used as tools to study their immunochemical relationships. These studies have shown that while DNA polymerases delta, epsilon and alpha are discrete proteins, they must share common structural features by virtue of the ability of several of our monoclonal antibodies to exhibit cross-reactivity. A second approach we have taken is the molecular cloning of human DNA polymerase delta and epsilon. We have cloned the DNA polymerase delta cDNA, and this has allowed us to compare its primary structure to those of human polymerase alpha and other members of this polymerase family. Multiple sequence alignments have revealed that human DNA polymerase delta is also closely related to the herpes virus family of DNA polymerases. In situ hybridization has shown that the human DNA polymerase delta gene is localized to chromosome 19 q13.3-q13.4. In order to further determine the functional regions of the DNA polymerase delta structure we are currently expressing human pol delta in E. coli and baculovirus systems. Other work in our laboratory is directed toward examining the expression of DNA polymerase delta during the cell cycle.

Amino Acid Sequence

Molecular Cloning, Recombinant Expression, and In Silico Structural Analysis of Cu/Zn-Superoxide Dismutase from Trachyspermum ammi.

Superoxide dismutase (SOD) is an essential antioxidant metalloenzyme that is critical for the cellular defense against oxidative damage, as it scavenges superoxide radicals and maintains the redox status. Cytosolic Cu/Zn-SOD is particularly important in the regulation of oxidative stress among different isoforms in higher plants. While Cu/Zn-SODs from several plant species have been characterized, molecular information is limited for Trachyspermum ammi, a medicinally important member of a family Apiaceae with antioxidant potential.In the present study, an integrated molecular and in silico approach has been taken to clone and analyze a Cu/Zn type SOD gene from T. ammi to get insight into its structural and evolutionary characteristics. PCR amplification yielded an open reading frame of 456 bp encoding a protein of 152 amino acids. Sequence analysis showed that plant Cu/Zn-SODs, especially those from Daucus carota, were highly similar to one another (about 90-95%).Multiple sequence alignment confirmed the presence of conserved catalytic motifs and metal-binding histidine residues, both of which are crucial for enzymatic function. Physicochemical analysis predicted the protein to be stable, hydrophilic and compatible with cytosolic localization. The analysis of secondary structure indicated a predominance of β-strands, consistent with the conserved β-barrel architecture of plant Cu/Zn-SODs.The three-dimensional structure was built by homology modeling using a closely related plant Cu/Zn-SOD template with high sequence identity. Structural validation demonstrated an acceptable stereochemical quality with 86.3% residues in the favored region of Ramachandran plot, satisfactory ERRAT and Verify3D scores, and a low RMSD value of 0.104 Å on structural superimposition. Phylogenetic analysis placed the enzyme in the Apiaceae lineage, suggesting evolutionary conservation among related plant species. In conclusion, this study presents the first molecular and structural characterization of Cu/Zn-SOD from T. ammi and confirms the existence of a conserved structural framework typical of plant Cu/Zn-SODs. These results provide a basis for further studies concerning recombinant expression, enzymatic validation and potential relevance in antioxidant and plant stress biology.

Cloning, Molecular

Sequence and localization of human NASP: conservation of a Xenopus histone-binding protein.

In this study the sequence and localization of human testicular NASP (nuclear autoantigenic sperm protein) are reported. NASP cDNA contains 2561 nt encoding a protein of 787 amino acids. The open reading frame contains 2446 nt followed by an ochre stop codon (TAA) and 104 nucleotides of untranslated sequence containing a poly(A) addition signal 10 bases upstream of the poly(A) tail. Northern blot analysis of human testis poly(A) mRNA indicates a message of approximately 3.2 kb. Multiple sequence alignment (MSA) analysis of the encoded human NASP amino acid sequence with the sequence for the Xenopus histone-binding protein N1/N2 and the rabbit NASP amino acid sequence demonstrates that the human sequence and the Xenopus sequence have extensive amino acid homology upstream of the rabbit initiation codon. Significantly, there is an 85% identity between the human and the rabbit NASP sequences when the alignment starts at the N-terminal of the rabbit sequence and at amino acid 101 of the human sequence. The nuclear translocation signal found in N1/N2 and rabbit NASP is completely conserved in human NASP. The first histone-binding domain of Xenopus is 70% identical and 90% similar to the human NASP domain. The second histone-binding domain of Xenopus is 48% identical and 71% similar to the human NASP domain. MSA analysis of the three sequences generated an unrooted ancestral tree with two branches, indicating that fewer amino acid changes have occurred between the Xenopus and the human sequences than between the Xenopus and the rabbit sequences. In the human testis, NASP is localized predominantly in primary spermatocytes and round spermatids. Spermatogonia, Sertoli cells, Leydig cells, peritubular cells, and other somatic cells do not stain. Human spermatozoa contain NASP in the acrosomal region. Following the acrosome reaction, some NASP remains in the equatorial and postacrosomal regions. We propose that mammalian testes and sperm contain a histone-binding protein which may play a role in regulating the early events of spermatogenesis.

Amino Acid Sequence

Prediction of domain organisation and secondary structure of thyroid peroxidase, a human autoantigen involved in destructive thyroiditis.

Organ specific autoimmune diseases are relatively common immunological disorders in man which include thyroid autoimmune disease, insulin-dependent diabetes mellitus and myasthenia gravis. The target autoantigens in some of these diseases have recently been characterised. In thyroid autoimmune disease this includes the key enzyme, thyroid peroxidase (TPO), which is involved in the generation of thyroid hormone. Structural knowledge about autoantigens such as thyroid peroxidase will allow a greater understanding of the interaction between autoantigens and the aberrant immune response, and facilitate the development of strategies for antigen-specific therapeutic manipulation. We report here a prediction of the secondary structure of thyroid peroxidase, together with the results of circular dichroic spectroscopy of a homologous purified enzyme. A combination of 3 secondary structure prediction programs has been used, following multiple sequence alignment, and TPO has been found to consist mainly of alpha-helical conformation, with little beta-sheet present. This structure prediction, together with knowledge of the exon-intron boundaries allows a model for the domain organisation of the TPO molecule to be proposed.

Amino Acid Sequence

A structure-derived sequence pattern for the detection of type I copper binding domains in distantly related proteins.

A structure-based approach to the definition of sequence patterns characteristic of protein domains is presented by example. The approach requires a multiple sequence alignment of a family (or set of related families) as well as at least one three-dimensional structure. The pattern derived does not merely summarize the information in the known sequences but attempts to generalize the pattern specifications based on structural insight. In this example, the pattern-driven database search identified correctly most of the known type I copper-binding domains and detected the presence of a homologous domain in a previously unknown case (CopA protein). The significance of these results is discussed.

Amino Acid Sequence

Conservation analysis and structure prediction of the SH2 family of phosphotyrosine binding domains.

Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.

Amino Acid Sequence

Analysis of insertions/deletions in protein structures.

An analysis of insertions and deletions (indels) occurring in a databank of multiple sequence alignments based on protein tertiary structure is reported. Indels prefer to be short (1 to 5 residues). The average intervening sequence length between them versus the percentage of residue identity in pairwise alignments shows an exponential behaviour, suggesting a stochastic process such that nearly every loop in an ancestral structure is a possible target for indels during evolution. The results also suggest a limit to the average size of indels accommodated by protein structures. The preferred indel conformations are reverse turn and coil as are the preferred conformations at the indel edges (N- and C-terminal sides). Interruptions in helices and strands were observed as very rare events.

Amino Acid Sequence

Structural analysis of homologous repeated domains in alpha-actinin and spectrin.

The amino acid sequences of chick and slime mould alpha-actinin each contain four repeats of approximately 122 residues. These repeats are homologous to the 18-22 repeats, each of approximately 106 residues, found in the alpha and beta subunits of spectrin and fodrin, and to the multiple repeats of approximately 110 residues found in the Duchenne muscular dystrophy protein (dystrophin). The repeats correspond to the elongated rod-like portion of these molecules. We present a multiple sequence alignment of 21 repeats from this superfamily (8 alpha-actinin and 13 spectrin/fodrin), based on optimal pairwise alignments, from which a characteristic consensus pattern of amino acid types is deduced. Trp 46 is invariant in all but one repeat, and physicochemical classes of amino acids are conserved at 25 other positions. Secondary structure prediction on both the alpha-actinin and spectrin repeats taken together with the distribution of proline residues in the sequences, strongly suggest that each repeated domain consists of a four-helix structure. Our predictions differ significantly from previous three-helix models based on analyses of fewer sequences. To determine possible interdomain regions, sites of limited proteolysis of the native chick alpha-actinin dimer were determined and located in the amino acid sequence. The majority of these sites were in corresponding positions in different repeats within a segment predicted as a long helix. We propose a model, consistent with the overall dimensions of the rod-like portions of the molecules, in which these long, probably interrupted helices, link adjacent domains.

Actinin

A proteolytically sensitive region common to several rat liver cytochromes P450: effect of cleavage on substrate binding.

Limited proteolysis of rat liver microsomes was used to probe the topography and structure of cytochrome P450 bound to the endoplasmic reticulum. Three cytochromes P450 from two families were examined. Monoclonal antibodies to cytochrome P450 forms 1A1, 2B1, and 2E1 were used to immunopurify these proteolyzed cytochromes P450 from microsomes from rats treated with 3-methylcholanthrene, phenobarbital, and acetone, respectively. Electrophoretic and immunoblot analysis of tryptic fragments revealed a highly sensitive cleavage site in all three cytochromes P450. N-Terminal sequencing was performed on the fragments after transfer onto poly(vinylidene difluoride) membranes and showed that this preferential cleavage site is at amino acid position 298 of P450 1A1, position 277 of P450 2B1, and position 278 of P450 2E1. Multiple sequence alignment revealed that these positions are at the amino terminal of a highly conserved region of these cytochromes P450. The important functional role implied by primary sequence conservation along with the proteolytic sensitivity at its amino terminal suggests that this region is a protein domain. Comparison with the known structure of the bacterial cytochrome P450cam predicts that this proteolytically sensitive site is within an interhelical turn region connected to the distal helix that partially encompasses the heme-containing active site. Substrate binding to the cleaved cytochromes P450 was examined in order to determine whether the newly added conformational freedom near the cleavage site functionally altered these cytochromes P450. Cleavage of P450 2B1 abolished benzphetamine binding, which indicates that the cleavage site contains an important structural determinant for binding this substrate. However, cleavage did not affect benzo[a]pyrene binding to P450 1A1.

Amino Acid Sequence

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking," Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana, contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a swapped domain topology; the other has no disulfide bonds and possesses two distinct, tandem domains. To explore the structural evolution of bicycle proteins, we attempted to predict bicycle protein structures with Alphafold2 (AF2) and other deep learning programs. While AF2 did not recover the two experimental structures using existing databases, it succeeded when provided with multiple sequence alignments (MSAs) of protein sequences from newly sequenced closely related species. Using this approach, we generated 2,400 high-confidence bicycle protein predictions from seven aphid species. While all aphid bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance. Our extension of AF2 with custom MSAs of proteins from closely related species provides a generalizable, powerful approach for predicting structures of rapidly evolving protein families.

Animals

Structural model of the nucleotide-binding conserved component of periplasmic permeases.

The amino acid sequences of 17 bacterial membrane proteins that are components of periplasmic permeases and function in the uptake of a variety of small molecules and ions are highly homologous to each other and contain sequence motifs characteristic of nucleotide-binding proteins. These proteins are known to bind ATP and are postulated to be the energy-coupling components of the permeases. Several medically important eukaryotic proteins, including the multidrug-resistance transporters and the protein encoded by the cystic fibrosis gene, are also homologous to this family. By multiple sequence alignment of these 17 proteins, the consensus sequence, secondary structure, and surface exposure were predicted. The secondary structural motifs that are conserved among nucleotide-binding proteins were identified in adenylate kinase, p21ras, and elongation factor Tu by superposition of their known tertiary structures. The equivalent secondary structural elements in the predicted conserved component were located. These, together with sequence information, served as guides for alignment with adenylate kinase. A model for the structure of the ATP-binding domain of the permease proteins is proposed by analogy to the adenylate kinase structure. The characteristics of several permease mutations and biochemical data lend support to the model.

Adenylate Kinase

GTPase domains of ras p21 oncogene protein and elongation factor Tu: analysis of three-dimensional structures, sequence families, and functional sites.

GTPase domains are functional and structural units employed as molecular switches in a variety of important cellular functions, such as growth control, protein biosynthesis, and membrane traffic. Amino acid sequences of more than 100 members of different subfamilies are known, but crystal structures of only mammalian ras p21 and bacterial elongation factor Tu have been determined. After optimal superposition of these remarkably similar structures, careful multiple sequence alignment, and calculation of residue-residue interactions, we analyzed the two subfamilies in terms of structural conservation, sequence conservation, and residue contact strength. There are three main results. (i) A structure-based alignment of p21 and elongation factor Tu. (ii) The definition of a common conserved structural core that may be useful as the basis of model building by homology of the three-dimensional structure of any GTPase domain. (iii) Identification of sequence regions, other than the effector loop and the nucleotide binding site, that may be involved in the functional cycle: they are loop L4, known to change conformation after GTP hydrolysis; helix alpha 2, especially Arg-73 and Met-67 in ras p21; loops L8 and L10, including ras p21 Arg-123, Lys-147, and Leu-120; and residues located spatially near the N and C termini. These regions are candidate sites for interaction either with the GTP/GDP exchange factor, with a GTPase-affected function, or with a molecule delivered to a destination site with the aid of the GTPase domain.

Amino Acid Sequence

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Inference of Cytochrome P450 Evolutionary History Using Structural and Physicochemical Metrics.

Cytochrome P450s are a superfamily of heme-binding monooxygenases involved with the detoxification of intrinsic and extrinsic toxins. They are near ubiquitous within biological domains and are found in all domains. Members of families within the superfamily are defined based on amino acid identity thresholds, with thresholds as low as 40% in some families. Relationships among Cytochrome P450 families have proven elusive due to sub-Twilight Zone interfamily identities (<30%) that result in poor multiple sequence alignment quality and thus low levels of support for downstream phylogenetic reconstructions. Despite the low identities, Cytochrome P450 structures are remarkably well conserved both within and among families. In such cases, structural phylogenetics has the potential to unveil elusive relationships because the selectively favored physicochemical properties giving rise to the structure and function of the proteins persist despite sequence-level divergence. Recently, in two separate publications, we demonstrated that by utilizing physicochemical vectors, dynamic time warping, and hierarchical clustering (PCDTW), large swaths of protein domain families and betacoronavirus receptor-binding domain clades were congruent with validated functional/structural relationships. These were important findings because anomalous sequence alignment-based maximum likelihood phylogenetic findings, which were not congruent with the known functional relationships, were resolved. That also validated the use of physicochemical vectors in making inferences about structural/functional homology. Additionally, it illuminated that the same methods might be applied to other protein families with relationships that are difficult to resolve from sequence data alone. Herein, we used Molecular Weight and Hydrophobicity Physicochemical Dynamic Time Warping (MWHP PCDTW) along with structural and sequence alignment-based phylogenetic methodologies to analyze all of the Cytochrome P450s found both in the high-fidelity Structural Classificaction of Proteins (SCOP) database and the reviewed sequences with both experimentally resolved and de novo predicted structures in the Protein Data Bank and the AlphaFold (AF) Protein Structure Database, respectively. We compared the resulting phylogenetic topologies and found that in some cases, structure-based methods may be less able to resolve random/convergent similarity than physicochemical and sequence-based methodologies. This finding agrees with previous findings that demonstrate the usefulness of physicochemical properties in resolving both random structural similarity and potentially convergent relationships.

Cytochrome P-450 Enzyme System

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Molecular cloning of the cDNA for the catalytic subunit of human DNA polymerase delta.

The cDNA of human DNA polymerase delta was cloned. The cDNA had a length of 3.5 kb and encoded a protein of 1107 amino acid residues with a calculated molecular mass of 124 kDa. Northern blot analysis showed that the cDNA hybridized to a mRNA of 3.4 kb. Monoclonal and polyclonal antibodies to the C-terminal 20 residues specifically immunoblotted the human pol delta catalytic polypeptide. A multiple sequence alignment was constructed. This showed that human pol delta is closely related to yeast pol delta and the herpes virus DNA polymerases. The levels of pol delta message were found to be induced concomitantly with DNA pol delta activity and DNA synthesis in serum restimulated proliferating IMR90 cultured cells. The human pol delta gene was localized to chromosome 19 by Southern blotting of EcoRI digested DNA from a panel of rodent/human cell hybrids.

Amino Acid Sequence

Homology modelling and protein engineering strategy of subtilases, the family of subtilisin-like serine proteinases.

Subtilases are members of the family of subtilisin-like serine proteases. Presently, greater than 50 subtilases are known, greater than 40 of which with their complete amino acid sequences. We have compared these sequences and the available three-dimensional structures (subtilisin BPN', subtilisin Carlsberg, thermitase and proteinase K). The mature enzymes contain up to 1775 residues, with N-terminal catalytic domains ranging from 268 to 511 residues, and signal and/or activation-peptides ranging from 27 to 280 residues. Several members contain C-terminal extensions, relative to the subtilisins, which display additional properties such as sequence repeats, processing sites and membrane anchor segments. Multiple sequence alignment of the N-terminal catalytic domains allows the definition of two main classes of subtilases. A structurally conserved framework of 191 core residues has been defined from a comparison of the four known three-dimensional structures. Eighteen of these core residues are highly conserved, nine of which are glycines. While the alpha-helix and beta-sheet secondary structure elements show considerable sequence homology, this is less so for peptide loops that connect the core secondary structure elements. These loops can vary in length by greater than 150 residues. While the core three-dimensional structure is conserved, insertions and deletions are preferentially confined to surface loops. From the known three-dimensional structures various predictions are made for the other subtilases concerning essential conserved residues, allowable amino acid substitutions, disulphide bonds, Ca(2+)-binding sites, substrate-binding site residues, ionic and aromatic interactions, proteolytically susceptible surface loops, etc. These predictions form a basis for protein engineering of members of the subtilase family, for which no three-dimensional structure is known.

Amino Acid Sequence