Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Inference of Cytochrome P450 Evolutionary History Using Structural and Physicochemical Metrics.

Cytochrome P450s are a superfamily of heme-binding monooxygenases involved with the detoxification of intrinsic and extrinsic toxins. They are near ubiquitous within biological domains and are found in all domains. Members of families within the superfamily are defined based on amino acid identity thresholds, with thresholds as low as 40% in some families. Relationships among Cytochrome P450 families have proven elusive due to sub-Twilight Zone interfamily identities (<30%) that result in poor multiple sequence alignment quality and thus low levels of support for downstream phylogenetic reconstructions. Despite the low identities, Cytochrome P450 structures are remarkably well conserved both within and among families. In such cases, structural phylogenetics has the potential to unveil elusive relationships because the selectively favored physicochemical properties giving rise to the structure and function of the proteins persist despite sequence-level divergence. Recently, in two separate publications, we demonstrated that by utilizing physicochemical vectors, dynamic time warping, and hierarchical clustering (PCDTW), large swaths of protein domain families and betacoronavirus receptor-binding domain clades were congruent with validated functional/structural relationships. These were important findings because anomalous sequence alignment-based maximum likelihood phylogenetic findings, which were not congruent with the known functional relationships, were resolved. That also validated the use of physicochemical vectors in making inferences about structural/functional homology. Additionally, it illuminated that the same methods might be applied to other protein families with relationships that are difficult to resolve from sequence data alone. Herein, we used Molecular Weight and Hydrophobicity Physicochemical Dynamic Time Warping (MWHP PCDTW) along with structural and sequence alignment-based phylogenetic methodologies to analyze all of the Cytochrome P450s found both in the high-fidelity Structural Classificaction of Proteins (SCOP) database and the reviewed sequences with both experimentally resolved and de novo predicted structures in the Protein Data Bank and the AlphaFold (AF) Protein Structure Database, respectively. We compared the resulting phylogenetic topologies and found that in some cases, structure-based methods may be less able to resolve random/convergent similarity than physicochemical and sequence-based methodologies. This finding agrees with previous findings that demonstrate the usefulness of physicochemical properties in resolving both random structural similarity and potentially convergent relationships.

Cytochrome P-450 Enzyme System

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Analysis of genetic diversity in cytotoxin-producing and non-cytotoxin-producing Helicobacter pylori strains.

Analysis of 32 Helicobacter pylori strains indicated a strong association between the presence of the cagA gene and a specific type of vacA allele found predominantly in cytotoxin-producing strains (P < .001). To determine whether tox+/CagA+ and tox-/CagA- strains constituted two separate noncombining lineages, sequences of the H. pylori ureC gene, cysS homologue, and the intergenic region between cysS and vacA were determined for multiple strains. The mean levels of nucleotide identity in the three regions were 96.7% +/- 0.5%, 95.0% +/- 1.0%, and 89.0% +/- 2.9%, respectively. Multiple sequence alignments and dendrograms based on these three regions failed to identify two clonal populations of organisms for which cagA and vacA genotypes were markers. The presence of a 63- to 64-bp insertion in the cysS-vacA intergenic region was unrelated to the vacA genotype of the strains. These data suggest that recombination between Helicobacter genomes may occur in vivo.

Base Sequence

Molecular cloning of the cDNA for the catalytic subunit of human DNA polymerase delta.

The cDNA of human DNA polymerase delta was cloned. The cDNA had a length of 3.5 kb and encoded a protein of 1107 amino acid residues with a calculated molecular mass of 124 kDa. Northern blot analysis showed that the cDNA hybridized to a mRNA of 3.4 kb. Monoclonal and polyclonal antibodies to the C-terminal 20 residues specifically immunoblotted the human pol delta catalytic polypeptide. A multiple sequence alignment was constructed. This showed that human pol delta is closely related to yeast pol delta and the herpes virus DNA polymerases. The levels of pol delta message were found to be induced concomitantly with DNA pol delta activity and DNA synthesis in serum restimulated proliferating IMR90 cultured cells. The human pol delta gene was localized to chromosome 19 by Southern blotting of EcoRI digested DNA from a panel of rodent/human cell hybrids.

Amino Acid Sequence

Molecular evolution of SRP cycle components: functional implications.

Signal recognition particle (SRP) is a cytoplasmic ribonucleoprotein that targets a subset of nascent presecretory proteins to the endoplasmic reticulum membrane. We have considered the SRP cycle from the perspective of molecular evolution, using recently determined sequences of genes or cDNAs encoding homologs of SRP (7SL) RNA, the Srp54 protein (Srp54p), and the alpha subunit of the SRP receptor (SR alpha) from a broad spectrum of organisms, together with the remaining five polypeptides of mammalian SRP. Our analysis provides insight into the significance of structural variation in SRP RNA and identifies novel conserved motifs in protein components of this pathway. The lack of congruence between an established phylogenetic tree and size variation in 7SL homologs implies the occurrence of several independent events that eliminated more than half the sequence content of this RNA during bacterial evolution. The apparently non-essential structures are domain I, a tRNA-like element that is constant in archaea, varies in size among eucaryotes, and is generally missing in bacteria, and domain III, a tightly base-paired hairpin that is present in all eucaryotic and archeal SRP RNAs but is invariably absent in bacteria. Based on both structural and functional considerations, we propose that the conserved core of SRP consists minimally of the 54 kDa signal sequence-binding protein complexed with the loosely base-paired domain IV helix of SRP RNA, and is also likely to contain a homolog of the Srp68 protein. Comparative sequence analysis of the methionine-rich M domains from a diverse array of Srp54p homologs reveals an extended region of amino acid identity that resembles a recently identified RNA recognition motif. Multiple sequence alignment of the G domains of Srp54p and SR alpha homologs indicates that these two polypeptides exhibit significant similarity even outside the four GTPase consensus motifs, including a block of nine contiguous amino acids in a location analogous to the binding site of the guanine nucleotide dissociation stimulator (GDS) for E. coli EF-Tu. The conservation of this sequence, in combination with the results of earlier genetic and biochemical studies of the SRP cycle, leads us to hypothesize that a component of the Srp68/72p heterodimer serves as the GDS for both Srp54p and SR alpha. Using an iterative alignment procedure, we demonstrate similarity between Srp68p and sequence motifs conserved among GDS proteins for small Ras-related GTPases. The conservation of SRP cycle components in organisms from all three major branches of the phylogenetic tree suggests that this pathway for protein export is of ancient evolutionary origin.

Amino Acid Sequence

The Genome Sequence DataBase version 1.0 (GSDB): from low pass sequences to complete genomes.

The Genome Sequence DataBase (GSDB) has completed its conversion to an improved relational database. The new database, GSDB 1.0, is fully operational and publicly available. Data contributions, including both original sequence submissions and community annotation, are being accomplished through the use of a graphical client-server interface tool, the GSDB Annotator, and via GIO (GSDB Input/Output) files. Data retrieval services are being provided through a new Web Query Tool and direct SQL. All methods of data contribution and data retrieval fully support the new data types that have been incorporated into GSDB, including discontiguous sequences, multiple sequence alignments, and community annotation.

Animals

Construction of a 3D model of cytochrome P450 2B4.

A three-dimensional structural model of rabbit phenobarbital-inducible cytochrome P450 2B4 (LM2) was constructed by homology modeling techniques previously developed for building and evaluating a 3D model of the cytochrome P450choP isozyme. Four templates with known crystal structures including cytochrome P450cam, terp, BM-3 and eryF were used in multiple sequence alignments and construction of the cytochrome P450 2B4 coordinates. The model was evaluated for its overall quality using available protein analysis programs and found to be satisfactory. The model structure was stable at room temperature during a 140 ps unconstrained full protein molecular dynamics simulation. A putative substrate access channel and binding site were identified. Two different substrates, benzphetamine and androstenedione, that are metabolized by cytochrome P450 2B4 with pronounced product specificity were docked into the putative binding site. Two orientations were found for each substrate that could lead to the observed preferred products. Using a geometric fit method three regions on the surface of the model cytochrome P450 structure were identified as possible sites for interaction with cytochrome b5, a redox partner of P450 2B4. Residues that may interact with the substrates and with cytochrome b5 have been identified and mutagenesis studies are currently in progress.

Amino Acid Sequence

Homology modelling and protein engineering strategy of subtilases, the family of subtilisin-like serine proteinases.

Subtilases are members of the family of subtilisin-like serine proteases. Presently, greater than 50 subtilases are known, greater than 40 of which with their complete amino acid sequences. We have compared these sequences and the available three-dimensional structures (subtilisin BPN', subtilisin Carlsberg, thermitase and proteinase K). The mature enzymes contain up to 1775 residues, with N-terminal catalytic domains ranging from 268 to 511 residues, and signal and/or activation-peptides ranging from 27 to 280 residues. Several members contain C-terminal extensions, relative to the subtilisins, which display additional properties such as sequence repeats, processing sites and membrane anchor segments. Multiple sequence alignment of the N-terminal catalytic domains allows the definition of two main classes of subtilases. A structurally conserved framework of 191 core residues has been defined from a comparison of the four known three-dimensional structures. Eighteen of these core residues are highly conserved, nine of which are glycines. While the alpha-helix and beta-sheet secondary structure elements show considerable sequence homology, this is less so for peptide loops that connect the core secondary structure elements. These loops can vary in length by greater than 150 residues. While the core three-dimensional structure is conserved, insertions and deletions are preferentially confined to surface loops. From the known three-dimensional structures various predictions are made for the other subtilases concerning essential conserved residues, allowable amino acid substitutions, disulphide bonds, Ca(2+)-binding sites, substrate-binding site residues, ionic and aromatic interactions, proteolytically susceptible surface loops, etc. These predictions form a basis for protein engineering of members of the subtilase family, for which no three-dimensional structure is known.

Amino Acid Sequence

Secondary structure prediction for modelling by homology.

An improved method of secondary structure prediction has been developed to aid the modelling of proteins by homology. Selected data from four published algorithms are scaled and combined as a weighted mean to produce consensus algorithms. Each consensus algorithm is used to predict the secondary structure of a protein homologous to the target protein and of known structure. By comparison of the predictions to the known structure, accuracy values are calculated and a consensus algorithm chosen as the optimum combination of the composite data for prediction of the homologous protein. This customized algorithm is then used to predict the secondary structure of the unknown protein. In this manner the secondary structure prediction is initially tuned to the required protein family before prediction of the target protein. The method improves statistical secondary structure prediction and can be incorporated into more comprehensive systems such as those involving consensus prediction from multiple sequence alignments. Thirty one proteins from five families were used to compare the new method to that of Garnier, Osguthorpe and Robson (GOR) and sequence alignment. The improvement over GOR is naturally dependent on the similarity of the homologous protein, varying from a mean of 3% to 7% with increasing alignment significance score.

Algorithms

Protein fold refinement: building models from idealized folds using motif constraints and multiple sequence data.

A general solution to the problem of directly incorporating data from multiple sequence alignments into the construction of molecular models was approached through the calculation of an estimated pairwise distance based on conserved hydrophobicity. A scaling method was developed that allowed the required bulk geometric properties of the estimated pair-wise distances (mean and mean squared) to mimic those expected in a globular protein. These properties were maintained independently of the composition, length, number or degree of conservation of the original sequences. Despite being a poor estimate for individual distances were found to be compatible with the native structure and could be weighted highly. While the estimated distances provided a general drive towards hydrophobic packing, more specific structure (including secondary structures and motifs) were induced by regularization towards an ideal form. These constraints were used to refine an outline starting structure (derived only from secondary structure axes) towards a compact form that was sufficiently protein-like for side chains to be added with almost no further adjustment of the alpha-carbon positions. This process allows rough folds based on abstract representations of protein architecture to be rapidly converted to a form where they can be analysed by the growing number of methods designed to assess molecular models.

Chemical Phenomena

Sequence divergence analysis for the prediction of seven-helix membrane protein structures: II. A 3-D model of human rhodopsin.

A three-dimensional (3-D) model of the transmembrane domain of human rhodopsin was predicted from the sequence divergence analysis of 42 sequences of rhodopsins and visual pigments without a template. The prediction steps include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, their secondary structure and packing shape in a helix bundle, prediction of side-chain conformations and structure refinement. The identification of the retinal binding site was assisted by its known covalent linkage with K296. The structural features of the predicted 3-D model are in good agreement with a low resolution electron density map of bovine rhodopsin and with residues in contact with retinal as determined experimentally.

Amino Acid Sequence

A simple procedure for assigning a sequence motif with an obscure pattern: application to the basic/helix-loop-helix motif.

We have developed a simple method to assign a sequence motif with an obscure pattern. Given a multiple sequence alignment for a region of protein that is known or strongly believed to have the same secondary and tertiary structures, the quantification method by principal component analysis is designed to find the regions most likely to have the same structure in a protein outside of the original set. The potential of this newly developed method was evaluated with reference to the known basic/helix-loop-helix (bHLH) motifs, and its characteristics were discussed with four obscure but well-defined motifs and compared with the other methods for searching sequence motifs. The method was also applied to assign the bHLH motif in Epstein-Barr virus nuclear antigen 1 (EBNA-1). This application revealed one candidate for the basic/helix 1 region and two candidates for the helix 2 region in the bHLH motif, within the region from amino acid residues 460 to 600, which is in good agreement with our previous experimental studies on the DNA binding region of EBNA-1. The basic/helix-loop-helix-loop-helix structure thus assigned suggests a function of EBNA-1 which is associated with both replication and transcription.

Amino Acid Sequence

Comparison of conservation within and between the Ser/Thr and Tyr protein kinase family: proposed model for the catalytic domain of the epidermal growth factor receptor.

The protein kinase family can be subdivided into two main groups based on their ability to phosphorylate Ser/Thr or Tyr substrates. In order to understand the basis of this functional difference, we have carried out a comparative analysis of sequence conservation within and between the Ser/Thr and Tyr protein kinases. A multiple sequence alignment of 86 protein kinase sequences was generated. For each position in the alignment we have computed the conservation of residue type in the Ser/Thr, in the Tyr and in both of the kinase subfamilies. To understand the structural and/or functional basis for the conservation, we have mapped these conservation properties onto the backbone of the recently determined structure of the cAMP-dependent Ser/Thr kinase. The results show that the kinase structure can be roughly segregated, based upon conservation, into three zones. The inner zone contains residues highly conserved in all the kinase family and describes the hydrophobic core of the enzyme together with residues essential for substrate and ATP binding and catalysis. The outer zone contains residues highly variable in all kinases and represents the solvent-exposed surface of the protein. The third zone is comprised of residues conserved in either the Ser/Thr or Tyr kinases or in both, but which are not conserved between them. These are sandwiched between the hydrophobic core and the solvent-exposed surface. In addition to analyzing overall conservation in the kinase family, we have also looked at conservation of its substrate and ATP binding sites. The ATP site is highly conserved throughout the kinases, whereas the substrate binding site is more variable. The active site contains several positions which differ between the Ser/Thr and Tyr kinases and may be responsible for discriminating between hydroxyl bearing side chains. Using this information we propose a model for Tyr substrate binding to the catalytic domain of the epidermal growth factor receptor (EGFR).

Amino Acid Sequence

A 3D model of the delta opioid receptor and ligand-receptor complexes.

A model for the 3D structure of the transmembrane domain of the delta opioid receptor was predicted from the sequence divergence analysis of 42 sequences of G-protein coupled peptide hormone receptors belonging to the opioid, somatostatin and angiotensin receptor families. No template was used in the prediction steps, which include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, prediction of their secondary structure, optimization of the packing shape in a helix bundle, prediction of side chain conformations and structural refinement. The general shape of the model is similar to that of the low resolution rhodopsin structure in that the TM3 and TM7 helices are most buried in the bundle and the TM1 and TM4 helices are most exposed to the lipid phase. An initial assessment of this model was made by determining to what extent a binding site identified using four structurally disparate high affinity delta opioid ligands was consistent with known mutational studies. With the assumption that the protonated amine nitrogen, a feature common to all delta opioid ligands, interacts with the highly conserved Asp127 in TM3, a pocket was found that satisfied the criteria of complementarity to the requirements for receptor recognition for these four diverse ligands, two delta selective antagonists (the fused ring naltrindole and the peptide Tyr-Tic-Phe-Phe-NH2) and the two agonists lofentanil and BW373U86 deduced from previous studies of the ligands alone. These ligands could be accommodated in a similar region of the receptor. The receptor binding site identified in the optimized complexes contained many residues in positions known to affect ligand binding in G-protein coupled receptors. These results also allowed identification of key residues as candidates for point mutations for further assessment and refinement of this model as well as preliminary indications of the requirements for recognition of this receptor.

Amino Acid Sequence

Structural and evolutionary relationships among the immunophilins: two ubiquitous families of peptidyl-prolyl cis-trans isomerases.

The immunophilins, protein receptors for the immunosuppressing drugs cyclosporin A and FK506 and related proteins from plants, fungi, and bacteria, have been analyzed structurally and evolutionarily. The cyclosporin A binding proteins (cyclophilins) represent one ubiquitous family of homologous proteins, and the FK506- and rapamycin-binding proteins (FKBPs) constitute a second, unrelated family. Multiple sequence alignments of members of each of these two protein families define the highly conserved residues that are likely to play important structural and functional roles, and mutations in representative members of these two families that abolish or alter function have been evaluated. FKBPs have undergone greater evolutionary divergence than the cyclophilins. Evolutionary trees were constructed using two distinct programs, and these trees establish the structural relationships that allow division of each of these families into subgroups. The results lead to the suggestion that several genes encoding isozymic forms of the FKBPs and possibly also of the cyclophilins existed in prokaryotes before the emergence of eukaryotes on earth and that representatives of these genes were transmitted to both kingdoms to give rise to current subfamilies of these proteins. By contrast, compartmentalization of both classes of immunophilins appears to have arisen independently in prokaryotes and eukaryotes, late in evolutionary history.

Amino Acid Isomerases

Genomic divergence of an HIV-2 from a German AIDS patient probably infected in Mali.

The complete nucleotide sequence of an HIV-2 isolate derived from a German AIDS patient with predominantly neurological symptoms is reported. The HIV-2BEN sequence is highly divergent from those of previously described HIV-2 and SIV strains. Evolutionary tree analysis of eight HIV-2 sequences reveals the existence of three HIV-2 groups. HIV-2BEN belongs to a group with two isolates from Ghana and The Gambia. Based on a comparison of HIV-2BEN with six HIV-2 isolates, SIVsmm and SIVmac, the variability of the structural env and gag proteins is similar within the HIV-2/SIVsmm/mac and HIV-1 groups. In contrast, the regulatory HIV-1 proteins are more highly conserved than those from HIV-2 strains. Multiple sequence alignments reveal that some domains of the envelope and regulatory proteins are well conserved among HIV-1, HIV-2/SIVsmm/mac, SIVagm and SIVmnd. The identification of conserved domains within the external glycoprotein could help to develop broadly active vaccines.

Acquired Immunodeficiency Syndrome

Phylogenetic analysis of gag genes from 70 international HIV-1 isolates provides evidence for multiple genotypes.

OBJECTIVE: To determine the extent of genetic variation among internationally collected HIV-1 isolates, to analyse phylogenetic relationships and the geographic distribution of different variants. DESIGN: Phylogenetic comparison of 70 HIV-1 isolates collected in 15 countries on four continents. METHODS: To sequence the complete gag genome of HIV-1 isolates, build multiple sequence alignments and construct phylogenetic trees using distance matrix methods and maximum parsimony algorithms. RESULTS: Phylogenetic tree analysis identified seven distinct genotypes. The seven genotypes were evident by both distance matrix methods and maximum parsimony analysis, and were strongly supported by bootstrap resampling of the data. The intra-genotypic gag distances averaged 7%, whereas the inter-genotypic distances averaged 14%. The geographic distribution of variants was complex. Some genotypes have apparently migrated to several continents and many areas harbor a mixture of genotypes. Related variants may cluster in certain areas, particularly isolates from a single city collected over a short time. CONCLUSIONS: The genetic variation among HIV-1 isolates is more extensive than previously appreciated. At least seven distinct HIV-1 genotypes can be identified. Diversification, migration and establishment of local, temporal 'blooms' of particular variants may all occur concomitantly.

Africa

Site-directed mutagenesis of the lipoate acetyltransferase of Escherichia coli.

Remote but significant similarities between the primary and predicted secondary structures of the chloramphenicol acetyltransferases (CAT) and lipoate acyltransferase subunits (LAT, E2) of the 2-oxo acid dehydrogenase complexes, have suggested that both types of enzyme may use similar catalytic mechanisms. Multiple sequence alignments for CAT and LAT have highlighted two conserved motifs that contain the active-site histidine and serine residues of CAT. Site-directed replacement of Ser550 in the E2p subunit (LAT) of the pyruvate dehydrogenase complex of Escherichia coli, deemed to be equivalent to the active-site Ser148 of CAT, supported the CAT-based model of LAT catalysis. The effects of other substitutions were also consistent with the predicted similarity in catalytic mechanism although specific details of active-site geometry may not be conserved.

Acetyltransferases