Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Molecular modelling of mammalian CYP2B isoforms and their interaction with substrates, inhibitors and redox partners.

1. The construction of three-dimensional models of CYP2B isozymes from rat (CYP2B1), rabbit (CYP2B4) and man (CYP2B6), based on a multiple sequence alignment with CYP102, a unique eukaryotic-like bacterial P450 (in terms of possessing an NADPH-dependent FAD- and FMN-containing oxidoreductase redox partner) of known crystal structure, is reported. 2. The enzyme models described are shown to be consistent with experimental evidence from site-directed mutagenesis studies, antibody recognition sites and amino acid residues identified as being associated with redox partner interactions, together with the location of a key serine residue (Ser-128) likely to be involved in protein kinaseA-mediated phosphorylation. 3. A substantial number of known substrates and inhibitors of CYP2B isozymes are shown to fit the putative active sites of the enzyme models in agreement with their reported position of metabolism or mode of inhibition respectively. In particular, there is complementarity between the characteristic non-planar geometries of CYP2B substrates and key groups in the enzymes' active sites. 4. Molecular modelling of CYP2B isozymes appears to rationalize a number of the reported findings from quantitative structure-activity relationship investigations on series of CYP2B substrates and inhibitors.

Amino Acid Sequence

A new family of plasma membrane polypeptides differentially regulated during plant development.

Two cDNAs encoding polypeptides identified in a tobacco leaf plasma membrane fraction prepared by phase partitioning were cloned. The deduced polypeptides, P16 and P17, exhibit a striking primary structure, similar to that of P19, a previously cloned plasma membrane polypeptide. Antibodies raised to the recombinant proteins were used to probe the cellular location of P16, P17 and P19 by means of western blotting of sucrose density-gradient fractions; all three polypeptides were found to be located solely at the plasma membrane. Furthermore, P19 antigen accumulated transiently at the time of floral induction while P16 and P17 antigens accumulated towards the end of the life-cycle. These results together with sequence database searches and multiple-sequence alignments suggest that we have identified a new family of plasma membrane polypeptides that (i) are putatively plant specific and (ii) are differentially regulated during plant development. These polypeptides are termed DREPPs for developmentally regulated plasma membrane polypeptides.

Amino Acid Sequence

HIV type 1 V3 domain serotyping and genotyping in Gauteng, Mpumalanga, KwaZulu-Natal, and Western Cape Provinces of South Africa.

More than 20.8 million people are living with HIV/AIDS in sub-Saharan Africa, with southern Africa the worst affected area and accounting for one of the fastest growing AIDS epidemics worldwide. Samples from 81 patients, including 25 from KwaZulu-Natal, 26 from Gauteng, 5 from Mpumalanga, and 25 from Western Cape Province, were serotyped using a competitive V3 peptide enzyme immunoassay (cPEIA). Viral RNA was also isolated from serum and the V3 region amplified by reverse transcriptase polymerase chain reaction (RT-PCR) to obtain a 240-bp product for direct sequencing of 29 samples. CLUSTAL W was used to make multiple sequence alignments. Distance calculation, tree construction methods, and bootstrap analysis were done using TREECON. Subtype C-like V3 loop sequences predominate in all provinces tested in South Africa. Discordant sero- and genotype results were observed in one patient only. The correlation between sero- and genotyping was 96% (24 of 25) in KwaZulu-Natal and 100% in Gauteng and Mpumalanga. In Western Cape Province 18% of patients were identified as sero/genotype B and 82% as sero/genotype C. Our data show that results of the second-generation V3 cPEIA correlated well with V3 sequencing and would be a rapid and affordable screening test to monitor the explosive southern African HIV-1 epidemic.

Adolescent

Statistical modeling, phylogenetic analysis and structure prediction of a protein splicing domain common to inteins and hedgehog proteins.

Inteins, introns spliced at the protein level, and the hedgehog family of proteins involved in eucaryotic development both undergo autocatalytic proteolysis. Here, a specific and sensitive hidden Markov model (HMM) of protein splicing domain shared by inteins and the hedgehog proteins has been trained and employed for further analysis. The HMM characterizes the common features of this domain including the position where a site-specific DNA endonuclease domain is inserted in the majority of the inteins. The HMM was used to identify several new putative inteins, such as that in the Methanococcus jannaschii klbA protein, and to generate a multiple sequence alignment of sequences possessing this domain. Phylogenetic analysis suggests that hedgehog proteins evolved from inteins. Secondary and tertiary structure predictions suggest that the domain has a structure similar to a beta-sandwich. Similarities between the serine protease cleavage mechanism and the protein splicing reaction mechanism are discussed. Examination of the locations of inteins indicates that they are not inserted randomly in an extein, but are often inserted at functionally important positions in the host proteins. A specific and sensitive HMM for a domain present in klbA proteins identified several additional bacterial and archaeal family members, and analysis of the site of insertion of the intein suggests residues that may be functionally important. This domain may play a role in formation of surface-associated protein complexes.

Algorithms

PHD--an automatic mail server for protein secondary structure prediction.

By the middle of 1993, > 30,000 protein sequences has been listed. For 1000 of these, the three-dimensional (tertiary) structure has been experimentally solved. Another 7000 can be modelled by homology. For the remaining 21,000 sequences, secondary structure prediction provides a rough estimate of structural features. Predictions in three states range between 35% (random) and 88% (homology modelling) overall accuracy. Using information about evolutionary conservation as contained in multiple sequence alignments, the secondary structure of 4700 protein sequences was predicted by the automatic e-mail server PHD. For proteins with at least one known homologue, the method has an expected overall three-state accuracy of 71.4% for proteins with at least one known homologue (evaluated on 126 unique protein chains).

Algorithms

SEQSEE: a comprehensive program suite for protein sequence analysis.

SEQSEE (SEQuence SEEker) is a multi-purpose, menu-driven suite of programs designed to provide a fully integrated, state-of-the-art package for the analysis and display of protein sequences and protein databases. It is currently configured to run on most UNIX-based machines including Sun, SGI and NeXT workstations with conversion to other architectures (e.g. Vax or Cray) being a relatively simple task. SEQSEE is capable of performing nearly all of the analytical and comparative tasks found in most comprehensive commercially available software packages. These include sequence/database searching, sequence retrieval, sequence entry and editing, statistical sequence analysis, multiple sequence alignment, flexible pattern matching, and secondary structure prediction. SEQSEE also integrates a number of unique databases which allow it to perform many additional functions such as structure-based sequence alignments and homology-based secondary structure prediction. Additional enhancements to many previously published algorithms have substantially improved the performance of SEQSEE over that found for most other commercial products. The source code, the documentation and all of the required databases for SEQSEE are freely available and may be obtained by anonymous ftp.

Algorithms

Comparison of side chain interactions performed by structurally equivalent residues in homologous protein structures.

The present work describes the computer program Hom-Bond, which allows to identify and compare intra-molecular interactions performed by side chain polar atoms as observed in a family of homologous protein structures with known and conserved 3-D conformation. For this purpose, the side chain to side chain and the side chain to main chain hydrogen bonds, the disulfide and the salt bridges are identified in each considered protein structure. Subsequently, the side chain interactions are displayed according to the multiple sequence alignment. The presented approach allows to easily identify bonds which are conserved in homologous proteins and to analyse rearrangements of the network of side chain interactions that characterize each protein structure.

Amino Acid Sequence

Sisyphus and prediction of protein structure.

The problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and 1D prediction have reopened the field. Mean-force-potentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue contacts (2D) can be detected by analysis of correlated mutations, albeit with low accuracy. Secondary structure, solvent accessibility and transmembrane helices (1D) can be predicted with significantly improved accuracy using multiple sequence alignments. Some of these new prediction methods have proven accurate and reliable enough to be useful in genome analysis, and in experimental structure determination. Moreover, the new generation of theoretical methods is increasingly influencing experiments in molecular biology.

Computers

Efficient discovery of conserved patterns using a pattern graph.

MOTIVATION: We have previously reported an algorithm for discovering patterns conserved in sets of related unaligned protein sequences. The algorithm was implemented in a program called Pratt. Pratt allows the user to define a class of patterns (e.g. the degree of ambiguity allowed and the length and number of gaps), and is then guaranteed to find the conserved patterns in this class scoring highest according to a defined fitness measure. In many cases, this version of Pratt was very efficient, but in other cases it was too time consuming to be applied. Hence, a more efficient algorithm was needed. RESULTS: In this paper, we describe a new and improved searching strategy that has two main advantages over the old strategy. First, it allows for easier integration with programs for multiple sequence alignment and data base search. Secondly, it makes it possible to use branch-and-bound search, and heuristics, to speed up the search. The new search strategy has been implemented in a new version of the Pratt program.

Algorithms

VHMPT: a graphical viewer and editor for helical membrane protein topologies.

MOTIVATION: Lacking structures resolved at atomic resolution, the great majority of membrane proteins have typically been depicted in a schematic two-dimensional (2D) topology consisting of putative transmembrane domains predicted from hydropathy plots. As more and more sequences of membrane proteins become available from genome projects, there is a need to automate the process of generating the schematic topology while allowing important information, such as the individual amino acid and the extent to which it is conserved in evolution, to be conveniently inspected. We addressed this need by developing a program called VHMPT. RESULTS: VHMPT (a graphical V iewer and editor for H elical line M embrane P rotein T opologies) can automatically generate a schematic 2D topology for a protein with transmembrane helices. Through an interactive graphical interface, VHMPT allows users to modify the layout of the generated topology, label specific amino acid or amino acid groups, and annotate with arrows and texts. Given a multiple sequence alignment file, VHMPT can also color code a normalized conservation score for each amino acid on the generated topology, allowing ready visual recognition of highly conserved (or variable) topological regions. VHMPT is written in Tcl/Tk and can run on platforms that have installed the Tcl/Tk interpreter. AVAILABILITY: The source code and a user manual for VHMPT are available for download at http://www. ibms.sinica.edu.tw/mjhwang/vhmpt. CONTACT: mjhwang@mail.ibms.sinica.edu.tw

Computational Biology

TOPAL: recombination detection in DNA and protein sequences.

UNLABELLED: TOPAL scans a multiple sequence alignment for evidence of recombinant sequences, prior to phylogenetic analysis. AVAILABILITY: The TOPAL package may be accessed at http://www.bioss.sari.ac.uk/grainne, and by anonymous ftp at ftp.bioss. sari.ac.uk in the directory pub/phylogeny/topal. CONTACT: grainne@bioss.sari.ac.uk

Computational Biology

Profile hidden Markov models.

The recent literature on profile hidden Markov model (profile HMM) methods and software is reviewed. Profile HMMs turn a multiple sequence alignment into a position-specific scoring system suitable for searching databases for remotely homologous sequences. Profile HMM analyses complement standard pairwise comparison methods for large-scale sequence analysis. Several software implementations and two large libraries of profile HMMs of common protein domains are available. HMM methods performed comparably to threading methods in the CASP2 structure prediction exercise.

Humans

Application of distance geometry to 3D visualization of sequence relationships.

SUMMARY: We describe the application of distance geometry methods to the three-dimensional visualization of sequence relationships, with examples for mumps virus SH gene cDNA and prion protein sequences. Sequence-sequence distance measures may be obtained from either a multiple sequence alignment or from sets of pairwise alignments. AVAILABILITY: C/Perl code and HTML/VRML files from http://www.nibsc.ac.uk/dg3dseq/

Algorithms

ENVIRON: a software package to compare protein three-dimensional structures with homologous sequences using local structural motifs.

This work presents a method to compare local clusters of interacting residues as observed in a known three-dimensional protein structure with corresponding clusters inferred from homologous protein sequences, assuming conserved protein folding. For this purpose the local environment of a selected residue in a known protein structure is defined as the ensemble of amino acids in contact with it in the folded state. Using a multiple sequence alignment to identify corresponding residues in homologous proteins, a detailed comparison can be performed between the local environment of a selected amino acid in the template protein structure and the expected local environments at the sets of equivalent residues, derived from the aligned protein sequences. The comparison makes it possible to detect conserved local features such as hydrogen bonding or complementarity in residue substitution. A global measure of environmental similarity is also defined, to search for conserved amino acid clusters subject to functional or structural constraints. The proposed approach is useful for investigating protein function as well as for site-directed mutagenesis experiments, where appropriate amino acid substitutions can be suggested by observing naturally occurring protein variants.

Algorithms

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Direct link between cytokine activity and a catalytic site for macrophage migration inhibitory factor.

Macrophage migration inhibitory factor (MIF) is a secreted protein that activates macrophages, neutrophils and T cells, and is implicated in sepsis, adult respiratory distress syndrome and rheumatoid arthritis. The mechanism of MIF function, however, is unknown. The three-dimensional structure of MIF is unlike that of any other cytokine, but bears striking resemblance to three microbial enzymes, two of which possess an N-terminal proline that serves as a catalytic base. Human MIF also possesses an N-terminal proline (Pro-1) that is invariant among all known homologues. Multiple sequence alignment of these MIF homologues reveals additional invariant residues that span the entire polypeptide but are in close proximity to the N-terminal proline in the folded protein. We find that p-hydroxyphenylpyruvate, a catalytic substrate of MIF, binds to the N-terminal region and interacts with Pro-1. Mutation of Pro-1 to a glycine substantially reduces the catalytic and cytokine activity of MIF. We suggest that the underlying biological activity of MIF may be based on an enzymatic reaction. The identification of the active site should facilitate the development of structure-based inhibitors.

Amino Acid Sequence

Inference of Cytochrome P450 Evolutionary History Using Structural and Physicochemical Metrics.

Cytochrome P450s are a superfamily of heme-binding monooxygenases involved with the detoxification of intrinsic and extrinsic toxins. They are near ubiquitous within biological domains and are found in all domains. Members of families within the superfamily are defined based on amino acid identity thresholds, with thresholds as low as 40% in some families. Relationships among Cytochrome P450 families have proven elusive due to sub-Twilight Zone interfamily identities (<30%) that result in poor multiple sequence alignment quality and thus low levels of support for downstream phylogenetic reconstructions. Despite the low identities, Cytochrome P450 structures are remarkably well conserved both within and among families. In such cases, structural phylogenetics has the potential to unveil elusive relationships because the selectively favored physicochemical properties giving rise to the structure and function of the proteins persist despite sequence-level divergence. Recently, in two separate publications, we demonstrated that by utilizing physicochemical vectors, dynamic time warping, and hierarchical clustering (PCDTW), large swaths of protein domain families and betacoronavirus receptor-binding domain clades were congruent with validated functional/structural relationships. These were important findings because anomalous sequence alignment-based maximum likelihood phylogenetic findings, which were not congruent with the known functional relationships, were resolved. That also validated the use of physicochemical vectors in making inferences about structural/functional homology. Additionally, it illuminated that the same methods might be applied to other protein families with relationships that are difficult to resolve from sequence data alone. Herein, we used Molecular Weight and Hydrophobicity Physicochemical Dynamic Time Warping (MWHP PCDTW) along with structural and sequence alignment-based phylogenetic methodologies to analyze all of the Cytochrome P450s found both in the high-fidelity Structural Classificaction of Proteins (SCOP) database and the reviewed sequences with both experimentally resolved and de novo predicted structures in the Protein Data Bank and the AlphaFold (AF) Protein Structure Database, respectively. We compared the resulting phylogenetic topologies and found that in some cases, structure-based methods may be less able to resolve random/convergent similarity than physicochemical and sequence-based methodologies. This finding agrees with previous findings that demonstrate the usefulness of physicochemical properties in resolving both random structural similarity and potentially convergent relationships.

Cytochrome P-450 Enzyme System

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95&#xa0;% or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics