Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

A model for human cytochrome P450 2D6 based on homology modeling and NMR studies of substrate binding.

The cytochrome P450 responsible for the debrisoquine/sparteine polymorphism (P450 2D6) has been produced in large quantities by expression of a modified cDNA in baculovirus. A polyhistidine extension was incorporated at the C-terminus of the expressed protein, which, after purification of the protein on a nickel-agarose column, could be removed proteolytically by treatment with thrombin. Purified yields of P450 2D6 were 2.4 mg from 700 mL of cell culture. The protein had a greater than 90% heme content and was fully active, having no residual absorbance at 420 nm in the reduced CO complex. The quantities produced allowed direct study of the interaction of the substrate codeine with the enzyme by paramagnetic relaxation effects on the NMR spectrum of the substrate. Distances between the heme iron atom and substrate protons were calculated from these experiments, and the orientation of the substrate in the binding pocket was determined. This showed that codeine was bound with the methoxy group of the molecule closest to the heme iron (iron-methyl proton distance of 3.1 +/- 0.1 A), consistent with the observed O-demethylation to morphine. A model of the complex Of P450 2D6 with codeine was built from a multiple sequence and structure alignment of the known crystal structures for P450s, incorporating the experimental constraints derived from the NMR studies. This showed that the overall fold Of P450 2D6 is more similar to that of P450 BM3 than to either P450 cam or P450 terp. Codeine binds to P450 2D6 so that the methoxy group is directly above the A ring of the heme, while the basic nitrogen interacts with the carboxylate of aspartate 301.

Amino Acid Sequence↗

Modelling the 2-kinase domain of 6-phosphofructo-2-kinase/fructose-2,6-bisphosphatase on adenylate kinase.

Simultaneous multiple alignment of available sequences of the bifunctional enzyme 6-phosphofructo-2-kinase/fructose-2,6-bisphosphatase revealed several segments of conserved residues in the 2-kinase domain. The sequence of the kinase domain was also compared with proteins of known three-dimensional structure. No similarity was found between the kinase domain of 6-phosphofructo-2-kinase and 6-phosphofructo-1-kinase. This questions the modelling of the 2-kinase domain on bacterial 6-phosphofructo-1-kinase that has previously been proposed [Bazan, Fletterick and Pilkis (1989) Proc. Natl. Acad. Sci. U.S.A. 86, 9642-9646]. However, sequence similarities were found between the 2-kinase domain and several nucleotide-binding proteins, the most similar being adenylate kinase. A structural model of the 2-kinase domain based on adenylate kinase is proposed. It accommodates all the results of site-directed mutagenesis studies carried out to date on residues in the 2-kinase domain. It also allows residues potentially involved in catalysis and/or substrate binding to be predicted.

Adenylate Kinase↗

Mutagenesis and the molecular modeling of the rat angiotensin II receptor (AT1).

The molecular interaction involved in the ligand binding of the rat angiotensin II receptor (AT1A) was studied by site-directed mutagenesis and receptor model building. The three-dimensional structure of AT1A was constructed on the basis of a multiple amino acid sequence alignment of seven transmembrane domain receptors and angiotensin II receptors and after the beta 2 adrenergic receptor model built on the template of the bacteriorhodopsin structure. These data indicated that there are conserved residues that are actively involved in the receptor-ligand interaction. Eleven conserved residues in AT1, His166, Arg167, Glu173, His183, Glu185, Lys199, Trp253, His256, Phe259, Thr260, and Asp263, were targeted individually for site-directed mutation to Ala. Using COS-7 cells transiently expressing these mutated receptors, we found that the binding of angiotensin II was not affected in three of the mutations in the second extracellular loop, whereas the ligand binding affinity was greatly reduced in mutants Lys199-->Ala, Trp253-->Ala, Phe259-->Ala, Asp263-->Ala, and Arg167-->Ala. These amino acid residues appeared to provide binding sites for Ang II. The molecular modeling provided useful structural information for the peptide hormone receptor AT1A. Binding of EXP985, a nonpeptide angiotensin II antagonist, was found to be involved with Arg167 but not Lys199.

Amino Acid Sequence↗

Distinct Ca2+ binding properties of novel C2 domains of plant phospholipase dalpha and beta.

Of the isoforms of plant phospholipase D (PLD) that have been cloned and characterized, PLDalpha requires millimolar levels of Ca(2+) for optimal activity, whereas PLDbeta is most active at micromolar concentrations of Ca(2+). Multiple amino acid sequence alignments suggest that PLDalpha and PLDbeta both contain a Ca(2+)-dependent phospholipid-binding C2 domain near their N termini. In the present study, we expressed and characterized the putative C2 domains of PLDalpha and PLDbeta, designated PLDalpha C2 and PLDbeta C2, by CD spectroscopy, isothermal titration calorimetry, and phospholipid binding assay. Both PLD C2 domains displayed CD spectra consistent with anticipated major beta-sheet structures but underwent spectral changes upon binding Ca(2+); the magnitude was larger for PLDbeta C2. These conformational changes, not shown by any of the previously characterized C2 domains of animal origin, occurred at micromolar Ca(2+) concentrations for PLDbeta C2 but at millimolar levels of the cation for PLDalpha C2. PLDbeta C2 exhibited three Ca(2+)-binding sites: one with a dissociation constant (K(d)) of 0.8 microm and the other two with a K(d) of 24 micrometer. In contrast, isothermal titration calorimetry data of PLDalpha C2 were consistent with 1-3 low affinity Ca(2+)-binding sites with K(d) in the range of 590-470 micrometer. The thermodynamics of Ca(2+) binding markedly differed for the two C2 domains. Likewise, PLDbeta C2 bound phosphatidylcholine (PC), the substrate of PLD, in the presence of submillimolar Ca(2+) concentrations, whereas PLDalpha C2 did so only in the presence of millimolar levels of the metal ion. Both C2 domains bound phosphatidylinoistol 4,5-bisphosphate, a regulator of PC hydrolysis by PLD. However, added Ca(2+) displaced the bound phosphatidylinoistol 4,5-bisphosphate. Ca(2+) and PC binding properties of PLDalpha C2 and PLDbeta C2 follow a trend similar to the Ca(2+) requirements of the whole enzymes, PLDalpha and PLDbeta, for PC hydrolysis. Taken together, the results suggest that the C2 domains of PLDalpha and PLDbeta have novel structural features and serve as handles by which Ca(2+) differentially regulates the activities of the isoforms.

Amino Acid Sequence↗

Sequence, distance, and accessibility are determinants of 5'-end-directed cleavages by retroviral RNases H.

The RNase H activity of reverse transcriptase is essential for retroviral replication. RNA 5'-end-directed cleavages represent a form of RNase H activity that is carried out on RNA/DNA hybrids that contain a recessed RNA 5'-end. Previously, the distance from the RNA 5'-end has been considered the primary determinant for the location of these cleavages. Employing model hybrid substrates and the HIV-1 and Moloney murine leukemia virus reverse transcriptases, we demonstrate that cleavage sites correlate with specific sequences and that the distance from the RNA 5'-end determines the extent of cleavage. An alignment of sequences flanking multiple RNA 5'-end-directed cleavage sites reveals that both enzymes strongly prefer A or U at the +1 position and C or G at the -2 position, and additionally for HIV-1, A is disfavored at the -4 position. For both enzymes, 5'-end-directed cleavages occurred when sites were positioned between the 13th and 20th nucleotides from the RNA 5'-end, a distance termed the cleavage window. In examining the importance of accessibility to the RNA 5'-end, it was found that the extent of 5'-end-directed cleavages observed in substrates containing a free recessed RNA 5'-end was most comparable to substrates with a gap of two or three bases between the upstream and downstream RNAs. Together these finding demonstrate that the selection of 5'-end-directed cleavage sites by retroviral RNases H results from a combination of nucleotide sequence, permissible distance, and accessibility to the RNA 5'-end.

Base Sequence↗

Testing a molecular clock without an outgroup: derivations of induced priors on branch-length restrictions in a Bayesian framework.

We propose a Bayesian method for testing molecular clock hypotheses for use with aligned sequence data from multiple taxa. Our method utilizes a nonreversible nucleotide substitution model to avoid the necessity of specifying either a known tree relating the taxa or an outgroup for rooting the tree. We employ reversible jump Markov chain Monte Carlo to sample from the posterior distribution of the phylogenetic model parameters and conduct hypothesis testing using Bayes factors, the ratio of the posterior to prior odds of competing models. Here, the Bayes factors reflect the relative support of the sequence data for equal rates of evolutionary change between taxa versus unequal rates, averaged over all possible phylogenetic parameters, including the tree and root position. As the molecular clock model is a restriction of the more general unequal rates model, we use the Savage-Dickey ratio to estimate the Bayes factors. The Savage-Dickey ratio provides a convenient approach to calculating Bayes factors in favor of sharp hypotheses. Critical to calculating the Savage-Dickey ratio is a determination of the prior induced on the modeling restrictions. We demonstrate our method on a well-studied mtDNA sequence data set consisting of nine primates. We find strong support against a global molecular clock, but do find support for a local clock among the anthropoids. We provide mathematical derivations of the induced priors on branch length restrictions assuming equally likely trees. These derivations also have more general applicability to the examination of prior assumptions in Bayesian phylogenetics.

Animals↗

Approximations to profile score distributions.

Profiles, which are summaries of multiple alignments of a sequence family, are used to find new instances of the family in databases. In this paper, we study the maximum score M obtained when the profile is aligned without indels at all possible positions of a random sequence. The main theorem gives an approximation to the distribution function of M with an explicit bound on the error. This theorem implies that M has a limiting extreme value distribution.

Computer Simulation↗

Using substitution probabilities to improve position-specific scoring matrices.

Each column of amino acids in a multiple alignment of protein sequences can be represented as a vector of 20 amino acid counts. For alignment and searching applications, the count vector is an imperfect representation of a position, because the observed sequences are an incomplete sample of the full set of related sequences. One general solution to this problem is to model unobserved sequences by adding artificial 'pseudo-counts' to the observed counts. We introduce a simple method for computing pseudo-counts that combines the diversity observed in each alignment position with amino acid substitution probabilities. In extensive empirical tests, this position-based method out-performed other pseudo-count methods and was a substantial improvement over the traditional average score method used for constructing profiles.

Amino Acid Sequence↗

STRAP: editor for STRuctural Alignments of Proteins.

STRAP is a comfortable and extensible tool for the generation and refinement of multiple alignments of protein sequences. Various sequence ordered input file formats are supported. These are the SwissProt-,GenBank-, EMBL-, DSSP- PDB-, MSF-, and plain ASCII text format. The special feature of STRAP is the simple visualization of spatial distances C(alpha)-atoms within the alignment. Thus structural information can easily be incorporated into the sequence alignment and can guide the alignment process in cases of low sequence similarities. Further STRAP is able to manage huge alignments comprising a lot of sequences. The protein viewers and modeling programs INSIGHT, RASMOL and WEBMOL are embedded into STRAP. STRAP is written in JAVA: The well-documented source code can be adapted easily to special requirements. STRAP may become the basis for complex alignment tools in the future.

Cysteine Endopeptidases↗

Probabilistic divergence measures for detecting interspecies recombination.

This paper proposes a graphical method for detecting interspecies recombination in multiple alignments of DNA sequences. A fixed-size window is moved along a given DNA sequence alignment. For every position, the marginal posterior probability over tree topologies is determined by means of a Markov chain Monte Carlo simulation. Two probabilistic divergence measures are plotted along the alignment, and are used to identify recombinant regions. The method is compared with established detection methods on a set of synthetic benchmark sequences and two real-world DNA sequence alignments.

Computational Biology↗

HyPhy: hypothesis testing using phylogenies.

UNLABELLED: The HyPhypackage is designed to provide a flexible and unified platform for carrying out likelihood-based analyses on multiple alignments of molecular sequence data, with the emphasis on studies of rates and patterns of sequence evolution. AVAILABILITY: http://www.hyphy.org CONTACT: muse@stat.ncsu.edu SUPPLEMENTARY INFORMATION: HyPhydocumentation and tutorials are available at http://www.hyphy.org.

Algorithms↗

MAGOS: multiple alignment and modelling server.

UNLABELLED: MAGOS is a web server allowing automated protein modelling coupled to the creation of a hierarchical and annotated multiple alignment of complete sequences. MAGOS is designed for an interactive approach of structural information within the framework of the evolutionary relevance of mined and predicted sequence information. AVAILABILITY: The web server is freely available at http://pig-pbil.ibcp.fr/magos.

Algorithms↗

Variation in evolutionary processes at different codon positions.

Evolutionary studies commonly model single nucleotide substitutions and assume that they occur as independent draws from a unique probability distribution across the sequence studied. This assumption is violated for protein-coding sequences, and we consider modeling approaches where codon positions (CPs) are treated as separate categories of sites because within each category the assumption is more reasonable. Such "codon-position" models have been shown to explain the evolution of codon data better than homogenous models in previous studies. This paper examines the ways in which codon-position models outperform homogeneous models and characterizes the differences in estimates of model parameters across CPs. Using the PANDIT database of multiple species DNA sequence alignments, we quantify the differences in the evolutionary processes at the 3 CPs in a systematic and comprehensive manner, characterizing previously undescribed features of protein evolution. We relate our findings to the functional constraints imposed by the genetic code, protein function, and the types of mutation that cause synonymous and nonsynonymous codon changes. The results increase our understanding of selective constraints and could be incorporated into phylogenetic analyses or gene-finding techniques in the future. The methods used are extended to an overlapping reading frame data set, and we discover that overlapping reading frames do not necessarily cause more stringent evolutionary constraints.

Base Sequence↗

RPG: the Ribosomal Protein Gene database.

RPG (http://ribosome.miyazaki-med.ac.jp/) is a new database that provides detailed information about ribosomal protein (RP) genes. It contains data from humans and other organisms, including Drosophila melanogaster, Caenorhabditis elegans, Saccharo myces cerevisiae, Methanococcus jannaschii and Escherichia coli. Users can search the database by gene name and organism. Each record includes sequences (genomic, cDNA and amino acid sequences), intron/exon structures, genomic locations and information about orthologs. In addition, users can view and compare the gene structures of the above organisms and make multiple amino acid sequence alignments. RPG also provides information on small nucleolar RNAs (snoRNAs) that are encoded in the introns of RP genes.

Amino Acid Sequence↗

CaspR: a web server for automated molecular replacement using homology modelling.

Molecular replacement (MR) is the method of choice for X-ray crystallography structure determination when structural homologues are available in the Protein Data Bank (PDB). Although the success rate of MR decreases sharply when the sequence similarity between template and target proteins drops below 35% identical residues, it has been found that screening for MR solutions with a large number of different homology models may still produce a suitable solution where the original template failed. Here we present the web tool CaspR, implementing such a strategy in an automated manner. On input of experimental diffraction data, of the corresponding target sequence and of one or several potential templates, CaspR executes an optimized molecular replacement procedure using a combination of well-established stand-alone software tools. The protocol of model building and screening begins with the generation of multiple structure-sequence alignments produced with T-COFFEE, followed by homology model building using MODELLER, molecular replacement with AMoRe and model refinement based on CNS. As a result, CaspR provides a progress report in the form of hierarchically organized summary sheets that describe the different stages of the computation with an increasing level of detail. For the 10 highest-scoring potential solutions, pre-refined structures are made available for download in PDB format. Results already obtained with CaspR and reported on the web server suggest that such a strategy significantly increases the fraction of protein structures which may be solved by MR. Moreover, even in situations where standard MR yields a solution, pre-refined homology models produced by CaspR significantly reduce the time-consuming refinement process. We expect this automated procedure to have a significant impact on the throughput of large-scale structural genomics projects. CaspR is freely available at http://igs-server.cnrs-mrs.fr/Caspr/.

Escherichia coli Proteins↗

CAMPO, SCR_FIND and CHC_FIND: a suite of web tools for computational structural biology.

The identification of evolutionarily conserved features of protein structures can provide insights into their functional and structural properties. Three methods have been developed and implemented as WWW tools, CAMPO, SCR_FIND and CHC_FIND, to analyze evolutionarily conserved residues (ECRs), structurally conserved regions (SCRs) and conserved hydrophobic contacts (CHCs) in protein families and superfamilies, on the basis of their 3D structures and the homologous sequences available. The programs identify protein segments that conserve a similar main-chain conformation, compute residue-to-residue hydrophobic contacts involving only apolar atoms common to all the 3D structures analyzed and allow the identification of conserved amino-acid sites among protein structures and their homologous sequences. The programs also allow the visualization of SCRs, CHCs and ECRs directly on the superposed structures and their multiple structural and sequence alignments. Tools and tutorials explaining their usage are available at http://schubert.bio.uniroma1.it/SCR_FIND, http://schubert.bio.uniroma1.it/CHC_FIND and http://schubert.bio.uniroma1.it/CAMPO.

Computational Biology↗

SCOPPI: a structural classification of protein-protein interfaces.

SCOPPI, the structural classification of protein-protein interfaces, is a comprehensive database that classifies and annotates domain interactions derived from all known protein structures. SCOPPI applies SCOP domain definitions and a distance criterion to determine inter-domain interfaces. Using a novel method based on multiple sequence and structural alignments of SCOP families, SCOPPI presents a comprehensive geometrical classification of domain interfaces. Various interface characteristics such as number, type and position of interacting amino acids, conservation, interface size, and permanent or transient nature of the interaction are further provided. Proteins in SCOPPI are annotated with Gene Ontology terms, and the ontology can be used to quickly browse SCOPPI. Screenshots are available for every interface and its participating domains. Here, we describe contents and features of the web-based user interface as well as the underlying methods used to generate SCOPPI's data. In addition, we present a number of examples where SCOPPI becomes a useful tool to analyze viral mimicry of human interface binding sites, gene fusion events, conservation of interface residues and diversity of interface localizations. SCOPPI is available at http://www.scoppi.org.

Binding Sites↗