Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

Designing gene libraries from protein profiles for combinatorial protein experiments.

Protein combinatorial libraries provide new ways to probe the determinants of folding and to discover novel proteins. Such libraries are often constructed by expressing an ensemble of partially random gene sequences. Given the intractably large number of possible sequences, some limitation on diversity must be imposed. A non-uniform distribution of nucleotides can be used to reduce the number of possible sequences and encode peptide sequences having a predetermined set of amino acid probabilities at each residue position, i.e., the amino acid sequence profile. Such profiles can be determined by inspection, multiple sequence alignment or physically-based computational methods. Here we present a computational method that takes as input a desired sequence profile and calculates the individual nucleotide probabilities among partially random genes. The calculated gene library can be readily used in the context of standard DNA synthesis to generate a protein library with essentially the desired profile. The fidelity between the desired profile and the calculated one coded by these partially random genes is quantitatively evaluated using the linear correlation coefficient and a relative entropy, each of which provides a measure of profile agreement at each position of the sequence. On average, this method of identifying such codon frequencies performs as well or better than other methods with regard to fidelity to the original profile. Importantly, the method presented here provides much better yields of complete sequences that do not contain stop codons, a feature that is particularly important when all or large fractions of a gene are subject to combinatorial mutation.

Amino Acids↗

SNP2CAPS: a SNP and INDEL analysis tool for CAPS marker development.

With the influx of various SNP genotyping assays in recent years, there has been a need for an assay that is robust, yet cost effective, and could be performed using standard gel-based procedures. In this context, CAPS markers have been shown to meet these criteria. However, converting SNPs to CAPS markers can be a difficult process if done manually. In order to address this problem, we describe a computer program, SNP2CAPS, that facilitates the computational conversion of SNP markers into CAPS markers. 413 multiple aligned sequences derived from barley ESTs were analysed for the presence of polymorphisms in 235 distinct restriction sites. 282 (90%) of 314 alignments that contain sequence variation due to SNPs and InDels revealed at least one polymorphic restriction site. After reducing the number of restriction enzymes from 235 to 10, 31% of the polymorphic sites could still be detected. In order to demonstrate the usefulness of this tool for marker development, we experimentally validated some of the results predicted by SNP2CAPS.

Base Sequence↗

Do introns favor or avoid regions of amino acid conservation?

Are intron positions correlated with regions of high amino acid conservation? For a set of ancient conserved proteins, with intronless prokaryotic but intron-containing eukaryotic homologs, multiple sequence alignments identified residues invariant throughout evolution. Intron positions between codons show no preferences. However, introns lying after the first base of a codon prefer conserved regions, markedly in glycines. Because glycines are in excess in conserved regions, this behavior could reflect phase-one introns entering glycine residues randomly in the ancestral sequences. Examination of intron positions within codons of evolutionarily invariable amino acids showed that roughly 50% of these introns are bordered by guanines at both 5'- and 3'-ends, 25% have a G only before the intron, and 5% have a G only after the intron, whereas about 20% are bordered by nonguanine bases.

Alternative Splicing↗

Construction of a 3D model of cytochrome P450 2B4.

A three-dimensional structural model of rabbit phenobarbital-inducible cytochrome P450 2B4 (LM2) was constructed by homology modeling techniques previously developed for building and evaluating a 3D model of the cytochrome P450choP isozyme. Four templates with known crystal structures including cytochrome P450cam, terp, BM-3 and eryF were used in multiple sequence alignments and construction of the cytochrome P450 2B4 coordinates. The model was evaluated for its overall quality using available protein analysis programs and found to be satisfactory. The model structure was stable at room temperature during a 140 ps unconstrained full protein molecular dynamics simulation. A putative substrate access channel and binding site were identified. Two different substrates, benzphetamine and androstenedione, that are metabolized by cytochrome P450 2B4 with pronounced product specificity were docked into the putative binding site. Two orientations were found for each substrate that could lead to the observed preferred products. Using a geometric fit method three regions on the surface of the model cytochrome P450 structure were identified as possible sites for interaction with cytochrome b5, a redox partner of P450 2B4. Residues that may interact with the substrates and with cytochrome b5 have been identified and mutagenesis studies are currently in progress.

Amino Acid Sequence↗

Interactions underlying subunit association in cholinesterases.

Cholinesterases occur in a family of molecular forms, both as homo-oligomers of catalytic subunits, which can be either soluble, amphiphilic or lipid-anchored to the membrane; and hetero-oligomers of catalytic subunits and structural subunits. The structural subunits afford a method for precise localization of cholinesterases for specific function. A number of mutagenesis studies suggest that the C-terminal region of one alternatively spliced form of cholinesterase is involved in association of catalytic subunits into tetramers and in the association of these tetramers with structural subunits, however, there is currently no structural information about this region. In addition, none of the mutagenesis studies have clearly defined the residues important in these interactions. Here, multiple sequence alignment, structure prediction techniques and analysis of three-dimensional structural data are combined with a re-examination of mutagenesis and biochemical data. Three-dimensional models for the C-terminal region and for soluble tetrameric cholinesterase are proposed, and a set of rules governing subunit association are formulated. The simple model for association of catalytic and structural subunits presented is consistent with data for all known cholinesterases from species as divergent as nematode and man.

Amino Acid Sequence↗

3D modeling, ligand binding and activation studies of the cloned mouse delta, mu; and kappa opioid receptors.

Refined 3D models of the transmembrane domains of the cloned delta, mu and kappa opioid receptors belonging to the superfamily of G-protein coupled receptors (GPCRs) were constructed from a multiple sequence alignment using the alpha carbon template of rhodopsin recently reported. Other key steps in the procedure were relaxation of the 3D helix bundle by unconstrained energy optimization and assessment of the stability of the structure by performing unconstrained molecular dynamics simulations of the energy optimized structure. The results were stable ligand-free models of the TM domains of the three opioid receptors. The ligand-free delta receptor was then used to develop a systematic and reliable procedure to identify and assess putative binding sites that would be suitable for similar investigation of the other two receptors and GPCRs in general. To this end, a non-selective, 'universal' antagonist, naltrexone, and agonist, etorphine, were used as probes. These ligands were first docked in all sites of the model delta opioid receptor which were sterically accessible and to which the protonated amine of the ligands could be anchored to a complementary proton-accepting residue. Using these criteria, nine ligand-receptor complexes with different binding pockets were identified and refined by energy minimization. The properties of all these possible ligand-substrate complexes were then examined for consistency with known experimental results of mutations in both opioid and other GPCRs. Using this procedure, the lowest energy agonist-receptor and antagonist-receptor complexes consistent with these experimental results were identified. These complexes were then used to probe the mechanism of receptor activation by identifying differences in receptor conformation between the agonist and the antagonist complex during unconstrained dynamics simulation. The results lent support to a possible activation mechanism of the mouse delta opioid receptor similar to that recently proposed for several other GPCRs. They also allowed the selection of candidate sites for future mutagenesis experiments.

Amino Acid Sequence↗

Molecular modeling of the amyloid-beta-peptide using the homology to a fragment of triosephosphate isomerase that forms amyloid in vitro.

The main component of the amyloid senile plaques found in Alzheimer's brain is the amyloid-beta-peptide (A beta), a proteolytic product of a membrane precursor protein. Previous structural studies have found different conformations for the A beta peptide depending on the solvent and pH used. In general, they have suggested an alpha-helix conformation at the N-terminal domain and a beta-sheet conformation for the C-terminal domain. The structure of the complete A beta peptide (residues 1-40) solved by NMR has revealed that only helical structure is present in A beta. However, this result cannot explain the large beta-sheet A beta aggregates known to form amyloid under physiological conditions. Therefore, we investigated the structure of A beta by molecular modeling based on extensive homology using the Smith and Waterman algorithm implemented in the MPsrch program (Blitz server). The results showed a mean value of 23% identity with selected sequences. Since these values do not allow a clear homology to be established with a reference structure in order to perform molecular modeling studies, we searched for detailed homology. A 28% identity with an alpha/beta segment of a triosephosphate isomerase (TIM) from Culex tarralis with an unsolved three-dimensional structure was obtained. Then, multiple sequence alignment was performed considering A beta, TIM from C.tarralis and another five TIM sequences with known three-dimensional structures. We found a TIM segment with secondary structure elements in agreement with previous experimental data for A beta. Moreover, when a synthetic peptide from this TIM segment was studied in vitro, it was able to aggregate and to form amyloid fibrils, as established by Congo red binding and electron microscopy. The A beta model obtained was optimized by molecular dynamics considering ionizable side chains in order to simulate A beta in a neutral pH environment. We report here the structural implications of this study.

Algorithms↗

Analysis of structural and physico-chemical parameters involved in the specificity of binding between alpha-amylases and their inhibitors.

Enzyme-inhibitor specificity was studied for alpha-amylases and their inhibitors. We purified and cloned the cDNAs of two different alpha-amylase inhibitors from the common bean (Phaseolus vulgaris) and have recently cloned the cDNA of an alpha-amylase of the Mexican bean weevil (Zabrotes subfasciatus), which is inhibited by alpha-amylase inhibitor 2 but not by alpha-amylase inhibitor 1. The crystal structure of AI-1 complexed with pancreatic porcine alpha-amylase allowed us to model the structure of AI-2. The structure of Zabrotes subfasciatus alpha-amylase was modeled based on the crystal structure of Tenebrio molitor alpha-amylase. Pairwise AI-1 and AI-2 with PPA and ZSA complexes were modeled. For these complexes we first identified the interface forming residues. In addition, we identified the hydrogen bonds, ionic interactions and loss of hydrophobic surface area resulting from complex formation. The parameters we studied provide insight into the general scheme of binding, but fall short of explaining the specificity of the inhibition. We also introduce three new tools-software packages STING, HORNET and STINGPaint-which efficiently determine the interface forming residues and the ionic interaction data, the hydrogen bond net as well as aid in interpretation of multiple sequence alignment, respectively.

Amino Acid Sequence↗

Homology modeling and identification of serine 160 as nucleophile of the active site in a thermostable carboxylesterase from the archaeon Archaeoglobus fulgidus.

The hyperthermophilic Archaeon Archaeoglobus fulgidus has a gene (AF1763) which encodes a thermostable carboxylesterase belonging to the hormone-sensitive lipase (HSL)-like group of the esterase/lipase family. Based on secondary structure predictions and a secondary structure-driven multiple sequence alignment with remote homologous proteins of known three-dimensional structure, we previously hypothesized for this enzyme the alpha/beta-hydrolase fold typical of several lipases and esterases and identified Ser160, Asp 255 and His285 as the putative members of the catalytic triad. In this paper we report the building of a 3D model for this enzyme based on the structure of the homologous brefeldin A esterase from Bacillus subtilis whose structure has been recently elucidated. The model reveals the topological organization of the fold corroborating our predictions. As regarding the active-site residues, Ser160, Asp255 and His285 are located close each other at hydrogen bond distances. The catalytic role of Ser160 as the nucleophilic member of the triad is demonstrated by the [(3)H]diisopropylphosphofluoridate (DFP) active-site labeling and sequencing of a radioactive peptide containing the signature sequence GDSAGG.

Affinity Labels↗

On the spatial disposition of the fifth transmembrane helix and the structural integrity of the transmembrane binding site in the opioid and ORL1 G protein-coupled receptor family.

Evidence from statistical cluster analyses of a multiple sequence alignment of G protein-coupled receptor seven-helix folds supports the existence of structurally conserved transmembrane (TM) ligand binding sites in the opioid/opioid receptor-like (ORL1) and amine receptor families. Based on the expectation that functionally conserved regions in homologous proteins will display locally higher levels of sequence identity compared with global sequence similarities that pertain to the overall fold, this approach may have wider applications in functional genomics to annotate sequence data. Binding sites in models of the kappa-opioid receptor seven-helix bundle built from the rhodopsin templates of Baldwin et al. (1997) [J. Mol. Biol., 272, 144-164] and Herzyk and Hubbard (1998) [J. Mol. Biol., 281, 742-751] are compared. The Herzyk and Hubbard template is found to be in better accord with experimental studies of amine, opioid and rhodopsin receptors owing to the reduced physical separation of the extracellular parts of TM helices V and VI and differences in the rotational orientation of the N-terminal of helix V that reveal side chain accessibilities in the Baldwin et al. structure to be out of phase with relative alkylation rates of engineered cysteine residues in the TM binding site of the alpha(2A)-adrenergic receptor. TM helix V in the Baldwin et al. template has been remodelled with a different proline kink to satisfy experimental constraints. A recent proposal that rotation of helix V is associated with receptor activation is critically discussed.

Binding Sites↗

Homology modelling and molecular dynamics studies of human placental tissue protein 13 (galectin-13).

The primary structure of the newly sequence analysed placental tissue protein 13 (PP13) was highly homologous to several members of the beta-galactoside-binding S-type lectin (galectin) family. By homology modelling, the three-dimensional structure of PP13 was built based on high-resolution crystal structures of homologues and also their characteristic 'jellyroll' fold was found in the case of PP13. Our model has been deposited in the Brookhaven Protein Data Bank. By multiple sequence alignment and structure-based secondary structure prediction, we underlined the structural similarity of PP13 with its homologues. The secondary structure of PP13 was identical with 'proto-type' galectins consisting of a five- and a six-stranded beta-sheet, joined by two alpha-helices, and galectins' highly conserved carbohydrate-recognition domain (CRD) was also present in PP13. Of the eight consensus residues in the CRD, four identical and three conservatively substituted were shared by PP13. By docking simulations PP13 possessed sugar-binding activity with highest affinity to N-acetyllactosamine and lactose typical of most galectins. All ligands were docked into the putative CRD of PP13. Based on several lines of evidence discussed in this paper demonstrating that PP13 is a novel galectin, PP13 was also designated galectin-13. These computational results provide some new insights into the possible role and importance of PP13 in various processes of the human body and can be of help in the initial steps of further functional research.

Amino Acid Sequence↗

Predicting the transmembrane secondary structure of ligand-gated ion channels.

Recent mutational analyses of ligand-gated ion channels (LGICs) have demonstrated a plausible site of anesthetic action within their transmembrane domains. Although there is a consensus that the transmembrane domain is formed from four membrane-spanning segments, the secondary structure of these segments is not known. We utilized 10 state-of-the-art bioinformatics techniques to predict the transmembrane topology of the tetrameric regions within six members of the LGIC family that are relevant to anesthetic action. They are the human forms of the GABA alpha 1 receptor, the glycine alpha 1 receptor, the 5HT3 serotonin receptor, the nicotinic AChR alpha 4 and alpha 7 receptors and the Torpedo nAChR alpha 1 receptor. The algorithms utilized were HMMTOP, TMHMM, TMPred, PHDhtm, DAS, TMFinder, SOSUI, TMAP, MEMSAT and TOPPred2. The resulting predictions were superimposed on to a multiple sequence alignment of the six amino acid sequences created using the CLUSTAL W algorithm. There was a clear statistical consensus for the presence of four alpha helices in those regions experimentally thought to span the membrane. The consensus of 10 topology prediction techniques supports the hypothesis that the transmembrane subunits of the LGICs are tetrameric bundles of alpha helices.

Amino Acid Sequence↗

Homology modelling and protein engineering strategy of subtilases, the family of subtilisin-like serine proteinases.

Subtilases are members of the family of subtilisin-like serine proteases. Presently, greater than 50 subtilases are known, greater than 40 of which with their complete amino acid sequences. We have compared these sequences and the available three-dimensional structures (subtilisin BPN', subtilisin Carlsberg, thermitase and proteinase K). The mature enzymes contain up to 1775 residues, with N-terminal catalytic domains ranging from 268 to 511 residues, and signal and/or activation-peptides ranging from 27 to 280 residues. Several members contain C-terminal extensions, relative to the subtilisins, which display additional properties such as sequence repeats, processing sites and membrane anchor segments. Multiple sequence alignment of the N-terminal catalytic domains allows the definition of two main classes of subtilases. A structurally conserved framework of 191 core residues has been defined from a comparison of the four known three-dimensional structures. Eighteen of these core residues are highly conserved, nine of which are glycines. While the alpha-helix and beta-sheet secondary structure elements show considerable sequence homology, this is less so for peptide loops that connect the core secondary structure elements. These loops can vary in length by greater than 150 residues. While the core three-dimensional structure is conserved, insertions and deletions are preferentially confined to surface loops. From the known three-dimensional structures various predictions are made for the other subtilases concerning essential conserved residues, allowable amino acid substitutions, disulphide bonds, Ca(2+)-binding sites, substrate-binding site residues, ionic and aromatic interactions, proteolytically susceptible surface loops, etc. These predictions form a basis for protein engineering of members of the subtilase family, for which no three-dimensional structure is known.

Amino Acid Sequence↗

Secondary structure prediction for modelling by homology.

An improved method of secondary structure prediction has been developed to aid the modelling of proteins by homology. Selected data from four published algorithms are scaled and combined as a weighted mean to produce consensus algorithms. Each consensus algorithm is used to predict the secondary structure of a protein homologous to the target protein and of known structure. By comparison of the predictions to the known structure, accuracy values are calculated and a consensus algorithm chosen as the optimum combination of the composite data for prediction of the homologous protein. This customized algorithm is then used to predict the secondary structure of the unknown protein. In this manner the secondary structure prediction is initially tuned to the required protein family before prediction of the target protein. The method improves statistical secondary structure prediction and can be incorporated into more comprehensive systems such as those involving consensus prediction from multiple sequence alignments. Thirty one proteins from five families were used to compare the new method to that of Garnier, Osguthorpe and Robson (GOR) and sequence alignment. The improvement over GOR is naturally dependent on the similarity of the homologous protein, varying from a mean of 3% to 7% with increasing alignment significance score.

Algorithms↗

Protein fold refinement: building models from idealized folds using motif constraints and multiple sequence data.

A general solution to the problem of directly incorporating data from multiple sequence alignments into the construction of molecular models was approached through the calculation of an estimated pairwise distance based on conserved hydrophobicity. A scaling method was developed that allowed the required bulk geometric properties of the estimated pair-wise distances (mean and mean squared) to mimic those expected in a globular protein. These properties were maintained independently of the composition, length, number or degree of conservation of the original sequences. Despite being a poor estimate for individual distances were found to be compatible with the native structure and could be weighted highly. While the estimated distances provided a general drive towards hydrophobic packing, more specific structure (including secondary structures and motifs) were induced by regularization towards an ideal form. These constraints were used to refine an outline starting structure (derived only from secondary structure axes) towards a compact form that was sufficiently protein-like for side chains to be added with almost no further adjustment of the alpha-carbon positions. This process allows rough folds based on abstract representations of protein architecture to be rapidly converted to a form where they can be analysed by the growing number of methods designed to assess molecular models.

Chemical Phenomena↗

Sequence divergence analysis for the prediction of seven-helix membrane protein structures: II. A 3-D model of human rhodopsin.

A three-dimensional (3-D) model of the transmembrane domain of human rhodopsin was predicted from the sequence divergence analysis of 42 sequences of rhodopsins and visual pigments without a template. The prediction steps include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, their secondary structure and packing shape in a helix bundle, prediction of side-chain conformations and structure refinement. The identification of the retinal binding site was assisted by its known covalent linkage with K296. The structural features of the predicted 3-D model are in good agreement with a low resolution electron density map of bovine rhodopsin and with residues in contact with retinal as determined experimentally.

Amino Acid Sequence↗

A simple procedure for assigning a sequence motif with an obscure pattern: application to the basic/helix-loop-helix motif.

We have developed a simple method to assign a sequence motif with an obscure pattern. Given a multiple sequence alignment for a region of protein that is known or strongly believed to have the same secondary and tertiary structures, the quantification method by principal component analysis is designed to find the regions most likely to have the same structure in a protein outside of the original set. The potential of this newly developed method was evaluated with reference to the known basic/helix-loop-helix (bHLH) motifs, and its characteristics were discussed with four obscure but well-defined motifs and compared with the other methods for searching sequence motifs. The method was also applied to assign the bHLH motif in Epstein-Barr virus nuclear antigen 1 (EBNA-1). This application revealed one candidate for the basic/helix 1 region and two candidates for the helix 2 region in the bHLH motif, within the region from amino acid residues 460 to 600, which is in good agreement with our previous experimental studies on the DNA binding region of EBNA-1. The basic/helix-loop-helix-loop-helix structure thus assigned suggests a function of EBNA-1 which is associated with both replication and transcription.

Amino Acid Sequence↗

Comparison of conservation within and between the Ser/Thr and Tyr protein kinase family: proposed model for the catalytic domain of the epidermal growth factor receptor.

The protein kinase family can be subdivided into two main groups based on their ability to phosphorylate Ser/Thr or Tyr substrates. In order to understand the basis of this functional difference, we have carried out a comparative analysis of sequence conservation within and between the Ser/Thr and Tyr protein kinases. A multiple sequence alignment of 86 protein kinase sequences was generated. For each position in the alignment we have computed the conservation of residue type in the Ser/Thr, in the Tyr and in both of the kinase subfamilies. To understand the structural and/or functional basis for the conservation, we have mapped these conservation properties onto the backbone of the recently determined structure of the cAMP-dependent Ser/Thr kinase. The results show that the kinase structure can be roughly segregated, based upon conservation, into three zones. The inner zone contains residues highly conserved in all the kinase family and describes the hydrophobic core of the enzyme together with residues essential for substrate and ATP binding and catalysis. The outer zone contains residues highly variable in all kinases and represents the solvent-exposed surface of the protein. The third zone is comprised of residues conserved in either the Ser/Thr or Tyr kinases or in both, but which are not conserved between them. These are sandwiched between the hydrophobic core and the solvent-exposed surface. In addition to analyzing overall conservation in the kinase family, we have also looked at conservation of its substrate and ATP binding sites. The ATP site is highly conserved throughout the kinases, whereas the substrate binding site is more variable. The active site contains several positions which differ between the Ser/Thr and Tyr kinases and may be responsible for discriminating between hydroxyl bearing side chains. Using this information we propose a model for Tyr substrate binding to the catalytic domain of the epidermal growth factor receptor (EGFR).

Amino Acid Sequence↗