Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Analysis and assessment of ab initio three-dimensional prediction, secondary structure, and contacts prediction.

CASP3 saw a substantial increase in the volume of ab initio 3D prediction data, with 507 datasets for fifteen selected targets and sixty-one groups participating. As with CASP2, methods ranged from computationally intensive strategies that attempt to recreate the physical and chemical forces involved in protein folding to the more recent knowledge-based approaches. These exploit information from the structure databases, extracting potentially similar fragments and/or distance constraints derived from multiple sequence alignments. The knowledge-based approaches generally gave more consistently successful predictions across the range of targets, particularly that of the Baker group (Bystroff and Baker, J Mol Biol 1998;281:565-577; Simons et al. Proteins Suppl 1999;3:171-176), which used a fragment library. In the secondary structure prediction category, the most successful approaches built on the concepts used in PHD (Rost et al. Comput Appl Biosci 1994;10:53-60), an accepted standard in this field. Like PHD, they exploit neural networks but have different strategies for incorporating multiple sequence data or position-dependent weight matrices for training the networks. Analysis of the contact data, for which only six groups participated, suggested that as yet this data provides a rather weak signal. However, in combination with other types of prediction data it can sometimes be a useful constraint for identifying the correct structure.

Animals↗

Relation between weight matrix and substitution matrix: motif search by similarity.

MOTIVATION: The discovery of patterns shared by several sequences that differ greatly is a basic task in sequence analysis, and still a challenge. Several methods have been developed for detecting patterns. Methods commonly used for motif search include the Gibbs sampler, Expectation-Maximization (EM) algorithm and some intuitive greedy approaches. One cannot guarantee the optimality of the result produced by the Gibbs sampler in a single run. The deterministic EM methods tend to get trapped by local optima. Solutions found by greedy approaches are rarely sufficiently good. RESULTS: A simple model describing a motif or a portion of local multiple sequence alignment is the weight matrix model, in which a motif is characterized with position-specific probabilities. Two substitution matrices are proposed to relate the sequence similarity with the weight matrix. Combining the substitution matrix and weight matrix, we examine three typical sets of protein sequences with increasing complexity. At a low score threshold for pair similarity, sliding windows are compared with a seed window to find the score sum, which provides a measure of statistical significance for multiple sequence comparison. Such a similarity analysis reveals many aspects of motifs. Blocks determined by similarity can be used to deduce a primary weight matrix or an improved substitution matrix. The algorithm successfully obtains the optimal solution for the test sets by just greedy iteration.

Algorithms↗

Mutational analysis of the OprM outer membrane component of the MexA-MexB-OprM multidrug efflux system of Pseudomonas aeruginosa.

OprM is the outer membrane component of the MexA-MexB-OprM efflux system of Pseudomonas aeruginosa. Multiple-sequence alignment of this protein and its homologues identified several regions of high sequence conservation that were targeted for site-directed mutagenesis. Of several deletions which were stably expressed, two, spanning residues G199 to A209 and A278 to N286 of the mature protein, were unable to restore antibiotic resistance in OprM-deficient strains of P. aeruginosa. Still, mutation of several conserved residues within these regions did not adversely affect OprM function. Mutation of the highly conserved N-terminal cysteine residue, site of acylation of this presumed lipoprotein, also did not affect expression or activity of OprM. Similarly, substitution of the OprM lipoprotein signal, including consensus lipoprotein box, with the signal peptide of OprF, the major porin of this organism, failed to impact on expression or activity. Apparently, acylation is not essential for OprM function. A large deletion at the N terminus, from A12 to R98, compromised OprM expression to some extent, although the deletion derivative did retain some activity. Several deletions failed to yield an OprM protein, including one lacking an absolutely conserved LGGGW sequence near the C terminus of the protein. The pattern of permissive and nonpermissive deletions was used to test a topology model for OprM based on the recently published crystal structure of the OprM homologue, TolC (V. Koronakis, A. Sharff, E. Koronakis, B. Luisi, and C. Hughes, Nature 405:914-919, 2000). The data are consistent with OprM monomer existing as a substantially periplasmic protein with four outer membrane-spanning regions.

Acylation↗

MAPPER: a search engine for the computational identification of putative transcription factor binding sites in multiple genomes.

BACKGROUND: Cis-regulatory modules are combinations of regulatory elements occurring in close proximity to each other that control the spatial and temporal expression of genes. The ability to identify them in a genome-wide manner depends on the availability of accurate models and of search methods able to detect putative regulatory elements with enhanced sensitivity and specificity. RESULTS: We describe the implementation of a search method for putative transcription factor binding sites (TFBSs) based on hidden Markov models built from alignments of known sites. We built 1,079 models of TFBSs using experimentally determined sequence alignments of sites provided by the TRANSFAC and JASPAR databases and used them to scan sequences of the human, mouse, fly, worm and yeast genomes. In several cases tested the method identified correctly experimentally characterized sites, with better specificity and sensitivity than other similar computational methods. Moreover, a large-scale comparison using synthetic data showed that in the majority of cases our method performed significantly better than a nucleotide weight matrix-based method. CONCLUSION: The search engine, available at http://mapper.chip.org, allows the identification, visualization and selection of putative TFBSs occurring in the promoter or other regions of a gene from the human, mouse, fly, worm and yeast genomes. In addition it allows the user to upload a sequence to query and to build a model by supplying a multiple sequence alignment of binding sites for a transcription factor of interest. Due to its extensive database of models, powerful search engine and flexible interface, MAPPER represents an effective resource for the large-scale computational analysis of transcriptional regulation.

Algorithms↗

Amino acid encoding schemes from protein structure alignments: multi-dimensional vectors to describe residue types.

Bioinformatic software has used various numerical encoding schemes to describe amino acid sequences. Orthogonal encoding, employing 20 numbers to describe the amino acid type of one protein residue, is often used with artificial neural network (ANN) models. However, this can increase the model complexity, thus leading to difficulty in implementation and poor performance. Here, we use ANNs to derive encoding schemes for the amino acid types from protein three-dimensional structure alignments. Each of the 20 amino acid types is characterized with a few real numbers. Our schemes are tested on the simulation of amino acid substitution matrices. These simplified schemes outperform the orthogonal encoding on small data sets. Using one of these encoding schemes, we generate a colouring scheme for the amino acids in which comparable amino acids are in similar colours. We expect it to be useful for visual inspection and manual editing of protein multiple sequence alignments.

Algorithms↗

Consensus patterns in DNA.

Matrices can provide realistic representations of protein/DNA specificity. In many cases simple mononucleotide-based matrices are adequate representations, but more complex matrices may be needed for other cases. Unlike simple consensus sequences, matrices allow for different penalties to be assessed for different changes to a binding site, a property that is essential for accurate description of a binding site pattern. When only a collection of binding site sequences is known, the best representation for the pattern is an information content formulation, based on both thermodynamic and statistical considerations. Quantitative data on relative binding affinities may be used to determine matrices that provide a best fit to the data. Matrix representations also provide an efficient method of aligning multiple sequences to identify binding site patterns that they have in common.

Base Sequence↗

Conservation and prediction of solvent accessibility in protein families.

Currently, the prediction of three-dimensional (3D) protein structure from sequence alone is an exceedingly difficult task. As an intermediate step, a much simpler task has been pursued extensively: predicting 1D strings of secondary structure. Here, we present an analysis of another 1D projection from 3D structure: the relative solvent accessibility of each residue. We show that solvent accessibility is less conserved in 3D homologues than is secondary structure, and hence is predicted less accurately from automatic homology modeling; the correlation coefficient of relative solvent accessibility between 3D homologues is only 0.77, and the average accuracy of predictions based on sequence alignments is only 0.68. The latter number provides an effective upper limit on the accuracy of predicting accessibility from sequence when homology modeling is not possible. We introduce a neural network system that predicts relative solvent accessibility (projected onto ten discrete states) using evolutionary profiles of amino acid substitutions derived from multiple sequence alignments. Evaluated in a cross-validation test on 238 unique proteins, the correlation between predicted and observed relative accessibility is 0.54. Interpreted in terms of a three-state (buried, intermediate, exposed) description of relative accessibility, the fraction of correctly predicted residue states is about 58%. In absolute terms this accuracy appears poor, but given the relatively low conservation of accessibility in 3D families, the network system is not far from its likely optimal performance. The most reliably predicted fraction of the residues (50%) is predicted as accurately as by automatic homology modeling. Prediction is best for buried residues, e.g., 86% of the completely buried sites are correctly predicted as having 0% relative accessibility.

Biological Evolution↗

Molecular modeling, affinity labeling, and site-directed mutagenesis define the key points of interaction between the ligand-binding domain of the vitamin D nuclear receptor and 1 alpha,25-dihydroxyvitamin D3.

We have combined molecular modeling and classical structure-function techniques to define the interactions between the ligand-binding domain (LBD) of the vitamin D nuclear receptor (VDR) and its natural ligand, 1alpha,25-dihydroxyvitamin D(3) [1alpha,25-(OH)(2)D(3)]. The affinity analogue 1alpha,25-(OH)(2)D(3)-3-bromoacetate exclusively labeled Cys-288 in the VDR-LBD. Mutation of C288 to glycine abolished this affinity labeling, whereas the VDR-LBD mutants C337G and C369G (other conserved cysteines in the VDR-LBD) were labeled similarly to the wild-type protein. These results revealed that the A-ring 3-OH group docks next to C288 in the binding pocket. We further mutated M284 and W286 (separately creating M284A, M284S, W286A, and W286F) and caused severe loss of ligand binding, indicating the crucial role played by the contiguous segment between M284 and C288. Alignment of the VDR-LBD sequence with the sequences of nuclear receptor LBDs of known 3-D structure positioned M284 and W286 in the presumed beta-hairpin of the molecule, thereby identifying it as the region contacting the A-ring of 1alpha, 25-(OH)(2)D(3). From the multiple sequence alignment, we developed a homologous extension model of the VDR-LBD. The model has a canonical nuclear receptor fold with helices H1-H12 and a single beta hairpin but lacks the long insert (residues 161-221) between H2 and H3. We docked the alpha-conformation of the A-ring into the binding pocket first so as to incorporate the above-noted interacting residues. The model predicts hydrogen bonding contacts between ligand and protein at S237 and D299 as well as at the site of the natural mutation R274L. Mutation of S237 or D299 to alanine largely abolished ligand binding, whereas changing K302, a nonligand-contacting residue, to alanine left binding unaffected. In the "activation" helix 12, the model places V418 closest to the ligand, and, consistent with this prediction, the mutation V418S abolished ligand binding. The studies together have enabled us to identify 1alpha,25-(OH)(2)D(3)-binding motifs in the ligand-binding pocket of VDR.

Affinity Labels↗

Approximate multiple protein structure alignment using the sum-of-pairs distance.

An algorithm is presented to compute a multiple structure alignment for a set of proteins and to generate a consensus (pseudo) protein for the set. The algorithm is a heuristic in that it computes an approximation to the optimal multiple structure alignment that minimizes the sum of the pairwise distances between the protein structures. The algorithm chooses an input protein as the initial consensus and computes a correspondence between the protein structures (which are represented as sets of unit vectors) using an approach analogous to the center-star method for multiple sequence alignment. From this correspondence, a set of rotation matrices (optimal for the given correspondence) is derived to align the structures and derive the new consensus. The process is iterated until the sum of pairwise distances converges. The computation of the optimal rotations is itself an iterative process that both makes use of the current consensus and generates simultaneously a new one. This approach is based on an interesting result that allows the sum of all pairwise distances to be represented compactly as distances to the consensus. Experimental results on several protein families are presented, showing that the algorithm converges quite rapidly.

Algorithms↗

Revealing the set of mutually correlated positions for the protein families of immunoglobulin fold.

In this study, I explain the observation that a rather limited number of residues (about 10) establishes the immunoglobulin fold for the sequences of about 100 residues. Immunoglobulin fold proteins (IgF) comprise SCOP protein superfamilies with rather different functions and with less than 10% sequence identity; their alignment can be accomplished only taking into account the 3D structure. Therefore, I believe that discovering the additional common features of the sequences is necessary to explain the existence of a common fold for these SCOP superfamilies. We propose a method for analysis of pair-wise interconnections between residues of the multiple sequence alignment which helps us to reveal the set of mutually correlated positions, inherent to almost every superfamily of this protein fold. Hence, the set of constant positions (comprising the hydrophobic common core) and the set of variable but mutually correlated ones can serve as a basis of having the common 3D structure for rather distinct protein sequences.

Algorithms↗

A method for the improvement of threading-based protein models.

A new method for the homology-based modeling of protein three-dimensional structures is proposed and evaluated. The alignment of a query sequence to a structural template produced by threading algorithms usually produces low-resolution molecular models. The proposed method attempts to improve these models. In the first stage, a high-coordination lattice approximation of the query protein fold is built by suitable tracking of the incomplete alignment of the structural template and connection of the alignment gaps. These initial lattice folds are very similar to the structures resulting from standard molecular modeling protocols. Then, a Monte Carlo simulated annealing procedure is used to refine the initial structure. The process is controlled by the model's internal force field and a set of loosely defined restraints that keep the lattice chain in the vicinity of the template conformation. The internal force field consists of several knowledge-based statistical potentials that are enhanced by a proper analysis of multiple sequence alignments. The template restraints are implemented such that the model chain can slide along the template structure or even ignore a substantial fraction of the initial alignment. The resulting lattice models are, in most cases, closer (sometimes much closer) to the target structure than the initial threading-based models. All atom models could easily be built from the lattice chains. The method is illustrated on 12 examples of target/template pairs whose initial threading alignments are of varying quality. Possible applications of the proposed method for use in protein function annotation are briefly discussed.

Amino Acid Sequence↗

A method for alpha-helical integral membrane protein fold prediction.

Integral membrane proteins (of the alpha-helical class) are of central importance in a wide variety of vital cellular functions. Despite considerable effort on methods to predict the location of the helices, little attention has been directed toward developing an automatic method to pack the helices together. In principle, the prediction of membrane proteins should be easier than the prediction of globular proteins: there is only one type of secondary structure and all helices pack with a common alignment across the membrane. This allows all possible structures to be represented on a simple lattice and exhaustively enumerated. Prediction success lies not in generating many possible folds but in recognizing which corresponds to the native. Our evaluation of each fold is based on how well the exposed surface predicted from a multiple sequence alignment fits its allocated position. Just as exposure to solvent in globular proteins can be predicted from sequence variation, so exposure to lipid can be recognized by variable-hydrophobic (variphobic) positions. Application to both bacteriorhodopsin and the eukaryotic rhodopsin/opsin families revealed that the angular size of the lipid-exposed faces must be predicted accurately to allow selection of the correct fold. With the inherent uncertainties in helix prediction and parameter choice, this accuracy could not be guaranteed but the correct fold was typically found in the top six candidates. Our method provides the first completely automatic method that can proceed from a scan of the protein sequence databanks to a predicted three-dimensional structure with no intervention required from the investigator. Within the limited domain of the seven helix bundle proteins, a good chance can be given of selecting the correct structure. However, the limited number of sequences available with a corresponding known structure makes further characterization of the method difficult.

Amino Acid Sequence↗

Genetic diversity and recombination of NA-PRRSV field strains in Vietnam: Implications for vaccine efficacy.

Porcine reproductive and respiratory syndrome (PRRS) causes severe reproductive losses in pregnant sows and piglets, resulting in substantial economic impact on the swine industry worldwide. However, due to the significant genetic diversity and rapid evolutionary changes of the pathogen, continuous surveillance and detailed genetic analysis of circulating strains are essential. The current study aimed to evaluate the genetic diversity of the hypervariable (HV) region of non-structural protein 2 (nsp2) among North American PRRSV strains isolated from swine farms in Vietnam. Phylogenetic analysis and multiple sequence alignment were conducted to determine subtype classification and assess genetic variability. A total of 48 field isolates were obtained, of which 12.5% belonged to classical NA-PRRSV, 16.6% to NADC30-like and 70.9% to HP-PRRSV, primarily distributed across sublineages 1.4, 5.1, 8.7 and 8.9. Amino acid comparisons found multiple insertions, deletions and substitutions at various positions within the hypervariable region of nsp2. The study revealed substantial genetic variation in the HV region of nsp2 among NA-PRRSV field strains, largely associated with recombination and immune escape. These findings highlight epidemiological risks to vaccine efficacy and underscore the need for continuous molecular surveillance to support effective PRRSV control in Vietnam.

PRRSV↗

Comparison of three classes of snake neurotoxins by homology modeling and computer simulation graphics.

We present a systematic structure comparison of three major classes of postsynaptic snake toxins, which include short and long chain alpha-type neurotoxins plus one angusticeps-type toxin of black mamba snake family. Two novel alpha-type neurotoxins isolated from Taiwan cobra (Naja naja atra) possessing distinct primary sequences and different postsynaptic neurotoxicities were taken as exemplars for short and long chain neurotoxins and compared with the major lethal short-chain neurotoxin in the same venom, i.e., cobrotoxin, based on the derived three-dimensional structure of this toxin in solution by NMR spectroscopy. A structure comparison among these two alpha-neurotoxins and angusticeps-type toxin (denoted as FS2) was carried out by the secondary-structure prediction together with computer homology-modeling based on multiple sequence alignment of their primary sequences and established NMR structures of cobrotoxin and FS2. It is of interest to find that upon pairwise superpositions of these modeled three-dimensional polypeptide chains, distinct differences in the overall peptide flexibility and interior microenvironment between these toxins can be detected along the three constituting polypeptide loops, which may reflect some intrinsic differences in the surface hydrophobicity of several hydrophobic peptide segments present on the surface loops of these toxin molecules as revealed by hydropathy profiles. Construction of a phylogenetic tree for these structurally related and functionally distinct toxins corroborates that all long and short toxins present in diverse snake families are evolutionarily related to each other, supposedly derived from an ancestral polypeptide by gene duplication and subsequent mutational substitutions leading to divergence of multiple three-loop toxin peptides.

Amino Acid Sequence↗

Molecular phylogeny in 3-D.

Molecular phylogenetic trees are constructed in three dimensions relative to the distribution of MW and pl classes and immunocrossreactivity against polyclonal antibodies to lens crystallins, as well as multiple sequence alignment between amino acid sequences, coding nucleotide sequences and the gene nucleotide sequences for beta-globin. Euclidian distances are estimated to position species in x, y, z space by multidimensional scaling and merged with bootstrap-tested branching pattern of Fitch & Margoliash plots to obtain 3-D phylogenetic tree. Compared to single attributes, phylogenetic trees based on multiple parameters allow significant repositioning of rodents, chiroptera and primates.

Animals↗

Proximal and distal histidines in thyroid peroxidase: relation to the alternatively spliced form, TPO-2.

The distal and proximal histidines in thyroid peroxidase (TPO), located by amino acid sequence alignment with their known counterparts in myeloperoxidase, are His 239 and His 494, respectively. These histidines lie outside the 57 amino acid peptide (residues 533-589) that is absent in the alternatively spliced form, TPO-2. However, asparagine 579, which very likely forms a stabilizing hydrogen bond with the proximal histidine in TPO, lies within the missing peptide region. The absence of Asn 579 from TPO-2 may be at least partially responsible for the reported lack of activity of this form of the enzyme. Formation of TPO compound I may also depend on Arg 396, based on analogy with the catalytic mechanism previously proposed for the more widely studied plant and fungal peroxidases. A multiple sequence alignment prepared with five mammalian and five invertebrate peroxidases shows complete conservation of Arg 396, as well as residues corresponding to His 239, His 494, and Asn 579 in TPO. The animal peroxidases comprise a family of homologous proteins that differ markedly from the plant/fungal/bacterial peroxidases in primary, secondary, and tertiary structure, yet share with them a common function. Animal peroxidases probably arose independently of the plant/fungal/bacterial peroxidase superfamily and most likely belong to a different gene family. The relation between animal and nonanimal peroxidases may represent an example of convergent evolution to a common enzymatic mechanism.

Alternative Splicing↗

Molecular characterization of the Cryptosporidium cervine genotype.

In this study, the 27 kDa immunodominant antigen (CP23), 70 kDa heat shock protein (HSP70), actin and beta-tubulin genes were amplified and sequenced for the first time from human isolates of Cryptosporidium cervine genotype. New primers were designed from reported sequences of other Cryptosporidium species and genotypes as well as the whole genome sequences of C. parvum and C. hominis, which enabled novel gene sequences and regions extending beyond those deposited in GenBank to be determined. In comparison with other species in the Cryptosporidium genus, multiple sequence alignment and phylogenetic analysis revealed that the Cryptosporidium cervine genotype isolates from humans clustered most closely with Cryptosporidium deer mouse genotype and C. suis (n. sp. formerly pig genotype I). The complete coding sequence of CP23 was determined to reveal low (72.4% and 68.0-69.8% respectively) identity to C. parvum and C. hominis sequences and the presence of a unique multiple proline-alanine-proline-valine (PAPV) repeat region.

Actins↗

Development and epidemiological investigation of a TaqMan-based multiplex real-time quantitative PCR assay for simultaneous detection of five bovine viruses (BVDV, AKAV, BNoV, BEV, and BCoV).

INRODUCTION: Infectious diseases caused by bovine viral diarrhea virus (BVDV), Akabane virus (AKAV), bovine norovirus (BNoV), bovine enterovirus (BEV), and bovine coronavirus (BCoV) significantly threaten the cattle industry, resulting in substantial economic losses. These pathogens often present similar clinical signs, such as diarrhea, vomiting, and reproductive disorders in pregnant cattle, and frequent covert or mixed infections further complicate accurate diagnosis. Therefore, rapid, sensitive, and field‑deployable diagnostic methods are essential for effective disease surveillance and control in the cattle industry. METHODS: In this study, we report for the first time the establishment of a TaqMan‑based real‑time quantitative PCR (qPCR) assay that enables simultaneous detection of these five bovine viruses. Multiple sequence alignment of conserved genomic regions was performed, and virus‑specific primers and probes were designed and optimized using Beacon Designer 7 software. Subsequently, a TaqMan‑based multiplex real‑time qPCR assay was established for simultaneous detection of BVDV, AKAV, BNoV, BEV, and BCoV. The established detection method was applied to 200 clinical samples collected from 10 farms in multiple regions of Jilin Province. RESULTS: The results showed that the detection rates for BVDV, AKAV, BNoV, BEV, and BCoV were 33.50%, 0.50%, 4.50%, 7.50%, and 12.00%, respectively. Mixed infections were detected in 9 samples co‑infected with two of the five pathogens, with an overall mixed infection rate of 4.50%. Compared with conventional PCR, coincidence rates were 100% for BVDV, AKAV, BNoV, BEV, and BCoV. DISCUSSION: These findings indicate that the TaqMan multiplex real‑time qPCR assay developed here demonstrates favorable specificity, sensitivity, and reproducibility. This assay enables efficient detection and surveillance of bovine viruses, offering a reliable technical tool for the diagnosis and control of corresponding viral diseases in cattle.

Akabane virus (AKAV)↗