Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Potential for dramatic improvement in sequence alignment against structures of remote homologous proteins by extracting structural information from multiple structure alignment.

A novel method has been developed for acquiring the correct alignment of a query sequence against remotely homologous proteins by extracting structural information from profiles of multiple structure alignment. A systematic search algorithm combined with a group of score functions based on sequence information and structural information has been introduced in this procedure. A limited number of top solutions (15,000) with high scores were selected as candidates for further examination. On a test-set comprising 301 proteins from 75 protein families with sequence identity less than 30%, the proportion of proteins with completely correct alignment as first candidate was improved to 39.8% by our method, whereas the typical performance of existing sequence-based alignment methods was only between 16.1% and 22.7%. Furthermore, multiple candidates for possible alignment were provided in our approach, which dramatically increased the possibility of finding correct alignment, such that completely correct alignments were found amongst the top-ranked 1000 candidates in 88.3% of the proteins. With the assistance of a sequence database, completely correct alignment solutions were achieved amongst the top 1000 candidates in 94.3% of the proteins. From such a limited number of candidates, it would become possible to identify more correct alignment using a more time-consuming but more powerful method with more detailed structural information, such as side-chain packing and energy minimization, etc. The results indicate that the novel alignment strategy could be helpful for extending the application of highly reliable methods for fold identification and homology modeling to a huge number of homologous proteins of low sequence similarity. Details of the methods, together with the results and implications for future development are presented.

Algorithms↗

Improvement in prediction of solvent accessibility by probability profiles.

The capability of predicting folding and conformation of a protein from its primary structure is probably one of the main goals of modern biology. An accurate prediction of solvent accessibility is an intermediate step along this way. A new method for predicting solvent accessibility from single sequence and multiple alignment data is described. The method is based on probability profiles calculated on an amino acid sequence centred on the residue whose accessibility has to be predicted. A profile is constructed for each exposure category considered so as to calculate the probability of a sequence being generated by the different profiles. Prediction accuracy was tested on a variety of protein sets with two- and three-state models. Different thresholds were used according to those adopted by the authors proposing the data sets. The prediction accuracy is significantly improved over existing methods.

Algorithms↗

Applications of hidden Markov models for characterization of homologous DNA sequences with a common gene.

Identifying and characterizing the structure in genome sequences is one of the principal challenges in modern molecular biology, and comparative genomics offers a powerful tool. In this paper, we introduce a hidden Markov model that allows a comparative analysis of multiple sequences related by a phylogenetic tree, and we present an efficient method for estimating the parameters of the model. The model integrates structure prediction methods for one sequence, statistical multiple alignment methods, and phylogenetic information. This unified model is particularly useful for a detailed characterization of DNA sequences with a common gene. We illustrate the model on a variety of homologous sequences.

Agrobacterium tumefaciens↗

Molecular modelling studies on G protein-coupled receptors: from sequence to structure?

In pharmaceutical research G protein-coupled receptors (GPCR) emerged as a superfamily of prominent drug targets. Extensive protein sequence analyses on GPCRs revealed a common protein topology consisting of a membrane-spanning seven-helix bundle, which is believed to accommodate the binding site for low-molecular weight ligands. Enormous efforts are undertaken to generate GPCR structural models by means of molecular modelling, since these have already been shown to aid the process of lead structure finding and optimisation in that they provide atomistic models for structure-based drug design approaches. One of the most critical steps in modelling the transmembrane domains of GPCRs is the assignment of the putative transmembrane sequence stretches from multiple protein sequence alignment analyses. This study focuses on the comparative evaluation of protein sequence analysis tools, such as periodicity analyses, multiple sequence analyses or directional helix descriptors, especially developed for modelling the 7TM domains of GPCRs. In this context we will demonstrate that from application of different methods contradictory results can be obtained for the identification of the putative transmembrane sequence stretches of peptide-binding GPCRs, as exemplified with a comprehensive protein sequence analysis study based on the most prominent members of that receptor class (angiotensin II, CCK/gastrin, interleukin 8, endothelin, etc.).

Animals↗

SMART: identification and annotation of domains from signalling and extracellular protein sequences.

SMART is a simple modular architecture research tool and database that provides domain identification and annotation on the WWW (http://coot.embl-heidelberg.de/SMART). The tool compares query sequences with its databases of domain sequences and multiple alignments whilst concurrently identifying compositionally biased regions such as signal peptide, transmembrane and coiled coil segments. Annotated and unannotated regions of the sequence can be used as queries in searches of sequence databases. The SMART alignment collection represents more than 250 signalling and extracellular domains. Each alignment is curated to assign appropriate domain boundaries and to ensure its quality. In addition, each domain is annotated extensively with respect to cellular localisation, species distribution, functional class, tertiary structure and functionally important residues.

Amino Acid Sequence↗

Accuracy of sequence alignment and fold assessment using reduced amino acid alphabets.

Reduced or simplified amino acid alphabets group the 20 naturally occurring amino acids into a smaller number of representative protein residues. To date, several reduced amino acid alphabets have been proposed, which have been derived and optimized by a variety of methods. The resulting reduced amino acid alphabets have been applied to pattern recognition, generation of consensus sequences from multiple alignments, protein folding, and protein structure prediction. In this work, amino acid substitution matrices and statistical potentials were derived based on several reduced amino acid alphabets and their performance assessed in a large benchmark for the tasks of sequence alignment and fold assessment of protein structure models, using as a reference frame the standard alphabet of 20 amino acids. The results showed that a large reduction in the total number of residue types does not necessarily translate into a significant loss of discriminative power for sequence alignment and fold assessment. Therefore, some definitions of a few residue types are able to encode most of the relevant sequence/structure information that is present in the 20 standard amino acids. Based on these results, we suggest that the use of reduced amino acid alphabets may allow to increasing the accuracy of current substitution matrices and statistical potentials for the prediction of protein structure of remote homologs.

Amino Acid Sequence↗

4SALE--a tool for synchronous RNA sequence and secondary structure alignment and editing.

BACKGROUND: In sequence analysis the multiple alignment builds the fundament of all proceeding analyses. Errors in an alignment could strongly influence all succeeding analyses and therefore could lead to wrong predictions. Hand-crafted and hand-improved alignments are necessary and meanwhile good common practice. For RNA sequences often the primary sequence as well as a secondary structure consensus is well known, e.g., the cloverleaf structure of the t-RNA. Recently, some alignment editors are proposed that are able to include and model both kinds of information. However, with the advent of a large amount of reliable RNA sequences together with their solved secondary structures (available from e.g. the ITS2 Database), we are faced with the problem to handle sequences and their associated secondary structures synchronously. RESULTS: 4SALE fills this gap. The application allows a fast sequence and synchronous secondary structure alignment for large data sets and for the first time synchronous manual editing of aligned sequences and their secondary structures. This study describes an algorithm for the synchronous alignment of sequences and their associated secondary structures as well as the main features of 4SALE used for further analyses and editing. 4SALE builds an optimal and unique starting point for every RNA sequence and structure analysis. CONCLUSION: 4SALE, which provides an user-friendly and intuitive interface, is a comprehensive toolbox for RNA analysis based on sequence and secondary structure information. The program connects sequence and structure databases like the ITS2 Database to phylogeny programs as for example the CBCAnalyzer. 4SALE is written in JAVA and therefore platform independent. The software is freely available and distributed from the website at http://4sale.bioapps.biozentrum.uni-wuerzburg.de.

Algorithms↗

Analysis and prediction of functional sub-types from protein sequence alignments.

The increasing number and diversity of protein sequence families requires new methods to define and predict details regarding function. Here, we present a method for analysis and prediction of functional sub-types from multiple protein sequence alignments. Given an alignment and set of proteins grouped into sub-types according to some definition of function, such as enzymatic specificity, the method identifies positions that are indicative of functional differences by comparison of sub-type specific sequence profiles, and analysis of positional entropy in the alignment. Alignment positions with significantly high positional relative entropy correlate with those known to be involved in defining sub-types for nucleotidyl cyclases, protein kinases, lactate/malate dehydrogenases and trypsin-like serine proteases. We highlight new positions for these proteins that suggest additional experiments to elucidate the basis of specificity. The method is also able to predict sub-type for unclassified sequences. We assess several variations on a prediction method, and compare them to simple sequence comparisons. For assessment, we remove close homologues to the sequence for which a prediction is to be made (by a sequence identity above a threshold). This simulates situations where a protein is known to belong to a protein family, but is not a close relative of another protein of known sub-type. Considering the four families above, and a sequence identity threshold of 30 %, our best method gives an accuracy of 96 % compared to 80 % obtained for sequence similarity and 74 % for BLAST. We describe the derivation of a set of sub-type groupings derived from an automated parsing of alignments from PFAM and the SWISSPROT database, and use this to perform a large-scale assessment. The best method gives an average accuracy of 94 % compared to 68 % for sequence similarity and 79 % for BLAST. We discuss implications for experimental design, genome annotation and the prediction of protein function and protein intra-residue distances.

Adenylyl Cyclases↗

Pairwise local structural alignment of RNA sequences with sequence similarity less than 40%.

MOTIVATION: Searching for non-coding RNA (ncRNA) genes and structural RNA elements (eleRNA) are major challenges in gene finding today as these often are conserved in structure rather than in sequence. Even though the number of available methods is growing, it is still of interest to pairwise detect two genes with low sequence similarity, where the genes are part of a larger genomic region. RESULTS: Here we present such an approach for pairwise local alignment which is based on foldalign and the Sankoff algorithm for simultaneous structural alignment of multiple sequences. We include the ability to conduct mutual scans of two sequences of arbitrary length while searching for common local structural motifs of some maximum length. This drastically reduces the complexity of the algorithm. The scoring scheme includes structural parameters corresponding to those available for free energy as well as for substitution matrices similar to RIBOSUM. The new foldalign implementation is tested on a dataset where the ncRNAs and eleRNAs have sequence similarity <40% and where the ncRNAs and eleRNAs are energetically indistinguishable from the surrounding genomic sequence context. The method is tested in two ways: (1) its ability to find the common structure between the genes only and (2) its ability to locate ncRNAs and eleRNAs in a genomic context. In case (1), it makes sense to compare with methods like Dynalign, and the performances are very similar, but foldalign is substantially faster. The structure prediction performance for a family is typically around 0.7 using Matthews correlation coefficient. In case (2), the algorithm is successful at locating RNA families with an average sensitivity of 0.8 and a positive predictive value of 0.9 using a BLAST-like hit selection scheme. AVAILABILITY: The program is available online at http://foldalign.kvl.dk/

Algorithms↗

Cloning of a Leishmania major gene encoding for an antigen with extensive homology to ribosomal protein S3a.

Following purification by affinity chromatography, a Leishmania major S-hexylglutathione- binding protein of molecular mass 66kDa was isolated. The immune serum against the parasite 66kDa polypeptide when used to screen a L. major cDNA library could identify clones encoding for the human v-fos transformation effector homologue, namely ribosomal protein S3a, and thus was named LmS3a-related protein (LmS3arp). A 1027bp cDNA fragment was found to contain the entire parasite gene encoding for a highly basic protein of 30kDa calculated molecular mass sharing homology to various ribosomal S3a proteins from different species. Using computer methods for a multiple alignment and sequence motif search, we found that LmS3arp shares a sequence homology to class theta glutathione S-transferase mainly in a segment containing critical residues involved in glutathione binding. These new findings are discussed in the light of recent published data showing multiple function(s) of the ribosomal proteins S3a.

Amino Acid Sequence↗

[Identification of the gene sequence of telomerase catalytic subunit in Plasmodium falciparum].

INTRODUCTION: The enzyme telomerase regulates telomere length by synthesis of telomeric repeats to compensate for telomeric loss in each DNA replication cycle. Therefore, telomerase is a potential target to block growth of cells with high replication rates. In Plasmodium falciparum, telomerase activity has been documented, but little information on its structure and role. METHODS: Herein, alignment of multiple sequences was undertaken comparing telomerase catalytic subunit sequences as found in existing databases. A consensus sequence was compared with the sequences in the P. falciparum genome project and as a result, a candidate sequence for a portion of the telomerase gene was recovered. Primer sets were designed for DNA and RNA amplifications. RESULTS: DNA fragments corresponding to telomerase conserved domains were amplified by using reverse transcription and PCR of cDNA. With a combination of bioinformatics and sequencing methods, the sequence of telomerase catalytic subunit gene (TERT) in P. falciparum was discovered, and its presence and transcription demonstrated.

Animals↗

Molecular analysis of a 444 bp fragment of the bovine leukaemia virus gp51 env gene reveals a high frequency of non-silent point mutations and suggests the presence of two subgroups of BLV in Chile.

With the aim of achieve a better understanding of the epidemiology and distribution of bovine leukaemia virus (BLV) infection in Chile, we assessed the suitability of using DNA isolated from the leukocyte fraction of bulk milk samples to carry out PCR-RFLP and DNA sequence analysis. The env fragment of BLV was successfully amplified from 33 serologically positive bulk milk samples collected from different geographical areas in the south of Chile. Restriction analysis allowed to classify 17 isolates within the Australian subgroup and 16 within the Belgium subgroup. DNA sequence and multiple alignment analysis of eight Chilean isolates showed a significantly higher frequency of single and double nucleotide substitutions. Most of these mutations were non-silent, resulting in changes at the protein level in several important epitopes of gp51. The Chilean sequences and 59 BLV env sequences available at GenBank, were subjected to a phylogenetic analysis, resulting in four different clusters. The groups identified were not related to those previously defined by restriction analysis. Chilean isolates were included in two different clusters and were genetically not related to isolates collected from neighbouring countries. Considering our results we can conclude: (i) bulk milk samples are suitable to identify the presence of BLV allowing epidemiological and genetic studies to be conducted on large geographical areas; (ii) at least four different genetic groups of BLV were identified by phylogenetic analysis, with Chilean isolates included in two different sub clusters.

Amino Acid Sequence↗

Human endogenous retrovirus family HERV-K(HML-5): status, evolution, and reconstruction of an ancient betaretrovirus in the human genome.

The human genome harbors numerous distinct families of so-called human endogenous retroviruses (HERV) which are remnants of exogenous retroviruses that entered the germ line millions of years ago. We describe here the hitherto little-characterized betaretrovirus HERV-K(HML-5) family (named HERVK22 in Repbase) in greater detail. Out of 139 proviruses, only a few loci represent full-length proviruses, and many lack gag protease and/or env gene regions. We generated a consensus sequence from multiple alignment of 62 HML-5 loci that displays open reading frames for the four major retroviral proteins. Four HML-5 long terminal repeat (LTR) subfamilies were identified that are associated with monophyletic proviral bodies, implying different evolution of HML-5 LTRs and genes. Sequence analysis indicated that the proviruses formed approximately 55 million years ago. Accordingly, HML-5 proviral sequences were detected in Old World and New World primates but not in prosimians. No recent activity is associated with this HERV family. We also conclude that the HML-5 consensus sequence primer binding site is identical to methionine tRNA. Therefore, the family should be designated HERV-M. Our study provides important insights into the structure and evolution of the oldest betaretrovirus in the primate genome known to date.

Animals↗

Evolutionary relationship between the TonB-dependent outer membrane transport proteins: nucleotide and amino acid sequences of the Escherichia coli colicin I receptor gene.

The nucleotide sequence of the Escherichia coli colicin I receptor gene (cir) has been determined. The predicted mature protein consists of 599 amino acids and has a molecular weight of 67,169. Several previously noted characteristics of other E. coli outer membrane protein sequences were also identified in the sequence of Cir. These include an overall acidic nature, the absence of long hydrophobic stretches of amino acids, and a lack of predicted alpha-helical secondary structure. Because two classes of outer membrane proteins (the TonB-dependent transport proteins and the porins) share some structural features, protein sequences from both of these groups were aligned pairwise and scored for sequence similarity. Statistical evidence suggested that the porins were not related to the proteins in the TonB-dependent group; however, there was a significant relationship between the proteins in the TonB-dependent group. On the basis of the multiple progressive sequence alignment and the similarity scores derived from it, a tree representing evolutionary distance between five TonB-dependent outer membrane transport proteins was generated.

Amino Acid Sequence↗

Identification of a novel 4-aminomethylpiperidine class of M3 muscarinic receptor antagonists and structural insight into their M3 selectivity.

Identification of a novel class of potent and highly selective M(3) muscarinic antagonists is described. First, the structure-activity relationship in the cationic amine core of our previously reported triphenylpropionamide class of M(3) selective antagonists was explored by a small diamine library constructed in solid phase. This led to the identification of M(3) antagonists with a novel piperidine pharmacophore and significantly improved subtype selectivity from a previously reported class. Successive modification on the terminal triphenylpropionamide part of the newly identified class gave 14a as a potent M(3) selective antagonist that had >100-fold selectivity versus the M(1), M(2), M(4), and M(5) receptors (M(3): K(i) = 0.30 nM, M(1)/M(3) = 570-fold, M(2)/M(3) = 1600-fold, M(4)/M(3) = 140-fold, M(5)/M(3) = 12000-fold). The possible rationale for its extraordinarily higher subtype selectivity than reported M(3) antagonists was hypothesized by sequence alignment of multiple muscarinic receptors and a computational docking of 14a into transmembrane domains of M(3) receptors.

Amino Acid Sequence↗

A multiple alignment of the capsid protein sequences of nepoviruses and comoviruses suggests a common structure.

The amino acid sequences of the regions encoding the structural proteins of eleven nepoviruses and five comoviruses, two genera of the family Comoviridae, have been aligned. The properties predicted by computer analysis (three-dimensional-3D-structure, hydrophobicity) are also correlated along this alignment, and aligned to the experimentally determined 3D structure of two comoviruses. It can thus be assumed that the 3D structure of the unique nepovirus coat protein matches that of the bipartite protomer found in the comovirus particles. In this model, the spatial locations of two amino-acid motifs characteristic of nepoviruses are in close vicinity, at the external surface of the virion. The coat proteins of nepoviruses and comoviruses may thus share a common evolutionary origin. A phylogenetic analysis was made using the multiple alignment, allowing a better understanding of the molecular relationships between these two groups of viruses.

Amino Acid Sequence↗

WAViS server for handling, visualization and presentation of multiple alignments of nucleotide or amino acids sequences.

Web Alignment Visualization Server contains a set of web-tools designed for quick generation of publication-quality color figures of multiple alignments of nucleotide or amino acids sequences. It can be used for identification of conserved regions and gaps within many sequences using only common web browsers. The server is accessible at http://wavis.img.cas.cz.

Computer Graphics↗

Minding the gap: frequency of indels in mtDNA control region sequence data and influence on population genetic analyses.

Insertions and deletions (indels) result in sequences of various lengths when homologous gene regions are compared among individuals or species. Although indels are typically phylogenetically informative, occurrence and incorporation of these characters as gaps in intraspecific population genetic data sets are rarely discussed. Moreover, the impact of gaps on estimates of fixation indices, such as F(ST), has not been reviewed. Here, I summarize the occurrence and population genetic signal of indels among 60 published studies that involved alignments of multiple sequences from the mitochondrial DNA (mtDNA) control region of vertebrate taxa. Among 30 studies observing indels, an average of 12% of both variable and parsimony-informative sites were composed of these sites. There was no consistent trend between levels of population differentiation and the number of gap characters in a data block. Across all studies, the average influence on estimates of PhiST was small, explaining only an additional 1.8% of among population variance (range 0.0-8.0%). Studies most likely to observe an increase in PhiST with the inclusion of gap characters were those with < 20 variable sites, but a near equal number of studies with few variable sites did not show an increase. In contrast to studies at interspecific levels, the influence of indels for intraspecific population genetic analyses of control region DNA appears small, dependent upon total number of variable sites in the data block, and related to species-specific characteristics and the spatial distribution of mtDNA lineages that contain indels.

Animals↗