Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

The megaprior heuristic for discovering protein sequence patterns.

Several computer algorithms for discovering patterns in groups of protein sequences are in use that are based on fitting the parameters of a statistical model to a group of related sequences. These include hidden Markov model (HMM) algorithms for multiple sequence alignment, and the MEME and Gibbs sampler algorithms for discovering motifs. These algorithms are sometimes prone to producing models that are incorrect because two or more patients have been combined. The statistical model produced in this situation is a convex combination (weighted average) of two or more different models. This paper presents a solution to the problem of convex combinations in the form of a heuristic based on using extremely low variance Dirichlet mixture priors as part of the statistical model. This heuristic, which we call the megaprior heuristic, increase the strength (i.e., decreases the variance) of the prior in proportion to the size of the sequence dataset. This causes each column in the final model to strongly resemble the mean of a single component of the prior, regardless of the size of the dataset. We describe the cause of the convex combination problem, analyze it mathematically, motivate and describe the implementation of the megaprior heuristic, and show how it can effectively eliminate the problem of convex combinations in protein sequence pattern discovery.

Algorithms↗

Prediction of the structure of the replication initiator protein DnaA.

The secondary structure of DnaA protein and its interaction with DNA and ribonucleotides has been predicted using biochemical, biophysical techniques, and prediction methods based on multiple-sequence alignment and neural networks. The core of all proteins from the DnaA family consists of an "open twisted alpha/beta structure," containing five alpha-helices alternating with five beta-strands. In our proposed structural model the interior of the core is formed by a parallel beta-sheet, whereas the alpha-helices are arranged on the surface of the core. The ATP-binding motif is located within the core, in a loop region following the first beta-strand. The N-terminal domain (80 aa) is composed of two alpha-helices, the first of which contains a potential leucine zipper motif for mediating protein-protein interaction, followed by a beta-strand and an additional alpha-helix. The N-terminal domain and the alpha/beta core region of DnaA are connected by a variable loop (45-70 aa); major parts of the loop region can be deleted without loss of protein activity. The C-terminal DNA-binding domain (94 aa) is mostly alpha-helical and contains a potential helix-loop-helix motif. DnaA protein does not dimerize in solution; instead, the two longest C-terminal alpha-helices could interact with each other, forming an internal "coiled coil" and exposing highly basic residues of a small loop region on the surface, probably responsible for DNA backbone contacts.

Amino Acid Sequence↗

Structure and genomic organization of a second class of immunoglobulin light chain genes in the channel catfish.

Earlier studies distinguished two classes of catfish light (L) chain (designated F and G). The cDNA structure and genomic organization of G L chain gene clusters has also been characterized previously. In this study, full length cDNA encoding F L chain was derived using PCR strategies based on the determined amino-terminal protein sequence. The encoded V region is readily delineated into framework regions (FR) and complementarity-determining regions (CDR). Multiple sequence alignments indicate that the F V(L) is closely related to kappa gene families. The F C(L) cannot be generally classified but it is structurally distinct from the C(L) regions of G: the amino acid sequence similarity is <35%. cDNA sequences representing processed sterile F transcripts of different loci were identified. Each sequence begins within the J(L) recombination signal sequence and extends downstream through the I(L)-C(L) segments. Genomic blots hybridized with C(L) probes indicate that there are at least 50 different C(L) segments. Based upon V(L) hybridization studies, different families of V(L) segments appear to be associated with closely related F C(L) segments. In characterized genomic clones, F gene segments are arranged in closely linked clusters with single copies of V(L), J(L), and C(L) segments within each cluster. The V(L) segments are located in opposite transcriptional polarity relative to the J(L) and C(L) segments, which indicates that V(L) segments rearrange by inversion. These combined studies establish that two structurally distinct classes of L chains are present in teleost fish and that both of the L chain classes evolved within a common organizational pattern of clustered segmental genes.

Amino Acid Sequence↗

A model for the nucleotide-binding domains of ABC transporters based on the large domain of aspartate aminotransferase.

ABC transporters are a large superfamily of integral membrane proteins involved inATP-dependent transport across biological membranes. Members of this superfamily play roles in a number of phenomena of biomedical interest, including cystic fibrosis (CFTR) and multidrug resistance (P-glycoprotein, MRP). Most ABC transporters are predicted to consist of four domains, two membrane-spanning domains and two cytoplasmic domains. The latter contain conserved nucleotide-binding motifs. Attempts to determine the structure of ABC transporters and of their separate domains are in progress but have not yet been successful. To aid structure determination and possibly learn more about the domain boundaries, we set out to model nucleotide-binding domains (NBDs) of ABC transporters based on a known structure. Previous attempts to predict the 3D structure of NBDs were based solely on sequence similarity with known nucleotide-binding folds. We have analyzed the sequences of a number of nucleotide-binding domains with the algorithm THREADER, developed by D.T. Jones, and a possible fold was found in the structure of aspartate aminotransferase. We present a model for the N-terminal NBD of CFTR, based on the large domain of the A chain of aspartate aminotransferase. The model is refined using multiple sequence alignment, secondary structure prediction, and 3D-1D profiles. Our model seems to be in good agreement with known properties of nucleotide-binding domains and has some appealing characteristics compared with the previous models.

ATP-Binding Cassette Transporters↗

Isolation and characterization of a skate retinal GABA transporter cDNA.

PURPOSE: The inhibitory neurotransmitter gamma-aminobutyric acid (GABA) is believed to play a crucial role in the processing of information within the vertebrate retina. Extracellular concentrations of GABA are thought to be tightly regulated by carrier-mediated transport proteins in neurons and glial cells. The purpose of this work was to isolate the gene that encodes one of these transport proteins in the skate retina. METHODS: cDNA clones were isolated from a skate retinal cDNA library using a mouse retinal GABA transporter (GAT1) cDNA as a probe. The PCR technique was used to fill sequence gaps, and 5' and 3' RACE were employed to amplify the 5' and 3' untranslated regions. The amplified fragments were subcloned into a T-vector. Blots containing RNA from 10 different tissues were probed to determine the size of the transcript and the tissue distribution. RESULTS: Sequence analysis revealed that the skate retinal GABA transporter cDNA shared 72% identity with the mouse GABA transporter-1 at the DNA level and 80% identity at the amino acid level. Multiple sequence alignments showed that our sequence is closest to the Torpedo GABA transporter-1. Two transcripts, 4.5 and 7 kb, were detected in retina and possibly brain by RNA blot analysis. Fourteen introns were detected in the skate GABA transporter gene. CONCLUSIONS: We successfully isolated a full length GABA transporter cDNA from the retina of the skate. The size of the full length sequence of the skate retinal GABA transporter is in agreement with the size of the smaller transcript detected on RNA blots. The larger transcript observed on the RNA blot may be the result of either alternative splicing or utilization of a downstream poly A signal.

Animals↗

AAA+: A class of chaperone-like ATPases associated with the assembly, operation, and disassembly of protein complexes.

Using a combination of computer methods for iterative database searches and multiple sequence alignment, we show that protein sequences related to the AAA family of ATPases are far more prevalent than reported previously. Among these are regulatory components of Lon and Clp proteases, proteins involved in DNA replication, recombination, and restriction (including subunits of the origin recognition complex, replication factor C proteins, MCM DNA-licensing factors and the bacterial DnaA, RuvB, and McrB proteins), prokaryotic NtrC-related transcription regulators, the Bacillus sporulation protein SpoVJ, Mg2+, and Co2+ chelatases, the Halobacterium GvpN gas vesicle synthesis protein, dynein motor proteins, TorsinA, and Rubisco activase. Alignment of these sequences, in light of the structures of the clamp loader delta' subunit of Escherichia coli DNA polymerase III and the hexamerization component of N-ethylmaleimide-sensitive fusion protein, provides structural and mechanistic insights into these proteins, collectively designated the AAA+ class. Whole-genome analysis indicates that this class is ancient and has undergone considerable functional divergence prior to the emergence of the major divisions of life. These proteins often perform chaperone-like functions that assist in the assembly, operation, or disassembly of protein complexes. The hexameric architecture often associated with this class can provide a hole through which DNA or RNA can be thread; this may be important for assembly or remodeling of DNA-protein complexes.

Adenosine Triphosphatases↗

Multiple DNA and protein sequence alignment on a workstation and a supercomputer.

This paper describes a multiple alignment method using a workstation and supercomputer. The method is based on the alignment of a set of aligned sequences with the new sequence, and uses a recursive procedure of such alignment. The alignment is executed in a reasonable computation time on diverse levels from a workstation to a supercomputer, from the viewpoint of alignment results and computational speed by parallel processing. The application of the algorithm is illustrated by several examples of multiple alignment of 12 amino acid and DNA sequences of HIV (human immunodeficiency virus) env genes. Colour graphic programs on a workstation and parallel processing on a supercomputer are discussed.

Algorithms↗

CRASP: a program for analysis of coordinated substitutions in multiple alignments of protein sequences.

Recent results suggest that during evolution certain substitutions at protein sites may occur in a coordinated manner due to interactions between amino acid residues. Information on these coordinated substitutions may be useful for analysis of protein structure and function. CRASP is an Internet-available software tool for the detection and analysis of coordinated substitutions in multiple alignments of protein sequences. The approach is based on estimation of the correlation coefficient between the values of a physicochemical parameter at a pair of positions of sequence alignment. The program enables the user to detect and analyze pairwise relationships between amino acid substitutions at protein sequence positions, estimate the contribution of the coordinated substitutions to the evolutionary invariance or variability in integral protein physicochemical characteristics such as the net charge of protein residues and hydrophobic core volume. The CRASP program is available at http://wwwmgs.bionet.nsc.ru/mgs/programs/crasp/.

Amino Acid Substitution↗

On the dependence structure of sequence alignment scores calculated with multiple scoring matrices.

A common practice in protein sequence alignment is to try several scoring matrices until "something interesting'' is found. This leads to a multiple testing problem making p- and E-values hard to interpret. We focus on local alignment and propose to use logistic copula functions to model explicitly the dependence structure of scores obtained using different scoring matrices. By doing this, we obtain p-value correction factors when using more than one scoring matrix on the same sequences. Furthermore the parameter of the logistic copula can be interpreted as measure of dependence, providing insight concerning the relatedness of the scores from different matrices.

Journal Article↗

Automatic discovery of sub-molecular sequence domains in multi-aligned sequences: a dynamic programming algorithm for multiple alignment segmentation.

Automatic identification of sub-structures in multi-aligned sequences is of great importance for effective and objective structural/functional domain annotation, phylogenetic treeing and other molecular analyses. We present a segmentation algorithm that optimally partitions a given multi-alignment into a set of potentially biologically significant blocks, or segments. This algorithm applies dynamic programming and progressive optimization to the statistical profile of a multi-alignment in order to optimally demarcate relatively homogenous sub-regions. Using this algorithm, a large multi-alignment of eukaryotic 16S rRNA was analyzed. Three types of sequence patterns were identified automatically and efficiently: shared conserved domain; shared variable motif; and rare signature sequence. Results were consistent with the patterns identified through independent phylogenetic and structural approaches. This algorithm facilitates the automation of sequence-based molecular structural and evolutionary analyses through statistical modeling and high performance computation.

Algorithms↗

Protein sequence alignments: a strategy for the hierarchical analysis of residue conservation.

An algorithm is described for the systematic characterization of the physico-chemical properties seen at each position in a multiple protein sequence alignment. The new algorithm allows questions important in the design of mutagenesis experiments to be quickly answered since positions in the alignment that show unusual or interesting residue substitution patterns may be rapidly identified. The strategy is based on a flexible set-based description of amino acid properties, which is used to define the conservation between any group of amino acids. Sequences in the alignment are gathered into subgroups on the basis of sequence similarity, functional, evolutionary or other criteria. All pairs of subgroups are then compared to highlight positions that confer the unique features of each subgroup. The algorithm is encoded in the computer program AMAS (Analysis of Multiply Aligned Sequences) which provides a textual summary of the analysis and an annotated (boxed, shaded and/or coloured) multiple sequence alignment. The algorithm is illustrated by application to an alignment of 67 SH2 domains where patterns of conserved hydrophobic residues that constitute the protein core are highlighted. The analysis of charge conservation across annexin domains identifies the locations at which conserved charges change sign. The algorithm simplifies the analysis of multiple sequence data by condensing the mass of information present, and thus allows the rapid identification of substitutions of structural and functional importance.

Algorithms↗

The CHAOS/DIALIGN WWW server for multiple alignment of genomic sequences.

Cross-species sequence comparison is a powerful approach to analyze functional sites in genomic sequences and many discoveries have been made based on genomic alignments. Herein, we present a WWW-based software system for multiple alignment of large genomic sequences. Our server utilizes the previously developed combination of CHAOS and DIALIGN to achieve both speed and alignment accuracy. CHAOS is a fast database search tool that creates a list of local sequence similarities. These are used by DIALIGN as anchor points to speed up the final alignment procedure. The resulting alignment is returned to the user in different formats together with a list of anchor points found by CHAOS. The CHAOS/DIALIGN software is freely available at http://dialign.gobics.de/chaos-dialign-submission.

Genomics↗

DIALIGN: multiple DNA and protein sequence alignment at BiBiServ.

DIALIGN is a widely used software tool for multiple DNA and protein sequence alignment. The program combines local and global alignment features and can therefore be applied to sequence data that cannot be correctly aligned by more traditional approaches. DIALIGN is available online through Bielefeld Bioinformatics Server (BiBiServ). The downloadable version of the program offers several new program features. To compare the output of different alignment programs, we developed the program AltAVisT. Our software is available at http://bibiserv.TechFak.Uni-Bielefeld.DE/dialign/.

Algorithms↗

Multiple sequence threading: conditional gap placement.

Preliminary work to improve on the gap placement in a novel multiple sequence threading method is presented here. Using the globin family, we construct measures for the assessment of gaps in a multiple sequence threading alignment based on the structural comparison of two of the proteins in the family. We take into account information from multiple sequence alignments on both the structure and sequence side of the problem. This work shows the parameterization of the problem allowing the foundation to optimize and test a gap placement weight or penalty. Four states were considered: deleted structure, inserted sequence, gap ends in structure, and broken ends in sequence. Each of these states was analyzed for exposure, occupancy and secondary structure. These measures enable us to gain insight into the placement of gaps in a multiple sequence threading alignment. We analyzed the most extreme violations of these properties and found in these cases that most secondary structures are broken by gaps. However, sequence inserts in structure were never found in deeply buried positions. Most end separations were 3-4 A in excess of the minimum 6 A, although some were larger. We show that the maximum amount of observed secondary structure found in inserts was about half of the predicted structure (typically 5 and 10%, respectively). A similar trend occurred with observed and predicted exposure in inserts. Our eventual aim is to devise a weight incorporating these measures of gap placement to further refine our algorithm for the threading of sequence on structure.

Amino Acid Sequence↗

Multiple DNA and protein sequence alignment based on segment-to-segment comparison.

In this paper, a new way to think about, and to construct, pairwise as well as multiple alignments of DNA and protein sequences is proposed. Rather than forcing alignments to either align single residues or to introduce gaps by defining an alignment as a path running right from the source up to the sink in the associated dot-matrix diagram, we propose to consider alignments as consistent equivalence relations defined on the set of all positions occurring in all sequences under consideration. We also propose constructing alignments from whole segments exhibiting highly significant overall similarity rather than by aligning individual residues. Consequently, we present an alignment algorithm that (i) is based on segment-to-segment comparison instead of the commonly used residue-to-residue comparison and which (ii) avoids the well-known difficulties concerning the choice of appropriate gap penalties: gaps are not treated explicity, but remain as those parts of the sequences that do not belong to any of the aligned segments. Finally, we discuss the application of our algorithm to two test examples and compare it with commonly used alignment methods. As a first example, we aligned a set of 11 DNA sequences coding for functional helix-loop-helix proteins. Though the sequences show only low overall similarity, our program correctly aligned all of the 11 functional sites, which was a unique result among the methods tested. As a by-product, the reading frames of the sequences were identified. Next, we aligned a set of ribonuclease H proteins and compared our results with alignments produced by other programs as reported by McClure et al. [McClure, M. A., Vasi, T. K. & Fitch, W. M. (1994) Mol. Biol. Evol. 11, 571-592]. Our program was one of the best scoring programs. However, in contrast to other methods, our protein alignments are independent of user-defined parameters.

Algorithms↗

Improvement of TRANSFAC matrices using multiple local alignment of transcription factor binding site sequences.

This paper describes a novel approach to constructing Position-Specific Weight Matrices (PWMs) based on the transcription factor binding site (TFBS) data provide by the TRANSFAC database and comparison of the newly generated PWMs with the original TRANSFAC matrices. Multiple local sequence alignment was performed on the TFBSs of each transcription factor. Several different alignment programs were tested and their matrices were compared to the original TRANSFAC matrices. One of the alignment programs, GLAM, produced comparable matrices in terms of the average ranking of true positive sites across the whole test set of sequences.

Algorithms↗

ARCS: an aggregated related column scoring scheme for aligned sequences.

MOTIVATION: Biologists frequently align multiple biological sequences to determine consensus sequences and/or search for predominant residues and conserved regions. Particularly, determining conserved regions in an alignment is one of the most important activities. Since protein sequences are often several-hundred residues or longer, it is difficult to distinguish biologically important conserved regions (motifs or domains) from others. The widely used tools, Logos, Al2co, Confind, and the entropy-based method, often fail to highlight such regions. Thus a computational tool that can highlight biologically important regions accurately will be highly desired. RESULTS: This paper presents a new scoring scheme ARCS (Aggregated Related Column Score) for aligned biological sequences. ARCS method considers not only the traditional character similarity measure but also column correlation. In an extensive experimental evaluation using 533 PROSITE patterns, ARCS is able to highlight the motif regions with up to 77.7% accuracy corresponding to the top three peaks. AVAILABILITY: The source code is available on http://bio.informatics.indiana.edu/projects/arcs and http://goldengate.case.edu/projects/arcs

Algorithms↗

PhyloGibbs: a Gibbs sampling motif finder that incorporates phylogeny.

A central problem in the bioinformatics of gene regulation is to find the binding sites for regulatory proteins. One of the most promising approaches toward identifying these short and fuzzy sequence patterns is the comparative analysis of orthologous intergenic regions of related species. This analysis is complicated by various factors. First, one needs to take the phylogenetic relationship between the species into account in order to distinguish conservation that is due to the occurrence of functional sites from spurious conservation that is due to evolutionary proximity. Second, one has to deal with the complexities of multiple alignments of orthologous intergenic regions, and one has to consider the possibility that functional sites may occur outside of conserved segments. Here we present a new motif sampling algorithm, PhyloGibbs, that runs on arbitrary collections of multiple local sequence alignments of orthologous sequences. The algorithm searches over all ways in which an arbitrary number of binding sites for an arbitrary number of transcription factors (TFs) can be assigned to the multiple sequence alignments. These binding site configurations are scored by a Bayesian probabilistic model that treats aligned sequences by a model for the evolution of binding sites and "background" intergenic DNA. This model takes the phylogenetic relationship between the species in the alignment explicitly into account. The algorithm uses simulated annealing and Monte Carlo Markov-chain sampling to rigorously assign posterior probabilities to all the binding sites that it reports. In tests on synthetic data and real data from five Saccharomyces species our algorithm performs significantly better than four other motif-finding algorithms, including algorithms that also take phylogeny into account. Our results also show that, in contrast to the other algorithms, PhyloGibbs can make realistic estimates of the reliability of its predictions. Our tests suggest that, running on the five-species multiple alignment of a single gene's upstream region, PhyloGibbs on average recovers over 50% of all binding sites in S. cerevisiae at a specificity of about 50%, and 33% of all binding sites at a specificity of about 85%. We also tested PhyloGibbs on collections of multiple alignments of intergenic regions that were recently annotated, based on ChIP-on-chip data, to contain binding sites for the same TF. We compared PhyloGibbs's results with the previous analysis of these data using six other motif-finding algorithms. For 16 of 21 TFs for which all other motif-finding methods failed to find a significant motif, PhyloGibbs did recover a motif that matches the literature consensus. In 11 cases where there was disagreement in the results we compiled lists of known target genes from the literature, and found that running PhyloGibbs on their regulatory regions yielded a binding motif matching the literature consensus in all but one of the cases. Interestingly, these literature gene lists had little overlap with the targets annotated based on the ChIP-on-chip data. The PhyloGibbs code can be downloaded from http://www.biozentrum.unibas.ch/~nimwegen/cgi-bin/phylogibbs.cgi or http://www.imsc.res.in/~rsidd/phylogibbs. The full set of predicted sites from our tests on yeast are available at http://www.swissregulon.unibas.ch.

Algorithms↗