Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals↗

A method for simultaneous alignment of multiple protein structures.

Here, we present MultiProt, a fully automated highly efficient technique to detect multiple structural alignments of protein structures. MultiProt finds the common geometrical cores between input molecules. To date, most methods for multiple alignment start from the pairwise alignment solutions. This may lead to a small overall alignment. In contrast, our method derives multiple alignments from simultaneous superpositions of input molecules. Further, our method does not require that all input molecules participate in the alignment. Actually, it efficiently detects high scoring partial multiple alignments for all possible number of molecules in the input. To demonstrate the power of MultiProt, we provide a number of case studies. First, we demonstrate known multiple alignments of protein structures to illustrate the performance of MultiProt. Next, we present various biological applications. These include: (1) a partial alignment of hinge-bent domains; (2) identification of functional groups of G-proteins; (3) analysis of binding sites; and (4) protein-protein interface alignment. Some applications preserve the sequence order of the residues in the alignment, whereas others are order-independent. It is their residue sequence order-independence that allows application of MultiProt to derive multiple alignments of binding sites and of protein-protein interfaces, making MultiProt an extremely useful structural tool.

Algorithms↗

Neighboring base composition and transversion/transition bias in a comparison of rice and maize chloroplast noncoding regions.

The correspondence between the transversion/transition ratio and the neighboring base composition in chloroplast DNA is examined. For 18 noncoding regions of the chloroplast genome, alignments between rice (Oryza sativa) and maize (Zea mays) were generated by two different methods. Difficulties of aligning noncoding DNA are discussed, and the alignments are analyzed in a manner that reduces alignment artifacts. Sequence divergence is < 10%, so multiple substitutions at a site are assumed to be rare. Observed substitutions were analyzed with respect to the A+T content of the two immediately flanking bases. It is shown that as this content increases, the proportion of transversions also increases. When both the 5'- and 3'-flanking nucleotides are G or C (A+T content of 0), only 25% of the observed substitutions are transversions. However, when both the 5'- and 3'-flanking nucleotides are A or T (A+T content of 2), 57% of the observed substitutions are transversions. Therefore, the influence of flanking base composition on substitutions, previously reported for a single noncoding region, is a general feature of the chloroplast genome.

Algorithms↗

Extensive sequence conservation among insect, nematode, and vertebrate vitellogenins reveals ancient common ancestry.

The eggs of most oviparous animals are provisioned with a class of protein called vitellogenin (Vg) which is stored as the major component of yolk. Until recently, deduced amino acid sequences were available only from vertebrate and nematode Vgs, which proved to be homologous. The sequences of several insect Vgs are now known, but early attempts at pairwise alignments with vertebrate and nematode Vgs have been problematic, leading to conflicting conclusions about how closely insect Vgs are related to the others. In this paper we demonstrate that insect Vg sequences can be confidently aligned with one another along their entire lengths and with multiple vertebrate and nematode Vg sequences along most of their spans. Although divergence is high, conservation among insect, vertebrate, and nematode Vg sequences is widespread with a preponderance of glycine, proline, and cysteine residues among strictly conserved amino acids, establishing conclusively that Vgs from the three phyla are homologous. Areas of least-certain alignment are primarily in and around insect and vertebrate polyserine domains which are not homologous. Phylogenetic reconstructions of Vgs based on sequence identities indicate that the insect lineage is the most diverged and that the mammalian serum protein, apolipoprotein B-100, arose from a Vg ancestor after the nematode/vertebrate divergence.

Amino Acid Sequence↗

T-cell-epitope mapping of the idiotypic monoclonal IgG heavy and light chains in multiple myeloma.

The idiotypic structures of the myeloma protein might be regarded as tumor-specific antigens. The present study was designed to map T-cell epitopes of the idiotypic myeloma protein to prove the existence of naturally occurring major-histocompatibility-complex-dependent idiotype (peptide)-specific T cells in multiple myeloma. The fine specificity of idiotype-reactive, interferon-gamma-producing blood T cells of a patient with multiple myeloma stage I was characterized by identification of idiotype (heavy and light chains)-derived MHC-restricted T-cell epitopes. T cells specifically reacting with peptides corresponding to each of the 3 complementarity-determining regions (CDRs) of the heavy-chain variable part (V(H)) of the autologous idiotype were found. In contrast, none of the peptides corresponding to the 3 CDRs of the light chain (V(L)) induced a specific T-cell response. The idiotype amino-acid sequence corresponding to the junction of the V(H), diversity (D), and joining (J) gene segments of the VH appeared to be an important target for T cells, since the sequence expressed MHC-class-I- as well as MHC-class-II-restricted epitopes. The study provides further support for the existence of MHC-restricted idiotype-specific T cells, which may target immunogenic CDR peptides in multiple myeloma. Such T cells could be an important part of the specific anti-tumor immune responses induced in idiotype vaccination protocols.

Aged↗

Evolutionary correlation between linker histones and microtubular structures.

Histones of the H1 group (linker histones) are abundant components of chromatin in eukaryotes, occurring on average at one molecule per nucleosome. The recent reports on the lack of a clear phenotypic effect of knock-out mutations as well as overexpression of histone H1 genes in different organisms have seriously undermined the long-held view that linker histones are essential for the basic functions of eukaryotic cells. In an attempt to resolve the paradox of an abundant conserved protein without a clear function, we re-examined the molecular and phylogenetic data on linker histones to see if they could reveal any correlation between the features of H1 and the functional or morphological characteristics of cells or organisms. Because of an earlier demonstration that in sea urchin the chromatin-type histone H1 is also found in the flagellar microtubules (Multigner et al. 1992), we focused on the correlation between the features of H1 and those of microtubular structures. A phylogenetic tree based on multiple alignment of over 100 available HI sequences suggests that the first divergence of the globular domain of H1 (GH1) resulted in branching into separate types characteristic for plants/Dictyostelium and for animals/ascomycetes, respectively. The GH1s of these two types differ by a short region (usually 5 amino acids) placed at a specific location within the C-terminal wing subdomain of GH1. Evolutionary analysis of the diversification of H1 mRNA into cell-cycle-dependent (polyA-) and independent (polyA+) forms showed a mosaic occurrence of these two forms in plants and animals, despite the fact that the H1 proteins of plants and animals belong to two well-distinguished groups. However, among organisms from both animal and plant kingdom, only those with H1 mRNA of a polyA- type have flagellated gametes. This correlation as well as the demonstration that in Volvox carteri the accumulation of polyA- mRNA of H1 occurs concurrently with the production of new flagella (Lindauer et al. 1993), suggests a direct link between polyA- phenotype of histone H1 mRNA and flagellogenesis.

Animals↗

The role of N286 and D320 in the reaction mechanism of human dihydrolipoamide dehydrogenase (E3) center domain.

According to the multiple alignment of various dihydrolipoamide dehydrogenases (E3s) sequences, three human mutant E3s of the conserved residues in the center domain, N286D, N286Q, and D320N were created, over-expressed and purified. We characterized these mutants to investigate the reaction mechanism of human dihydrolipoamide dehydrogenases. The specific activities of N286D, N286Q, and D320N are 30.84%, 24.57% and 48.60% to that of the wild-type E3 respectively. The FAD content analysis indicated that these mutant E3s about 96.0%, 99.4% and 82.7% of FAD content compared to that of wild-type E3 respectively. The molecular weight analysis showed that these three mutant proteins form the dimer. Kinetic's data demonstrated that the K(cat) of both forward and reverse reactions of these mutant proteins were decreased. These results suggest that N286 and D320 play a role in the catalytic function of the E3.

Amino Acid Sequence↗

Immune cascade of Spodoptera litura: cloning, expression, and characterization of inducible prophenol oxidase.

Haemolymph associated phenol oxidase is a critical component of invertebrate immune reaction and cuticle sclerotization. Phenol oxidase catalyses the conversion of mono-phenols to diphenols and quinones which finally leads to melanin formation. We have cloned the c-DNA encoding phenol oxidase from the haemocytes of Spodoptera litura and expressed it in Escherichia coli. The encoding gene is 2452bp with an open reading frame of 2091 bp translating into a 697 amino acid protein. Multiple alignment analysis of the predicted protein sequence shows close homology to other lepidopeteran PPOII type genes. The transcription of the gene is induced upon microbial challenge of 6th instar larvae with E. coli and is unresponsive to injury. Cloning of the ORF of SLPPO in-frame in the E. coli expression vector pQE30 resulted in its expression. Enzymatic analysis of the recombinant protein reveals that the recombinant protein is catalytically active on 4-methyl pyrocatechol upon activation by cetyl pyridinium chloride.

Amino Acid Sequence↗

Improving profile HMM discrimination by adapting transition probabilities.

Profile hidden Markov models (HMMs) are used to model protein families and for detecting evolutionary relationships between proteins. Such a profile HMM is typically constructed from a multiple alignment of a set of related sequences. Transition probability parameters in an HMM are used to model insertions and deletions in the alignment. We show here that taking into account unrelated sequences when estimating the transition probability parameters helps to construct more discriminative models for the global/local alignment mode. After normal HMM training, a simple heuristic is employed that adjusts the transition probabilities between match and delete states according to observed transitions in the training set relative to the unrelated (noise) set. The method is called adaptive transition probabilities (ATP) and is based on the HMMER package implementation. It was benchmarked in two remote homology tests based on the Pfam and the SCOP classifications. Compared to the HMMER default procedure, the rate of misclassification was reduced significantly in both tests and across all levels of error rate.

Algorithms↗

Characterization of a recombinant immunodiagnostic antigen (NIE) from Strongyloides stercoralis L3-stage larvae.

Due to the process of internal autoinfection, even chronic asymptomatic infections with Strongyloides stercoralis have the potential to become severe disseminated disease with fatal outcome. Intermittent and scanty larval excretion makes parasitologic diagnosis difficult. Serodiagnosis is helpful, but antigen preparation from infective larvae requires access to patients or immunosuppressed experimental animals. For these reasons, attention has turned to recombinant antigens for immunodiagnosis. A 31-kDa candidate antigen (NIE) derived from an L3 cDNA library is described in this report. Multiple alignment of the deduced amino acid sequence of NIE showed approximately 12-18% identity with various other organisms, including 17.9% of Asp1 of Ancylostoma caninum, 12.6% of Hemonchus contortus, and 17.6% of insect venom allergen 5 of yellow jacket. By ELISA, antibodies to the purified recombinant NIE antigen were demonstrated in 87.5% of 48 sera from strongyloides-infected patients and in only 6.5% of sera from presumed normal controls. Immunoreactivity of purified NIE antigen with parasite-specific IgE from sera of strongyloides-infected patients indicated its potential use as an immediate sensitivity skin test antigen. This application of the NIE antigen was supported by its capacity to trigger release of histamine upon in vitro exposure to blood from strongyloides-infected patients and its failure to produce histamine release from blood of normal controls.

Amino Acid Sequence↗

Involvement of Gln937 of Streptococcus downei GTF-I glucansucrase in transition-state stabilization.

Multiple alignment of deduced amino-acid sequences of glucansucrases (glucosyltransferases and dextransucrases) from oral streptococci and Leuconostoc mesenteroides has shown them to share a well-conserved catalytic domain. A portion of this domain displays homology to members of the alpha-amylase family (glycoside hydrolase family 13), which all have a (beta/alpha)8 barrel structure. In the glucansucrases, however, the alpha-helix and beta-strand elements are circularly permuted with respect to the order in family 13. Previous work has shown that amino-acid residues contributing to the active site of glucansucrases are situated in structural elements that align with those of family 13. In alpha-amylase and cyclodextrin glucanotransferase, a histidine residue has been identified that acts to stabilize the transition state, and a histidine is conserved at the corresponding position in all other members of family 13. In all the glucansucrases, however, the aligned position is occupied by glutamine. Mutants of glucosyltransferase I were constructed in which this glutamine, Gln937, was changed to histidine, glutamic acid, aspartic acid, asparagine or alanine. The effects on specific activity, ability to form glucan and ability to transfer glucose to a maltose acceptor were examined. Only histidine could substitute for glutamine and maintain Michaelis-Menten kinetics, albeit at a greatly reduced kcat, showing that Gln937 plays a functionally equivalent role to the histidine in family 13. This provides additional evidence in support of the proposed alignment of the (beta/alpha)8 barrel structures. Mutation at position 937 altered the acceptor reaction with maltose, and resulted in the synthesis of novel gluco-oligosaccharides in which alpha1,3-linked glucosyl units are joined sequentially to maltose.

Enzyme Stability↗

Changing a single amino acid residue switches processive and non-processive behavior of Aspergillus niger endopolygalacturonase I and II.

Processivity, also known as multiple attack on a single chain, is a feature commonly encountered only in enzymes in which the substrate binds in a tunnel. However, of the seven Aspergillus niger endopolygalacturonases, which have an open substrate binding cleft, four enzymes show processive behavior, whereas the other endopolygalacturonases are randomly acting enzymes. In a previous study (Benen, J.A.E., Kester, H.C.M., and Visser, J. (1999) Eur. J. Biochem. 259, 577-585) we proposed that the high affinity for the substrate of subsite -5 of processive endopolygalacturonase I constitutes the origin of the multiple attack behavior. Based on primary sequence alignments of A. niger endopolygalacturonases and three-dimensional structure analysis of endopolygalacturonase II, an arginine residue was identified in the processive enzymes at a position commensurate with subsite -5, whereas a serine residue was present at this position in the non-processive enzymes. In endopolygalacturonase I mutation R95S was introduced, and in endopolygalacturonase II mutation S91R was introduced. Product progression analysis on polymer substrate and bond cleavage frequency studies using oligogalacturonides of defined chain length for the mutant enzymes revealed that processive/non-processive behavior is indeed interchangeable by one single amino acid substitution at subsite -5, Arg-->Ser or Ser-->Arg.

Amino Acid Sequence↗

Detection of BCR-ABL mutations and resistance to imatinib mesylate.

The major mechanism of imatinib resistance for patients with chronic myeloid leukemia (CML) is clonal expansion of leukemic cells with mutations in the Bcr-Abl fusion tyrosine kinase that reduce the capacity of imatinib to inhibit kinase activity. The early detection of such mutations may allow timely treatment intervention to prevent or overcome resistance. Direct sequencing of the BCR-ABL kinase domain is relatively rapid and allows detection of emerging mutations at a sensitivity of approx 20%. Mutations have been detected over a range of 242 amino acids, which spans the entire kinase domain. For optimal sensitivity, the kinase domain of the abnormal gene should be isolated by reverse-transcription (RT) polymerase chain reaction (PCR) amplification using primers that hybridize to the BCR and ABL genes. The quality of the RNA is assessed by real-time quantitative PCR prior to analysis, and BCR-ABL levels are determined. Only RNA of adequate quality is used to ensure accurate and reproducible mutation analysis. Depending on the level of BCR-ABL transcripts, a one- or two-step PCR is required to amplify the kinase domain. Direct sequencing with dye terminator chemistry is performed using PCR-purified products. The sequence is compared to an ABL kinase domain reference sequence using sequencing analysis software, which aligns the sequences and highlights single or multiple mutations.

Amino Acid Substitution↗

The deferred path heuristic for the generalized tree alignment problem.

Many multiple alignment methods implicitly or explicitly try to minimize the amount of biological change implied by an alignment. At the level of sequences, biological change is measured along a phylogenetic tree, a structure frequently being predicted only after the multiple alignment instead of together with it. The Generalized Tree Alignment problem addresses both questions simultaneously. It can formally be viewed as a Steiner tree problem in sequence space and our approach merges a path heuristic for the construction of a Steiner tree with a clustering method as usually applied only to distance data. This combination is achieved using sequence graphs, a data structure for efficient representation of similar sequences. Although somewhat slower in practice than an earlier method by Hein (1989) the current approach achieves significantly better results in terms of the underlying scoring function. Furthermore, a variant of the algorithm is introduced that maintains a guaranteed error bound of (2 - 2/n) for n sequences.

Algorithms↗

Improved genetic algorithm-based protein structure comparisons: pairwise and multiple superpositions.

Three major improvements to a previously described method for automatic protein structure comparison are described. First, a limit to translations for the rigid-body superposition is now assigned according to the dimensions of the structures being compared. Second, examination of the effect of the gap penalty on the derivation of a sequence alignment corresponding to a given structure superposition has led to a method to evaluate alternative structure-based sequence alignments. Third, the pairwise procedure has been generalized to multiple structure alignment. This implementation of rigid-body superposition can recognize well documented distant relationships which hitherto have required consideration of additional features and properties as well as those relationships between proteins of different sizes. A much larger common scaffold or framework between six globins can be extracted than that obtained using a standard algorithm for multiple structure superposition.

Algorithms↗

Finding the most significant common sequence and structure motifs in a set of RNA sequences.

We present a computational scheme to locally align a collection of RNA sequences using sequence and structure constraints. In addition, the method searches for the resulting alignments with the most significant common motifs, among all possible collections. The first part utilizes a simplified version of the Sankoff algorithm for simultaneous folding and alignment of RNA sequences, but maintains tractability by constructing multi-sequence alignments from pairwise comparisons. The algorithm finds the multiple alignments using a greedy approach and has similarities to both CLUSTAL and CONSENSUS, but the core algorithm assures that the pairwise alignments are optimized for both sequence and structure conservation. The choice of scoring system and the method of progressively constructing the final solution are important considerations that are discussed. Example solutions, and comparisons with other approaches, are provided. The solutions include finding consensus structures identical to published ones.

Algorithms↗

BLAST Filter and GraphAlign: rule-based formation and analysis of sets of related DNA and protein sequences.

BLAST Filter and GraphAlign are web-based tools that offer novel methods for building and analyzing sets of related (i.e. similar) DNA and protein sequences. They can be used separately or together. BLAST Filter generates related sequence sets in an automated, objective and reproducible way based on an input query sequence. Sequences matched by BLAST are filtered through a set of 15 user-configurable rules based on full-length query/subject comparisons, high-scoring segment pair statistics and the level of redundancy in the sequence set. Such sets can be used for multiple alignments, profile hidden Markov models and other bioinformatics applications, including GraphAlign, which provides several novel methods for analyzing global query/subject alignments along with graphical representations of sequence similarities. These services are available at the following URLs: http://darwin.nmsu.edu/cgi-bin/blast_filter.cgi and http://darwin.nmsu.edu/cgi-bin/graph_align.cgi.

Computational Biology↗

The phylogenetic analysis of variable-length sequence data: elongation factor-1alpha introns in European populations of the parasitoid wasp genus Pauesia (Hymenoptera: Braconidae: Aphidiinae).

Elongation factor-1alpha (EF-1alpha) is a highly conserved nuclear coding gene that can be used to investigate recent divergences due to the presence of rapidly evolving introns. However, a universal feature of intron sequences is that even closely related species exhibit insertion and deletion events, which cause variation in the lengths of the sequences. Indels are frequently rich in evolutionary information, but most investigators ignore sites that fall within these variable regions, largely because the analytical tools and theory are not well developed. We examined this problem in the taxonomically problematic parasitoid wasp genus Pauesia (Hymenoptera: Braconidae: Aphidiinae) using congruence as a criterion for assessing a range of methods for aligning such variable-length EF-1alpha intron sequences. These methods included distance- and parsimony-based multiple-alignment programs (CLUSTAL W and MALIGN), direct optimization (POY), and two "by eye" alignment strategies. Furthermore, with one method (CLUSTAL W) we explored in detail the robustness of results to changes in the gap cost parameters. Phenetic-based alignments ("by eye" and CLUSTAL W) appeared, under our criterion, to perform as well as more readily defensible, but computationally more demanding, methods. In general, all of our alignment and tree-building strategies recovered the same basic topological structure, which means that an underlying phylogenetic signal remained regardless of the strategy chosen. However, several relationships between clades were sensitive both to alignment and to tree-building protocol. Further alignments, considering only sequences belonging to the same group, allowed us to infer a range of phylogenetic relationships that were highly robust to tree-building protocol. By comparing these topologies with those obtained by varying the CLUSTAL parameters, we generated the distribution area of congruence and taxonomic compatibility. Finally, we present the first robust estimate of the European Pauesia phylogeny by using two EF-1alpha introns and 38 taxa (plus 3 outgroups). This estimate conflicts markedly with the traditional subgeneric classification. We recommend that this classification be abandoned, and we propose a series of monophyletic species groups.

Animals↗