Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

A novel approach to phylogeny reconstruction from protein sequences.

The reliable reconstruction of tree topology from a set of homologous sequences is one of the main goals in the study of molecular evolution. If consistent estimators of distances from a multiple sequence alignment are known, the distance method is attractive because the tree reconstruction is consistent. To obtain a distance estimate d, the observed proportion of differences p (p-distance) is usually "corrected" for multiple and back substitutions by means of a functional relationship d = f(p). In this paper the conditions under which this correction of p-distances will not alter the selection of the tree topology are specified. When these conditions are not fulfilled the selection of the tree topology may depend on the correction function applied. A novel method which includes estimates of distances not only between sequence pairs, but between triplets, quadruplets, etc., is proposed to strengthen the proper selection of correction function and tree topology. A "super" tree that includes all tree topologies as special cases is introduced.

Animals↗

Multiple protein structure alignment.

A method was developed to compare protein structures and to combine them into a multiple structure consensus. Previous methods of multiple structure comparison have only concatenated pairwise alignments or produced a consensus structure by averaging coordinate sets. The current method is a fusion of the fast structure comparison program SSAP and the multiple sequence alignment program MULTAL. As in MULTAL, structures are progressively combined, producing intermediate consensus structures that are compared directly to each other and all remaining single structures. This leads to a hierarchic "condensation," continually evaluated in the light of the emerging conserved core regions. Following the SSAP approach, all interatomic vectors were retained with well-conserved regions distinguished by coherent vector bundles (the structural equivalent of a conserved sequence position). Each bundle of vectors is summarized by a resultant, whereas vector coherence is captured in an error term, which is the only distinction between conserved and variable positions. Resultant vectors are used directly in the comparison, which is weighted by their error values, giving greater importance to the matching of conserved positions. The resultant vectors and their errors can also be used directly in molecular modeling. Applications of the method were assessed by the quality of the resulting sequence alignments, phylogenetic tree construction, and databank scanning with the consensus. Visual assessment of the structural superpositions and consensus structure for various well-characterized families confirmed that the consensus had identified a reasonable core.

Amino Acid Sequence↗

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural (3-D) and sequence (1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in SwissProt using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 29% of all SwissProt-stored sequences.

Amino Acid Sequence↗

Local multiple alignment of numerical sequences: detection of subtle motifs from protein sequences and structures.

This paper presents a new method to find motifs from multiple protein sequences and multiple protein structures. The method consists of two parts: quantification and local multiple alignment. In the former part, protein sequences and protein structures are transformed into sequences of real numbers and real vectors respectively. In the latter part, fixed length regions having similar shapes are located. A Gibbs sampling algorithm for sequences of real numbers/vectors is newly developed for finding common regions. The results of the comparison with a standard Gibbs sampling program show that the method is particularly useful when structural information is available.

Algorithms↗

SnapDRAGON: a method to delineate protein structural domains from sequence data.

We describe a method to identify protein domain boundaries from sequence information alone based on the assumption that hydrophobic residues cluster together in space. SnapDRAGON is a suite of programs developed to predict domain boundaries based on the consistency observed in a set of alternative ab initio three-dimensional (3D) models generated for a given protein multiple sequence alignment. This is achieved by running a distance geometry-based folding technique in conjunction with a 3D-domain assignment algorithm. The overall accuracy of our method in predicting the number of domains for a non-redundant data set of 414 multiple alignments, representing 185 single and 231 multiple-domain proteins, is 72.4 %. Using domain linker regions observed in the tertiary structures associated with each query alignment as the standard of truth, inter-domain boundary positions are delineated with an accuracy of 63.9 % for proteins comprising continuous domains only, and 35.4 % for proteins with discontinuous domains. Overall, domain boundaries are delineated with an accuracy of 51.8 %. The prediction accuracy values are independent of the pair-wise sequence similarities within each of the alignments. These results demonstrate the capability of our method to delineate domains in protein sequences associated with a wide variety of structural domain organisation.

Algorithms↗

Identification of multiple distinct Snf2 subfamilies with conserved structural motifs.

The Snf2 family of helicase-related proteins includes the catalytic subunits of ATP-dependent chromatin remodelling complexes found in all eukaryotes. These act to regulate the structure and dynamic properties of chromatin and so influence a broad range of nuclear processes. We have exploited progress in genome sequencing to assemble a comprehensive catalogue of over 1300 Snf2 family members. Multiple sequence alignment of the helicase-related regions enables 24 distinct subfamilies to be identified, a considerable expansion over earlier surveys. Where information is known, there is a good correlation between biological or biochemical function and these assignments, suggesting Snf2 family motor domains are tuned for specific tasks. Scanning of complete genomes reveals all eukaryotes contain members of multiple subfamilies, whereas they are less common and not ubiquitous in eubacteria or archaea. The large sample of Snf2 proteins enables additional distinguishing conserved sequence blocks within the helicase-like motor to be identified. The establishment of a phylogeny for Snf2 proteins provides an opportunity to make informed assignments of function, and the identification of conserved motifs provides a framework for understanding the mechanisms by which these proteins function.

Adenosine Triphosphatases↗

An expectation maximization algorithm for training hidden substitution models.

We derive an expectation maximization algorithm for maximum-likelihood training of substitution rate matrices from multiple sequence alignments. The algorithm can be used to train hidden substitution models, where the structural context of a residue is treated as a hidden variable that can evolve over time. We used the algorithm to train hidden substitution matrices on protein alignments in the Pfam database. Measuring the accuracy of multiple alignment algorithms with reference to BAliBASE (a database of structural reference alignments) our substitution matrices consistently outperform the PAM series, with the improvement steadily increasing as up to four hidden site classes are added. We discuss several applications of this algorithm in bioinformatics.

Algorithms↗

The HSSP database of protein structure-sequence alignments and family profiles.

HSSP (http: //www.sander.embl-ebi.ac.uk/hssp/) is a derived database merging structure (3-D) and sequence (1-D) information. For each protein of known 3D structure from the Protein Data Bank (PDB), we provide a multiple sequence alignment of putative homologues and a sequence profile characteristic of the protein family, centered on the known structure. The list of homologues is the result of an iterative database search in SWISS-PROT using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed putative homologues are very likely to have the same 3D structure as the PDB protein to which they have been aligned. As a result, the database not only provides aligned sequence families, but also implies secondary and tertiary structures covering 33% of all sequences in SWISS-PROT.

Computer Communication Networks↗

Improved alignment of nucleosome DNA sequences using a mixture model.

DNA sequences that are present in nucleosomes have a preferential approximately 10 bp periodicity of certain dinucleotide signals, but the overall sequence similarity of the nucleosomal DNA is weak, and traditional multiple sequence alignment tools fail to yield meaningful alignments. We develop a mixture model that characterizes the known dinucleotide periodicity probabilistically to improve the alignment of nucleosomal DNAs. We assume that a periodic dinucleotide signal of any type emits according to a probability distribution around a series of 'hot spots' that are equally spaced along nucleosomal DNA with 10 bp period, but with a 1 bp phase shift across the middle of the nucleosome. We model the three statistically most significant dinucleotide signals, AA/TT, GC and TA, simultaneously, while allowing phase shifts between the signals. The alignment is obtained by maximizing the likelihood of both Watson and Crick strands simultaneously. The resulting alignment of 177 chicken nucleosomal DNA sequences revealed that all 10 distinct dinucleotides are periodic, however, with only two distinct phases and varying intensity. By Fourier analysis, we show that our new alignment has enhanced periodicity and sequence identity compared with center alignment. The significance of the nucleosomal DNA sequence alignment is evaluated by comparing it with that obtained using the same model on non-nucleosomal sequences.

Algorithms↗

The HSSP database of protein structure-sequence alignments.

HSSP is a derived database merging structural three dimensional (3-D) and sequence one dimensional(1-D) information. For each protein of known 3-D structure from the Protein Data Bank (PDB), the database has a multiple sequence alignment of all available homologues and a sequence profile characteristic of the family. The list of homologues is the result of a database search in Swissprot using a position-weighted dynamic programming method for sequence profile alignment (MaxHom). The database is updated frequently. The listed homologues are very likely to have the same 3-D structure as the PDB protein to which they have been aligned. As a result, the database is not only a database of aligned sequence families, but also a database of implied secondary and tertiary structures covering 27% of all Swissprot-stored sequences.

Amino Acid Sequence↗

PANDIT: an evolution-centric database of protein and associated nucleotide domains with inferred trees.

PANDIT is a database of homologous sequence alignments accompanied by estimates of their corresponding phylogenetic trees. It provides a valuable resource to those studying phylogenetic methodology and the evolution of coding-DNA and protein sequences. Currently in version 17.0, PANDIT comprises 7738 families of homologous protein domains; for each family, DNA and corresponding amino acid sequence multiple alignments are available together with high quality phylogenetic tree estimates. Recent improvements include expanded methods for phylogenetic tree inference, assessment of alignment quality and a redesigned web interface, available at the URL http://www.ebi.ac.uk/goldman-srv/pandit.

Databases, Nucleic Acid↗

Assessing the odd secondary structural properties of nuclear small subunit ribosomal RNA sequences (18S) of the twisted-wing parasites (Insecta: Strepsiptera).

We report the entire sequence (2864 nts) and secondary structure of the nuclear small subunit ribosomal RNA (SSU rRNA) gene (18S) from the twisted-wing parasite Caenocholax fenyesi texensis Kathirithamby & Johnston (Strepsiptera: Myrmecolacidae). The majority of the base pairings in this structural model map on to the SSU rRNA secondary and tertiary helices that were previously predicted with comparative analysis. These regions of the core rRNA were unambiguously aligned across all Arthropoda. In contrast, many of the variable regions, as previously characterized in other insect taxa, had very large insertions in C. f. texensis. The helical base pairs in these regions were predicted with a comparative analysis of a multiple sequence alignment (that contains C. f. texensis and 174 published arthropod 18S rRNA sequences, including eleven strepsipterans) and thermodynamic-based algorithms. Analysis of our structural alignment revealed four unusual insertions in the core rRNA structure that are unique to animal 18S rRNA and in general agreement with previously proposed insertion sites for strepsipterans. One curious result is the presence of a large insertion within a hairpin loop of a highly conserved pseudoknot helix in variable region 4. Despite the extraordinary variability in sequence length and composition, this insertion contains the conserved sequences 5'-AUUGGCUUAAA-3' and 5'-GAC-3' that immediately flank a putative helix at the 5'- and 3'-ends, respectively. The longer sequence has the potential to form a nine base pair helix with a sequence in the variable region 2, consistent with a recent study proposing this tertiary interaction. Our analysis of a larger set of arthropod 18S rRNA sequences has revealed possible errors in some of the previously published strepsipteran 18S rRNA sequences. Thus we find no support for the previously recovered heterogeneity in the 18S molecules of strepsipterans. Our findings lend insight to the evolution of RNA structure and function and the impact large insertions pose on genome size. We also provide a novel alignment template that will improve the phylogenetic placement of the Strepsiptera among other insect taxa.

Animals↗

Detecting overlapping coding sequences in virus genomes.

BACKGROUND: Detecting new coding sequences (CDSs) in viral genomes can be difficult for several reasons. The typically compact genomes often contain a number of overlapping coding and non-coding functional elements, which can result in unusual patterns of codon usage; conservation between related sequences can be difficult to interpret--especially within overlapping genes; and viruses often employ non-canonical translational mechanisms--e.g. frameshifting, stop codon read-through, leaky-scanning and internal ribosome entry sites--which can conceal potentially coding open reading frames (ORFs). RESULTS: In a previous paper we introduced a new statistic--MLOGD (Maximum Likelihood Overlapping Gene Detector)--for detecting and analysing overlapping CDSs. Here we present (a) an improved MLOGD statistic, (b) a greatly extended suite of software using MLOGD, (c) a database of results for 640 virus sequence alignments, and (d) a web-interface to the software and database. Tests show that, from an alignment with just 20 mutations, MLOGD can discriminate non-overlapping CDSs from non-coding ORFs with a typical accuracy of up to 98%, and can detect CDSs overlapping known CDSs with a typical accuracy of 90%. In addition, the software produces a variety of statistics and graphics, useful for analysing an input multiple sequence alignment. CONCLUSION: MLOGD is an easy-to-use tool for virus genome annotation, detecting new CDSs--in particular overlapping or short CDSs--and for analysing overlapping CDSs following frameshift sites. The software, web-server, database and supplementary material are available at http://guinevere.otago.ac.nz/mlogd.html.

Algorithms↗

Identification and phylogenetic comparison of Salem virus, a novel paramyxovirus of horses.

A virus that could not be identified as a previously known equine virus was isolated from the mononuclear cells of a horse. Electron microscopy revealed enveloped virions with nucleocapsid structures characteristic of viruses in the Paramyxoviridae family. The virus failed to hemabsorb chicken or guinea pig red blood cells and lacked neuraminidase activity. Two viral genes were isolated from a cDNA expression library. Multiple sequence alignments of one gene indicated an average identity of 45% as compared to Morbillivirus N protein sequences. A weaker relationship was found with Tupaia paramyxovirus (TPMV) and Hendra virus (HeV) N proteins. In the second gene, multiple open reading frames (ORFs) were identified, corresponding to the arrangement of the P, V, and C ORFs in the Morbillivirus and Respirovirus viruses. Short stretches in the C-terminal regions of the P and C proteins showed limited homologies to viruses in the Morbillivirus genus but no obvious relationship to viruses in other genera. The V ORF translation product contained a highly conserved, cysteine-rich domain that is common to most viruses in the Paramyxovirinae subfamily. Sequencing of P gene cDNA clones confirmed the use of a cotranscriptional editing mechanism for the regulation of P/V expression. Based on the location of its origin it has been named Salem virus (SalV).

Amino Acid Sequence↗

Rfam: annotating non-coding RNAs in complete genomes.

Rfam is a comprehensive collection of non-coding RNA (ncRNA) families, represented by multiple sequence alignments and profile stochastic context-free grammars. Rfam aims to facilitate the identification and classification of new members of known sequence families, and distributes annotation of ncRNAs in over 200 complete genome sequences. The data provide the first glimpses of conservation of multiple ncRNA families across a wide taxonomic range. A small number of large families are essential in all three kingdoms of life, with large numbers of smaller families specific to certain taxa. Recent improvements in the database are discussed, together with challenges for the future. Rfam is available on the Web at http://www.sanger.ac.uk/Software/Rfam/ and http://rfam.wustl.edu/.

Animals↗

ZmDB, an integrated database for maize genome research.

Zea mays DataBase (ZmDB) seeks to provide a comprehensive view of maize (corn) genetics by linking genomic sequence data with gene expression analysis and phenotypes of mutant plants. ZmDB originated in 1999 as the Web portal for a large project of maize gene discovery, sequencing and phenotypic analysis using a transposon tagging strategy and expressed sequence tag (EST) sequencing. Recently, ZmDB has broadened its scope to include all public maize ESTs, genome survey sequences (GSSs), and protein sequences. More than 170 000 ESTs are currently clustered into approximately 20 000 contigs and about an equal number of apparent singlets. These clusters are continuously updated and annotated with respect to potential encoded protein products. More than 100 000 GSSs are similarly assembled and annotated by spliced alignment with EST and protein sequences. The ZmDB interface provides quick access to analytical tools for further sequence analysis. Every sequence record is linked to several display options and similarity search tools, including services for multiple sequence alignment, protein domain determination and spliced alignment. Furthermore, ZmDB provides web-based ordering of materials generated in the project, including ESTs, ordered collections of genomic sequences tagged with the RescueMu transposon and microarrays of amplified ESTs. ZmDB can be accessed at http://zmdb.iastate.edu/.

DNA Transposable Elements↗

Amino acid sequence analysis of the annexin super-gene family of proteins.

The annexins are a widespread family of calcium-dependent membrane-binding proteins. No common function has been identified for the family and, until recently, no crystallographic data existed for an annexin. In this paper we draw together 22 available annexin sequences consisting of 88 similar repeat units, and apply the techniques of multiple sequence alignment, pattern matching, secondary structure prediction and conservation analysis to the characterisation of the molecules. The analysis clearly shows that the repeats cluster into four distinct families and that greatest variation occurs within the repeat 3 units. Multiple alignment of the 88 repeats shows amino acids with conserved physicochemical properties at 22 positions, with only Gly at position 23 being absolutely conserved in all repeats. Secondary structure prediction techniques identify five conserved helices in each repeat unit and patterns of conserved hydrophobic amino acids are consistent with one face of a helix packing against the protein core in predicted helices a, c, d, e. Helix b is generally hydrophobic in all repeats, but contains a striking pattern of repeat-specific residue conservation at position 31, with Arg in repeats 4 and Glu in repeats 2, but unconserved amino acids in repeats 1 and 3. This suggests repeats 2 and 4 may interact via a buried saltbridge. The loop between predicted helices a and b of repeat 3 shows features distinct from the equivalent loop in repeats 1, 2 and 4, suggesting an important structural and/or functional role for this region. No compelling evidence emerges from this study for uteroglobin and the annexins sharing similar tertiary structures, or for uteroglobin representing a derivative of a primordial one-repeat structure that underwent duplication to give the present day annexins. The analyses performed in this paper are re-evaluated in the Appendix, in the light of the recently published X-ray structure for human annexin V. The structure confirms most of the predictions and shows the power of techniques for the determination of tertiary structural information from the amino acid sequences of an aligned protein family.

Algorithms↗

Two novel alpha-neurotoxins isolated from Taiwan cobra: sequence characterization and phylogenetic comparison of homologous neurotoxins.

Two novel postsynaptic neurotoxins (alpha-neurotoxins) isolated and purified from the Taiwan cobra venom (Naja naja atra) possess distinct primary sequences and different neurotoxicities as compared with the most abundant and lethal component in the venom, i.e., cobrotoxin characterized before from the same venom. The complete sequences of two neurotoxin analogues were determined by N-terminal Edman degradation and comparison of amino acid compositions of proteolytic toxin fragments with other homologous toxins of known sequences. The short-chain neurotoxin consists of 61 amino acid residues with eight conserved cysteine residues and is found to show 78% sequence identity with cobrotoxin. The other toxin, consisting of 65 residues with ten cysteines, belongs to the family of long-chain neurotoxins. It is the first long-chain alpha-neurotoxin reported from the Taiwan cobra. The lethal toxicities of these two novel neurotoxins were much lower than cobrotoxin, albeit with close structural homology among the three toxins in terms of their primary sequences and tertiary structure predicted by homology modeling. Multiple sequence alignment and comparison coupled with construction of a phylogenetic tree for various alpha-neurotoxins of Naja and closely related genuses have established that all nicotinic alpha-neurotoxins present in the snake family of Elapidae are closely related to each other, presumably derived from an ancestral polypeptide by gene duplication and subsequent multiple mutational substitutions.

Amino Acid Sequence↗