Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

[Pattern recognition in the computer analysis of nucleotide sequences].

We have used an algorithm from the pattern recognition theory "generalized portrait" to find a distinguishing vector for Escherichia coli promoters. We have made an attempt to solve closely linked problems for choosing significant signs of that signal, multiple alignment and for calculation of the recognition vector (matrix). The promoters with known strength have been ranged with this vector. The analysis of the occurrence of predicted promoters has been carried out. The promoters search program for IBM-compatible computers is available from the authors.

Base Sequence

Ulysses transposable element of Drosophila shows high structural similarities to functional domains of retroviruses.

We have determined the DNA structure of the Ulysses transposable element of Drosophila virilis and found that this transposon is 10,653 bp and is flanked by two unusually large direct repeats 2136 bp long. Ulysses shows the characteristic organization of LTR-containing retrotransposons, with matrix and capsid protein domains encoded in the first open reading frame. In addition, Ulysses contains protease, reverse transcriptase, RNase H and integrase domains encoded in the second open reading frame. Ulysses lacks a third open reading frame present in some retrotransposons that could encode an env-like protein. A dendrogram analysis based on multiple alignments of the protease, reverse transcriptase, RNase H, integrase and tRNA primer binding site of all known Drosophila LTR-containing retrotransposon sequences establishes a phylogenetic relationship of Ulysses to other retrotransposons and suggests that Ulysses belongs to a new family of this type of elements.

Amino Acid Sequence

Alignment-free integration of single-nucleus ATAC-seq across species with sPYce.

Changes in gene regulation largely contribute to differences in cellular identities and phenotypes between species. Single-nucleus assays for transposase-accessible chromatin with sequencing (snATAC-seq) are an efficient strategy to identify putative gene regulatory elements and provide new insight into evolutionary divergence of regulatory programmes. However, no dedicated framework exists to integrate and compare snATAC-seq data across species, while methods designed for single-cell gene expression data have serious limitations. Here we present sPYce, a cross-species snATAC-seq integration method that relies on sequence composition similarities through k-mer histograms of regulatory regions, removing the need for genome alignments to anchor data from different species. sPYce can embed datasets from multiple species into the same mathematical space and permits further downstream analysis steps. We benchmarked sPYce against existing approaches on two publicly available datasets spanning more than 160 myr of evolution, showing that it successfully uncovers conserved cellular programmes while preserving biologically relevant species-specific differences. By comparing cerebellar development in mice and opossums, sPYce identifies regulatory divergence in granule cell differentiation programmes, particularly driven by nuclear factor 1. As an easy-to-use, alignment-free cross-species snATAC-seq integration approach, sPYce opens new perspectives to compare gene regulatory evolution across species.

Animals

MAFin: motif detection in multiple alignment files.

MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.

Software

Evolutionary relationships among aminotransferases. Tyrosine aminotransferase, histidinol-phosphate aminotransferase, and aspartate aminotransferase are homologous proteins.

A data base was compiled containing the amino acid sequences of 12 aspartate aminotransferases and 11 other aminotransferases. A comparison of these sequences by a standard alignment method confirmed the previously reported homology of all aspartate aminotransferases and Escherichia coli tyrosine aminotransferase. However, no significant similarity between these proteins and any of the other aminotransferases was detected. A more rigorous analysis, focusing on short sequence segments rather than the total polypeptide chain, revealed that rat tyrosine aminotransferase and Saccharomyces cerevisiae and Escherichia coli histidinol-phosphate aminotransferase share several homologous sequence segments with aspartate aminotransferases. For comparison of the complete sequences, a multiple sequence editor was developed to display the whole set of amino acid sequences in parallel on a single work-sheet. The editor allows gaps in individual sequences or a set of sequences to be introduced and thus facilitates their parallel analysis and alignment. Several clusters of invariant residues at corresponding positions in the amino acid sequences became evident, clearly establishing that the cytosolic and the mitochondrial isoenzyme of vertebrate aspartate aminotransferase, E. coli aspartate aminotransferase, rat and E. coli tyrosine aminotransferase, and S. cerevisiae and E. coli histidinol-phosphate aminotransferase are homologous proteins. Only 12 amino acid residues out of a total of about 400 proved to be invariant in all sequences compared; they are either involved in the binding of pyridoxal 5'-phosphate and the substrate, or appear to be essential for the conformation of the enzymes.

Amino Acid Sequence

Detection of Caenorhabditis transposon homologs in diverse organisms.

Although transposons that move via DNA intermediates are common in bacteria, invertebrates, and plants, none have been clearly documented in vertebrates and certain other classes of organisms. One such family of transposons includes invertebrate elements related to Caenorhabditis elegans Tc1. Blocks of aligned protein segments derived from this family were used to search a nucleotide sequence databank. Among the relatives detected were known bacterial insertion elements, revealing the ancient origin of the family. Furthermore, a Tc1-like homolog was detected in a catfish, raising the possibility that this valuable tool of C. elegans genetics can be used with vertebrate genomes. This study illustrates the use of multiple protein blocks for detection and evaluation of distant relationships.

Amino Acid Sequence

Recognition of distantly related protein sequences using conserved motifs and neural networks.

A sensitive technique for protein sequence motif recognition based on neural networks has been developed. It involves three major steps. (1) At each appropriate alignment position of a set of N matched sequences, a set of N aligned oligopeptides is specified with preselected window length. N neural nets are subsequently and successively trained on N-1 amino acid spans after eliminating each ith oligopeptide. A test for recognition of each of the ith spans is performed. The average neural net recognition over N such trials is used as a measure of conservation for the particular windowed region of the multiple alignment. This process is repeated for all possible spans of given length in the multiple alignment. (2) The M most conserved regions are regarded as motifs and the oligopeptides within each are used to train intensively M individual neural networks. (3) The M networks are then applied in a search for related primary structures in a databank of known protein sequences. The oligopeptide spans in the database sequence with strongest neural net output for each of the M networks are saved and then scored according to the output signals and the proper combination that follows the expected N- to C-terminal sequence order. The motifs from the database with highest similarity scores can then be used to retrain the M neural nets, which can be subsequently utilized for further searches in the databank, thus providing even greater sensitivity to recognize distant familial proteins. This technique was successfully applied to the integrase, DNA-polymerase and immunoglobulin families.

Aldehyde Dehydrogenase

Structural predictions for the central domain of dystrophin.

The amino acid sequence of dystrophin indicates that the molecule has globular N- and C-terminal domains separated by a long central rod domain. The central rod contains multiple repeats, about 100 amino acids long and of variable length. These diverge sufficiently in sequence that, in previous studies, only 14 of the most similar repeats have been aligned and analysed in any detail. We show here that a heptad pattern of hydrophobic residues is preserved across all repeats. Using the heptad pattern together with a consensus sequence template, we identified and aligned 25 repeats in the dystrophin rod sequence. Each repeat consists of a constant-length core helix of 54 residues, coupled via a short linker to a weakly conserved variable-length helix, and then via a second linker to the next core. The variable-length helix appears truncated in repeats 10 and 13 and extended in repeats 4 and 20. The extension of repeat 20 is particularly interesting since it corresponds to a hotspot of dystrophy-inducing mutations. Detailed modelling suggests that the classical Speicher-Marchesi [(1984) Nature 311, 177-180] model for spectrin may not be appropriate to dystrophin without some modification. We propose that whilst the repeating structural motif in dystrophin is probably a bead of triple coiled coil, this bead is twice as massive as, and out of phase with, those proposed for spectrin. Our model raises the possibility that the rod domain of dystrophin may confer elasticity on the molecule. Deletions which truncate this region would then reduce the extensibility of the molecule without affecting actin crosslinking, consistent with their typically producing the relatively benign Becker phenotype of muscular dystrophy.

Actinin

PanDelos-plus: A parallel algorithm for computing sequence homology in pangenomic analysis.

The identification of homologous gene families across multiple genomes is a central task in bacterial pangenomics traditionally requiring computationally demanding all-against-all comparisons. PanDelos addresses this challenge with an alignment-free and parameter-free approach based on k-mer profiles, combining high speed, ease of use, and competitive accuracy with state-of-the-art methods. However, the increasing availability of genomic data requires tools that can scale efficiently to larger datasets. To address this need, we present PanDelos-plus, a fully parallel, gene-centric redesign of PanDelos. The algorithm parallelizes the most computationally intensive phases (Best Hit detection and Bidirectional Best Hit extraction) through data decomposition and a thread pool strategy, while employing lightweight data structures to reduce memory usage. Benchmarks on synthetic datasets show that PanDelos-plus achieves up to 14x faster execution and reduces memory usage by up to 96%, while maintaining consistency with the original algorithm. These improvements allow the PanDelos methodology to be applied to population-scale comparative genomics, thus enabling more precise characterisation of pangenome structure and dynamics. PanDelos-plus is available at github.com/synbionics/PanDelos-plus.

Journal Article

A database of protein structure families with common folding motifs.

The availability of fast and robust algorithms for protein structure comparison provides an opportunity to produce a database of three-dimensional comparisons, called families of structurally similar proteins (FSSP). The database currently contains an extended structural family for each of 154 representative (below 30% sequence identity) protein chains. Each data set contains: the search structure; all its relatives with 70-30% sequence identity, aligned structurally; and all other proteins from the representative set that contain substructures significantly similar to the search structure. Very close relatives (above 70% sequence identity) rarely have significant structural differences and are excluded. The alignments of remote relatives are the result of pairwise all-against-all structural comparisons in the set of 154 representative protein chains. The comparisons were carried out with each of three novel automatic algorithms that cover different aspects of protein structure similarity. The user of the database has the choice between strict rigid-body comparisons and comparisons that take into account interdomain motion or geometrical distortions; and, between comparisons that require strictly sequential ordering of segments and comparisons, which allow altered topology of loop connections or chain reversals. The data sets report the structurally equivalent residues in the form of a multiple alignment and as a list of matching fragments to facilitate inspection by three-dimensional graphics. If substructures are ignored, the result is a database of structure alignments of full-length proteins, including those in the twilight zone of sequence similarity.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

DNA sequence determinants for binding of transformed Ah receptor to a dioxin-responsive enhancer.

We have utilized gel retardation analysis and DNA mutagenesis to examine the specific interaction of transformed guinea pig hepatic cytosolic TCDD.AhR complex with a dioxin-responsive element (DRE). Sequence alignment of the mouse CYPIA1 upstream DREs has identified a common invariant "core" consensus sequence of TNGCGTG flanked by several variable nucleotides. Competitive gel retardation analysis using a series of DRE oligonucleotides containing single or multiple base substitutions has allowed identification of those nucleotides important for TCDD.AhR.DRE complex formation. A putative TCDD.AhR DNA-binding consensus sequence of GCGTGNNA/TNNNC/G has been derived. The four core nucleotides, CGTG, appear to be critical for TCDD-inducible protein-DNA complex formation since their substitution decreased AhR binding affinity by 100-800-fold; the remaining conserved bases are also important, albeit to a lesser degree (3-5-fold). The 5'-ward thymine, present in the invariant core sequence of all the DREs identified to date, appears not to be involved in DNA binding of the AhR. The results obtained here indicate that although the primary interaction of the TCDD.AhR complex with the DRE occurs with the conserved "core" sequence, nucleotides flanking the core also contribute to the specificity of DRE binding.

Animals

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans

A survey of multiple sequence comparison methods.

Multiple sequence comparison refers to the search for similarity in three or more sequences. This article presents a survey of the exhaustive (optimal) and heuristic (possibly sub-optimal) methods developed for the comparison of multiple macromolecular sequences. Emphasis is given to the different approaches of the heuristic methods. Four distance measures derived from information engineering and genetic studies are introduced for the comparison between two alignments of sequences. The use of entropy, which plays a central role in information theory as measures of information, choice and uncertainty, is proposed as a simple measure for the evaluation of the optimality of an alignment in the absence of any a priori knowledge about the structures of the sequences being compared. This article also gives two examples of comparison between alternative alignments of the same set of 5SRNAs as obtained by several different heuristic methods.

Algorithms

Nucleotide sequence, secondary structure and evolution of the 5S ribosomal RNA from five bacterial species.

The nucleotide sequences of the 5S ribosomal RNAs of the bacteria Agrobacterium tumefaciens, Alcaligenes faecalis, Pseudomonas cepacia, Aquaspirillum serpens and Acinetobacter calcoaceticus have been determined. The sequences fit in a generally accepted model for 5S RNA secondary structure. However, a closer comparative examination of these and other bacterial 5S RNA primary structures reveals the potential of additional base pairing and of multiple equilibria between a set of slightly different alternative secondary structures in one area of the molecule. The phylogenetic position of the examined bacteria is derived from a 5S RNA sequence alignment by a clustering method and compared with the position derived on the basis of 16S ribosomal RNA oligonucleotide catalogs.

Bacteria

Patellofemoral joint: evaluation during active flexion with ultrafast spoiled GRASS MR imaging.

An ultrafast spoiled gradient-recalled acquisition in the steady state pulse sequence was developed that permits multiple images to be obtained at a temporal resolution suitable for examining the patellofemoral joint during active flexion. This pulse sequence was used to perform kinematic magnetic resonance imaging of patellar alignment and tracking in five healthy subjects and seven patients with a provisional clinical diagnosis of abnormal patellofemoral joints.

Female

A pattern of partially homologous recombination in mouse L cells.

Herpes simplex virus thymidine kinase gene and pBR322 DNA (in large excess to the thymidine kinase gene) were introduced into mouse L cells by calcium phosphate DNA-mediated gene transfer. DNA fragments encompassing six junctions between the exogenous DNAs have been cloned and their nucleotide sequences determined. Analysis of these sequences has shown that stretches of partial homology involving from 20-50 base pairs are present near the points at which joining occurs between the donor molecules. The structure of the junction sequences suggests that the recombination event involves the alignment of the two donor DNA molecules at partially homologous regions followed by staggered cutting and joining. One donor molecule is always cut in the region of partial homology, while the second is cut at some distance that is a small multiple of 13.5 +/- 0.5 base pairs away (at 0, 14, 27, 39, 41, and 54 base pairs). In the three junctions where the second cut is far from the region of homology, a 17- to 19-base-pair segment of DNA separates the donor sequences. In all cases the origin of this "filler" DNA appears to be oligonucleotides derived from pBR322.

Animals

Identification of a functionally conserved surface region of rat cytochromes P450IA.

A region of rat cytochrome P450IA1 at residues 294-301 (Gln-Asp-Arg-Arg-Leu-Asp-Glu-Asn), equivalent to a proinhibitory region of cytochrome P450IA2, was identified by sequence alignment. Anti-peptide antibodies were successfully raised when the peptide was coupled through either its N- or its C-terminus to carrier protein, but no antibodies were produced against the so-called multiple peptide antigen, which consisted of eight copies of the peptide attached through its C-terminus to a synthetic base. Both of the anti-peptide antibodies bound specifically to cytochrome P450IA1 in the rat, as shown by e.l.i.s.a. and immunoblotting. They inhibited microsomal aryl hydrocarbon hydroxylase activity and the mutagenic activation of 2-acetylaminofluorene (these reactions are catalysed by cytochrome P450IA1), but not high-affinity phenacetin O-de-ethylation activity, which is catalysed by cytochrome P450IA2. However, there was differences in the properties of the two antisera in their binding to cytochromes P450IA1 in species other than the rat, their relative binding to the multiple peptide antigen, the yield of antibody following affinity purification using peptide coupled through its N-terminus to CNBr-activated Sepharose, and the binding of the purified preparations to N- and C-terminal-coupled peptide conjugates. These observations indicated that the antibodies were directed to the region of the peptide opposite to the end which was coupled to the carrier protein. Nevertheless, both of the antibody preparations bound equally well to the target cytochrome P450, thus indicating that, in the native protein, the whole of the peptide region is exposed on the surface of cytochrome P450IA1 and is available for binding by the antibodies. The role of this region appears to be the same in both cytochromes P450IA1 and P450IA2, despite the difference in its primary structure in the two cytochromes P450.

Amino Acid Sequence

The product of unr, the highly conserved gene upstream of N-ras, contains multiple repeats similar to the cold-shock domain (CSD), a putative DNA-binding motif.

We show that the open reading frame transcribed from the unr gene (immediately upstream of N-ras) in mammals consists of multiple repeats similar to the cold-shock domain (CSD), a putative DNA-binding motif found in prokaryotic cold-shock proteins, and eukaryotic DNA-binding proteins. Alignment of the CSD sequences of unr with those from other proteins reveals a core of similarity for which a consistent secondary structure prediction can be derived. This prediction suggests that the CSD consists primarily of beta-sheet, in contrast to most known eukaryotic DNA-binding proteins. Sequence analysis of the 3' end of the guinea pig unr gene shows that the core of one CSD repeat is encoded in a single exon, consistent with the modular assembly of the gene from ancestral CSD-coding units.

Amino Acid Sequence