Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Hierarchical method to align large numbers of biological sequences.

The method presented here is intended as a compromise between finding a good overall alignment and the time taken to do so. Many multiple alignment algorithms spend an excessively large amount of effort trying to find the best global alignment. This time is often ill spent because the results of the standard dynamic programming alignment algorithm are dominated by the choice of gap penalty and the form of the score matrix, both of which have a poor theoretical foundation. Nonetheless, it is important that savings in time do not compromise the quality of the alignment. By using the consensus sequence approach, this danger is largely avoided as the conserved features of the sequences are quickly identified and preserved through further cycles. In the alignment of existing alignments, which is one of the more novel aspects of the method, each alignment was treated as an averaged consensus sequence with gaps making no contribution. This gives rise to the advantageous property that gaps will have a greater propensity to be inserted where there are already gaps and is equivalent to a local change in the gap penalty. This type of behavior represents a transition away from the homogeneous scoring schemes used in aligning two sequences toward a scoring scheme that depends on position in the sequence. The alignment of consensus sequences thus forms a bridge between simple pair alignment and the alignment of discrete patterns in which sequence features and allowed gap locations are exaggerated. To complete this transition the program described above has been integrated into the earlier pattern matching (template) program. Such templates can reliably locate sequence similarities that are too weak or scattered to be found by the more standard alignment methods and should therefore produce a further condensation of the sequence data bank. Only by continually extending our knowledge of the relationships between sequences to increasingly distant similarities can we hope to avoid being overwhelmed by the increasing amount of data.

Algorithms

Sequence-directed mutagenesis: evidence from a phylogenetic history of human alpha-interferon genes.

We have studied the potential contribution of template-dependent events to genetic variation in mammals by examining the sequence alterations that have occurred in the recent evolution of human interferon genes. Fifteen members of the human alpha-interferon gene family were aligned, and a phylogenetic history was inferred. Many multiple events are inferred to have occurred in the evolution of the interferon genes and for the majority of these local DNA sequences were present that were capable of serving as templates for their occurrence. We conclude that the DNA sequence has the potential to explain many of the inferred spontaneous events and to explain complex alterations to sequences--i.e., the joint occurrence of base substitutions and insertions/deletions. Thus, such a mechanism would often cause multiple sequence changes as a result of a single mutational event and would provide additional genetic variation for evolution. Sequence-directed mutations would depend upon the local DNA sequences and, hence, would not be random at the DNA level.

Base Sequence

tRNA-rRNA sequence homologies: evidence for an ancient modular format shared by tRNAs and rRNAs.

Homologies between tRNAs and rRNAs are identified in searches using various combinations of Escherichia coli, yeast, Halobacterium volcanii and bovine mitochondrial sequences. As in previously reported comparisons, the homologies are too frequent and long to be attributed to coincidence, and similar frequencies from inter- and intraspecies comparisons preclude evolutionary convergence as an explanation. In contrast to the earlier studies, patterns in the positioning of the homologies are now described. Graphing the positions of the homologies along orthogonal axes that represent numbers of bases in tRNA and rRNA shows recurring patterns in the alignments. Preferred spacings of integral multiples of 9 bases are found, suggesting a periodicity in the ancestral structure from which the tRNAs and rRNAs were derived. The periodicity also suggests persistence of a modular format in both classes of molecules that survived changes in sequence that occurred during evolution. A model is proposed for the generation of the ancestral molecule and the early evolution of the coding mechanism. Elongation by self-priming and self-templating gave a hairpin with a 9 base stem. Two additional cycles gave a 70-80 base tRNA-like structure. Additional cycles yielded a tandem repeat of this unit, roughly equivalent in size to the combined rRNAs of prokaryotes. The larger RNA would contain the information and materials for generating the smaller RNAs. It is proposed that multiple recombination among such molecules gave composite structures, presumed progenitors of today's t- and rRNAs. The distribution of the conserved domains among today's species argues for the existence of the ancestral molecule prior to divergence of lines leading to the various kingdoms. Their presence in the different nucleic acids suggests the existence of a nucleic acid with multiple functions prior to partitioning of these functions among the nucleic acids that exist today. The occurrence of overlaps, overlays and consensus alignments among the homologies provides the means for identifying contiguous and neighboring conserved regions and holds promise for the reconstruction of the sequence of an ancestral molecule.

Animals

Multiple alignment using simulated annealing: branch point definition in human mRNA splicing.

A method for the simultaneous alignment of a very large number of sequences using simulated annealing is presented. The total running time of the algorithm does not depend explicitly on the number of sequences treated. The method has been used for the simultaneous alignment of 1462 human intron sequences upstream of the intron-exon boundary. The consensus sequence of the aligned set together with a calculation of the Shannon information clearly shows that several sequence motives are conserved: (i) a previously undetected guanosine rich region, (ii) the branch point and (iii) the polypyrimidine tract. The nucleotide frequencies at each position of the branch point consensus sequence qualitatively reproduce the frequencies of the experimentally determined branch points.

Algorithms

The sequence of squash NADH:nitrate reductase and its relationship to the sequences of other flavoprotein oxidoreductases. A family of flavoprotein pyridine nucleotide cytochrome reductases.

Nucleotide sequences were determined for cDNA clones for squash NADH:nitrate oxidoreductase (EC 1.6.6.1), which is one of the most completely characterized forms of this higher plant enzyme. An open reading frame of 2754 nucleotides began at the first ATG. The deduced amino acid sequence contains 918 residues, with a predicted Mr = 103,376. The amino acid sequence is very similar to sequences deduced for other higher plant nitrate reductases. The squash sequence has significant similarity to the amino acid sequences of sulfite oxidase, cytochrome b5, and NADH:cytochrome b5 reductase. Alignment of these sequences with that of squash defines domains of nitrate reductase that appear to bind its 3 prosthetic groups (molybdopterin, heme-iron, and FAD). The amino acid sequence of the FAD domain of squash nitrate reductase was aligned with FAD domain sequences of other NADH:nitrate reductases, NADH:cytochrome b5 reductases, NADPH:nitrate reductases, ferredoxin:NADP+ reductases, NADPH:cytochrome P-450 reductases, NADPH:sulfite reductase flavoproteins, and Bacillus megaterium cytochrome P-450BM-3. In this multiple alignment, 14 amino acid residues are invariant, which suggests these proteins are members of a family of flavoenzymes. Secondary structure elements of the structural model of spinach ferredoxin:NADP+ reductase were used to predict the secondary structure of squash nitrate reductase and the other related flavoenzymes in this family. We suggest that this family of flavoenzymes, nearly all of which reduce a hemoprotein, be called "flavoprotein pyridine nucleotide cytochrome reductases."

Amino Acid Sequence

Sequence and linkage analysis of the Coxiella burnetii citrate synthase-encoding gene.

The nucleotide (nt) sequence of the Coxiella burnetii citrate synthase-encoding gene (gltA), previously cloned in Escherichia coli, was determined. The nt sequence analysis revealed an open reading frame (ORF) of 1290 bp capable of coding for a protein of 430 amino acids (aa) with a deduced Mr of 48,633. Preceding an ATG start codon, a possible transcription start point (tsp) with homology to the E. coli promoter consensus was detected. A poly-purine-rich region occurred immediately upstream from the gltA reading frame and potentially serves as a ribosome-binding site. Additionally, a G + C-rich region of dyad symmetry 3' to the translational stop codon was found that could possibly function as a Rho-independent transcriptional termination signal. A large, nearly perfect, inverted repeat was identified upstream from the gltA tsp and was shown by Southern analysis to be present in multiple copies in the C. burnetii genome. The deduced aa sequence of C. burnetii GltA was optimally aligned with enzymes from various prokaryotic sources and one eukaryotic source (pig heart). Using perfect aa identity, the C. burnetii enzyme demonstrated the greatest homology with GltA from Acinetobacter anitratum (65%). Although only 26% aa identity was seen with the pig heart enzyme, many of the residues identified in ligand binding appear to be conserved. Sequencing studies of a region centered approx. 5.6 kb upstream from gltA revealed an ORF read with opposite polarity that encodes a peptide highly homologous to the C terminus of the flavoprotein subunit of E. coli succinate dehydrogenase. This report represents the first nt sequence analysis of a gene of known function from the obligate intracellular parasite, C. burnetii.

Amino Acid Sequence

Genome-wide SNP data support species boundaries in sympatric Polylepis Ruiz & Pav. (Rosaceae) species from Bolivia and Ecuador.

Species delimitation in the South American genus Polylepis is notoriously challenging due to high morphological similarity and phenotypic plasticity, likely driven by hybridization and gene flow. Previous phylogenetic studies suggested that genetic structure aligns more strongly with geography than with taxonomy, questioning existing species concepts and hampering conservation efforts. We used double-digest RAD sequencing (ddRADseq) to generate genome-wide SNP data for 11 Polylepis species sampled across multiple localities in Bolivia and Ecuador. Population genetic analyses, phylogenetic inference, and network approaches were combined to assess whether genetic structure aligns more closely with taxonomy or geography. Morphologically defined species formed largely cohesive genetic lineages across regions, with species identity explaining substantially more genetic variation than locality. While localized admixture and reticulation were detected among closely related taxa, widespread species showed strong genetic cohesion and clear separation from congeners. Our results indicate that the sampled Polylepis species from Bolivia and Ecuador maintain distinct genetic identities despite localized signals consistent with gene flow. This genome-wide support for current taxonomy highlights Polylepis as a valuable model for studying speciation under gene flow and indicates that multiple geographic sampling will be essential in reconstructing a robust phylogeny of the genus, with important implications for conservation planning in Andean montane forests.

Bolivia

Efficient sequence alignment algorithms.

Sequence alignments are becoming more important with the increase of nucleic acid data. Fitch and Smith have recently given an example where multiple insertion/deletions (rather than a series of adjacent single insertion/deletions) are necessary to achieve the correct alignment. Multiple insertion/deletions are known to increase computation time from O(n2) to O(n3) although Gotoh has presented an O(n2) algorithm in the case the multiple insertion/deletion weighting function is linear. It is argued in this paper that it could be desirable to use concave weighting functions. For that case, an algorithm is derived that is conjectured to be O(n2).

Base Sequence

Sex hormone-binding globulin, androgen-binding protein, and vitamin K-dependent protein S are homologous to laminin A, merosin, and Drosophila crumbs protein.

Androgen-binding protein (ABP) and sex hormone-binding globulin (SHBG) are extracellular steroid-binding proteins that are homologous to the COOH-terminal domain of vitamin K-dependent protein S, a protein important in blood clotting. We find that the sequences of ABP, SHBG, and protein S are also similar to two basement membrane proteins, laminin and merosin, and to an integral membrane protein, Drosophila crumbs protein. These latter three proteins have important roles in regulating differentiation and development. The sequence similarity corresponds to the G domain of laminin A chain, which binds heparin and type IV collagen. Analysis of a multiple alignment of these proteins reveals one well-conserved segment corresponding to the part of SHBG that binds to its membrane receptor and another corresponding to the part of protein S that binds to C4b-binding protein. The similarities suggest that ABP, SHBG, and protein S may also have functions related to that of laminin and merosin.

Amino Acid Sequence

Inching toward reality: an improved likelihood model of sequence evolution.

Our previous evolutionary model is generalized to permit approximate treatment of multiple-base insertions and deletions as well as regional heterogeneity of substitution rates. Parameter estimation and alignment procedures that incorporate these generalizations are developed. Simulations are used to assess the accuracy of the parameter estimation procedure and an example of an inferred alignment is included.

Animals

Similarities between putative transport proteins of plant viruses.

The nucleic acids of many plant viruses encode proteins with one or more of the following properties: an Mr of approximately 30,000, localization in the cell wall of the infected plant and a demonstrated role in cell-to-cell transport of infection. A progressive alignment strategy, aligning first those sequences known to be similar, and then aligning the resulting groups of sequences, was used to examine further the relatedness of the amino acid sequences of putative transport proteins of caulimoviruses, of proteins similar to the putative transport protein of alfalfa mosaic virus (A1MV) and of those similar to the tobacco mosaic virus (TMV) 30K protein. The strategy first identified regions in which multiple dipeptides of one group were similar to those of another group. The regions of similarity were brought into alignment by the conservative introduction of gaps. The positions of the introduction of gaps were adjusted to optimize similarity. Statistical significances of the resulting alignments, determined both by comparison with shuffled amino acid sequences and with the sequence alignment off-set by 1 to 15 residues in each direction, suggest that the amino acid sequences of the three groups of viruses are distantly related. Nevertheless, significant relationships between members of the caulimoviral group of sequences and members of each of the A1MV-like and TMV-like groups were found. These relationships and the analysis of the number of insertions/deletions between present sequences and a hypothetical common ancestor suggest that the sequences of the caulimoviral proteins are less diverged from the ancestor than either the A1MV-like or TMV-like proteins. The alignment identified common regions of predicted secondary structure and regions of similar hydropathy, regions possibly crucial for proper functioning of the proteins.

Amino Acid Sequence

Rat liver carboxylesterase: cDNA cloning, sequencing, and evidence for a multigene family.

A cDNA clone was isolated from a rat liver lambda gt11 expression library by screening with polyclonal antibodies raised against a rat liver microsomal carboxylesterase. This clone of 1.8 kb contained an open reading frame encoding a mature protein of 531 amino acids with a predicted molecular weight of 58,084. The 5' portion of the clone coded for 9 amino acids of a putative signal peptide. The 3' end of the clone included an untranslated region and a poly (A) tail. Carboxylesterase active site regions, five potential N-linked glycosylation sites, and 2 postulated cystine disulfide bridges were found in the cDNA-deduced amino acid sequence. Sequences obtained from tryptic peptides and the NH2-terminus of the purified native carboxylesterase were aligned with the deduced amino acid sequence, and the overall identity was 84%. Southern blot analysis suggested the presence of multiple genes. Thus it is concluded that we have cloned a rat liver carboxylesterase, and that this enzyme is a member of a multigene family.

Amino Acid Sequence

Mitochondrial Impostors: Prevalence and Impacts of NUMTs on Genetic and Evolutionary Studies in Carnivora.

Nuclear mitochondrial pseudogenes are mitochondria-derived DNA sequences integrated into the nuclear genome, which can introduce errors in species identification, phylogenetic inference, and population genetics. Although nuclear mitochondrial pseudogene contamination has been reported in some Carnivora species, a systematic investigation into the prevalence and impacts of nuclear mitochondrial pseudogenes across an order is still lacking. In this study, 22,102 mitochondrial DNA sequences of 80 Carnivora species from 14 families and 54 genera were retrieved from the public National Center for Biotechnology Information database and further analyzed. Using alignment-based methods, 158 problematic sequences/sequence groups were identified and categorized into four types: nuclear mitochondrial pseudogenes, species misidentification or mislabeling, sequence errors, and anomalous sites. Among families, Felidae exhibited the highest rate of nuclear mitochondrial pseudogene contamination, particularly in species of the genus Panthera. In contrast, no nuclear mitochondrial pseudogene contamination was detected in members of Ursidae and Ailuridae. Phylogenetic analysis revealed multiple independent origins of nuclear mitochondrial pseudogene, with some tracing back to the common ancestor of Carnivora. To mitigate nuclear mitochondrial pseudogene-related errors, rigorous sequence verification strategies, such as sequence alignment and phylogenetic validation, should be implemented. In conclusion, our findings highlight the necessity of nuclear mitochondrial pseudogene awareness in genetic and evolutionary studies of Carnivora and other taxa.

Animals

High-accuracy SNV calling for bacterial isolates using deep learning with AccuSNV.

Accurate detection of mutations within bacterial species is critical for fundamental studies of microbial evolution, reconstruction of transmission events, and identification of antimicrobial resistance mutations. Although many tools have been developed to identify single-nucleotide variants (SNVs) from whole-genome sequencing, they often suffer from high false-positive rates owing to the complexity of bacterial genomes and the need for different filtering cutoffs across sample types and sequencing depths. As data sets increase in size, the manual filtering required for high accuracy presents a significant obstacle. Here, we present AccuSNV, a novel deep learning-based tool for high-precision and automated bacterial SNV calling. Unlike traditional methods that process one sample at a time, AccuSNV leverages a convolutional neural network (CNN) that integrates alignment information across multiple samples, enhancing precision through learned across-sample patterns. We evaluate AccuSNV against seven popular SNV-calling tools using simulated data from six bacterial species with varied sequencing depths, numbers of isolates, mutations, and divergence levels. To further validate its real-world utility, we test AccuSNV on multiple curated bacterial data sets containing reported SNVs. In both simulated and real-world scenarios, AccuSNV consistently achieves the best performance. Moreover, AccuSNV provides comprehensive user-friendly downstream analysis modules and outputs, including mutation annotation information, phylogenetic inference, d N/d S calculations, and optional manual filtering. Together with the automated deep learning-based calling, these features make AccuSNV broadly accessible to users with different levels of computational expertise.

Deep Learning

A comparison of several similarity indices used in the classification of protein sequences: a multivariate analysis.

The present work describes an attempt to identify reliable criteria which could be used as distance indices between protein sequences. Seven different criteria have been tested: i and ii) the scores of the alignments as given by the BESTFIT and the FASTA programs; iii) the ratio parameter, i.e. the BESTFIT score divided by the length of the aligned peptides; iv and v) the statistical significance (Z-scores) of the scores calculated by BESTFIT and FASTA, as obtained by comparison with shuffled sequences; vi) the Z-scores provided by the program RELATE which performs a segment-by-segment comparison of 2 sequences, and vii) an original distance index calculated by the program DOCMA from all the pairwise dotplots between the sequences. These 7 criteria have been tested against the aminoacid sequences of 39 globins and those of the 20 aminoacyl-tRNA synthetases from E. coli. The distances between the sequences were analyzed by the multivariate analysis techniques. The results show that the distances calculated from the scores of the pairwise alignments are not adequately sensitive. The Z-score from RELATE is not selective enough and too demanding in computer time. Three criteria gave a classification consistent with the known similarities between the sequences in the sets, namely the Z-scores from BESTFIT and FASTA and the multiple dotplot comparison distance index from DOCMA.

Algorithms

Ulysses transposable element of Drosophila shows high structural similarities to functional domains of retroviruses.

We have determined the DNA structure of the Ulysses transposable element of Drosophila virilis and found that this transposon is 10,653 bp and is flanked by two unusually large direct repeats 2136 bp long. Ulysses shows the characteristic organization of LTR-containing retrotransposons, with matrix and capsid protein domains encoded in the first open reading frame. In addition, Ulysses contains protease, reverse transcriptase, RNase H and integrase domains encoded in the second open reading frame. Ulysses lacks a third open reading frame present in some retrotransposons that could encode an env-like protein. A dendrogram analysis based on multiple alignments of the protease, reverse transcriptase, RNase H, integrase and tRNA primer binding site of all known Drosophila LTR-containing retrotransposon sequences establishes a phylogenetic relationship of Ulysses to other retrotransposons and suggests that Ulysses belongs to a new family of this type of elements.

Amino Acid Sequence

Alignment-free integration of single-nucleus ATAC-seq across species with sPYce.

Changes in gene regulation largely contribute to differences in cellular identities and phenotypes between species. Single-nucleus assays for transposase-accessible chromatin with sequencing (snATAC-seq) are an efficient strategy to identify putative gene regulatory elements and provide new insight into evolutionary divergence of regulatory programmes. However, no dedicated framework exists to integrate and compare snATAC-seq data across species, while methods designed for single-cell gene expression data have serious limitations. Here we present sPYce, a cross-species snATAC-seq integration method that relies on sequence composition similarities through k-mer histograms of regulatory regions, removing the need for genome alignments to anchor data from different species. sPYce can embed datasets from multiple species into the same mathematical space and permits further downstream analysis steps. We benchmarked sPYce against existing approaches on two publicly available datasets spanning more than 160 myr of evolution, showing that it successfully uncovers conserved cellular programmes while preserving biologically relevant species-specific differences. By comparing cerebellar development in mice and opossums, sPYce identifies regulatory divergence in granule cell differentiation programmes, particularly driven by nuclear factor 1. As an easy-to-use, alignment-free cross-species snATAC-seq integration approach, sPYce opens new perspectives to compare gene regulatory evolution across species.

Animals

MAFin: motif detection in multiple alignment files.

MOTIVATION: Whole Genome and Proteome Alignments, represented by the multiple alignment file format, have become a standard approach in comparative genomics and proteomics. These often require identifying conserved motifs, which is crucial for understanding functional and evolutionary relationships. However, current approaches lack a direct method for motif detection within MAF files. We present MAFin, a novel tool that enables efficient motif detection and conservation analysis in MAF files to address this gap, streamlining genomic and proteomic research. RESULTS: We developed MAFin, the first motif detection tool for Multiple Alignment Format files. MAFin enables the multithreaded search of conserved motifs using three approaches: (i) using user-specified k-mers to search the sequences. (ii) with regular expressions, in which case one or more patterns are searched, and (iii) with predefined Position Weight Matrices. Once the motif has been found, MAFin detects the motif instances and calculates the conservation across the aligned sequences. MAFin also calculates a conservation percentage, which provides information about the conservation levels of each motif across the aligned sequences, based on the number of matches relative to the length of the motif. A set of statistics enables the interpretation of each motif's conservation level, and the detected motifs are exported in JSON and CSV files for downstream analyses. AVAILABILITY AND IMPLEMENTATION: MAFin is offered as a Python package under the GPL license as a multi-platform application and is available at: https://github.com/Georgakopoulos-Soares-lab/MAFin.

Software