Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

A genetic analysis of HIV-1 from Punjab, India reveals the presence of multiple variants.

OBJECTIVE: To determine the extent of HIV-1 genetic variation in Indian patients. DESIGN: To avoid any bias in selecting viral variants, HIV-1 DNA was amplified directly from the peripheral blood mononuclear cells of patients and sequenced. Genetic similarity between Indian sequences and other geographic isolates was analysed by phylogenetic analysis algorithms. METHODS: A fragment encompassing the C2/V3-V5 regions of HIV-1 gp120 was amplified from the lymphocyte DNA of 12 Indian patients. Multiple clones from each patient were sequenced. Nucleotide sequences encompassing about 650 base pairs were aligned for the Indian and other geographically distinct isolates. Inter-isolate relationships were analysed by means of distance, parsimony and neighbour-joining algorithms. RESULTS: Nucleotide sequence comparisons showed low interpatient variation. Amino-acid comparisons revealed a high degree of homology between Indian sequences in this study and those studied earlier. On distance and parsimony trees, most of the Indian sequences clustered together as subtype C. However, sequences from three patients also showed significant homologies and phylogenetic clustering outside of subtype C. CONCLUSIONS: The predominant strain of HIV-1 in India belongs to subtype C and little interpatient nucleotide sequence divergence in the majority of cases suggests recent spread of HIV-1 in this region. This study also presents the first evidence for non-C subtypes in the Indian population with two epidemiologically linked samples remaining unclassified for any existing env subtype. The presence of variant subtypes in Indian patients sheds light on the transmission routes of HIV-1 to India and emphasizes the need to include these sequences in vaccine development strategies.

Acquired Immunodeficiency Syndrome

DNA sequence comparisons of the human, mouse, and rabbit immunoglobulin kappa gene.

A comparative analysis between human, mouse, and rabbit immunoglobulin (Ig) kappa-gene DNA sequences is presented. New formulas for determining the expected length and variance of the longest block identity (a succession of matching nucleotides) between multiple random sequences are given and are used to establish statistical criteria for ascertaining the significance of block identities shared in r out of s sequences. The statistically significant block identities within and between the Ig-kappa-gene sequences are ascertained, and alignment maps based on these similarities are constructed. The human and rabbit sequences (especially in the noncoding regions) and the human and mouse sequences (on the coding regions) show a similarity much stronger than that between the mouse and rabbit sequences. The existence of several highly significant shared oligonucleotides occurring in alignment with each other or with respect to the J- and C-gene segments suggests a configuration of multiple control sites. Discussion and interpretations of the form and distribution of the block identities are given.

Animals

A distance-based block searching algorithm.

We present in this paper an algorithm for the multiple comparison of a set of protein sequences. Our approach is that of peptide matching and consists in looking for all the words that occur approximatively in at least q of the sequences in the set, where q is a parameter. Words are compared by using a reference object called a model, that is itself a word over the alphabet of the amino acids, and the comparison between a model and a word is based on w-length words instead of single symbols. This idea is similar to the one used in the Blast program in the case of pairwise comparisons. Two w-length words are considered to be related if an alignment without gaps of the two using a similarity matrix has a score greater than a certain threshold value t. In our case, we say that a k-length word u is an occurrence of a model m of the same length if every w-length subword of u is related to the corresponding subword of m in the sense given above. If a model m has occurrences in at least q of the sequences of the set, m is said to occur in the set. In percentage terms, the value of q may correspond to something as small as 5% of the sequences (search for recurrent words in a set of non homologous proteins) or as high as 70-100% (establishment of a list of all similar words as a first step in a multiple alignment program). The algorithm presented here is an efficient and exact way of looking for all the models, of a fixed length k or of the greatest possible length kmax, that occur in a set of sequences. It can work with any kind of scoring matrix and an extension of the algorithm allows for the introduction of gaps between a model and its occurrences.

Algorithms

Further analysis of cDNA clones for maize phosphoenolpyruvate carboxylase involved in C4 photosynthesis. Nucleotide sequence of entire open reading frame and evidence for polyadenylation of mRNA at multiple sites in vivo.

Four clones of cDNA for phosphoenolpyruvate carboxylase [EC 4.1.1.31] were obtained from a maize green leaf cDNA library by colony hybridization. The largest cDNA was of full-length (3335 nucleotides), being 243 nucleotides longer than the cDNA cloned previously [(1986) Nucleic Acids Res. 14, 1615-1628]. Alignment of the sequence for the N-terminal coding region found in two of the four clones with the sequence reported previously, established the sequence of the entire coding region for the enzyme. The sequencing of 3'-untranslated region of the clones revealed that the poly(A) tract is attached at multiple sites in vivo.

Base Sequence

Primary structure of carboxypeptidase T: delineation of functionally relevant features in Zn-carboxypeptidase family.

The primary structure of carboxypeptidase T--a Zn-dependent extracellular enzyme of Thermoactinomyces vulgaris--was determined from the cloned cpT gene nucleotide sequence and compared to Zn-carboxypeptidases from various organisms. The compilation and analysis of multiple alignment accompanied by consideration of available tertiary structure data have shown that in the overall spatial structure and active site arrangement CpT is similar to other enzymes constituting the Zn-carboxypeptidase family. Nine of 16 amino acid residues found to be strictly invariant are presumably located close to the active site. The preservation of His69, Glu72, Asn144, Arg145, His196, Tyr248, and Glu270 identified previously as essential catalytic site participants implicates basically the same catalytic mechanism in the Zn-carboxypeptidase family. It is proposed that Pro205 and Asp256 should play an important role in proper S1'-pocket spatial arrangement. The comparative analysis of amino acid variations in S1'-pocket enabled us to reveal structural determinants of the Zn-carboxypeptidase primary specificity. The relatively reduced size of the pocket and negative charge of Asp253 are supposed to contribute correspondingly to A- and B-type substrate preferences of carboxypeptidase T endowed with dual primary specificity.

Amino Acid Sequence

Definition of general topological equivalence in protein structures. A procedure involving comparison of properties and relationships through simulated annealing and dynamic programming.

A protein is defined as an indexed string of elements at each level in the hierarchy of protein structure: sequence, secondary structure, super-secondary structure, etc. The elements, for example, residues or secondary structure segments such as helices or beta-strands, are associated with a series of properties and can be involved in a number of relationships with other elements. Element-by-element dissimilarity matrices are then computed and used in the alignment procedure based on the sequence alignment algorithm of Needleman & Wunsch, expanded by the simulated annealing technique to take into account relationships as well as properties. The utility of this method for exploring the variability of various aspects of protein structure and for comparing distantly related proteins is demonstrated by multiple alignment of serine proteinases, aspartic proteinase lobes and globins.

Amino Acid Sequence

Hierarchical method to align large numbers of biological sequences.

The method presented here is intended as a compromise between finding a good overall alignment and the time taken to do so. Many multiple alignment algorithms spend an excessively large amount of effort trying to find the best global alignment. This time is often ill spent because the results of the standard dynamic programming alignment algorithm are dominated by the choice of gap penalty and the form of the score matrix, both of which have a poor theoretical foundation. Nonetheless, it is important that savings in time do not compromise the quality of the alignment. By using the consensus sequence approach, this danger is largely avoided as the conserved features of the sequences are quickly identified and preserved through further cycles. In the alignment of existing alignments, which is one of the more novel aspects of the method, each alignment was treated as an averaged consensus sequence with gaps making no contribution. This gives rise to the advantageous property that gaps will have a greater propensity to be inserted where there are already gaps and is equivalent to a local change in the gap penalty. This type of behavior represents a transition away from the homogeneous scoring schemes used in aligning two sequences toward a scoring scheme that depends on position in the sequence. The alignment of consensus sequences thus forms a bridge between simple pair alignment and the alignment of discrete patterns in which sequence features and allowed gap locations are exaggerated. To complete this transition the program described above has been integrated into the earlier pattern matching (template) program. Such templates can reliably locate sequence similarities that are too weak or scattered to be found by the more standard alignment methods and should therefore produce a further condensation of the sequence data bank. Only by continually extending our knowledge of the relationships between sequences to increasingly distant similarities can we hope to avoid being overwhelmed by the increasing amount of data.

Algorithms

Molecular cloning, nucleotide sequence, and expression of a carboxypeptidase-encoding gene from the archaebacterium Sulfolobus solfataricus.

Mammalian metallocarboxypeptidases play key roles in major biological processes, such as digestive-protein degradation and specific proteolytic processing. A Sulfolobus solfataricus gene (cpsA) encoding a recently described zinc carboxypeptidase with an unusually broad substrate specificity was cloned, sequenced, and expressed in Escherichia coli. Despite the lack of overall sequence homology with known carboxypeptidases, seven homology blocks, including the Zn-coordinating and catalytic residues, were identified by multiple alignment with carboxypeptidases A, B, and T. S. solfataricus carboxypeptidase expressed in E. coli was found to be enzymatically active, and both its substrate specificity and thermostability were comparable to those of the purified S. solfataricus enzyme.

Amino Acid Sequence

Sequence-directed mutagenesis: evidence from a phylogenetic history of human alpha-interferon genes.

We have studied the potential contribution of template-dependent events to genetic variation in mammals by examining the sequence alterations that have occurred in the recent evolution of human interferon genes. Fifteen members of the human alpha-interferon gene family were aligned, and a phylogenetic history was inferred. Many multiple events are inferred to have occurred in the evolution of the interferon genes and for the majority of these local DNA sequences were present that were capable of serving as templates for their occurrence. We conclude that the DNA sequence has the potential to explain many of the inferred spontaneous events and to explain complex alterations to sequences--i.e., the joint occurrence of base substitutions and insertions/deletions. Thus, such a mechanism would often cause multiple sequence changes as a result of a single mutational event and would provide additional genetic variation for evolution. Sequence-directed mutations would depend upon the local DNA sequences and, hence, would not be random at the DNA level.

Base Sequence

tRNA-rRNA sequence homologies: evidence for an ancient modular format shared by tRNAs and rRNAs.

Homologies between tRNAs and rRNAs are identified in searches using various combinations of Escherichia coli, yeast, Halobacterium volcanii and bovine mitochondrial sequences. As in previously reported comparisons, the homologies are too frequent and long to be attributed to coincidence, and similar frequencies from inter- and intraspecies comparisons preclude evolutionary convergence as an explanation. In contrast to the earlier studies, patterns in the positioning of the homologies are now described. Graphing the positions of the homologies along orthogonal axes that represent numbers of bases in tRNA and rRNA shows recurring patterns in the alignments. Preferred spacings of integral multiples of 9 bases are found, suggesting a periodicity in the ancestral structure from which the tRNAs and rRNAs were derived. The periodicity also suggests persistence of a modular format in both classes of molecules that survived changes in sequence that occurred during evolution. A model is proposed for the generation of the ancestral molecule and the early evolution of the coding mechanism. Elongation by self-priming and self-templating gave a hairpin with a 9 base stem. Two additional cycles gave a 70-80 base tRNA-like structure. Additional cycles yielded a tandem repeat of this unit, roughly equivalent in size to the combined rRNAs of prokaryotes. The larger RNA would contain the information and materials for generating the smaller RNAs. It is proposed that multiple recombination among such molecules gave composite structures, presumed progenitors of today's t- and rRNAs. The distribution of the conserved domains among today's species argues for the existence of the ancestral molecule prior to divergence of lines leading to the various kingdoms. Their presence in the different nucleic acids suggests the existence of a nucleic acid with multiple functions prior to partitioning of these functions among the nucleic acids that exist today. The occurrence of overlaps, overlays and consensus alignments among the homologies provides the means for identifying contiguous and neighboring conserved regions and holds promise for the reconstruction of the sequence of an ancestral molecule.

Animals

Multiple alignment using simulated annealing: branch point definition in human mRNA splicing.

A method for the simultaneous alignment of a very large number of sequences using simulated annealing is presented. The total running time of the algorithm does not depend explicitly on the number of sequences treated. The method has been used for the simultaneous alignment of 1462 human intron sequences upstream of the intron-exon boundary. The consensus sequence of the aligned set together with a calculation of the Shannon information clearly shows that several sequence motives are conserved: (i) a previously undetected guanosine rich region, (ii) the branch point and (iii) the polypyrimidine tract. The nucleotide frequencies at each position of the branch point consensus sequence qualitatively reproduce the frequencies of the experimentally determined branch points.

Algorithms

The sequence of squash NADH:nitrate reductase and its relationship to the sequences of other flavoprotein oxidoreductases. A family of flavoprotein pyridine nucleotide cytochrome reductases.

Nucleotide sequences were determined for cDNA clones for squash NADH:nitrate oxidoreductase (EC 1.6.6.1), which is one of the most completely characterized forms of this higher plant enzyme. An open reading frame of 2754 nucleotides began at the first ATG. The deduced amino acid sequence contains 918 residues, with a predicted Mr = 103,376. The amino acid sequence is very similar to sequences deduced for other higher plant nitrate reductases. The squash sequence has significant similarity to the amino acid sequences of sulfite oxidase, cytochrome b5, and NADH:cytochrome b5 reductase. Alignment of these sequences with that of squash defines domains of nitrate reductase that appear to bind its 3 prosthetic groups (molybdopterin, heme-iron, and FAD). The amino acid sequence of the FAD domain of squash nitrate reductase was aligned with FAD domain sequences of other NADH:nitrate reductases, NADH:cytochrome b5 reductases, NADPH:nitrate reductases, ferredoxin:NADP+ reductases, NADPH:cytochrome P-450 reductases, NADPH:sulfite reductase flavoproteins, and Bacillus megaterium cytochrome P-450BM-3. In this multiple alignment, 14 amino acid residues are invariant, which suggests these proteins are members of a family of flavoenzymes. Secondary structure elements of the structural model of spinach ferredoxin:NADP+ reductase were used to predict the secondary structure of squash nitrate reductase and the other related flavoenzymes in this family. We suggest that this family of flavoenzymes, nearly all of which reduce a hemoprotein, be called "flavoprotein pyridine nucleotide cytochrome reductases."

Amino Acid Sequence

Sequence and linkage analysis of the Coxiella burnetii citrate synthase-encoding gene.

The nucleotide (nt) sequence of the Coxiella burnetii citrate synthase-encoding gene (gltA), previously cloned in Escherichia coli, was determined. The nt sequence analysis revealed an open reading frame (ORF) of 1290 bp capable of coding for a protein of 430 amino acids (aa) with a deduced Mr of 48,633. Preceding an ATG start codon, a possible transcription start point (tsp) with homology to the E. coli promoter consensus was detected. A poly-purine-rich region occurred immediately upstream from the gltA reading frame and potentially serves as a ribosome-binding site. Additionally, a G + C-rich region of dyad symmetry 3' to the translational stop codon was found that could possibly function as a Rho-independent transcriptional termination signal. A large, nearly perfect, inverted repeat was identified upstream from the gltA tsp and was shown by Southern analysis to be present in multiple copies in the C. burnetii genome. The deduced aa sequence of C. burnetii GltA was optimally aligned with enzymes from various prokaryotic sources and one eukaryotic source (pig heart). Using perfect aa identity, the C. burnetii enzyme demonstrated the greatest homology with GltA from Acinetobacter anitratum (65%). Although only 26% aa identity was seen with the pig heart enzyme, many of the residues identified in ligand binding appear to be conserved. Sequencing studies of a region centered approx. 5.6 kb upstream from gltA revealed an ORF read with opposite polarity that encodes a peptide highly homologous to the C terminus of the flavoprotein subunit of E. coli succinate dehydrogenase. This report represents the first nt sequence analysis of a gene of known function from the obligate intracellular parasite, C. burnetii.

Amino Acid Sequence

Genome-wide SNP data support species boundaries in sympatric Polylepis Ruiz & Pav. (Rosaceae) species from Bolivia and Ecuador.

Species delimitation in the South American genus Polylepis is notoriously challenging due to high morphological similarity and phenotypic plasticity, likely driven by hybridization and gene flow. Previous phylogenetic studies suggested that genetic structure aligns more strongly with geography than with taxonomy, questioning existing species concepts and hampering conservation efforts. We used double-digest RAD sequencing (ddRADseq) to generate genome-wide SNP data for 11 Polylepis species sampled across multiple localities in Bolivia and Ecuador. Population genetic analyses, phylogenetic inference, and network approaches were combined to assess whether genetic structure aligns more closely with taxonomy or geography. Morphologically defined species formed largely cohesive genetic lineages across regions, with species identity explaining substantially more genetic variation than locality. While localized admixture and reticulation were detected among closely related taxa, widespread species showed strong genetic cohesion and clear separation from congeners. Our results indicate that the sampled Polylepis species from Bolivia and Ecuador maintain distinct genetic identities despite localized signals consistent with gene flow. This genome-wide support for current taxonomy highlights Polylepis as a valuable model for studying speciation under gene flow and indicates that multiple geographic sampling will be essential in reconstructing a robust phylogeny of the genus, with important implications for conservation planning in Andean montane forests.

Bolivia

Efficient sequence alignment algorithms.

Sequence alignments are becoming more important with the increase of nucleic acid data. Fitch and Smith have recently given an example where multiple insertion/deletions (rather than a series of adjacent single insertion/deletions) are necessary to achieve the correct alignment. Multiple insertion/deletions are known to increase computation time from O(n2) to O(n3) although Gotoh has presented an O(n2) algorithm in the case the multiple insertion/deletion weighting function is linear. It is argued in this paper that it could be desirable to use concave weighting functions. For that case, an algorithm is derived that is conjectured to be O(n2).

Base Sequence

Sex hormone-binding globulin, androgen-binding protein, and vitamin K-dependent protein S are homologous to laminin A, merosin, and Drosophila crumbs protein.

Androgen-binding protein (ABP) and sex hormone-binding globulin (SHBG) are extracellular steroid-binding proteins that are homologous to the COOH-terminal domain of vitamin K-dependent protein S, a protein important in blood clotting. We find that the sequences of ABP, SHBG, and protein S are also similar to two basement membrane proteins, laminin and merosin, and to an integral membrane protein, Drosophila crumbs protein. These latter three proteins have important roles in regulating differentiation and development. The sequence similarity corresponds to the G domain of laminin A chain, which binds heparin and type IV collagen. Analysis of a multiple alignment of these proteins reveals one well-conserved segment corresponding to the part of SHBG that binds to its membrane receptor and another corresponding to the part of protein S that binds to C4b-binding protein. The similarities suggest that ABP, SHBG, and protein S may also have functions related to that of laminin and merosin.

Amino Acid Sequence

Inching toward reality: an improved likelihood model of sequence evolution.

Our previous evolutionary model is generalized to permit approximate treatment of multiple-base insertions and deletions as well as regional heterogeneity of substitution rates. Parameter estimation and alignment procedures that incorporate these generalizations are developed. Simulations are used to assess the accuracy of the parameter estimation procedure and an example of an inferred alignment is included.

Animals

Similarities between putative transport proteins of plant viruses.

The nucleic acids of many plant viruses encode proteins with one or more of the following properties: an Mr of approximately 30,000, localization in the cell wall of the infected plant and a demonstrated role in cell-to-cell transport of infection. A progressive alignment strategy, aligning first those sequences known to be similar, and then aligning the resulting groups of sequences, was used to examine further the relatedness of the amino acid sequences of putative transport proteins of caulimoviruses, of proteins similar to the putative transport protein of alfalfa mosaic virus (A1MV) and of those similar to the tobacco mosaic virus (TMV) 30K protein. The strategy first identified regions in which multiple dipeptides of one group were similar to those of another group. The regions of similarity were brought into alignment by the conservative introduction of gaps. The positions of the introduction of gaps were adjusted to optimize similarity. Statistical significances of the resulting alignments, determined both by comparison with shuffled amino acid sequences and with the sequence alignment off-set by 1 to 15 residues in each direction, suggest that the amino acid sequences of the three groups of viruses are distantly related. Nevertheless, significant relationships between members of the caulimoviral group of sequences and members of each of the A1MV-like and TMV-like groups were found. These relationships and the analysis of the number of insertions/deletions between present sequences and a hypothetical common ancestor suggest that the sequences of the caulimoviral proteins are less diverged from the ancestor than either the A1MV-like or TMV-like proteins. The alignment identified common regions of predicted secondary structure and regions of similar hydropathy, regions possibly crucial for proper functioning of the proteins.

Amino Acid Sequence