Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

[Genes of the lipase family: comparison of nucleic and proteinic sequences].

Vertebrates' plasmatic apolipoproteins and a few number of lipases in their metabolism present sequence homologies. They are grouped in genes families. The four exons apolipoproteins gene family includes nine human genes: the divergence rate of their sequences allows to place the first ancestral gene very high in the phylogenetic tree of the evolution. However, a more recent duplication of apolipoprotein C-I gene dating from 40 millions years, may be a phylogenetic marker for the radiation of Monkeys. Pancreatic lipase and isoforms, lipoprotein-lipase and hepatic triacylglycerol-lipase form by their homologies a "superfamily" of genes, which also includes yolk proteins of Dipterians eggs. Sequence homologies of PL, LPL and HL are analysed and compared with multiple alignments of amino-acids and nucleotides on spreadsheets. From these comparisons we may characterize four classes of phylogenetic markers: 1) repetitive DNA sequence (Alu, B1, PRE-1) appeared during Mammals evolution, 2) short insertions or deletions (within N-terminal domain) and a gene conversion in guinea-pig lineage, 3) a progressive reduction of intron number during the lipases evolution, 4) several duplications of genes which have produced the five genes of this superfamily currently known in the human genome.

Amino Acid Sequence

Phylogenetic analyses of 55 retroelements on the basis of the nucleotide and product amino acid sequences of the pol gene.

Comparisons of pol gene nucleotide and reverse transcriptase (RT) amino acid sequences of 47 retroviruses, 3 caulimoviruses, and 5 hepadnaviruses showed that approximately one-third of the gene at the 5' end is much more conserved than other pol regions. The most conserved regions on both the nucleotide and amino acid sequences were chosen for construction of phylogenetic trees. The maximum-parsimony and distance-matrix methods were used for analyses of aligned amino acid sequences; these two methods, and the compatibility method, were used to analyze the aligned nucleotide sequences. Essentially identical majority-rule consensus trees were produced by these different methods from both the pol gene nucleotide and RT amino acid sequences, which divided the 55 retroelements into six major groups. The reliability of the phylogenetic trees was probed with the bootstrapping of 100 replicates of the original sequence alignments. The grouping results were shown to be statistically significant by multiple comparisons with the least-significant-difference procedure.

Amino Acid Sequence

Complement components C1r/C1s, bone morphogenic protein 1 and Xenopus laevis developmentally regulated protein UVS.2 share common repeats.

Property patterns were constructed, based on an alignment of related domains in human complement subcomponents C1r and C1s as well as in the sea urchin protein uEGF. This kind of consensus pattern was able to identify similar domains in a human bone morphogenic protein, in a Xenopus laevis embryonal protein involved in dorsoanterior development and in a calcium-dependent serine protease secreted from malignant hamster embryo fibroblast cells. Because of the high level of overall sequence homology this protease may be the hamsters' equivalent of the human complement subcomponent C1s. The resulting multiple alignment of all studied domains suggests functionally and structurally important regions.

Amino Acid Sequence

Sensitive methods for determining the relatedness of proteins with limited sequence homology.

Recently, considerable advances have been made in attempts to determine the relatedness of protein sequences distant in evolution, when little or no knowledge is available concerning the corresponding tertiary architectures. Several improvements have been made to existing techniques, and these include better amino acid substitution weights contained in scoring matrices, better understanding of the effect of different gap penalty values in delineating the optimal alignment of two sequences, improved assessment of the significance of suggested sequence similarities, consideration of high scoring alternative alignments, and advances in searching entire sequence databases with the profile technique utilizing multiple-sequence information. New approaches that search for similarity a query sequence against large data banks rely on highly conserved segmental motifs defined from an aligned family of sequences. A sensitive algorithm to find distant repeats within one primary structure has also been developed recently. Solution of the inverse protein-folding problem, which involves an estimation of the ability of a sequence to take on a known main-chain tertiary topology (despite little homology with the known sequence), is being facilitated by the recent explosion in the number of new algorithms.

Amino Acid Sequence

An efficient and reliable method for cloning PCR-amplification products: a survey of point mutations in integrin cDNA.

A highly efficient, non-labor-intensive method for cloning DNA fragments produced by PCR amplification was used to carry out a rapid survey of potential point mutations in integrin alpha 6 cDNA from 17 different cell-type sources. The method includes glass powder purification of the PCR reaction mixture, followed by simultaneous treatment with T4 polynucleotide kinase and DNA polymerase I, and another glass powder purification. Sequences from multiple subclones of each cell type were readily generated, aligned and checked for mismatches. Several commonly used alternative procedures were compared for cloning efficiency and size-fidelity of inserted DNA fragments.

Amino Acid Sequence

Molecular mimicry in T cell-mediated autoimmunity: viral peptides activate human T cell clones specific for myelin basic protein.

Structural similarity between viral T cell epitopes and self-peptides could lead to the induction of an autoaggressive T cell response. Based on the structural requirements for both MHC class II binding and TCR recognition of an immunodominant myelin basic protein (MBP) peptide, criteria for a data base search were developed in which the degeneracy of amino acid side chains required for MHC class II binding and the conservation of those required for T cell activation were considered. A panel of 129 peptides that matched the molecular mimicry motif was tested on seven MBP-specific T cell clones from multiple sclerosis patients. Seven viral and one bacterial peptide efficiently activated three of these clones. Only one peptide could have been identified as a molecular mimic by sequence alignment. The observation that a single T cell receptor can recognize quite distinct but structurally related peptides from multiple pathogens has important implications for understanding the pathogenesis of autoimmunity.

Amino Acid Sequence

A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given.

Among the fundamental problems in molecular evolution and in the analysis of homologous sequences are alignment, phylogeny reconstruction, and the reconstruction of ancestral sequences. This paper presents a fast, combined solution to these problems. The new algorithm gives an approximation to the minimal history in terms of a distance function on sequences. The distance function on sequences is a minimal weighted path length constructed from substitutions and insertions-deletions of segments of any length. Substitutions are weighted with an arbitrary metric on the set of nucleotides or amino acids, and indels are weighted with a gap penalty function of the form gk = a + (bxk), where k is the length of the indel and a and b are two positive numbers. A novel feature is the introduction of the concept of sequence graphs and a generalization of the traditional dynamic sequence comparison algorithm to the comparison of sequence graphs. Sequence graphs ease several computational problems. They are used to represent large sets of sequences that can then be compared simultaneously. Furthermore, they allow the handling of multiple, equally good, alignments, where previous methods were forced to make arbitrary choices. A program written in C implemented this method; it was tested first on 22 5S RNA sequences.

Algorithms

Sequence analysis of firefly luciferase family reveals a conservative sequence motif.

A conservative sequence motif was extracted from an alignment of the firefly luciferase family. AngR derived from a pathogenic bacterium and acetyl-CoA synthetases derived from two ascomycete fungi were identified as members of the firefly luciferase family by a homology search with the motif and other sequence comparison analyses. The motif sequence shares several characteristics with the phosphate-binding sites of phosphoproteins and nucleotide-binding proteins. A multiple alignment and an unrooted phylogenetic tree were constructed for the investigation of evolutionary relationships within the firefly luciferase family.

Amino Acid Sequence

PATMAT: a searching and extraction program for sequence, pattern and block queries and databases.

A program has been developed that provides molecular biologists with multiple tools for searching databases, yet uses a very simple interface. PATMAT can use protein or (translated) DNA sequences, patterns or blocks of aligned proteins as queries of databases consisting of amino acid or nucleotide sequences, patterns or blocks. The ability to search databases of blocks by 'on-the-fly' conversion to scoring matrices provides a new tool for detection and evaluation of distant relationships. PATMAT uses a pull-down, menu-driven interface to carry out its multiple searching, extraction and viewing functions. Each query or database type is recognized, reported, and the appropriate search carried out, with matches and alignments reported in windows as they occur. Any of the high scoring matches can be exported to a file, viewed and recalled as a query using only a few keystrokes or mouse selections. Searches of multiple database files are carried out by user selection within a window. PATMAT runs under DOS; the searching engine also runs under UNIX.

Amino Acid Sequence

A homologue of the mammalian multidrug resistance gene (mdr) is functionally expressed in the intestine of Xenopus laevis.

P-glycoprotein is an integral membrane protein that functions in multidrug resistance (MDR) cells as a drug efflux pump to maintain intracellular concentrations of antitumor drugs below cytotoxic levels. A homologue of the mammalian mdr gene has been isolated and characterized from Xenopus laevis (Xe-mdr). The cDNA was isolated from a tadpole cDNA library using the full length mouse mdrlb cDNA as a probe. The Xe-mdr encodes a protein that is 66% identical to the mouse mdrlb and 68% identical to the human mdrl. The predicted structure of the Xe-mdr gene product identifies twelve membrane spanning domains and two ATP binding sites both of which are the hallmark of the ABC (ATP binding cassette) transporters. Xe-mdr mRNA is expressed as a single message of 4.5 kb and is found predominantly in the intestine. Xe-mdr message is increased 3- to 4-fold in the ileum compared to the rest of the small intestine. In situ hybridization of sequential sections from the small intestine localized the expression of the Xe-mdr to the cells lining the lumenal epithelium. Brush border membrane vesicles prepared from the small intestine of Xenopus laevis effluxed vinblastine in an ATP-dependent manner. Efflux was decreased by verapamil, a known inhibitor of P-glycoprotein function. These studies indicate that the structure of Xe-mdr has been conserved and suggest that the protein has a role in maintaining the function of the normal intestine in Xenopus.

ATP Binding Cassette Transporter, Subfamily B, Mem

Paramyosin gene (unc-15) of Caenorhabditis elegans. Molecular cloning, nucleotide sequence and models for thick filament structure.

Paramyosin is a major structural component of thick filaments isolated from many invertebrate muscles. The Caenorhabditis elegans paramyosin gene (unc-15) was identified by screening with specific antibodies an "exon-expression" library containing lacZ/nematode gene fusions. Short probes recovered from the library were used to identify bacteriophage lambda and cosmid clones that encompass the entire paramyosin (unc-15) gene. From these clones, numerous subclones containing epitopes reacting with anti-paramyosin sera were obtained, providing strong evidence that the initial cloned fragment was, in fact, derived from the structural gene for paramyosin. The complete nucleotide sequence of a 12 x 10(3) base-pair region spanning the gene was obtained. The gene is composed of ten short exons encoding a protein of 866 [corrected] amino acid residues. Paramyosin is highly similar to residues 267 to 1089 of myosin heavy chain rods. For most of its length, paramyosin appears to form an alpha-helical coiled-coil and shows the expected heptad repeat of hydrophobic amino acid residues and the 28-residue repeat of charged amino acids characteristic of myosin heavy chain rods. However, paramyosin differs from myosin in having non-helical extensions at both the N and C termini and an additional "skip" residue that interrupts the 28-residue repeat. The distribution of charges along the length of the paramyosin rod is also significantly different from that of myosin heavy chain rods. Potential charge-mediated interactions between paramyosin rods and between paramyosin and myosin rods were calculated using a model successfully applied previously to the analysis of the myosin rod sequences. Myosin rods aligned in parallel show optimal charge-charge interactions at multiples of 98 residue staggers (i.e. at axial displacements of multiples of 143 A). Paramyosin rods, in contrast, appear to interact optimally at parallel staggers of 493 residues (i.e. at axial displacements of 720 A) but show only weak interaction peaks at 98 or 296 residues. Similar calculations suggest optimal interactions between paramyosin molecules and myosin rods and in their anti-parallel alignments. The implications of these results for the structure of the bare zone and the assembly of nematode thick filaments are discussed.

Animals

Cooperation of transposable elements to endow global networks of initiators of hybrid assembly pathways of endogenous multiprotein complexes.

Mechanisms governing initiation steps of the assembly of endogenous multi-protein complexes (EMC) remain incompletely understood. Here, multiple lines of observations are reported describing the function-aligned initiation sequence of hybrid assembly pathways (HAP) of EMC. The first step of HAP-guided chain reactions of protein-protein interactions (PPI) of EMC assemblies constitutes the creation of cell type-specific pools of hetero and homo dimers. The molecular anatomy of HAP was elucidated by defining qualitative and quantitative characteristics of protein binding to a compendium of 200,393 distinct genomic regulatory elements (GRE), including 49,667 sequences representing control sets of genomic loci as well as 150,726 GRE of different evolutionary origins. The consensus sequence of HAP actions consists of: a) Initiation on genomic DNA of the formation of metastable hetero- and homodimers of EMCs' protein constituents; b) Release of dimers from DNA templates for delivery to the EMC assembly compartments; c) Assembly of defined EMC by sequential on demand addition of proteins to preformed dimers serving as attractors of EMC-specific ensembles of monomers. Chromosome-naïve DNA scaffolds facilitating creation of intracellular dimer pools engage networks of ~700 transcription factors (TFs), 534 of which manifest region-specific patterns of significantly enriched expression in 1358 brain regions. HAP initiators appear to operate within nucleosome-depleted islands of transposable elements (TE) - derived sequences within heterochromatin. PPI assembly lines of EMCs operate in 2 concurrent modes: TF-TF PPI cascade and PPI HUB protein cascade. Regardless of the number of DNA-bound initiator TFs (ranging from one to 716 TFs), both modes of operations reached the equilibrium at the PPI constituents saturation levels of ~245 proteins for TF-TF PPI modes and of ~351 proteins for PPI HUB protein modes. Distinct panels of DNA-bound initiator TFs and proteins of PPI cascade ensembles are enriched in either defined sets of neuroanatomical structures (TF-TF mode) or among structural-functional constituents of synapses (HUB proteins mode). Thus, these bifurcated cascades appear biologically congruent: TF-TF constituents map to transcriptional signatures of hundreds of brain regions, whereas HUB constituents map to synaptogenesis and synaptic structures, suggesting the unified logic of genomic functions coordinating region identity and connectivity. Evidence-supported examples of default operations of PPI-guided assemblies of hetero- and homodimers of Yamanaka factors, neurogenesis constituents, and protein components of postsynaptic density of excitatory and inhibitory synaptogenesis are reported with detailed analytical focus on human Claustrum. The foundational set of observations reported in this contribution should facilitate experimental and theoretical explorations of TE-seeded genomic codes for initiators of PPI chain reactions of protein dimerization creating pools of attractors to guide and accelerate the EMC assemblies.

Humans

Maximum entropy weighting of aligned sequences of proteins or DNA.

In a family of proteins or other biological sequences like DNA the various subfamilies are often very unevenly represented. For this reason a scheme for assigning weights to each sequence can greatly improve performance at tasks such as database searching with profiles or other consensus models based on multiple alignments. A new weighting scheme for this type of database search is proposed. In a statistical description of the searching problem it is derived from the maximum entropy principle. It can be proved that, in a certain sense, it corrects for uneven representation. It is shown that finding the maximum entropy weights is an easy optimization problem for which standard techniques are applicable.

Amino Acid Sequence

Sequence homology and absence of mRNA defines a possible pseudogene member of the Trypanosoma cruzi gp85/sialidase multigene family.

A genomic clone, pTt21, containing DNA apparently transcribed specifically in Trypanosoma cruzi trypomastigotes, was obtained by differentially screening a genomic library with trypomastigote and epimastigote cDNA. This 3444-bp clone contained open reading frames at each end, separated by a 1.8-kb non-coding region. The translated polypeptide from the 3' open reading frame (ORF2) of 1037 bp had 25-30% identity with 5 recently published T. cruzi gp85/sialidase sequences, and 20-25% identity with bacterial sialidases. Rabbit antiserum raised against an Escherichia coli fusion protein derived from the 5' open reading frame (ORF1) identified a surface antigen of 160 kDa, specifically expressed in trypomastigotes. A probe containing the first 211 bp from ORF1 was used to obtain a complete copy (c1821) of a gene that was closely related to ORF1, and encoded another member of the gp85/sialidase family. c1821 encodes a protein of 897 amino acids, but assignment of the N-terminus of the polypeptide was not possible. The 5'-most start codon is an unfavourable context to act as a translation initiator, it does not align with the initiator methionines of other gp85/sialidase sequences, nor is it followed by a signal peptide sequence characteristically found in other gp85/sialidase sequences. Although homology with the 5' ends of other gp85/sialidase sequences decays towards the 5' end of c1821, alignment of c1821 with 4 other gp85/sialidases indicated that the coding sequence should extend upstream at least 160 amino acids. In this region of c1821 there are multiple stop codons in each frame. The presence of the stop codons, the alignment data and our inability to amplify reverse transcribed mRNA using four internal primers, suggest that c1821 may not be present as a mature mRNA and is a pseudogene. Comparison of the apparently non-repetitive 3' coding domain of c1821 with the corresponding repetitive domains of two other members of the gp85/sialidase family revealed a high degree of similarity in nucleotide but not in amino acid sequence, and c1821 may thus represent an evolutionary intermediate between sub-families of the gp85/sialidase superfamily.

Amino Acid Sequence

Structure and expression of the human apolipoprotein A-IV gene.

We have isolated the human apolipoprotein (apo) A-IV gene from a cosmid library and determined its complete nucleotide sequence. The gene contains three exons of 162, 127, and 1180 nucleotides separated by two introns of 357 and 777 nucleotides. A sequence polymorphism has been identified in the 3' noncoding portion of the third exon. The human apoA-IV gene lacks an intron in the area encoding the 5' nontranslated region of its mRNA, which distinguishes it from all the other human apolipoprotein genes whose sequences are known. Comparison matrix analysis of the human apoA-IV gene sequence revealed evidence for an ancestral 11-nucleotide repeat unit that spans the third exon. These repeated sequences are much more highly conserved than those present in either rat apoA-IV or in any other human apolipoprotein. Optimal alignments of the 5' flanking regions of the rat and human apoA-IV genes disclosed multiple deletions in the rat sequence as well as a highly conserved region of 90 nucleotides (90% sequence identity) located within 170 nucleotides of the start site of transcription. The 5' flanking regions of the human and rat apoA-IV genes were ligated to the bacterial chloramphenicol acetyltransferase gene, then transfected into different cultured cells. The apoA-IV gene sequences elicited preferential expression of chloramphenicol acetyltransferase activity when introduced into intestinally derived Caco-2 cells and liver-derived Hep-G2 cells, consistent with the tissue specificity of the native gene. Analysis of deletion mutants of the human apoA-IV 5' flanking region indicated that regions from -293 to -233 and from -127 to -60 upstream of the transcription start site contain sequences required for maximum gene expression. These findings on the structure and expression of rat and human apoA-IV should prove useful in studying the control of the apoA-IV gene.

Amino Acid Sequence

A genetic analysis of HIV-1 from Punjab, India reveals the presence of multiple variants.

OBJECTIVE: To determine the extent of HIV-1 genetic variation in Indian patients. DESIGN: To avoid any bias in selecting viral variants, HIV-1 DNA was amplified directly from the peripheral blood mononuclear cells of patients and sequenced. Genetic similarity between Indian sequences and other geographic isolates was analysed by phylogenetic analysis algorithms. METHODS: A fragment encompassing the C2/V3-V5 regions of HIV-1 gp120 was amplified from the lymphocyte DNA of 12 Indian patients. Multiple clones from each patient were sequenced. Nucleotide sequences encompassing about 650 base pairs were aligned for the Indian and other geographically distinct isolates. Inter-isolate relationships were analysed by means of distance, parsimony and neighbour-joining algorithms. RESULTS: Nucleotide sequence comparisons showed low interpatient variation. Amino-acid comparisons revealed a high degree of homology between Indian sequences in this study and those studied earlier. On distance and parsimony trees, most of the Indian sequences clustered together as subtype C. However, sequences from three patients also showed significant homologies and phylogenetic clustering outside of subtype C. CONCLUSIONS: The predominant strain of HIV-1 in India belongs to subtype C and little interpatient nucleotide sequence divergence in the majority of cases suggests recent spread of HIV-1 in this region. This study also presents the first evidence for non-C subtypes in the Indian population with two epidemiologically linked samples remaining unclassified for any existing env subtype. The presence of variant subtypes in Indian patients sheds light on the transmission routes of HIV-1 to India and emphasizes the need to include these sequences in vaccine development strategies.

Acquired Immunodeficiency Syndrome