Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

The evolution of hexamerins and the phylogeny of insects.

The evolutionary relationships among arthropod hemocyanins and insect hexamerins were investigated. A multiple sequence alignment of 12 hemocyanin and 31 hexamerin subunits was constructed and used for studying sequence conservation and protein phylogeny. Although hexamerins and hemocyanins belong to a highly divergent protein superfamily and only 18 amino acid positions are identical in all the sequences, the core structures of the three protein domains are well conserved. Under the assumption of maximum parsimony, a phylogenetic tree was obtained that matches perfectly the assumed phylogeny of the insect orders. An interesting common clade of the hymenopteran and coleopteran hexamerins was observed. In most insect orders, several paralogous hexamerin subclasses were identified that diversified after the splitting of the major insect orders. The dipteran arylphorin/LSP-1-like hexamerins were subject to closer examination, demonstrating hexamerin gene amplification and gene loss in the brachyceran Diptera. The hexamerin receptors, which belong to the hexamerin/hemocyanin superfamily, diverged early in insect evolution, before the radiation of the winged insects. After the elimination of some rapidly or slowly evolving sequences, a linearized phylogenetic tree of the hexamerins was constructed under the assumption of a molecular clock. The inferred time scale of hexamerin evolution, which dates back to the Carboniferous, agrees with the available paleontological data and reveals some previously unknown divergence times among and within the insect orders.

Amino Acid Sequence

Cloning of the homogentisate 1,2-dioxygenase gene, the key enzyme of alkaptonuria in mouse.

We determined 48 amino acid residues from five peptides from the homogeneous monomer of homogentisate 1,2-dioxygenase (HGO; E.C. 1.13. 11.15) of mouse liver. After digestion with trypsin, peptides were separated by reversed phase chromatography and amino acid sequenced. The deduced codon sequence of three peptides was used to derive degenerated oligomeres. By combining these oligos, we were able to amplify fragments from 100 to 300 bases (b) from mouse liver cDNA by polymerase chain reaction after reverse transcription (RT-PCR). A fragment of 200 b was cloned and used as a probe to screen a mouse liver cDNA library. One clone from this library contained the complete cDNA-insert for HGO as determined by sequencing. The cDNA encodes for a protein of 50 kDa, as predicted. The cDNA of mouse HGO has an overall identity of 41% to the corresponding gene hmgA from Aspergillus. Sequence similarities to human expressed sequence tags (EST) clones ranged from 70% to 20%. The positions of 122 conserved amino acids could be determined by multiple sequence alignment. We identified one first intron of 928 b in the mouse gene. The gene for HGO seems to be expressed in various tissues, as shown by RT-PCR on different cDNAs. FISH experiments with the whole murine cDNA as probe clearly revealed signals at the human chromosomal band 3q13. 3-q21. This corresponds well to the previous assignment of the locus for the human alkaptonuria gene (AKU) to the same chromosomal region by multipoint linkage analysis. We therefore conclude that the HGO cDNA encodes the gene responsible for alkaptonuria.

Alkaptonuria

High sequence similarity within ras exons 1 and 2 in different mammalian species and phylogenetic divergence of the ras gene family.

We have determined the canine and feline N-, K-, and H-ras gene sequences from position +23 to +270 covering exons I and II which contain the mutational hot spot codons 12, 13, and 61. The results were used to assess the degree of similarity between ras gene DNA regions containing the critical domains affected in neoplastic disorders in different mammalian species. The comparative analyses performed included human, canine, feline, murine, rattine, and, whenever possible, bovine, leporine (rabbit), porcelline (guinea pig), and mesocricetine (hamster) ras gene sequences within the region of interest. Comparison of feline and canine nucleotide sequences with the corresponding regions in human DNA revealed a sequence similarity greater than 85% to the human sequence. Contemporaneous analysis of previously published ras DNA sequences from other mammalian species showed a similar degree of homology to human DNA. Most nucleotide differences observed represented synonymous changes without effect on the amino acid sequence of the respective proteins. For assessment of the phylogenetic evolution of ras gene family, a maximum parsimony dendrogram based on multiple sequence alignment of the common region of exons I and II in the N-, K-, and H-ras genes was constructed. Interestingly, a higher substitution rate among the H-ras genes became apparent, indicating accelerated sequence evolution within this particular clade. The most parsimonious tree clearly shows that the duplications giving rise to the three ras genes must have occurred before the mammalian radiation.

Animals

A helix-turn-helix DNA-binding motif predicted for transposases of DNA transposons.

A helix-turn-helix (HTH) DNA-binding motif is identified in transposase sequences in Tc1, mariner and pogo DNA transposum. The findings are supported by results of various sequence analysis methods. Tc1 transposases are also predicted to contain another DNA-binding region. These findings are in accord with experimental evidence obtained from Tc1A, Tc3A and pogo transposases. The pogo family transposases, but not the pogo-type transcription factors, contain the HTH motif, suggesting that HTH structures are essential for Tc1/mariner/pogo transposition. Analysis of multiple sequence alignments enabled the identification of the HTH motif in distantly related protein sequences.

Amino Acid Sequence

Molecular characterization of morphologically typical human calicivirus Sapporo.

Human calicivirus Sapporo (SV) has typical calicivirus morphology and causes acute gastroenteritis in children. The nucleotide sequence of 3.2 kb of the 3' end of SV was determined from a cloned cDNA. The 3' end of the SV genome is predicted to encode the RNA-dependent RNA polymerase region, the capsid protein and two small open reading frames. The nonstructural and capsid protein coding sequences in the SV genome are fused in a single open reading frame. The organization of these proteins in the SV sequence is similar to that of rabbit hemorrhagic disease virus and the recently described Manchester virus, and distinct from the genome organization of the prototype human calicivirus, Norwalk virus, that lacks typical calicivirus morphology and has been described as a small round structured virus (SRSV). Sequence analysis of the predicted capsid region showed that the SV capsid is longer by approximately 30 amino acids than the capsid of any of the SRSVs, and multiple sequence alignments showed that these additional amino acids are located in the variable region of the capsid protein. Expression of the capsid protein of SV in insect cells resulted in the self-assembly of virus-like particles that have a morphology similar to that of the native virus. This result shows that calicivirus morphology is determined by the primary sequence of the capsid protein.

Amino Acid Sequence

A histidine gene cluster of the hyperthermophile Thermotoga maritima: sequence analysis and evolutionary significance.

The sequences of histidine operon genes in hyperthermophiles are informative for understanding high protein thermostability and the evolution of metabolic pathways. Therefore, a cluster of eight his genes from the hyperthermophilic and phylogenetically early bacterium Thermotoga maritima was cloned and sequenced. The cluster has the gene order hisDCBdHAFI-E, lacking only hisG and hisBp, and does not contain intercistronic regions. This compact organization of his genes resembles the his operon of enterobacteria. Sequence analysis downstream of the stop codon of hisI-E identifies a region with a significantly higher cytosine over guanosine content, which is indicative of a rho-dependent termination of transcription of the his operon. Multiple sequence alignments of N1-((5'-phosphoribosyl)-formimino)-5-aminoimidazole-4-carboxyam ide ribonucleotide isomerase (HisA) and of the cycloligase moiety of imidazoleglycerol phosphate synthase (HisF) support the previous assignment of the (beta alpha)8-barrel fold to these proteins. The alignments also reveal a second phosphate-binding motif located in the first halves of both enzymes and thereby support the hypothesis that HisA and HisF have evolved by a sequence of two gene duplication events. Comparison of the amino acid compositions of HisA and HisF from mesophiles and thermophiles shows that the thermostable variants of both enzymes contain a significantly increased number of charged amino acid residues and may therefore be stabilized by additional salt bridges.

Aldose-Ketose Isomerases

Molecular Cloning, Recombinant Expression, and In Silico Structural Analysis of Cu/Zn-Superoxide Dismutase from Trachyspermum ammi.

Superoxide dismutase (SOD) is an essential antioxidant metalloenzyme that is critical for the cellular defense against oxidative damage, as it scavenges superoxide radicals and maintains the redox status. Cytosolic Cu/Zn-SOD is particularly important in the regulation of oxidative stress among different isoforms in higher plants. While Cu/Zn-SODs from several plant species have been characterized, molecular information is limited for Trachyspermum ammi, a medicinally important member of a family Apiaceae with antioxidant potential.In the present study, an integrated molecular and in silico approach has been taken to clone and analyze a Cu/Zn type SOD gene from T. ammi to get insight into its structural and evolutionary characteristics. PCR amplification yielded an open reading frame of 456 bp encoding a protein of 152 amino acids. Sequence analysis showed that plant Cu/Zn-SODs, especially those from Daucus carota, were highly similar to one another (about 90-95%).Multiple sequence alignment confirmed the presence of conserved catalytic motifs and metal-binding histidine residues, both of which are crucial for enzymatic function. Physicochemical analysis predicted the protein to be stable, hydrophilic and compatible with cytosolic localization. The analysis of secondary structure indicated a predominance of β-strands, consistent with the conserved β-barrel architecture of plant Cu/Zn-SODs.The three-dimensional structure was built by homology modeling using a closely related plant Cu/Zn-SOD template with high sequence identity. Structural validation demonstrated an acceptable stereochemical quality with 86.3% residues in the favored region of Ramachandran plot, satisfactory ERRAT and Verify3D scores, and a low RMSD value of 0.104 Å on structural superimposition. Phylogenetic analysis placed the enzyme in the Apiaceae lineage, suggesting evolutionary conservation among related plant species. In conclusion, this study presents the first molecular and structural characterization of Cu/Zn-SOD from T. ammi and confirms the existence of a conserved structural framework typical of plant Cu/Zn-SODs. These results provide a basis for further studies concerning recombinant expression, enzymatic validation and potential relevance in antioxidant and plant stress biology.

Cloning, Molecular

Sequence and localization of human NASP: conservation of a Xenopus histone-binding protein.

In this study the sequence and localization of human testicular NASP (nuclear autoantigenic sperm protein) are reported. NASP cDNA contains 2561 nt encoding a protein of 787 amino acids. The open reading frame contains 2446 nt followed by an ochre stop codon (TAA) and 104 nucleotides of untranslated sequence containing a poly(A) addition signal 10 bases upstream of the poly(A) tail. Northern blot analysis of human testis poly(A) mRNA indicates a message of approximately 3.2 kb. Multiple sequence alignment (MSA) analysis of the encoded human NASP amino acid sequence with the sequence for the Xenopus histone-binding protein N1/N2 and the rabbit NASP amino acid sequence demonstrates that the human sequence and the Xenopus sequence have extensive amino acid homology upstream of the rabbit initiation codon. Significantly, there is an 85% identity between the human and the rabbit NASP sequences when the alignment starts at the N-terminal of the rabbit sequence and at amino acid 101 of the human sequence. The nuclear translocation signal found in N1/N2 and rabbit NASP is completely conserved in human NASP. The first histone-binding domain of Xenopus is 70% identical and 90% similar to the human NASP domain. The second histone-binding domain of Xenopus is 48% identical and 71% similar to the human NASP domain. MSA analysis of the three sequences generated an unrooted ancestral tree with two branches, indicating that fewer amino acid changes have occurred between the Xenopus and the human sequences than between the Xenopus and the rabbit sequences. In the human testis, NASP is localized predominantly in primary spermatocytes and round spermatids. Spermatogonia, Sertoli cells, Leydig cells, peritubular cells, and other somatic cells do not stain. Human spermatozoa contain NASP in the acrosomal region. Following the acrosome reaction, some NASP remains in the equatorial and postacrosomal regions. We propose that mammalian testes and sperm contain a histone-binding protein which may play a role in regulating the early events of spermatogenesis.

Amino Acid Sequence

Prediction of domain organisation and secondary structure of thyroid peroxidase, a human autoantigen involved in destructive thyroiditis.

Organ specific autoimmune diseases are relatively common immunological disorders in man which include thyroid autoimmune disease, insulin-dependent diabetes mellitus and myasthenia gravis. The target autoantigens in some of these diseases have recently been characterised. In thyroid autoimmune disease this includes the key enzyme, thyroid peroxidase (TPO), which is involved in the generation of thyroid hormone. Structural knowledge about autoantigens such as thyroid peroxidase will allow a greater understanding of the interaction between autoantigens and the aberrant immune response, and facilitate the development of strategies for antigen-specific therapeutic manipulation. We report here a prediction of the secondary structure of thyroid peroxidase, together with the results of circular dichroic spectroscopy of a homologous purified enzyme. A combination of 3 secondary structure prediction programs has been used, following multiple sequence alignment, and TPO has been found to consist mainly of alpha-helical conformation, with little beta-sheet present. This structure prediction, together with knowledge of the exon-intron boundaries allows a model for the domain organisation of the TPO molecule to be proposed.

Amino Acid Sequence

A structure-derived sequence pattern for the detection of type I copper binding domains in distantly related proteins.

A structure-based approach to the definition of sequence patterns characteristic of protein domains is presented by example. The approach requires a multiple sequence alignment of a family (or set of related families) as well as at least one three-dimensional structure. The pattern derived does not merely summarize the information in the known sequences but attempts to generalize the pattern specifications based on structural insight. In this example, the pattern-driven database search identified correctly most of the known type I copper-binding domains and detected the presence of a homologous domain in a previously unknown case (CopA protein). The significance of these results is discussed.

Amino Acid Sequence

Conservation analysis and structure prediction of the SH2 family of phosphotyrosine binding domains.

Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.

Amino Acid Sequence

Pattern recognition and self-correcting distance geometry calculations applied to myohemerythrin.

A topological list, consisting of segments of regular secondary structures and a list of buried and solvent accessible residues, is automatically predicted from multiple aligned sequences in a protein family. This topological list is translated into geometric constraints for distance geometry calculation in torsion angle space. A new self-correcting distance geometry method detects and eliminates false distance constraints. In an application to the four-helix bundle protein, myohem-erythrin, the right-handed global fold was correctly reproduced with a root-mean-square deviation of 2.6 A, when the topological list was derived from the X-ray structure. A predicted topological list, coupled with constraints from the residues in the active site of myohemerythrin, predicted the correct fold with a root-mean-square deviation of 4 A for backbone atoms.

Amino Acid Sequence

Specificity of the cytochrome P-450 interaction with cytochrome b5.

The specificity of the interaction of cytochrome b5 with different forms of cytochrome P-450 was examined. Immunopurification of cytochromes P-450 1A1, 2B1 and 2E1 from rat liver microsomes resulted in co-purification of cytochrome b5 with cytochrome P-450 forms 2B1 and 2E1 but not 1A1. This specificity was evaluated in conjunction with multiple sequence alignment of the three cytochrome P-450s and a molecular model of the cytochrome P-450-cytochrome b5 complex [(1989) Biochemistry 28, 8201-8205]. These analyses suggest two basic residues in the arginine cluster region of P-450, which are present in P-450s 2B1 and 2E1 but are absent in P-450 1A1, as potential binding sites for cytochrome b5.

Amino Acid Sequence

Hydrophobic cluster analysis and secondary structure predictions revealed that major and minor structural subunits of K88-related adhesins of Escherichia coli share a common overall fold and differ structurally from other fimbrial subunits.

The structural relatedness of K88-related major and minor subunits was deduced from their sequences by hydrophobic cluster analysis (HCA) and secondary structure predictions produced by the profile neural network prediction program (PHD) on multiple sequence alignments. Although the weak residue identity between major and minor subunits is evidence of a high evolutionary distance, an overall structural similarity was observed In addition, clear amphipathic conformations were conserved in predicted secondary structure. On the basis of this predicted structural similarity, a schematic 2D model of ClpG subunit was developed.

Amino Acid Sequence

Identification of additional homologues of subunits VII and VIII of the ubiquinol-cytochrome c oxidoreductase enables definition of consensus sequences.

The Candida utilis QCR7 gene encoding subunit VII of the ubiquinol-cytochrome c oxidoreductase was isolated by functional complementation of the Saccharomyces cerevisiae subunit VII-null mutant. Several other subunit VII homologues as well as homologues for subunit VIII were identified by screening the GenBank database. Some of these homologues for subunit VII could only be identified as such using a consensus sequence that was derived from the multiple sequence alignment. Definition of the consensus should facilitate further analysis of structure/function relationships in this protein.

Amino Acid Sequence

Ribosomal protein L22 from Thermus thermophilus: sequencing, overexpression and crystallisation.

The gene for the ribosomal protein L22 from Thermus thermophilus has been sequenced and overexpressed in Escherichia coli. A multiple sequence alignment was carried out for all proteins of the L22 family reported so far. The recombinant protein was purified and crystallized. The crystals belong to the space group P2(1)2(1)2(1), with cell parameters of a = 32.6 A, b = 66.0 A, c = 67.8 A.

Amino Acid Sequence

Spectroscopic study of an HIV-1 capsid protein major homology region peptide analog.

The capsid (CA) domain of retroviral Gag proteins possesses one subdomain, the major homology region (MHR), which is conserved among nearly all avian and mammalian retroviruses. While it is known that the mutagenesis of residues in the MHR will impair virus infectivity, the precise structure and function of the MHR is not known. In order to obtain further information on the MHR, we have examined the structure of a synthetic peptide encompassing the MHR of human immunodeficiency virus type I (HIV-1) CA protein. Multiple sequence alignment and secondary structure prediction indicate that the peptide could form 50% alpha-helix and 10% beta-sheet. In addition, circular dichroism studies indicate that, in the presence of 50% trifluoroethanol (TFE), the peptide adopts an alpha-helical structure over half of its length. Further analysis by proton nuclear magnetic resonance spectroscopy suggests that the C-terminal portion of the MHR forms a helix in aqueous solution. Upon the addition of TFE, the position of the helix remains nearly constant, but the magnitude of the changes in H alpha chemical shifts of the residues indicate a more stable helix. These results suggest that a helical C-terminus of retroviral MHRs may be integral to the function of this region.

Amino Acid Sequence

Cloning and sequencing of cDNA clones encoding chicken lamins A and B1 and comparison of the primary structures of vertebrate A- and B-type lamins.

Nuclear lamins are intermediate-filament-type proteins forming a fibrillar meshwork underlying the inner nuclear membrane. The existence of multiple isoforms of lamin proteins in vertebrates is believed to reflect functional specializations during cell division and differentiation. Although biochemical criteria may be used to classify many lamin isoforms into A- and B-type subfamilies, the structural features distinguishing the members of these subfamilies remain to be characterized fully. Here, we report the complete primary structures of chicken lamins A and B1, as they are deduced from cloned cDNAs; in the accompanying paper we present the complete sequence of lamin B2, a second avian B-type lamin. Comparisons of the chicken lamin sequences with each other and with those of other lamins allow us to establish structural features that are common to members of both subfamilies. Conversely, multiple sequence alignments make it possible to identify a number of structural motifs that clearly differentiate B-type lamins from A-type lamins. With this information at hand, we attempt to correlate different biochemical properties of A- and B-type lamins with the presence or absence of specific sequence motifs.

Amino Acid Sequence