Search PubMedSearch

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Predicted structure of the adenovirus DNA binding protein.

The DNA sequence of a portion of the MAV1 SmaI-D fragment coding for the C-terminal 147 amino acids of the adenoviral DNA-binding protein (DBP) has been determined. A multiple sequence alignment was constructed of the MAV1 fragment and the DBPs of Ad.2, 4, 5, 7, 12, 40, and 41 to examine the degree of conservation of features that have been mapped on the Ad.2 DBP and to identify further conserved features. The less conserved N-terminal segment of the protein contains two nuclear localization signals and two acidic regions, the host range region, and all of the 11 phosphorylation sites. The highly conserved C-terminal segment contains a potential leucine zipper and zinc finger motifs. These sequence features were mapped onto a predicted secondary structure of the Ad.2 DBP.

Adenoviridae

Sequence of the avian adenovirus FAV 1 (CELO) DNA encoding the hexon-associated protein pVI and hexon.

The genomic region of the avian adenovirus FAV1 (CELO) encoding the precursor to virion structural protein VI (p VI) and the major capsid protein hexon has been sequenced. The 223-unit sequence of the CELO pVI protein has two potential Ad endoproteinase cleavage sites and a conserved C-terminal sequence including the Cys residue supposedly involved in endoproteinase activation. The CELO hexon gene sequence predicts a 942-residue protein (106.7 kDa). Multiple sequence alignment with other six known hexon protein sequences (human, bovine, murine, and avian) reveals high overall homology. The identity is highest in the regions corresponding to the pedestals which from the base of the hexon, and lowest in the regions corresponding to the loops which are exposed on the outer surface of the virion.

Adenoviridae

Compared chemical properties of dermonecrotic and lethal toxins from spiders of the genus Loxosceles (Araneae).

Loxosceles spider venom usually causes a typical dermonecrotic lesion in bitten patients, but it may also cause systemic effects that may be lethal. Gel filtration on Sephadex G-100 of Loxosceles gaucho, L. laeta, or L. intermedia spider venoms resulted in three fractions (A, containing higher molecular mass components. B containing intermediate molecular mass components, and C with lower molecular mass components). The dermonecrotic and lethal activities were detected exclusively in fraction A of all three species. Analysis by SDS-PAGE showed that the major protein contained in fraction A has molecular weight approximately 35 kDa in L. gaucho and L. intermedia, but 32 kDa in L. laeta venom. These toxins were isolated from venoms of L. gaucho, L. laeta, and L. intermedia by SDS-PAGE followed by blotting to PVDF membrane and sequencing. A database search showed a high level of identity between each toxin and a fragment of the L. reclusa (North American spider) toxin. A multiple sequence alignment of the Loxosceles toxins showed many common identical residues in their N-terminal sequences. Identities ranged from 50.0% (L. gaucho and L. reclusa) to 61.1% (L. intermedia and L. reclusa). The purified toxins were also submitted to capillary electrophoresis peptide mapping after in situ partial hydrolysis of the blotted samples. The results obtained suggest that L. intermedia protein is more similar to L. laeta toxin than L. gaucho toxin and revealed a smaller homology between L. intermedia and L. gaucho. Altogether these findings suggest that the toxins responsible for most important activities of venoms of Loxosceles species have a molecular mass of 32-35 kDa and are probably homologous proteins.

Amino Acid Sequence

The evolution of rhodopsins and neurotransmitter receptors.

Rhodopsins share a limited number of amino acid identities with a variety of other integral membrane proteins. Most of these proteins have seven putative transmembrane segments and are likely to play a role in transmembrane signaling. We have undertaken a systematic series of comparisons of primary and secondary structure in order to clarify the functional and evolutionary significance of these sequence similarities. On the basis of consistently high similarity scores, we find that the most internally consistent definition of the rhodopsin gene family would include vertebrate rhodopsins, alpha- and beta-adrenergic receptors, M1 and M2 muscarinic acetylcholine receptors, substance K receptors, and insect rhodopsins, while excluding bacteriorhodopsin, the mass human oncogene, vertebrate and insect nicotinic acetylcholine receptors, and the yeast STE2 and STE3 peptide receptors. The rhodopsin gene family is highly diverged at the primary sequence level but has maintained a conserved secondary structure, including a previously unidentified hierarchy of transmembrane segment hydrophobicity. We have developed new computer algorithms for progressive multiple sequence alignment and the analysis of local conservation of protein domains, and we have used these algorithms to examine the phylogeny of the rhodopsin gene family and the changing domains of sequence conservation. The results show striking differences and similarities in the conserved domains in each of the three main branches of the rhodopsin gene family, and indicate that color vision arose independently in the lines of descent leading to modern humans and fruit flies.

Algorithms

Common origin of arthropod tyrosinase, arthropod hemocyanin, insect hexamerin, and dipteran arylphorin receptor.

Dipteran arylphorin receptors, insect hexamerins, cheliceratan and crustacean hemocyanins, and crustacean and insect tyrosinases display significant sequence similarities. We have undertaken a systematic comparison of primary and secondary structures of these proteins. On the basis of multiple sequence alignments the phylogeny of these proteins was investigated. Hexamerin subunits, hemocyanin subunits, and tyrosinases share extensive similarities throughout the entire amino acid sequence. Our studies suggest the origin of arthropod hemocyanins from ancient tyrosinase-like proteins. Insect hexamerins likely evolved from hemocyanins of ancient crustaceans, supporting the proposed sister-group position of these subphyla. Arylphorin receptors, responsible for incorporation of hexamerins into the larval fat body of diptera, are related to hexamerins, hemocyanins, and tyrosinase. The receptor sequences display extensive similarities to the first and third domains of hemocyanins and hexamerins. In the middle region only limited amino acid conservation was observed. Elements important for hexamer formation are deleted in the receptors. Phylogenetic analysis indicated that dipteran arylphorin receptors diverged from ancient hexamerins, probably early in insect evolution.

Amino Acid Sequence

Structural and functional relationships of human DNA polymerases.

A continuing theme of our laboratory has been the understanding of human DNA polymerases at the structural level. We have purified DNA polymerases delta, epsilon and alpha from human placenta. Monoclonal antibodies to these polymerases were isolated and used as tools to study their immunochemical relationships. These studies have shown that while DNA polymerases delta, epsilon and alpha are discrete proteins, they must share common structural features by virtue of the ability of several of our monoclonal antibodies to exhibit cross-reactivity. A second approach we have taken is the molecular cloning of human DNA polymerase delta and epsilon. We have cloned the DNA polymerase delta cDNA, and this has allowed us to compare its primary structure to those of human polymerase alpha and other members of this polymerase family. Multiple sequence alignments have revealed that human DNA polymerase delta is also closely related to the herpes virus family of DNA polymerases. In situ hybridization has shown that the human DNA polymerase delta gene is localized to chromosome 19 q13.3-q13.4. In order to further determine the functional regions of the DNA polymerase delta structure we are currently expressing human pol delta in E. coli and baculovirus systems. Other work in our laboratory is directed toward examining the expression of DNA polymerase delta during the cell cycle.

Amino Acid Sequence

Hepatitis C virus encodes a selenium-dependent glutathione peroxidase gene. Implications for oxidative stress as a risk factor in progression to hepatocellular carcinoma.

AIM: Using structural bioinformatics methods, the aim is to assess the hypothesis that hepatitis C virus (HCV) encodes a glutathione peroxidase (GPx) gene in an overlapping reading frame, linking HCV expression and pathogenesis to the Se status and dietary oxidant/Antioxidant balance of the host. METHODS: The putative HCV GPx gene was identified by searching viral sequence databases, using conserved GPx active site sequences as probes, giving particular weight to the UGA (selenocysteine) codon. Multiple sequence alignments were generated and analyzed to validate the sequence similarity, and to establish the degree of conservation of the identified genomic features in HCV. Molecular modeling was used to assess the structural feasibility of the proposed homology. RESULTS: The GPx homology region overlaps the NS4 gene, and is well conserved in HCV. The sequence similarity of the conserved active site regions to a set of known GPx is high (4 to 6 SD greater than expected for similar random sequences). The computed strain energy of a molecular model of the HCV GPx is energetically favorable, comparable to the bovine GPx structure. CONCLUSIONS: By linking HCV replication and pathogenesis to the Se status and dietary oxidant/antioxidant balance of the host, the existence of a viral GPx gene could help to explain why HCV disease progression is accelerated by oxidant stresses such as alcoholism and iron overload.

Amino Acid Sequence

The evolution of hexamerins and the phylogeny of insects.

The evolutionary relationships among arthropod hemocyanins and insect hexamerins were investigated. A multiple sequence alignment of 12 hemocyanin and 31 hexamerin subunits was constructed and used for studying sequence conservation and protein phylogeny. Although hexamerins and hemocyanins belong to a highly divergent protein superfamily and only 18 amino acid positions are identical in all the sequences, the core structures of the three protein domains are well conserved. Under the assumption of maximum parsimony, a phylogenetic tree was obtained that matches perfectly the assumed phylogeny of the insect orders. An interesting common clade of the hymenopteran and coleopteran hexamerins was observed. In most insect orders, several paralogous hexamerin subclasses were identified that diversified after the splitting of the major insect orders. The dipteran arylphorin/LSP-1-like hexamerins were subject to closer examination, demonstrating hexamerin gene amplification and gene loss in the brachyceran Diptera. The hexamerin receptors, which belong to the hexamerin/hemocyanin superfamily, diverged early in insect evolution, before the radiation of the winged insects. After the elimination of some rapidly or slowly evolving sequences, a linearized phylogenetic tree of the hexamerins was constructed under the assumption of a molecular clock. The inferred time scale of hexamerin evolution, which dates back to the Carboniferous, agrees with the available paleontological data and reveals some previously unknown divergence times among and within the insect orders.

Amino Acid Sequence

Cloning of the homogentisate 1,2-dioxygenase gene, the key enzyme of alkaptonuria in mouse.

We determined 48 amino acid residues from five peptides from the homogeneous monomer of homogentisate 1,2-dioxygenase (HGO; E.C. 1.13. 11.15) of mouse liver. After digestion with trypsin, peptides were separated by reversed phase chromatography and amino acid sequenced. The deduced codon sequence of three peptides was used to derive degenerated oligomeres. By combining these oligos, we were able to amplify fragments from 100 to 300 bases (b) from mouse liver cDNA by polymerase chain reaction after reverse transcription (RT-PCR). A fragment of 200 b was cloned and used as a probe to screen a mouse liver cDNA library. One clone from this library contained the complete cDNA-insert for HGO as determined by sequencing. The cDNA encodes for a protein of 50 kDa, as predicted. The cDNA of mouse HGO has an overall identity of 41% to the corresponding gene hmgA from Aspergillus. Sequence similarities to human expressed sequence tags (EST) clones ranged from 70% to 20%. The positions of 122 conserved amino acids could be determined by multiple sequence alignment. We identified one first intron of 928 b in the mouse gene. The gene for HGO seems to be expressed in various tissues, as shown by RT-PCR on different cDNAs. FISH experiments with the whole murine cDNA as probe clearly revealed signals at the human chromosomal band 3q13. 3-q21. This corresponds well to the previous assignment of the locus for the human alkaptonuria gene (AKU) to the same chromosomal region by multipoint linkage analysis. We therefore conclude that the HGO cDNA encodes the gene responsible for alkaptonuria.

Alkaptonuria

High sequence similarity within ras exons 1 and 2 in different mammalian species and phylogenetic divergence of the ras gene family.

We have determined the canine and feline N-, K-, and H-ras gene sequences from position +23 to +270 covering exons I and II which contain the mutational hot spot codons 12, 13, and 61. The results were used to assess the degree of similarity between ras gene DNA regions containing the critical domains affected in neoplastic disorders in different mammalian species. The comparative analyses performed included human, canine, feline, murine, rattine, and, whenever possible, bovine, leporine (rabbit), porcelline (guinea pig), and mesocricetine (hamster) ras gene sequences within the region of interest. Comparison of feline and canine nucleotide sequences with the corresponding regions in human DNA revealed a sequence similarity greater than 85% to the human sequence. Contemporaneous analysis of previously published ras DNA sequences from other mammalian species showed a similar degree of homology to human DNA. Most nucleotide differences observed represented synonymous changes without effect on the amino acid sequence of the respective proteins. For assessment of the phylogenetic evolution of ras gene family, a maximum parsimony dendrogram based on multiple sequence alignment of the common region of exons I and II in the N-, K-, and H-ras genes was constructed. Interestingly, a higher substitution rate among the H-ras genes became apparent, indicating accelerated sequence evolution within this particular clade. The most parsimonious tree clearly shows that the duplications giving rise to the three ras genes must have occurred before the mammalian radiation.

Animals

A helix-turn-helix DNA-binding motif predicted for transposases of DNA transposons.

A helix-turn-helix (HTH) DNA-binding motif is identified in transposase sequences in Tc1, mariner and pogo DNA transposum. The findings are supported by results of various sequence analysis methods. Tc1 transposases are also predicted to contain another DNA-binding region. These findings are in accord with experimental evidence obtained from Tc1A, Tc3A and pogo transposases. The pogo family transposases, but not the pogo-type transcription factors, contain the HTH motif, suggesting that HTH structures are essential for Tc1/mariner/pogo transposition. Analysis of multiple sequence alignments enabled the identification of the HTH motif in distantly related protein sequences.

Amino Acid Sequence

Molecular characterization of morphologically typical human calicivirus Sapporo.

Human calicivirus Sapporo (SV) has typical calicivirus morphology and causes acute gastroenteritis in children. The nucleotide sequence of 3.2 kb of the 3' end of SV was determined from a cloned cDNA. The 3' end of the SV genome is predicted to encode the RNA-dependent RNA polymerase region, the capsid protein and two small open reading frames. The nonstructural and capsid protein coding sequences in the SV genome are fused in a single open reading frame. The organization of these proteins in the SV sequence is similar to that of rabbit hemorrhagic disease virus and the recently described Manchester virus, and distinct from the genome organization of the prototype human calicivirus, Norwalk virus, that lacks typical calicivirus morphology and has been described as a small round structured virus (SRSV). Sequence analysis of the predicted capsid region showed that the SV capsid is longer by approximately 30 amino acids than the capsid of any of the SRSVs, and multiple sequence alignments showed that these additional amino acids are located in the variable region of the capsid protein. Expression of the capsid protein of SV in insect cells resulted in the self-assembly of virus-like particles that have a morphology similar to that of the native virus. This result shows that calicivirus morphology is determined by the primary sequence of the capsid protein.

Amino Acid Sequence

A histidine gene cluster of the hyperthermophile Thermotoga maritima: sequence analysis and evolutionary significance.

The sequences of histidine operon genes in hyperthermophiles are informative for understanding high protein thermostability and the evolution of metabolic pathways. Therefore, a cluster of eight his genes from the hyperthermophilic and phylogenetically early bacterium Thermotoga maritima was cloned and sequenced. The cluster has the gene order hisDCBdHAFI-E, lacking only hisG and hisBp, and does not contain intercistronic regions. This compact organization of his genes resembles the his operon of enterobacteria. Sequence analysis downstream of the stop codon of hisI-E identifies a region with a significantly higher cytosine over guanosine content, which is indicative of a rho-dependent termination of transcription of the his operon. Multiple sequence alignments of N1-((5'-phosphoribosyl)-formimino)-5-aminoimidazole-4-carboxyam ide ribonucleotide isomerase (HisA) and of the cycloligase moiety of imidazoleglycerol phosphate synthase (HisF) support the previous assignment of the (beta alpha)8-barrel fold to these proteins. The alignments also reveal a second phosphate-binding motif located in the first halves of both enzymes and thereby support the hypothesis that HisA and HisF have evolved by a sequence of two gene duplication events. Comparison of the amino acid compositions of HisA and HisF from mesophiles and thermophiles shows that the thermostable variants of both enzymes contain a significantly increased number of charged amino acid residues and may therefore be stabilized by additional salt bridges.

Aldose-Ketose Isomerases

Molecular Cloning, Recombinant Expression, and In Silico Structural Analysis of Cu/Zn-Superoxide Dismutase from Trachyspermum ammi.

Superoxide dismutase (SOD) is an essential antioxidant metalloenzyme that is critical for the cellular defense against oxidative damage, as it scavenges superoxide radicals and maintains the redox status. Cytosolic Cu/Zn-SOD is particularly important in the regulation of oxidative stress among different isoforms in higher plants. While Cu/Zn-SODs from several plant species have been characterized, molecular information is limited for Trachyspermum ammi, a medicinally important member of a family Apiaceae with antioxidant potential.In the present study, an integrated molecular and in silico approach has been taken to clone and analyze a Cu/Zn type SOD gene from T. ammi to get insight into its structural and evolutionary characteristics. PCR amplification yielded an open reading frame of 456 bp encoding a protein of 152 amino acids. Sequence analysis showed that plant Cu/Zn-SODs, especially those from Daucus carota, were highly similar to one another (about 90-95%).Multiple sequence alignment confirmed the presence of conserved catalytic motifs and metal-binding histidine residues, both of which are crucial for enzymatic function. Physicochemical analysis predicted the protein to be stable, hydrophilic and compatible with cytosolic localization. The analysis of secondary structure indicated a predominance of β-strands, consistent with the conserved β-barrel architecture of plant Cu/Zn-SODs.The three-dimensional structure was built by homology modeling using a closely related plant Cu/Zn-SOD template with high sequence identity. Structural validation demonstrated an acceptable stereochemical quality with 86.3% residues in the favored region of Ramachandran plot, satisfactory ERRAT and Verify3D scores, and a low RMSD value of 0.104 Å on structural superimposition. Phylogenetic analysis placed the enzyme in the Apiaceae lineage, suggesting evolutionary conservation among related plant species. In conclusion, this study presents the first molecular and structural characterization of Cu/Zn-SOD from T. ammi and confirms the existence of a conserved structural framework typical of plant Cu/Zn-SODs. These results provide a basis for further studies concerning recombinant expression, enzymatic validation and potential relevance in antioxidant and plant stress biology.

Cloning, Molecular

Sequence and localization of human NASP: conservation of a Xenopus histone-binding protein.

In this study the sequence and localization of human testicular NASP (nuclear autoantigenic sperm protein) are reported. NASP cDNA contains 2561 nt encoding a protein of 787 amino acids. The open reading frame contains 2446 nt followed by an ochre stop codon (TAA) and 104 nucleotides of untranslated sequence containing a poly(A) addition signal 10 bases upstream of the poly(A) tail. Northern blot analysis of human testis poly(A) mRNA indicates a message of approximately 3.2 kb. Multiple sequence alignment (MSA) analysis of the encoded human NASP amino acid sequence with the sequence for the Xenopus histone-binding protein N1/N2 and the rabbit NASP amino acid sequence demonstrates that the human sequence and the Xenopus sequence have extensive amino acid homology upstream of the rabbit initiation codon. Significantly, there is an 85% identity between the human and the rabbit NASP sequences when the alignment starts at the N-terminal of the rabbit sequence and at amino acid 101 of the human sequence. The nuclear translocation signal found in N1/N2 and rabbit NASP is completely conserved in human NASP. The first histone-binding domain of Xenopus is 70% identical and 90% similar to the human NASP domain. The second histone-binding domain of Xenopus is 48% identical and 71% similar to the human NASP domain. MSA analysis of the three sequences generated an unrooted ancestral tree with two branches, indicating that fewer amino acid changes have occurred between the Xenopus and the human sequences than between the Xenopus and the rabbit sequences. In the human testis, NASP is localized predominantly in primary spermatocytes and round spermatids. Spermatogonia, Sertoli cells, Leydig cells, peritubular cells, and other somatic cells do not stain. Human spermatozoa contain NASP in the acrosomal region. Following the acrosome reaction, some NASP remains in the equatorial and postacrosomal regions. We propose that mammalian testes and sperm contain a histone-binding protein which may play a role in regulating the early events of spermatogenesis.

Amino Acid Sequence

Prediction of domain organisation and secondary structure of thyroid peroxidase, a human autoantigen involved in destructive thyroiditis.

Organ specific autoimmune diseases are relatively common immunological disorders in man which include thyroid autoimmune disease, insulin-dependent diabetes mellitus and myasthenia gravis. The target autoantigens in some of these diseases have recently been characterised. In thyroid autoimmune disease this includes the key enzyme, thyroid peroxidase (TPO), which is involved in the generation of thyroid hormone. Structural knowledge about autoantigens such as thyroid peroxidase will allow a greater understanding of the interaction between autoantigens and the aberrant immune response, and facilitate the development of strategies for antigen-specific therapeutic manipulation. We report here a prediction of the secondary structure of thyroid peroxidase, together with the results of circular dichroic spectroscopy of a homologous purified enzyme. A combination of 3 secondary structure prediction programs has been used, following multiple sequence alignment, and TPO has been found to consist mainly of alpha-helical conformation, with little beta-sheet present. This structure prediction, together with knowledge of the exon-intron boundaries allows a model for the domain organisation of the TPO molecule to be proposed.

Amino Acid Sequence

A structure-derived sequence pattern for the detection of type I copper binding domains in distantly related proteins.

A structure-based approach to the definition of sequence patterns characteristic of protein domains is presented by example. The approach requires a multiple sequence alignment of a family (or set of related families) as well as at least one three-dimensional structure. The pattern derived does not merely summarize the information in the known sequences but attempts to generalize the pattern specifications based on structural insight. In this example, the pattern-driven database search identified correctly most of the known type I copper-binding domains and detected the presence of a homologous domain in a previously unknown case (CopA protein). The significance of these results is discussed.

Amino Acid Sequence

Conservation analysis and structure prediction of the SH2 family of phosphotyrosine binding domains.

Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.

Amino Acid Sequence