Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple Sequence Alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Co-evolution of proteins with their interaction partners.

The divergent evolution of proteins in cellular signaling pathways requires ligands and their receptors to co-evolve, creating new pathways when a new receptor is activated by a new ligand. However, information about the evolution of binding specificity in ligand-receptor systems is difficult to glean from sequences alone. We have used phosphoglycerate kinase (PGK), an enzyme that forms its active site between its two domains, to develop a standard for measuring the co-evolution of interacting proteins. The N-terminal and C-terminal domains of PGK form the active site at their interface and are covalently linked. Therefore, they must have co-evolved to preserve enzyme function. By building two phylogenetic trees from multiple sequence alignments of each of the two domains of PGK, we have calculated a correlation coefficient for the two trees that quantifies the co-evolution of the two domains. The correlation coefficient for the trees of the two domains of PGK is 0. 79, which establishes an upper bound for the co-evolution of a protein domain with its binding partner. The analysis is extended to ligands and their receptors, using the chemokines as a model. We show that the correlation between the chemokine ligand and receptor trees' distances is 0.57. The chemokine family of protein ligands and their G-protein coupled receptors have co-evolved so that each subgroup of chemokine ligands has a matching subgroup of chemokine receptors. The matching subfamilies of ligands and their receptors create a framework within which the ligands of orphan chemokine receptors can be more easily determined. This approach can be applied to a variety of ligand and receptor systems.

Chemokines↗

Molecular cloning of the arylsulfate sulfotransferase gene and characterization of its product from Enterobacter amnigenus AR-37.

The gene encoding the Enterobacter amnigenus AR-37 arylsulfate sulfotransferase (ASST) was cloned, sequenced, and expressed in Escherichia coli NM522. Sequencing led to the identification of three contiguous open reading frames (ORFs) on the same strand. Based on amino acid sequence homology, ORF1, ORF2, and ORF3 are designated astA, dsbA, and dsbB, respectively. A multiple sequence alignment revealed conserved regions in ASST. An N-terminal amino acid sequence analysis of the purified ASST from E. coli NM522 (pEAST72) showed that it is subject to N-terminal processing. The specific activity of purified ASST is 436.5 U/mg of protein. The enzyme is a monomeric protein with a molecular mass of 64 kDa. Using phenol as an acceptor substrate, 4-methylumbelliferyl sulfate is the best donor substrate, followed by beta-naphthyl sulfate, p-nitrophenyl sulfate (PNS), and alpha-naphthyl sulfate. For PNS, alpha-naphthol is the best acceptor substrate, followed by phenol, resorcinol, p-acetaminophen, tyramine, and tyrosine. The enzyme has a different acceptor specificity than the enzyme purified from Eubacterium A-44. It is similar to Klebsiella K-36 and Haemophilus K-12. The apparent K(m) values for PNS using phenol as an acceptor and for phenol using PNS as a donor are 0.163 and 0.314 mM, respectively. The pI and optimum pH are 6.1 and 9.0, respectively.

Amino Acid Sequence↗

Adelaide river rhabdovirus expresses consecutive glycoprotein genes as polycistronic mRNAs: new evidence of gene duplication as an evolutionary process.

A 3914 nucleotide region of the Adelaide River virus (ARV) genome, located immediately downstream of the M2 gene, has been cloned and sequenced. The region contains two long open reading frames (ORFs). The first encodes a protein comprising 660 amino acids which shares extensive sequence homology with the virion G protein of bovine ephemeral fever virus (BEFV) and less but significant homology with other rhabdovirus glycoproteins. The size and structural characteristics of the product indicate that it represents the 90-kDa ARV virion G protein. The second ORF encodes a polypeptide of 609 residues with nine potential glycosylation sites which is most closely related to the BEFV non-structural glycoprotein (GNS). In infected mammalian cells, the ARV G and GNS genes are transcribed primarily as a polycistronic mRNA which appears to extend from the consensus sequence (AACAG) at the start of the G gene to the next recognized polyadenylation signal (CATG[A]7) located 697 nucleotides downstream of the GNS protein termination codon. Less abundant mRNAs which appeared to initiate at consensus sequences immediately preceding and following the GNS ORF and terminate at the same polyadenylation signal were also detected. Polyadenylation-like sequences at the end of each ORF do not appear to be recognized as transcription stop signals. Multiple sequence alignments and phylogenetic analyses indicated that the ARV G and GNS glycoproteins, like those of BEFV, are structurally related and appear to have evolved at different rates from a common ancestral gene. A copy-choice mechanism, involving upstream relocation of the polymerase during replication, is proposed to account for the evolution of the tandem glycoprotein genes.

Amino Acid Sequence↗

Evolutionary relationships among putative RNA-dependent RNA polymerases encoded by a mitochondrial virus-like RNA in the Dutch elm disease fungus, Ophiostoma novo-ulmi, by other viruses and virus-like RNAs and by the Arabidopsis mitochondrial genome.

The nucleotide sequence (2617 nucleotides) of virus-like double-stranded (ds) RNA 3a in a diseased isolate, Log1/3-8d2 (Ld), of the ascomycete fungus Ophiostoma novo-ulmi has been determined. One strand of the dsRNA contains an open reading frame (ORF) with the potential to encode a protein of 718 amino acids, and the complementary strand contains two smaller ORFs with the potential to encode proteins of 178 and 182 amino acids, respectively. The large ORF contains 12 UGA codons which code for tryptophan in ascomycete mitochondria and has a codon bias typical of mitochondrial genes, consistent with the localization of Ld dsRNAs within the mitochondria. The amino acid sequence contains motifs characteristic of RNA-dependent RNA polymerases (RdRps). This putative RdRp was shown to be related to putative RdRps of mitochondrial dsRNAs of another ascomycete and a basidiomycete fungus and also to a putative RdRp encoded by the mitochondrial genome of Arabidopsis thaliana. In multiple sequence alignments, the fungal mitochondrial dsRNA-encoded RdRp-like proteins formed a cluster, ancestrally related to the RdRps of the yeast 20S and 23S RNA replicons and of the positive-stranded RNA bacteriophages of the Leviviridae family, but distinct from RdRps of other families and genera of fungal RNA viruses and related plant and animal RNA viruses. Northern blot analysis with RNA 3a strand-specific probes indicated that nucleic acid extracts of Ld contain more single-stranded (positive-stranded) RNA than dsRNA, consistent with an evolutionary relationship between RNA 3a and positive-stranded RNA phages.

Amino Acid Sequence↗

Bioinformatics in protein analysis.

The chapter gives an overview of bioinformatic techniques of importance in protein analysis. These include database searches, sequence comparisons and structural predictions. Links to useful World Wide Web (WWW) pages are given in relation to each topic. Databases with biological information are reviewed with emphasis on databases for nucleotide sequences (EMBL, GenBank, DDBJ), genomes, amino acid sequences (Swissprot, PIR, TrEMBL, GenePept), and three-dimensional structures (PDB). Integrated user interfaces for databases (SRS and Entrez) are described. An introduction to databases of sequence patterns and protein families is also given (Prosite, Pfam, Blocks). Furthermore, the chapter describes the widespread methods for sequence comparisons, FASTA and BLAST, and the corresponding WWW services. The techniques involving multiple sequence alignments are also reviewed: alignment creation with the Clustal programs, phylogenetic tree calculation with the Clustal or Phylip packages and tree display using Drawtree, njplot or phylo_win. Finally, the chapter also treats the issue of structural prediction. Different methods for secondary structure predictions are described (Chou-Fasman, Garnier-Osguthorpe-Robson, Predator, PHD). Techniques for predicting membrane proteins, antigenic sites and postranslational modifications are also reviewed.

Computational Biology↗

Protein engineering in the alpha-amylase family: catalytic mechanism, substrate specificity, and stability.

Most starch hydrolases and related enzymes belong to the alpha-amylase family which contains a characteristic catalytic (beta/alpha)8-barrel domain. Currently known primary structures that have sequence similarities represent 18 different specificities, including starch branching enzyme. Crystal structures have been reported in three of these enzyme classes: the alpha-amylases, the cyclodextrin glucanotransferases, and the oligo-1,6-glucosidases. Throughout the alpha-amylase family, only eight amino acid residues are invariant, seven at the active site and a glycine in a short turn. However, comparison of three-dimensional models with a multiple sequence alignment suggests that the diversity in specificity arises by variation in substrate binding at the beta-->alpha loops. Designed mutations thus have enhanced transferase activity and altered the oligosaccharide product patterns of alpha-amylases, changed the distribution of alpha-, beta- and gamma-cyclodextrin production by cyclodextrin glucanotransferases, and shifted the relative alpha-1,4:alpha-1,6 dual-bond specificity of neopullulanase. Barley alpha-amylase isozyme hybrids and Bacillus alpha-amylases demonstrate the impact of a small domain B protruding from the (beta/alpha)8-scaffold on the function and stability. Prospects for rational engineering in this family include important members of plant origin, such as alpha-amylase, starch branching and debranching enzymes, and amylomaltase.

Amino Acid Sequence↗

Response regulators of bacterial signal transduction systems: selective domain shuffling during evolution.

Response regulators of bacterial sensory transduction systems generally consist of receiver module domains covalently linked to effector domains. The effector domains include DNA binding and/or catalytic units that are regulated by sensor kinase-catalyzed aspartyl phosphorylation within their receiver modules. Most receiver modules are associated with three distinct families of DNA binding domains, but some are associated with other types of DNA binding domains, with methylated chemotaxis protein (MCP) demethylases, or with sensor kinases. A few exist as independent entities which regulate their target systems by noncovalent interactions. In this study the molecular phylogenies of the receiver modules and effector domains of 49 fully sequenced response regulators and their homologues were determined. The three major, evolutionarily distinct, DNA binding domains found in response regulators were evaluated for their phylogenetic relatedness, and the phylogenetic trees obtained for these domains were compared with those for the receiver modules. Members of one family (family 1) of DNA binding domains are linked to large ATPase domains which usually function cooperatively in the activation of E. coli sigma 54-dependent promoters or their equivalents in other bacteria. Members of a second family (family 2) always function in conjunction with the E. coli sigma 70 or its equivalent in other bacteria. A third family of DNA binding domains (family 3) functions by an uncharacterized mechanism involving more than one sigma factor. These three domain families utilize distinct helix-turn-helix motifs for DNA binding. The phylogenetic tree of the receiver modules revealed three major and several minor clusters of these domains. The three major receiver module clusters (clusters 1, 2, and 3) generally function with the three major families of DNA binding domains (families 1, 2, and 3, respectively) to comprise three classes of response regulators (classes 1, 2, and 3), although several exceptions exist. The minor clusters of receiver modules were usually, but not always, associated with other types of effector domains. Finally, several receiver modules did not fit into a cluster. It was concluded that receiver modules usually diverged from common ancestral protein domains together with the corresponding effector domains, although domain shuffling, due to intragenic splicing and fusion, must have occurred during the evolution of some of these proteins. Multiple sequence alignments of the 49 receiver modules and their various types of effector domains, together with other homologous domains, allowed definition of regions of striking sequence similarity and degrees of conservation of specific residues. Sequence data were correlated with structure/function when such information was available.(ABSTRACT TRUNCATED AT 250 WORDS)

Adenosine Triphosphatases↗

The evolutionary divergence of neurotransmitter receptors and second-messenger pathways.

Members of the superfamily of G-protein-coupled neurotransmitter receptors have a conserved secondary structure, a moderate and reasonably steady rate of sequence change, and usually lack introns within the coding sequence. These properties are advantageous for evolutionary studies. The duplication and divergence of the genes in this gene family led to the formation of distinct neurotransmitter pathways and may have facilitated the evolution of complex nervous systems. I have analyzed this evolutionary divergence by quantitative multiple sequence alignment, bootstrap resampling, and statistical analysis of 49 adrenergic, muscarinic cholinergic, dopamine, and octopamine receptor sequences from 12 animal species. The results indicate that the first event to occur within this gene family was the divergence of the catecholamine receptors from the muscarinic acetylcholine receptors, which occurred prior to the divergence of the arthropod and vertebrate lineages. Subsequently, the ability to activate specific second-messenger pathways diverged independently in both the muscarinic and the catecholamine receptors. This appears to have occurred after the divergence of the arthropod and vertebrate lineages but before the divergence of the avian and mammalian lineages. However, the second-messenger pathways activated by adrenergic and dopamine receptors did not diverge independently. Rather, the ability of the catecholamine receptors to bind to specific ligands, such as epinephrine, norepinephrine, dopamine, or octopamine, was repeatedly modified in evolutionary history, and in some cases was modified after the divergence of the second-messenger pathways.

Adenylyl Cyclases↗

Comparison of lantibiotic gene clusters and encoded proteins.

Lantibiotics form a group of modified peptides with unique structures, containing post-translationally modified amino acids such as dehydrated and lanthionine residues. In the gram-positive bacteria that secrete these lantibiotics, the gene clusters flanking the structural genes for various linear (type A) lantibiotics have recently been characterized. The best studied representatives are those of nisin (nis), subtilin (spa), epidermin (epi), Pep5 (pep), cytolysin (cyl), lactocin S (las) and lacticin 481 (lct). Comparison of the lantibiotic gene clusters shows that they contain conserved genes that probably encode similar functions. The nis, spa, epi and pep clusters contain lanB and lanC genes that are presumed to code for two types of enzymes that have been implicated in the modification reactions characteristic of all lantibiotics, i.e. dehydration and thio-ether ring formation. The cyl, las and lct gene clusters have no homologue of the lanB gene, but they do contain a much larger lanM gene that is the lanC gene homologue. Most lantibiotic gene clusters contain a lanP gene encoding a serine protease that is presumably involved in the proteolytic processing of the prelantibiotics. All clusters contain a lanT gene encoding an ABC transporter likely to be involved in the export of (precursors of) the lantibiotics. The lanE, lanF and lanG genes in the nis, spa and epi clusters encode another transport system that is possibly involved in self-protection. In the nisin and subtilin gene clusters two tandem genes, lanR and lanK, have been located that code for a two-component regulatory system. Finally, non-homologous genes are found in some lantibiotic gene clusters. The nisI and spaI genes encode lipoproteins that are involved in immunity, the pepI gene encodes a membrane-located immunity protein, and epiD encodes an enzyme involved in a post-translational modification found only in the C-terminus of epidermin. Several genes of unknown function are also found in the las gene cluster. A database has been assembled for all putative gene products of type A lantibiotic gene clusters. Database searches, multiple sequence alignment and secondary structure prediction have been used to identify conserved sequence segments in the LanB, LanC, LanE, LanF, LanG, LanK, LanM, LanP, LanR and LanT gene products that may be essential for structure and function. This database allows for a rapid screening of newly determined sequences in lantibiotic gene clusters.

ATP-Binding Cassette Transporters↗

Identification of novel homologues of three low molecular weight subunits of the mitochondrial bc1 complex.

Large-scale random cDNA sequencing projects have been started for several organisms and are a valuable tool for the analysis of quantitative and qualitative aspects of gene expression. However, the reliability of the obtained data is limited as most of the clones are only partially analysed on one strand. As a consequence the sequence entries derived from random cDNA sequencing projects usually comprise incomplete open reading frames. They nevertheless define complete and reliable coding sequences, if two prerequisites are fulfilled: (i) the clones encode very small proteins, and (ii) the clones have a high frequency in the cDNA-banks. The present study describes the use of cDNA databases for the identification of homologues of three low-molecular-weight subunits of the mitochondrial bc1 complex, termed the QCR6, QCR9 and QCR10 proteins. These polypeptides are only characterized for a small number of organisms, have a scarcely defined function and exhibit a low degree of structural conservation if compared between different species. Several clones were identified for each polypeptide by searches with TBLASTN using the known sequences as probes. Most of the database entries contain complete open reading frames and sequencing queries could be excluded due to the abundancy of the clones. Multiple sequence alignments are presented for all three polypeptides and consensus sequences are given which may provide a basis for the investigation of the proteins by site-directed mutagenesis.

Animals↗

Taxonomic relationships between distinct potato virus Y isolates based on detailed comparisons of the viral coat proteins and 3'-nontranslated regions.

Detailed comparisons were made of the sequences of the coat protein (CP) cistrons and 3'-nontranslated regions (3'-NTR) of 21 (geographically) distinct isolates of potato virus Y (PVY) and a virus isolate initially described as pepper mottle virus (PepMoV). Multiple sequence alignments and phylogenetic relationships based on these alignments resulted into a subgrouping of virus isolates which largely corresponded with the historical strain differentiation based on biological criteria as host range, symptomatology and serology. Virus isolates belonging to the same subgroup shared a number of characteristic CP amino acid and 3'-NTR nucleotide residues indicating that, by using sequences from the 3'-terminal region of the potyvirus genome, a distinction could be made between different isolates of one virus species as well as between different virus species. RNA secondary structure analysis of the 3'-NTR of twelve PVY isolates revealed four major stem-loop structures of which, surprisingly, the loop sequences gave a similar clustering of isolates as resulting from the overall comparisons of CP and 3'-NTR sequences. This implies a biological significance of these structural elements.

Amino Acid Sequence↗

Predicted structure of the adenovirus DNA binding protein.

The DNA sequence of a portion of the MAV1 SmaI-D fragment coding for the C-terminal 147 amino acids of the adenoviral DNA-binding protein (DBP) has been determined. A multiple sequence alignment was constructed of the MAV1 fragment and the DBPs of Ad.2, 4, 5, 7, 12, 40, and 41 to examine the degree of conservation of features that have been mapped on the Ad.2 DBP and to identify further conserved features. The less conserved N-terminal segment of the protein contains two nuclear localization signals and two acidic regions, the host range region, and all of the 11 phosphorylation sites. The highly conserved C-terminal segment contains a potential leucine zipper and zinc finger motifs. These sequence features were mapped onto a predicted secondary structure of the Ad.2 DBP.

Adenoviridae↗

Sequence of the avian adenovirus FAV 1 (CELO) DNA encoding the hexon-associated protein pVI and hexon.

The genomic region of the avian adenovirus FAV1 (CELO) encoding the precursor to virion structural protein VI (p VI) and the major capsid protein hexon has been sequenced. The 223-unit sequence of the CELO pVI protein has two potential Ad endoproteinase cleavage sites and a conserved C-terminal sequence including the Cys residue supposedly involved in endoproteinase activation. The CELO hexon gene sequence predicts a 942-residue protein (106.7 kDa). Multiple sequence alignment with other six known hexon protein sequences (human, bovine, murine, and avian) reveals high overall homology. The identity is highest in the regions corresponding to the pedestals which from the base of the hexon, and lowest in the regions corresponding to the loops which are exposed on the outer surface of the virion.

Adenoviridae↗

Compared chemical properties of dermonecrotic and lethal toxins from spiders of the genus Loxosceles (Araneae).

Loxosceles spider venom usually causes a typical dermonecrotic lesion in bitten patients, but it may also cause systemic effects that may be lethal. Gel filtration on Sephadex G-100 of Loxosceles gaucho, L. laeta, or L. intermedia spider venoms resulted in three fractions (A, containing higher molecular mass components. B containing intermediate molecular mass components, and C with lower molecular mass components). The dermonecrotic and lethal activities were detected exclusively in fraction A of all three species. Analysis by SDS-PAGE showed that the major protein contained in fraction A has molecular weight approximately 35 kDa in L. gaucho and L. intermedia, but 32 kDa in L. laeta venom. These toxins were isolated from venoms of L. gaucho, L. laeta, and L. intermedia by SDS-PAGE followed by blotting to PVDF membrane and sequencing. A database search showed a high level of identity between each toxin and a fragment of the L. reclusa (North American spider) toxin. A multiple sequence alignment of the Loxosceles toxins showed many common identical residues in their N-terminal sequences. Identities ranged from 50.0% (L. gaucho and L. reclusa) to 61.1% (L. intermedia and L. reclusa). The purified toxins were also submitted to capillary electrophoresis peptide mapping after in situ partial hydrolysis of the blotted samples. The results obtained suggest that L. intermedia protein is more similar to L. laeta toxin than L. gaucho toxin and revealed a smaller homology between L. intermedia and L. gaucho. Altogether these findings suggest that the toxins responsible for most important activities of venoms of Loxosceles species have a molecular mass of 32-35 kDa and are probably homologous proteins.

Amino Acid Sequence↗

The evolution of rhodopsins and neurotransmitter receptors.

Rhodopsins share a limited number of amino acid identities with a variety of other integral membrane proteins. Most of these proteins have seven putative transmembrane segments and are likely to play a role in transmembrane signaling. We have undertaken a systematic series of comparisons of primary and secondary structure in order to clarify the functional and evolutionary significance of these sequence similarities. On the basis of consistently high similarity scores, we find that the most internally consistent definition of the rhodopsin gene family would include vertebrate rhodopsins, alpha- and beta-adrenergic receptors, M1 and M2 muscarinic acetylcholine receptors, substance K receptors, and insect rhodopsins, while excluding bacteriorhodopsin, the mass human oncogene, vertebrate and insect nicotinic acetylcholine receptors, and the yeast STE2 and STE3 peptide receptors. The rhodopsin gene family is highly diverged at the primary sequence level but has maintained a conserved secondary structure, including a previously unidentified hierarchy of transmembrane segment hydrophobicity. We have developed new computer algorithms for progressive multiple sequence alignment and the analysis of local conservation of protein domains, and we have used these algorithms to examine the phylogeny of the rhodopsin gene family and the changing domains of sequence conservation. The results show striking differences and similarities in the conserved domains in each of the three main branches of the rhodopsin gene family, and indicate that color vision arose independently in the lines of descent leading to modern humans and fruit flies.

Algorithms↗

Common origin of arthropod tyrosinase, arthropod hemocyanin, insect hexamerin, and dipteran arylphorin receptor.

Dipteran arylphorin receptors, insect hexamerins, cheliceratan and crustacean hemocyanins, and crustacean and insect tyrosinases display significant sequence similarities. We have undertaken a systematic comparison of primary and secondary structures of these proteins. On the basis of multiple sequence alignments the phylogeny of these proteins was investigated. Hexamerin subunits, hemocyanin subunits, and tyrosinases share extensive similarities throughout the entire amino acid sequence. Our studies suggest the origin of arthropod hemocyanins from ancient tyrosinase-like proteins. Insect hexamerins likely evolved from hemocyanins of ancient crustaceans, supporting the proposed sister-group position of these subphyla. Arylphorin receptors, responsible for incorporation of hexamerins into the larval fat body of diptera, are related to hexamerins, hemocyanins, and tyrosinase. The receptor sequences display extensive similarities to the first and third domains of hemocyanins and hexamerins. In the middle region only limited amino acid conservation was observed. Elements important for hexamer formation are deleted in the receptors. Phylogenetic analysis indicated that dipteran arylphorin receptors diverged from ancient hexamerins, probably early in insect evolution.

Amino Acid Sequence↗

Structural and functional relationships of human DNA polymerases.

A continuing theme of our laboratory has been the understanding of human DNA polymerases at the structural level. We have purified DNA polymerases delta, epsilon and alpha from human placenta. Monoclonal antibodies to these polymerases were isolated and used as tools to study their immunochemical relationships. These studies have shown that while DNA polymerases delta, epsilon and alpha are discrete proteins, they must share common structural features by virtue of the ability of several of our monoclonal antibodies to exhibit cross-reactivity. A second approach we have taken is the molecular cloning of human DNA polymerase delta and epsilon. We have cloned the DNA polymerase delta cDNA, and this has allowed us to compare its primary structure to those of human polymerase alpha and other members of this polymerase family. Multiple sequence alignments have revealed that human DNA polymerase delta is also closely related to the herpes virus family of DNA polymerases. In situ hybridization has shown that the human DNA polymerase delta gene is localized to chromosome 19 q13.3-q13.4. In order to further determine the functional regions of the DNA polymerase delta structure we are currently expressing human pol delta in E. coli and baculovirus systems. Other work in our laboratory is directed toward examining the expression of DNA polymerase delta during the cell cycle.

Amino Acid Sequence↗

Hepatitis C virus encodes a selenium-dependent glutathione peroxidase gene. Implications for oxidative stress as a risk factor in progression to hepatocellular carcinoma.

AIM: Using structural bioinformatics methods, the aim is to assess the hypothesis that hepatitis C virus (HCV) encodes a glutathione peroxidase (GPx) gene in an overlapping reading frame, linking HCV expression and pathogenesis to the Se status and dietary oxidant/Antioxidant balance of the host. METHODS: The putative HCV GPx gene was identified by searching viral sequence databases, using conserved GPx active site sequences as probes, giving particular weight to the UGA (selenocysteine) codon. Multiple sequence alignments were generated and analyzed to validate the sequence similarity, and to establish the degree of conservation of the identified genomic features in HCV. Molecular modeling was used to assess the structural feasibility of the proposed homology. RESULTS: The GPx homology region overlaps the NS4 gene, and is well conserved in HCV. The sequence similarity of the conserved active site regions to a set of known GPx is high (4 to 6 SD greater than expected for similar random sequences). The computed strain energy of a molecular model of the HCV GPx is energetically favorable, comparable to the bovine GPx structure. CONCLUSIONS: By linking HCV replication and pathogenesis to the Se status and dietary oxidant/antioxidant balance of the host, the existence of a viral GPx gene could help to explain why HCV disease progression is accelerated by oxidant stresses such as alcoholism and iron overload.

Amino Acid Sequence↗