Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Multiple sequence alignment”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Identification and characterization of mycobacterial proteins differentially expressed under standing and shaking culture conditions, including Rv2623 from a novel class of putative ATP-binding proteins.

The environmental signals that affect gene regulation in Mycobacterium tuberculosis remain largely unknown despite their importance to tuberculosis pathogenesis. Other work has shown that several promoters, including acr (also known as hspX) (alpha-crystallin homolog), are upregulated in shallow standing cultures compared with constantly shaking cultures. Each of these promoters is also induced to a similar extent within macrophages. The present study used two-dimensional gel electrophoresis and mass spectrometry to further characterize differences in mycobacterial protein expression during growth under standing and shaking culture conditions. Metabolic labeling of M. bovis BCG showed that at least 45 proteins were differentially expressed under standing and shaking culture conditions. Rv2623, CysA2-CysA3, Gap, and Acr were identified from each of four spots or gel bands that were specifically increased in bacteria from standing cultures. An additional standing-induced spot contained two comigrating proteins, GlcB and KatG. The greatest induction was observed with Rv2623, a 32-kDa protein of unknown function that was strongly expressed under standing conditions and absent in shaking cultures. Analysis using PROBE, a multiple sequence alignment and database mining tool, classified M. tuberculosis Rv2623 as a member of a novel class of ATP-binding proteins that may be involved in M. tuberculosis's response to environmental signals. These studies demonstrate the power of combined proteomic and computational approaches and demonstrate that subtle differences in bacterial culture conditions may have important implications for the study of gene expression in mycobacteria.

Amino Acid Sequence↗

Identification of a Mycobacterium tuberculosis putative classical nitroreductase gene whose expression is coregulated with that of the acr aene within macrophages, in standing versus shaking cultures, and under low oxygen conditions.

Tuberculosis remains a leading killer worldwide, and new approaches for its treatment and prevention are urgently needed. This effort will benefit greatly from a better understanding of gene regulation in Mycobacterium tuberculosis, particularly with respect to this pathogen's response to its host environment. We examined the behavior of two promoters from the divergently transcribed M. tuberculosis genes acr/hspX/Rv2031c (alpha-crystallin homolog) and Rv2032/acg (acr-coregulated gene) by using a promoter-GFP fusion assay in Mycobacterium bovis BCG. We found that Rv2032 is a novel macrophage-induced gene whose expression is coregulated with that of acr. Relative levels of intracellular induction for both promoters were significantly affected by shallow standing versus shaking bacterial culture conditions prior to macrophage infection, and both promoters were strongly induced under low oxygen conditions. Deletion analyses showed that DNA sequences within a 43-bp region were required for expression of these promoters under all conditions. Multiple sequence alignment and database searches performed with PROBE indicated that Rv2032 is one of eight M. tuberculosis genes of previously unknown function that belong to an unusual superfamily of classical nitroreductases, which may have a role for bacteria within the host environment. These findings show that mycobacterial culture conditions can greatly influence the results and interpretation of subsequent gene regulation experiments. We propose that these differences might be exploited for dissection of the regulatory factors that affect mycobacterial gene expression within the host.

Amino Acid Motifs↗

Membrane topology of Escherichia coli diacylglycerol kinase.

The topology of Escherichia coli diacylglycerol kinase (DAGK) within the cytoplasmic membrane was elucidated by a combined approach involving both multiple aligned sequence analysis and fusion protein experiments. Hydropathy plots of the five prokaryotic DAGK sequences available were uniform in their prediction of three transmembrane segments. The hydropathy predictions were experimentally tested genetically by fusing C-terminal deletion derivatives of DAGK to beta-lactamase and beta-galactosidase. Following expression, the enzymatic activities of the chimeric proteins were measured and used to determine the cellular location of the fusion junction. These studies confirmed the hydropathy predictions for DAGK with respect to the number and approximate sequence locations of the transmembrane segments. Further analysis of the aligned DAGK sequences detected probable alpha-helical N-terminal capping motifs and two amphipathic alpha-helices within the enzyme. The combined fusion and sequence data indicate that DAGK is a polytopic integral membrane protein with three transmembrane segments with the N terminus of the protein in the cytoplasm, the C terminus in the periplasmic space, and two amphipathic helices near the cytoplasmic surface.

Amino Acid Sequence↗

Cloning, sequencing, and disruption of a levanase gene of Bacillus polymyxa CF43.

The Bacillus polymyxa CF43 lelA gene, expressing both sucrose and fructan hydrolase activities, was isolated from a genomic library of B. polymyxa screened in Bacillus subtilis. The gene was detected as expressing sucrose hydrolase activity; B. subtilis transformants did not secrete the lelA gene product (LelA) into the extracellular medium. A 1.7-kb DNA fragment sufficient for lelA expression in Escherichia coli was sequenced. It contains a 548-codon open reading frame. The deduced amino acid sequence shows 54% identity with mature B. subtilis levanase and is similar to other fructanases and sucrases (beta-D-fructosyltransferases). Multiple-sequence alignment of 14 of these proteins revealed several previously unreported features. LelA appears to be a 512-amino-acid polypeptide containing no canonical signal peptide. The hydrolytic activities of LelA on sucrose, levan, and inulin were compared with those of B. subtilis levanase and sucrase, confirming that LelA is indeed a fructanase. The lelA gene in the chromosome of B. polymyxa was disrupted with a chloramphenicol resistance gene (cat) by "inter-gramic" conjugation: the lelA::cat insertion on a mobilizable plasmid was transferred from an E. coli transformant to B. polymyxa CF43, and B. polymyxa transconjugants containing the lelA::cat construct replacing the wild-type lelA gene in their chromosomes were selected directly. The growth of the mutant strain on levan, inulin, and sucrose was not affected.

Amino Acid Sequence↗

Biochemical characterization and sequence analysis of the gluconate:NADP 5-oxidoreductase gene from Gluconobacter oxydans.

Gluconate:NADP 5-oxidoreductase (GNO) from the acetic acid bacterium Gluconobacter oxydans subsp. oxydans DSM3503 was purified to homogeneity. This enzyme is involved in the nonphosphorylative, ketogenic oxidation of glucose and oxidizes gluconate to 5-ketogluconate. GNO was localized in the cytoplasm, had an isoelectric point of 4.3, and showed an apparent molecular weight of 75,000. In sodium dodecyl sulfate gel electrophoresis, a single band appeared corresponding to a molecular weight of 33,000, which indicated that the enzyme was composed of two identical subunits. The pH optimum of gluconate oxidation was pH 10, and apparent Km values were 20.6 mM for the substrate gluconate and 73 microM for the cosubstrate NADP. The enzyme was almost inactive with NAD as a cofactor and was very specific for the substrates gluconate and 5-ketogluconate. D-Glucose, D-sorbitol, and D-mannitol were not oxidized, and 2-ketogluconate and L-sorbose were not reduced. Only D-fructose was accepted, with a rate that was 10% of the rate of 5-ketogluconate reduction. The gno gene encoding GNO was identified by hybridization with a gene probe complementary to the DNA sequence encoding the first 20 N-terminal amino acids of the enzyme. The gno gene was cloned on a 3.4-kb DNA fragment and expressed in Escherichia coli. Sequencing of the gene revealed an open reading frame of 771 bp, encoding a protein of 257 amino acids with a predicted relative molecular mass of 27.3 kDa. Plasmid-encoded gno was functionally expressed, with 6.04 U/mg of cell-free protein in E. coli and with 6.80 U/mg of cell-free protein in G. oxydans, which corresponded to 85-fold overexpression of the G. oxydans wild-type GNO activity. Multiple sequence alignments showed that GNO was affiliated with the group II alcohol dehydrogenases, or short-chain dehydrogenases, which display a typical pattern of six strictly conserved amino acid residues.

Amino Acid Sequence↗

Structure-function relationship of bacterial prolipoprotein diacylglyceryl transferase: functionally significant conserved regions.

The structure-function relationship of bacterial prolipoprotein diacylgyceryl transferase (LGT) Has been investigated by a comparison of the primary structures of this enzyme in phylogenetically distant bacterial species, analysis of the sequences of mutant enzymes, and specific chemical modification of the Escherichia coli enzyme. A clone containing the gene for LGT, lgt, of the gram-positive species Staphylococcus aureus was isolated by complementation of the temperature-sensitive lgt mutant of E. coli (strain SK634) defective in LGT activity. In vivo and in vitro assays for prolipoprotein diacylglyceryl modification activity indicated that the complementing clone restored the prolipoprotein modification activity in the mutant strain. Sequence determination of the insert DNA revealed an open reading frame of 837 bp encoding a protein of 279 amino acids with a calculated molecular mass of 31.6 kDa. S. aureus LGT showed 24% identity and 47% similarity with E. coli, Salmonella typhimurium, and Haemophilus influenzae LGT.S. aureus LGT, while 12 amino acids shorter than the E. coli enzyme, had a hydropathic profile and a predicted pI (10.4) similar to those of the E. coli enzyme. Multiple sequence alignment among E. coli, S. typhimurium, H. influenzae, and S. aureus LGT proteins revealed regions of highly conserved amino acid sequences throughout the molecule. Three independent lgt mutant alleles from E. coli SK634, SK635, and SK636 and one lgt allele from S. typhimurium SE5221, all defective in LGT activity at the nonpermissive temperature, were cloned by PCR and sequenced. The mutant alleles were found to contain a single base alteration resulting in the substitution of a conserved amino acid. The longest set of identical amino acids without any gap was H-103-GGLIG-108 in LGT from these four microorganisms. In E. coli lgt mutant SK634, Gly-104 in this region was mutated to Ser, and the mutant organism was temperature sensitive in growth and exhibited low LGT activity in vitro. Diethylpyrocarbonate inactivated the E. coli LGT with a second-order rate constant of 18.6 M-1S-1, and the inactivation of LGT activity was reversed by hydroxylamine at pH 7. The inactivation kinetics were consistent with the modification of a single residue, His or Tyr, essential for LGT activity.

Amino Acid Sequence↗

The Alcaligenes eutrophus protein HoxN mediates nickel transport in Escherichia coli.

HoxN, an integral membrane protein with seven transmembrane helices and a molecular mass of 33.1 kDa, is involved in high-affinity nickel transport in Alcaligenes eutrophus H16. From genetic analyses, it has been concluded that HoxN is a single-component ion carrier. To investigate this assumption, hoxN was introduced into Escherichia coli. The recombinant strain showed significantly enhanced nickel uptake in a short-interval assay. Likewise, growth in the presence of 63NiCl2 yielded a more than 15-fold-increased cellular nickel content. The HoxN-based nickel transport activity could also be demonstrated in a physiological assay: an E. coli strain coexpressing hoxN and the urease operon of Klebsiella aerogenes exhibited urease activity 10-fold greater than that in the strain lacking a functional hoxN. These results strongly suggest that HoxN is sufficient to operate as a nickel permease. Multiple sequence alignment of HoxN and four other bacterial membrane proteins implicated in nickel metabolism revealed two conserved signatures which may play a role in the nickel translocation process.

Alcaligenes↗

Purification, characterization, and sequence analysis of 2-aminomuconic 6-semialdehyde dehydrogenase from Pseudomonas pseudoalcaligenes JS45.

2-Aminonumconic 6-semialdehyde is an unstable intermediate in the biodegradation of nitrobenzene and 2-aminophenol by Pseudomonas pseudoalcaligenes JS45. Previous work has shown that enzymes in cell extracts convert 2-aminophenol to 2-aminomuconate in the presence of NAD+. In the present work, 2-aminomuconic semialdehyde dehydrogenase was purified and characterized. The purified enzyme migrates as a single band on sodium dodecyl sulfate-polyacrylamide gel electrophoresis with a molecular mass of 57 kDa. The molecular mass of the native enzyme was estimated to be 160 kDa by gel filtration chromatography. The optimal pH for the enzyme activity was 7.3. The enzyme is able to oxidize several aldehyde analogs, including 2-hydroxymuconic semialdehyde, hexaldehyde, and benzaldehyde. The gene encoding 2-aminomuconic semialdehyde dehydrogenase was identified by matching the deduced N-terminal amino acid sequence of the gene with the first 21 amino acids of the purified protein. Multiple sequence alignment of various semialdehyde dehydrogenase protein sequences indicates that 2-aminomuconic 6-semialdehyde dehydrogenase has a high degree of identity with 2-hydroxymuconic 6-semialdehyde dehydrogenases.

Aldehyde Oxidoreductases↗

PHR1 and PHR2 of Candida albicans encode putative glycosidases required for proper cross-linking of beta-1,3- and beta-1,6-glucans.

PHR1 and PHR2 encode putative glycosylphosphatidylinositol-anchored cell surface proteins of the opportunistic fungal pathogen Candida albicans. These proteins are functionally related, and their expression is modulated in relation to the pH of the ambient environment in vitro and in vivo. Deletion of either gene results in a pH-conditional defect in cell morphology and virulence. Multiple sequence alignments demonstrated a distant relationship between the Phr proteins and beta-galactosidases. Based on this alignment, site-directed mutagenesis of the putative active-site residues of Phr1p and Phr2p was conducted and two conserved glutamate residues were shown to be essential for activity. By taking advantage of the pH-conditional expression of the genes, a temporal analysis of cell wall changes was performed following a shift of the mutants from permissive to nonpermissive pH. The mutations did not grossly affect the amount of polysaccharides in the wall but did alter their distribution. The most immediate alteration to occur was a fivefold increase in the rate of cross-linking between beta-1,6-glycosylated mannoproteins and chitin. This increase was followed shortly thereafter by a decline in beta-1,3-glucan-associated beta-1, 6-glucans and, within several generations, a fivefold increase in the chitin content of the walls. The increased accumulation of chitin-linked glucans was not due to a block in subsequent processing as determined by pulse-chase analysis. Rather, the results suggest that the glucans are diverted to chitin linkage due to the inability of the mutants to establish cross-links between beta-1,6- and beta-1,3-glucans. Based on these and previously published results, it is suggested that the Phr proteins process beta-1,3-glucans and make available acceptor sites for the attachment of beta-1,6-glucans.

Amino Acid Sequence↗

Insertion mutagenesis and membrane topology model of the Pseudomonas aeruginosa outer membrane protein OprM.

Pseudomonas aeruginosa OprM is a protein involved in multiple-antibiotic resistance as the outer membrane component for the MexA-MexB-OprM efflux system. Planar lipid bilayer experiments showed that OprM had channel-forming activity with an average single-channel conductance of only about 80 pS in 1 M KCl. The gene encoding OprM was subjected to insertion mutagenesis by cloning of a foreign epitope from the circumsporozoite form of the malarial parasite Plasmodium falciparum into 11 sites. In Escherichia coli, 8 of the 11 insertion mutant genes expressed proteins at levels comparable to those obtained with the wild-type gene and the inserted malarial epitopes were surface accessible as assessed by indirect immunofluorescence. When moved to a P. aeruginosa OprM-deficient strain, seven of the insertion mutant genes expressed proteins at variable levels comparable to that of wild-type OprM and three of these reconstituted MIC profiles resembling those of the wild-type protein, while the other mutant forms showed variable MIC results. Utilizing the data from these experiments, in conjunction with multiple sequence alignments and structure predictions, an OprM topology model with 16 beta strands was proposed.

Amino Acid Sequence↗

The VirB4 family of proposed traffic nucleoside triphosphatases: common motifs in plasmid RP4 TrbE are essential for conjugation and phage adsorption.

Proteins of the VirB4 family are encoded by conjugative plasmids and by type IV secretion systems, which specify macromolecule export machineries related to conjugation systems. The central feature of VirB4 proteins is a nucleotide binding site. In this study, we asked whether members of the VirB4 protein family have similarities in their primary structures and whether these proteins hydrolyze nucleotides. A multiple-sequence alignment of 19 members of the VirB4 protein family revealed striking overall similarities. We defined four common motifs and one conserved domain. One member of this protein family, TrbE of plasmid RP4, was genetically characterized by site-directed mutagenesis. Most mutations in trbE resulted in complete loss of its activities, which eliminated pilus production, propagation of plasmid-specific phages, and DNA transfer ability in Escherichia coli. Biochemical studies of a soluble derivative of RP4 TrbE and of the full-length homologous protein R388 TrwK revealed that the purified forms of these members of the VirB4 protein family do not hydrolyze ATP or GTP and behave as monomers in solution.

Acid Anhydride Hydrolases↗

The ubiquitous protein domain EAL is a cyclic diguanylate-specific phosphodiesterase: enzymatically active and inactive EAL domains.

The EAL domain (also known as domain of unknown function 2 or DUF2) is a ubiquitous signal transduction protein domain in the Bacteria. Its involvement in hydrolysis of the novel second messenger cyclic dimeric GMP (c-di-GMP) was demonstrated in vivo but not in vitro. The EAL domain-containing protein Dos from Escherichia coli was reported to hydrolyze cyclic AMP (cAMP), implying that EAL domains have different substrate specificities. To investigate the biochemical activity of EAL, the E. coli EAL domain-containing protein YahA and its individual EAL domain were overexpressed, purified, and characterized in vitro. Both full-length YahA and the EAL domain hydrolyzed c-di-GMP into linear dimeric GMP, providing the first biochemical evidence that the EAL domain is sufficient for phosphodiesterase activity. This activity was c-di-GMP specific, optimal at alkaline pH, dependent on Mg(2+) or Mn(2+), strongly inhibited by Ca(2+), and independent of protein oligomerization. Linear dimeric GMP was shown to be 5'pGpG. The EAL domain from Dos was overexpressed, purified, and found to function as a c-di-GMP-specific phosphodiesterase, not as a cAMP-specific phosphodiesterase, in contrast to previous reports. The EAL domains can hydrolyze 5'pGpG into GMP, however, very slowly, thus implying that this activity is irrelevant in vivo. Therefore, c-di-GMP is the exclusive substrate of EAL. Multiple-sequence alignment revealed two groups of EAL domains hypothesized to correspond to enzymatically active and inactive domains. The domains in the latter group have mutations in residues conserved in the active domains. The enzymatic inactivity of EAL domains may explain their coexistence with GGDEF domains in proteins possessing c-di-GMP synthase (diguanulate cyclase) activity.

3',5'-Cyclic-GMP Phosphodiesterases↗

Sinorhizobium meliloti dctA mutants with partial ability to transport dicarboxylic acids.

Sinorhizobium meliloti dctA encodes a transport protein needed for a successful nitrogen-fixing symbiosis between the bacteria and alfalfa. Using the toxicity of the DctA substrate fluoroorotic acid as a selective agent in an iterated selection procedure, four independent S. meliloti dctA mutants were isolated that retained some ability to transport dicarboxylates. Two mutations were located in a region called motif B located in a predicted transmembrane helix of the protein that has been shown in other members of the glutamate transporter family to be involved in cation binding. A G114D mutation was located in the third transmembrane helix, which had not previously been directly implicated in transport. Multiple sequence alignment of more than 60 members of the glutamate transporter family revealed a glycine at this position in nearly all members of the family. The fourth mutant was able to transport succinate at almost wild-type levels but was impaired in malate and fumarate transport. It contains two mutations: one in a periplasmic domain and the other predicted to be in the cytoplasm. Separation of the mutations showed that each contributed to the altered substrate preference. dctA deletion mutants that contain the mutant dctA alleles on a plasmid can proceed further in symbiotic development than null mutants of dctA, but none of the plasmids could support symbiotic nitrogen fixation, although they can transport dicarboxylates, some at relatively high levels.

Amino Acid Sequence↗

Rapid detection of Shiga toxin-producing bacteria in feces by multiplex PCR with molecular beacons on the smart cycler.

We have developed a rapid (1-h) real-time fluorescence-based PCR assay with the Smart Cycler thermal cycler (Cepheid, Sunnyvale, Calif.) for the detection of Shiga toxin-producing Escherichia coli (STEC), as well as other Shiga toxin-producing bacteria. Based on multiple-sequence alignments, we have designed two pairs of PCR primers that efficiently amplify all variants of the Shiga toxin genes stx(1) and stx(2), respectively. These primer pairs were combined for use in a multiplex assay. Two molecular beacons bearing different fluorophores were used as internal probes specific for each amplicon. Assays performed with purified genomic DNA from a variety of STEC strains (n = 23) from diverse geographic locations showed analytical sensitivities of about 10 genome copies per PCR. Non-STEC strains (n = 20) were also tested, and no amplification was observed. The PCR results correlated perfectly with the phenotypic characterization of toxin production in both STEC and non-STEC strains, thereby confirming the specificity of the assay. The assay was validated by testing 38 fecal samples obtained from 27 patients. Of these samples, 26 were PCR positive for stx(1) and/or stx(2). Compared with the culture results, both the sensitivity and the negative predictive value were 100%. The specificity was 92%, and the positive predictive value was 96%. Moreover, this assay detected STEC from a sample in which the STEC concentration was at the limit of detection of the conventional culture methods and from a sample in which STEC was not detected by the conventional culture methods. This real-time PCR assay is simple, rapid, sensitive, and specific and allows detection of all Shiga toxin-producing bacteria directly from fecal samples, irrespective of their serotypes.

DNA, Bacterial↗

Broadly reactive and highly sensitive assay for Norwalk-like viruses based on real-time quantitative reverse transcription-PCR.

We have developed an assay for the detection of Norwalk-like viruses (NLVs) based on reverse transcription-PCR (RT-PCR) that is highly sensitive to a broad range of NLVs. We isolated virus from 71 NLV-positive stool specimens from 37 outbreaks of nonbacterial acute gastroenteritis and sequenced the open reading frame 1 (ORF1)-ORF2 junction region, the most conserved region of the NLV genome. The data were subjected to multiple-sequence alignment analysis and similarity plot analysis. We used the most conserved sequences that react with diverse NLVs to design primers and TaqMan probes for the respective genogroups of NLV, GI and GII, for use in a real-time quantitative RT-PCR assay. Our method detected NLV in 99% (80 of 81) of the stool specimens that were positive by electron microscopy, a better detection rate than with the two available RT-PCR methods. Furthermore, our new method also detected NLV in 20 of 28 stool specimens from the same NLV-related outbreaks that were negative for virus by electron microscopy. Our new assay is free from carryover DNA contamination and detects low copy numbers of NLV RNA. It can be used as a routine assay for diagnosis as well as for elucidation of the epidemiology of NLV infections.

Base Sequence↗

Characterization of the hepatitis C virus NS2/3 processing reaction by using a purified precursor protein.

The NS2-NS3 region of the hepatitis C virus polyprotein encodes a proteolytic activity that is required for processing of the NS2/3 junction. Membrane association of NS2 and the autocatalytic nature of the NS2/3 processing event have so far constituted hurdles to the detailed investigation of this reaction. We now report the first biochemical characterization of the self-processing activity of a purified NS2/3 precursor. Using multiple sequence alignments, we were able to define a minimal domain, devoid of membrane-anchoring sequences, which was still capable of performing the processing reaction. This truncated protein was efficiently expressed and processed in Escherichia coli. The processing reaction could be significantly suppressed by growth in minimal medium in the absence of added zinc ions, leading to the accumulation of an unprocessed precursor protein in inclusion bodies. This protein was purified to homogeneity, refolded, and shown to undergo processing at the authentic NS2/NS3 cleavage site with rates comparable to those observed using an in vitro-translated full-length NS2/3 precursor. Size-exclusion chromatography and a dependence of the processing rate on the concentration of truncated NS2/3 suggested a functional multimerization of the precursor protein. However, we were unable to observe trans cleavage activity between cleavage-site mutants and active-site mutants. Furthermore, the cleavage reaction of the wild-type protein was not inhibited by addition of a mutant that was unable to undergo self-processing. Site-directed mutagenesis data and the independence of the processing rate from the nature of the added metal ion argue in favor of NS2/3 being a cysteine protease having Cys993 and His952 as a catalytic dyad. We conclude that a purified protein can efficiently reproduce processing at the NS2/3 site in the absence of additional cofactors.

Amino Acid Sequence↗

Reversible oxidative modification as a mechanism for regulating retroviral protease dimerization and activation.

Human immunodeficiency virus protease activity can be regulated by reversible oxidation of a sulfur-containing amino acid at the dimer interface. We show here that oxidation of this amino acid in human immunodeficiency virus type 1 protease prevents dimer formation. Moreover, we show that human T-cell leukemia virus type 1 protease can be similarly regulated through reversible glutathionylation of its two conserved cysteine residues. Based on the known three-dimensional structures and multiple sequence alignments of retroviral proteases, it is predicted that the majority of retroviral proteases have sulfur-containing amino acids at the dimer interface. The regulation of protease activity by the modification of a sulfur-containing amino acid at the dimer interface may be a conserved mechanism among the majority of retroviruses.

Amino Acid Sequence↗

In silico pattern-based analysis of the human cytomegalovirus genome.

More than 200 open reading frames (ORFs) from the human cytomegalovirus genome have been reported as potentially coding for proteins. We have used two pattern-based in silico approaches to analyze this set of putative viral genes. With the help of an objective annotation method that is based on the Bio-Dictionary, a comprehensive collection of amino acid patterns that describes the currently known natural sequence space of proteins, we have reannotated all of the previously reported putative genes of the human cytomegalovirus. Also, with the help of MUSCA, a pattern-based multiple sequence alignment algorithm, we have reexamined the original human cytomegalovirus gene family definitions. Our analysis of the genome shows that many of the coded proteins comprise amino acid combinations that are unique to either the human cytomegalovirus or the larger group of herpesviruses. We have confirmed that a surprisingly large portion of the analyzed ORFs encode membrane proteins, and we have discovered a significant number of previously uncharacterized proteins that are predicted to be G-protein-coupled receptor homologues. The analysis also indicates that many of the encoded proteins undergo posttranslational modifications such as hydroxylation, phosphorylation, and glycosylation. ORFs encoding proteins with similar functional behavior appear in neighboring regions of the human cytomegalovirus genome. All of the results of the present study can be found and interactively explored online (http://cbcsrv.watson.ibm.com/virus/).

Algorithms↗