Search PubMedSearch

Biomedical subjects

K Robison

Publications and source records attributed to K Robison.

10 recordsLinked to original sources

A comprehensive library of DNA-binding site matrices for 55 proteins applied to the complete Escherichia coli K-12 genome.

A major mode of gene regulation occurs via the binding of specific proteins to specific DNA sequences. The availability of complete bacterial genome sequences offers an unprecedented opportunity to describe networks of such interactions by correlating existing experimental data with computational predictions. Of the 240 candidate Escherichia coli DNA-binding proteins, about 55 have DNA-binding sites identified by DNA footprinting. We used these sites to construct recognition matrices, which we used to search for additional binding sites in the E. coli genomic sequence. Many of these matrices show a strong preference for non-coding DNA. Discrepancies are identified between matrices derived from natural sites and those derived from SELEX (Systematic Evolution of Ligands by Exponential enrichment) experiments. We have constructed a database of these proteins and binding sites, called DPInteract (available at http://arep.med.harvard.edu/dpinteract).

Bacterial Proteins

Comparing the predicted and observed properties of proteins encoded in the genome of Escherichia coli K-12.

Mining the emerging abundance of microbial genome sequences for hypotheses is an exciting prospect of "functional genomics". At the forefront of this effort, we compared the predictions of the complete Escherichia coli genomic sequence with the observed gene products by assessing 381 proteins for their mature N-termini, in vivo abundances, isoelectric points, molecular masses, and cellular locations. Two-dimensional gel electrophoresis (2-DE) and Edman sequencing were combined to sequence Coomassie-stained 2-DE spots representing the abundant proteins of wild-type E. coli K-12 strains. Greater than 90% of the abundant proteins in the E. coli proteome lie in a small isoelectric point and molecular mass window of 4-7 and 10-100 kDa, respectively. We identified several highly abundant proteins, YjbJ, YjbP, YggX, HdeA, and AhpC, which would not have been predicted from the genomic sequence alone. Of the 223 uniquely identified loci, 60% of the encoded proteins are proteolytically processed. As previously reported, the initiator methionine was efficiently cleaved when the penultimate amino acid was serine or alanine. In contrast, when the penultimate amino acid was threonine, glycine, or proline, cleavage was variable, and valine did not signal cleavage. Although signal peptide cleavage sites tended to follow predicted rules, the length of the putative signal sequence was occassionally greater than the consensus. For proteins predicted to be in the cytoplasm or inner membrane, the N-terminal amino acids were highly constrained compared to proteins localized to the periplasm or outer membrane. Although cytoplasmic proteins follow the N-end rule for protein stability, proteins in the periplasm or outer membrane do not follow this rule; several have N-terminal amino acids predicted to destabilize the proteins. Surprisingly, 18% of the identified 2-DE spots represent isoforms in which protein products of the same gene have different observed pI and M(r), suggesting they are post-translationally processed. Although most of the predicted and observed values for isoelectric point and molecular mass show reasonable concordance, for several proteins the observed values significantly deviate from the expected values. Such discrepancies may represent either highly processed proteins or misinterpretations of the genomic sequence. Our data suggest that AhpC, CspC, and HdeA exist as covalent homomultimers, and that IcdA exists as at least three isoforms even under conditions in which covalent modification is not predicted. We enriched for proteins based on subcellular location and found several proteins in unexpected subcellular locations.

Amino Acid Sequence

Multiplex sequencing of 1.5 Mb of the Mycobacterium leprae genome.

The nucleotide sequence of 1.5 Mb of genomic DNA from Mycobacterium leprae was determined using computer-assisted multiplex sequencing technology. This brings the 2.8-Mb M. leprae genome sequence to approximately 66% completion. The sequences, derived from 43 recombinant cosmids, contain 1046 putative protein-coding genes, 44 repetitive regions, 3 tRNAs, and 15 tRNAs. The gene density of one per 1.4 kb is slightly lower than that of Mycoplasma (1.2 kb). Of the protein coding genes, 44% have significant matches to genes with well-defined functions. Comparison of 1157 M. leprae and 1564 Mycobacterium tuberculosis proteins shows a complex mosaic of homologous genomic blocks with up to 22 adjacent proteins in conserved map order. Matches to known enzymatic, antigenic, membrane, cell wall, cell division, multidrug resistance, and virulence proteins suggest therapeutic and vaccine targets. Unusual features of the M. leprae genome include large polyketide synthase (pks) operons, inteins, and highly fragmented pseudogenes.

Amino Acid Sequence

The tigA gene is a transcriptional fusion of glycolytic genes encoding triose-phosphate isomerase and glyceraldehyde-3-phosphate dehydrogenase in oomycota.

Genes encoding triose-phosphate isomerase (TPI) and glyceraldehyde-3-phosphate dehydrogenase (GAPDH) are fused and form a single transcriptional unit (tigA) in Phytophthora species, members of the order Pythiales in the phylum Oomycota. This is the first demonstration of glycolytic gene fusion in eukaryotes and the first case of a TPI-GAPDH fusion in any organism. The tigA gene from Phytophthora infestans has a typical Oomycota transcriptional start point consensus sequence and, in common with most Phytophthora genes, has no introns. Furthermore, Southern and PCR analyses suggest that the same organization exists in other closely related genera, such as Pythium, from the same order (Oomycota), as well as more distantly related genera, Saprolegnia and Achlya, in the order Saprolegniales. Evidence is provided that in P. infestans, there is at least one other discrete copy of a GAPDH-encoding gene but not of a TPI-encoding gene. Finally, a phylogenetic analysis of TPI does not place Phytophthora within the assemblage of crown eukaryotes and suggests TPI may not be particularly useful for resolving relationships among major eukaryotic groups.

Amino Acid Sequence

Novel Gq alpha isoform is a candidate transducer of rhodopsin signaling in a Drosophila testes-autonomous pacemaker.

DGq is the alpha subunit of the heterotrimeric GTPase (G alpha), which couples rhodopsin to phospholipase C in Drosophila vision. We have uncovered three duplicated exons in dgq by scanning the GenBank data base for unrecognized coding sequences. These alternative exons encode sites involved in GTPase activity and G beta-binding, NorpA (phospholipase C)-binding, and rhodopsin-binding. We examined the in vivo splicing of dgq in adult flies and find that, in all but the male gonads, only two isoforms are expressed. One, dgqA, is the original visual isoform and is expressed in eyes, ocelli, brain, and male gonads. The other, dgqB, has the three novel exons and is widely expressed. Remarkably, all three nonvisual B exons are highly similar (82% identity at the amino acid level) to the Gq alpha family consensus, from Caenorhabditis elegans to human, but all three visual A exons are divergent (61% identity). Intriguingly, we have found a third isoform, dgqC, which is specifically and abundantly expressed in male gonads, and shares the divergent rhodopsin-binding exon of dgqA. We suggest that DGqC is a candidate for the light-signal transducer of a testes-autonomous photosensory clock. This proposal is supported by the finding that rhodopsin 2 and arrestin 1, two photoreceptor-cell-specific genes, are also expressed in male gonads.

Adult

Discovery of amphibian Tc1-like transposon families.

We have discovered transposase sequences in the bull frog (Rana catesbeiana) and in the clawed frog (Xenopus laevis), which demonstrates that there are DNA-mediated transposons in Amphibia. The DNA sequences of 11 new Xenopus elements describe two new vertebrate transposon families. Phylogenetic analysis, using these sequences along with previously defined vertebrate and invertebrate elements, reveals at least five families of Tc1-like elements in Vertebrata. Some of these families co-exist in the same genome. Furthermore, the grouping of one of the amphibian transposon families with a branch of the teleost transposons raises the possibility of horizontal transfer.

Amino Acid Sequence

Cloning and sequencing of thiol-specific antioxidant from mammalian brain: alkyl hydroperoxide reductase and thiol-specific antioxidant define a large family of antioxidant enzymes.

A cDNA corresponding to a thiol-specific antioxidant enzyme (TSA) was isolated from a rat brain cDNA library with the use of antibodies to bovine TSA. The cDNA clone encoded an open reading frame capable of encoding a 198-residue polypeptide. The rat and yeast TSA proteins show significant sequence homology to the 21-kDa component (AhpC) of Salmonella typhimurium alkyl hydroperoxide reductase, and we have found that AhpC exhibits TSA activity. AhpC and TSA define a family of > 25 different proteins present in organisms from all kingdoms. The similarity among the family members extends over the entire sequence and ranges between 23% and 98% identity. A majority of the members of the AhpC/TSA family contain two conserved cysteines. At least eight of the genes encoding AhpC/TSA-like polypeptides are found in proximity to genes encoding other oxidoreductase activities, and the expression of several of the homologs has been correlated with pathogenicity. We suggest that the AhpC/TSA family represents a widely distributed class of antioxidant enzymes. We also report that a second family of proteins, defined by the 57-kDa component (AhpF) of alkyl hydroperoxide reductase and by thioredoxin reductase, has expanded to include six additional members.

Amino Acid Sequence

Large scale bacterial gene discovery by similarity search.

DNA sequencing efforts frequently uncover genes other than the targeted ones. We have used rapid database scanning methods to search for undescribed eubacterial and archean protein coding frames in regions flanking known genes. By searching all prokaryotic DNA sequences not marked as coding for proteins or stable RNAs against the protein databases, we have identified more than 450 new examples of bacterial proteins, as well as a smaller number of possible revisions to known proteins, at a surprisingly high rate of one new protein or revision for every 24 initial DNA sequences or 8,300 nucleotides examined. Seven proteins are members of families which have not been described in prokaryotic sequences. We also describe 49 re-interpretations of existing sequence data of particular biological significance.

Amino Acid Sequence