Search PubMed⌕ Search

Biomedical subjects

G Wesley Hatfield

Publications and source records attributed to G Wesley Hatfield.

16 recordsLinked to original sources

Indirect recognition in sequence-specific DNA binding by Escherichia coli integration host factor: the role of DNA deformation energy.

Integration host factor (IHF) is a bacterial histone-like protein whose primary biological role is to condense the bacterial nucleoid and to constrain DNA supercoils. It does so by binding in a sequence-independent manner throughout the genome. However, unlike other structurally related bacterial histone-like proteins, IHF has evolved a sequence-dependent, high affinity DNA-binding motif. The high affinity binding sites are important for the regulation of a wide range of cellular processes. A remarkable feature of IHF is that it employs an indirect readout mechanism to bind and wrap DNA at both the nonspecific and high affinity (sequence-dependent) DNA sites. In this study we assessed the contributions of pre-formed and protein-induced DNA conformations to the energetics of IHF binding. Binding energies determined experimentally were compared with energies predicted for the IHF-induced deformation of the DNA helix (DNA deformation energy) in the IHF-DNA complex. Combinatorial sets of de novo DNA sequences were designed to systematically evaluate the influence of sequence-dependent structural characteristics of the conserved IHF recognition elements of the consensus DNA sequence. We show that IHF recognizes pre-formed conformational characteristics of the consensus DNA sequence at high affinity sites, whereas at all other sites relative affinity is determined by the deformational energy required for nearest-neighbor base pairs to adopt the DNA structure of the bound DNA-IHF complex.

Amino Acid Motifs↗

HB tag modules for PCR-based gene tagging and tandem affinity purification in Saccharomyces cerevisiae.

We have recently developed the HB tag as a useful tool for tandem-affinity purification under native as well as fully denaturing conditions. The HB tag and its derivatives consist of a hexahistidine tag and a bacterially-derived in vivo biotinylation signal peptide, which support sequential purification by Ni2+ -chelate chromatography and binding to immobilized streptavidin. To facilitate tagging of budding yeast proteins with HB tags, we have created a series of plasmids with various selectable markers. These plasmids allow single-step PCR-based tagging and expression under control of the endogenous promoters or the inducible GAL1 promoter. HB tagging of several budding yeast ORFs demonstrated efficient biotinylation of the HB tag in vivo by endogenous yeast biotin ligases. No adverse effects of the HB tag on protein function were observed. The HB tagging plasmids presented here are related to previously reported epitope-tagging plasmids, allowing PCR-based tagging with the same locus-specific primer sets that are used for other widely used epitope-tagging strategies. The Sequences for the described plasmids were submitted to GenBank under Accession Numbers DQ407918-pFA6a-HBH-kanMX6 DQ407927-pFA6a-RGS18H-kanMX6 DQ407919-pFA6a-HBH-hphMX4 DQ407928-pFA6a-RGS18H-hphMX4 DQ407920-pFA6a-HBH-TRP1 DQ407929-pFA6a-RGS18H-TRP1 DQ407921-pFA6a-HTB-kanMX6 DQ407930-pFA6a-kanMX6-PGAL1-HBH DQ407922-pFA6a-HTB-hphMX4 DQ407931-pFA6a-TRP1-PGAL1-HBH DQ407923-pFA6a-HTB-TRP1 DQ407924-pFA6a-BIO-kanMX6 DQ407925-pFA6a-BIO-hphMX4 DQ407926-pFA6a-BIO-TRP1.

Base Sequence↗

Application of a generalized MWC model for the mathematical simulation of metabolic pathways regulated by allosteric enzymes.

In our effort to elucidate the systems biology of the model organism, Escherichia coli, we have developed a mathematical model that simulates the allosteric regulation for threonine biosynthesis pathway starting from aspartate. To achieve this goal, we used kMech, a Cellerator language extension that describes enzyme mechanisms for the mathematical modeling of metabolic pathways. These mechanisms are converted by Cellerator into ordinary differential equations (ODEs) solvable by Mathematica. In this paper, we describe a more flexible model in Cellerator, which generalizes the Monod, Wyman, Changeux (MWC) model for enzyme allosteric regulation to allow for multiple substrate, activator and inhibitor binding sites. Furthermore, we have developed a model that describes the behavior of the bifunctional allosteric enzyme aspartate kinase I-homoserine dehydrogenase I (AKI-HDHI). This model predicts the partition of enzyme activities in the steady state which paves the way for a more generalized prediction of the behavior of bifunctional enzymes.

Algorithms↗

Global gene expression profiling in Escherichia coli K12: effects of oxygen availability and ArcA.

The ArcAB two-component system of Escherichia coli regulates the aerobic/anaerobic expression of genes that encode respiratory proteins whose synthesis is coordinated during aerobic/anaerobic cell growth. A genomic study of E. coli was undertaken to identify other potential targets of oxygen and ArcA regulation. A group of 175 genes generated from this study and our previous study on oxygen regulation (Salmon, K., Hung, S. P., Mekjian, K., Baldi, P., Hatfield, G. W., and Gunsalus, R. P. (2003) J. Biol. Chem. 278, 29837-29855), called our gold standard gene set, have p values <0.00013 and a posterior probability of differential expression value of 0.99. These 175 genes clustered into eight expression patterns and represent genes involved in a large number of cell processes, including small molecule biosynthesis, macromolecular synthesis, and aerobic/anaerobic respiration and fermentation. In addition, 119 of these 175 genes were also identified in our previous study of the fnr allele. A MEME/weight matrix method was used to identify a new putative ArcA-binding site for all genes of the E. coli genome. 16 new sites were identified upstream of genes in our gold standard set. The strict statistical analyses that we have performed on our data allow us to predict that 1139 genes in the E. coli genome are regulated either directly or indirectly by the ArcA protein with a 99% confidence level.

Alleles↗

A mathematical model for the branched chain amino acid biosynthetic pathways of Escherichia coli K12.

As a first step toward the elucidation of the systems biology of the model organism Escherichia coli, it was our goal to mathematically model a metabolic system of intermediate complexity, namely the well studied end product-regulated pathways for the biosynthesis of the branched chain amino acids L-isoleucine, L-valine, and L-leucine. This has been accomplished with the use of kMech (Yang, C.-R., Shapiro, B. E., Mjolsness, E. D., and Hatfield, G. W. (2005) Bioinformatics 21, in press), a Cellerator (Shapiro, B. E., Levchenko, A., Meyerowitz, E. M., Wold, B. J., and Mjolsness, E. D. (2003) Bioinformatics 19, 677-678) language extension that describes a suite of enzyme reaction mechanisms. Each enzyme mechanism is parsed by kMech into a set of fundamental association-dissociation reactions that are translated by Cellerator into ordinary differential equations. These ordinary differential equations are numerically solved by Mathematica. Any metabolic pathway can be simulated by stringing together appropriate kMech models and providing the physical and kinetic parameters for each enzyme in the pathway. Writing differential equations is not required. The mathematical model of branched chain amino acid biosynthesis in E. coli K12 presented here incorporates all of the forward and reverse enzyme reactions and regulatory circuits of the branched chain amino acid biosynthetic pathways, including single and multiple substrate (Ping Pong and Bi Bi) enzyme kinetic reactions, feedback inhibition (allosteric, competitive, and non-competitive) mechanisms, the channeling of metabolic flow through isozymes, the channeling of metabolic flow via transamination reactions, and active transport mechanisms. This model simulates the results of experimental measurements.

Acetolactate Synthase↗

Application of a generalized MWC model for the mathematical simulation of metabolic pathways regulated by allosteric enzymes.

In our effort to elucidate the systems biology of the model organism, Escherichia coli, we have developed a mathematical model that simulates the allosteric regulation for threonine biosynthesis pathway starting from aspartate. To achieve this goal, we used kMech, a Cellerator language extension that describes enzyme mechanisms for the mathematical modeling of metabolic pathways. These mechanisms are converted by Cellerator into ordinary differential equations (ODEs) solvable by Mathematica. In this paper, we describe a more flexible model in Cellerator, which generalizes the Monod, Wyman, Changeux (MWC) model for enzyme allosteric regulation to allow for multiple substrate, activator and inhibitor binding sites. Furthermore, we have developed a model that describes the behavior of the bifunctional allosteric enzyme aspartate Kinase I-Homoserine Dehydrogenase I (AKI-HDHI). This model predicts the partition of enzyme activities in the steady state which paves a way for a more generalized prediction of the behavior of bifunctional enzymes.

Algorithms↗

Identification of hair cycle-associated genes from time-course gene expression profile data by using replicate variance.

The hair-growth cycle is an example of a cyclic process that is well characterized morphologically but understood incompletely at the molecular level. As an initial step in discovering regulators in hair-follicle morphogenesis and cycling, we used DNA microarrays to profile mRNA expression in mouse back skin from eight representative time points. We developed a statistical algorithm to identify the set of genes expressed within skin that are associated specifically with the hair-growth cycle. The methodology takes advantage of higher replicate variance during asynchronous hair cycles in comparison with synchronous cycles. More than one-third of genes with detectable skin expression showed hair-cycle-related changes in expression, suggesting that many more genes may be associated with the hair-growth cycle than have been identified in the literature. By using a probabilistic clustering algorithm for replicated measurements, these genes were grouped into 30 time-course profile clusters, which fall into four major classes. Distinct genetic pathways were characteristic for the different time-course profile clusters, providing insights into the regulation of hair-follicle cycling and suggesting that this approach is useful for identifying hair follicle regulators. In addition to revealing known hair-related genes, we identified genes that were not previously known to be hair cycle-associated and confirmed their temporal and spatial expression patterns during the hair-growth cycle by quantitative real-time PCR and in situ hybridization. The same computational approach should be generally useful for identifying genes associated with cyclic processes from complex tissues.

Algorithms↗

An enzyme mechanism language for the mathematical modeling of metabolic pathways.

MOTIVATION: As a first step toward the elucidation of the systems biology of complex biological systems, it was our goal to mathematically model common enzyme catalytic and regulatory mechanisms that repeatedly appear in biological processes such as signal transduction and metabolic pathways. RESULTS: We describe kMech, a Cellerator language extension that describes a suite of enzyme mechanisms. Each enzyme mechanism is parsed by kMech into a set of fundamental association-dissociation reactions that are translated by Cellerator into ordinary differential equations that are numerically solved by Mathematica. In addition, we present methods that use commonly available kinetic measurements to estimate rate constants required to solve these differential equations.

Algorithms↗

Effects of exercise on gene expression in human peripheral blood mononuclear cells.

Exercise leads to increases in circulating levels of peripheral blood mononuclear cells (PBMCs) and to a simultaneous, seemingly paradoxical increase in both pro- and anti-inflammatory mediators. Whether this is paralleled by changes in gene expression within the circulating population of PBMCs is not fully understood. Fifteen healthy men (18-30 yr old) performed 30 min of constant work rate cycle ergometry (approximately 80% peak O2 uptake). Blood samples were obtained preexercise (Pre), end-exercise (End-Ex), and 60 min into recovery (Recovery), and gene expression was measured using microarray analysis (Affymetrix GeneChips). Significant differential gene expression was defined with a posterior probability of differential expression of 0.99 and a Bayesian P value of 0.005. Significant changes were observed from Pre to End-Ex in 311 genes, from End-Ex to Recovery in 552 genes, and from Pre to Recovery in 293 genes. Pre to End-Ex upregulation of PBMC genes related to stress and inflammation [e.g., heat shock protein 70 (3.70-fold) and dual-specificity phosphatase-1 (4.45-fold)] was followed by a return of these genes to baseline by Recovery. The gene for interleukin-1 receptor antagonist (an anti-inflammatory mediator) increased between End-Ex and Recovery (1.52-fold). Chemokine genes associated with inflammatory diseases [macrophage inflammatory protein-1alpha (1.84-fold) and -1beta (2.88-fold), and regulation-on-activation, normal T cell expressed and secreted (1.34-fold)] were upregulated but returned to baseline by Recovery. Exercise also upregulated growth and repair genes such as epiregulin (3.50-fold), platelet-derived growth factor (1.55-fold), and hypoxia-inducible factor-I (2.40-fold). A single bout of heavy exercise substantially alters PBMC gene expression characterized in many cases by a brisk activation and deactivation of genes associated with stress, inflammation, and tissue repair.

Adolescent↗

Use of RT-PCR and DNA microarrays to characterize RNA recovered by non-invasive tape harvesting of normal and inflamed skin.

We describe a non-invasive approach for recovering RNA from the surface of skin via a simple tape stripping procedure that permits a direct quantitative and qualitative assessment of pathologic and physiologic biomarkers. Using semi-quantitative RT-PCR we show that tape-harvested RNA is comparable in quality and utility to RNA recovered by biopsy. It is likely that tape-harvested RNA is derived from epidermal cells residing close to the surface and includes adnexal structures and present data showing that tape and biopsy likely recover different cell populations. We report the successful amplification of tape-harvested RNA for hybridization to DNA microarrays. These experiments showed no significant gene expression level differences between replicate sites on a subject and minimal differences between a male and female subject. We also compared the array generated RNA profiles between normal and 24 h 1% SLS-occluded skin and observed that SLS treatment resulted in statistically significant changes in the expression levels of more than 1,700 genes. These data establish the utility of tape harvesting as a non-invasive method for capturing RNA from human skin and support the hypothesis that tape harvesting is an efficient method for sampling the epidermis and identifying select differentially regulated epidermal biomarkers.

Actins↗

Activation of transcription initiation from a stable RNA promoter by a Fis protein-mediated DNA structural transmission mechanism.

The leuV operon of Escherichia coli encodes three of the four genes for the tRNA1Leu isoacceptors. Transcription from this and other stable RNA promoters is known to be affected by a cis-acting UP element and by Fis protein interactions with the carboxyl-terminal domain of the alpha-subunits of RNA polymerase. In this report, we suggest that transcription from the leuV promoter also is activated by a Fis-mediated, DNA supercoiling-dependent mechanism similar to the IHF-mediated mechanism described previously for the ilvP(G) promoter (S. D. Sheridan et al., 1998, J Biol Chem 273: 21298-21308). We present evidence that Fis binding results in the translocation of superhelical energy from the promoter-distal portion of a supercoiling-induced DNA duplex destabilized (SIDD) region to the promoter-proximal portion of the leuV promoter that is unwound within the open complex. A mutant Fis protein, which is defective in contacting the carboxyl-terminal domain of the alpha-subunits of RNA polymerase, remains competent for stimulating open complex formation, suggesting that this DNA supercoiling-dependent component of Fis-mediated activation occurs in the absence of specific protein interactions between Fis and RNA polymerase. Fis-mediated translocation of superhelical energy from upstream binding sites to the promoter region may be a general feature of Fis-mediated activation of transcription at stable RNA promoters, which often contain A+T-rich upstream sequences.

Base Sequence↗

Global gene expression profiling in Escherichia coli K12. The effects of oxygen availability and FNR.

The work presented here is a first step toward a long term goal of systems biology, the complete elucidation of the gene regulatory networks of a living organism. To this end, we have employed DNA microarray technology to identify genes involved in the regulatory networks that facilitate the transition of Escherichia coli cells from an aerobic to an anaerobic growth state. We also report the identification of a subset of these genes that are regulated by a global regulatory protein for anaerobic metabolism, FNR. Analysis of these data demonstrated that the expression of over one-third of the genes expressed during growth under aerobic conditions are altered when E. coli cells transition to an anaerobic growth state, and that the expression of 712 (49%) of these genes are either directly or indirectly modulated by FNR. The results presented here also suggest interactions between the FNR and the leucine-responsive regulatory protein (Lrp) regulatory networks. Because computational methods to analyze and interpret high dimensional DNA microarray data are still at an early stage, and because basic issues of data analysis are still being sorted out, much of the emphasis of this work is directed toward the development of methods to identify differentially expressed genes with a high level of confidence. In particular, we describe an approach for identifying gene expression patterns (clusters) obtained from multiple perturbation experiments based on a subset of genes that exhibit high probability for differential expression values.

Cell Division↗

Differential analysis of DNA microarray gene expression data.

Here, we review briefly the sources of experimental and biological variance that affect the interpretation of high-dimensional DNA microarray experiments. We discuss methods using a regularized t-test based on a Bayesian statistical framework that allow the identification of differentially regulated genes with a higher level of confidence than a simple t-test when only a few experimental replicates are available. We also describe a computational method for calculating the global false-positive and false-negative levels inherent in a DNA microarray data set. This method provides a probability of differential expression for each gene based on experiment-wide false-positive and -negative levels driven by experimental error and biological variance.

Analysis of Variance↗

Global gene expression profiling in Escherichia coli K12. The effects of leucine-responsive regulatory protein.

Leucine-responsive regulatory protein (Lrp) is a global regulatory protein that affects the expression of multiple genes and operons in bacteria. Although the physiological purpose of Lrp-mediated gene regulation remains unclear, it has been suggested that it functions to coordinate cellular metabolism with the nutritional state of the environment. The results of gene expression profiles between otherwise isogenic lrp(+) and lrp(-) strains of Escherichia coli support this suggestion. The newly discovered Lrp-regulated genes reported here are involved either in small molecule or macromolecule synthesis or degradation, or in small molecule transport and environmental stress responses. Although many of these regulatory effects are direct, others are indirect consequences of Lrp-mediated changes in the expression levels of other global regulatory proteins. Because computational methods to analyze and interpret high dimensional DNA microarray data are still an early stage, much of the emphasis of this work is directed toward the development of methods to identify differentially expressed genes with a high level of confidence. In particular, we describe a Bayesian statistical framework for a posterior estimate of the standard deviation of gene measurements based on a limited number of replications. We also describe an algorithm to compute a posterior estimate of differential expression for each gene based on the experiment-wide global false positive and false negative level for a DNA microarray data set. This allows the experimenter to compute posterior probabilities of differential expression for each individual differential gene expression measurement.

Base Sequence↗

DNA topology-mediated control of global gene expression in Escherichia coli.

Because the level of DNA superhelicity varies with the cellular energy charge, it can change rapidly in response to a wide variety of altered nutritional and environmental conditions. This is a global alteration, affecting the entire chromosome and the expression levels of all operons whose promoters are sensitive to superhelicity. In this way, the global pattern of gene expression may be dynamically tuned to changing needs of the cell under a wide variety of circumstances. In this article, we propose a model in which chromosomal superhelicity serves as a global regulator of gene expression in Escherichia coli, tuning expression patterns across multiple operons, regulons, and stimulons to suit the growth state of the cell. This model is illustrated by the DNA supercoiling-dependent mechanisms that coordinate basal expression levels of operons of the ilv regulon both with one another and with cellular growth conditions.

Base Sequence↗

The role of DNA deformation energy at individual base steps for the identification of DNA-protein binding sites.

We examine the use of deformation propensity at individual base steps for the identification of DNA-protein binding sites. We have previously demonstrated that estimates of the total energy to bend DNA to its bound conformation can partially explain indirect DNA-protein interactions. We now show that the deformation propensities at each base step are not equally informative for classifying a sequence as a binding site, and that applying non-uniform weights to the contribution of each base step to aggregate deformation propensity can greatly improve classification accuracy. We show that a perceptron can be trained to use the deformation propensity at each step in a sequence to generate such weights.

Binding Sites↗