Search PubMed⌕ Search

Biomedical subjects

Terence Hwa

Publications and source records attributed to Terence Hwa.

9 recordsLinked to original sources

Distinct changes of genomic biases in nucleotide substitution at the time of Mammalian radiation.

Differences in the regional substitution patterns in the human genome created patterns of large-scale variation of base composition known as genomic isochores. To gain insight into the origin of the genomic isochores, we develop a maximum-likelihood approach to determine the history of substitution patterns in the human genome. This approach utilizes the vast amount of repetitive sequence deposited in the human genome over the past approximately 250 Myr. Using this approach, we estimate the frequencies of seven types of substitutions: the four transversions, two transitions, and the methyl-assisted transition of cytosine in CpG. Comparing substitutional patterns in repetitive elements of various ages, we reconstruct the history of the base-substitutional process in the different isochores for the past 250 Myr. At around 90 MYA (around the time of the mammalian radiation), we find an abrupt fourfold to eightfold increase of the cytosine transition rate in CpG pairs compared with that of the reptilian ancestor. Further analysis of nucleotide substitutions in regions with different GC content reveals concurrent changes in the substitutional patterns. Although the substitutional pattern was dependent on the regional GC content in such ways that it preserved the regional GC content before the mammalian radiation, it lost this dependence afterward. The substitutional pattern changed from an isochore-preserving to an isochore-degrading one. We conclude that isochores have been established before the radiation of the eutherian mammals and have been subject to the process of homogenization since then.

CpG Islands↗

On schemes of combinatorial transcription logic.

Cells receive a wide variety of cellular and environmental signals, which are often processed combinatorially to generate specific genetic responses. Here we explore theoretically the potentials and limitations of combinatorial signal integration at the level of cis-regulatory transcription control. Our analysis suggests that many complex transcription-control functions of the type encountered in higher eukaryotes are already implementable within the much simpler bacterial transcription system. Using a quantitative model of bacterial transcription and invoking only specific protein-DNA interaction and weak glue-like interaction between regulatory proteins, we show explicit schemes to implement regulatory logic functions of increasing complexity by appropriately selecting the strengths and arranging the relative positions of the relevant protein-binding DNA sequences in the cis-regulatory region. The architectures that emerge are naturally modular and evolvable. Our results suggest that the transcription regulatory apparatus is a "programmable" computing machine, belonging formally to the class of Boltzmann machines. Crucial to our results is the ability to regulate gene expression at a distance. In bacteria, this can be achieved for isolated genes via DNA looping controlled by the dimerization of DNA-bound proteins. However, if adopted extensively in the genome, long-distance interaction can cause unintentional intergenic cross talk, a detrimental side effect difficult to overcome by the known bacterial transcription-regulation systems. This may be a key factor limiting the genome-wide adoption of complex transcription control in bacteria. Implications of our findings for combinatorial transcription control in eukaryotes are discussed.

Bacteria↗

Localization of denaturation bubbles in random DNA sequences.

We study the thermodynamic and dynamic behaviors of twist-induced denaturation bubbles in a long, stretched random sequence of DNA. The small bubbles associated with weak twist are delocalized. Above a threshold torque, the bubbles of several tens of bases or larger become preferentially localized to AT-rich segments. In the localized regime, the bubbles exhibit "aging" and move around subdiffusively with continuously varying dynamic exponents. These properties are derived by using results of large-deviation theory together with scaling arguments and are verified by Monte Carlo simulations.

Base Composition↗

Dynamics of competitive evolution on a smooth landscape.

We study competitive DNA sequence evolution directed by in vitro protein binding. The steady-state dynamics of this process is well described by a shape-preserving pulse which decelerates and eventually reaches equilibrium. We explain this dynamical behavior within a continuum mean-field framework. Analytical results obtained on the motion of the pulse agree with simulations. Furthermore, finite population correction to the mean-field results are found to be insignificant.

Binding, Competitive↗

Mechanically probing the folding pathway of single RNA molecules.

We study theoretically the denaturation of single RNA molecules by mechanical stretching, focusing on signatures of the (un)folding pathway in molecular fluctuations. Our model describes the interactions between nucleotides by incorporating the experimentally determined free energy rules for RNA secondary structure, whereas exterior single-stranded regions are modeled as freely jointed chains. For exemplary RNA sequences (hairpins and the Tetrahymena thermophila group I intron), we compute the quasiequilibrium fluctuations in the end-to-end distance as the molecule is unfolded by pulling on opposite ends. Unlike the average quasiequilibrium force-extension curves, these fluctuations reveal clear signatures from the unfolding of individual structural elements. We find that the resolution of these signatures depends on the spring constant of the force-measuring device, with an optimal value intermediate between very rigid and very soft. We compare and relate our results to recent experiments by Liphardt et al. (2001).

Animals↗

DNA sequence evolution with neighbor-dependent mutation.

We introduce a model of DNA sequence evolution which can account for biases in mutation rates that depend on the identity of the neighboring bases. An analytic solution for this class of models is developed by adopting well-known methods of nonlinear dynamics. Results are presented for the CpG-methylation-deamination process, which dominates point substitutions in vertebrates. The dinucleotide frequencies generated by the model (using empirically obtained mutation rates) match the overall pattern observed in noncoding DNA. A web-based tool has been constructed to compute single- and dinucleotide frequencies for arbitrary neighbor-dependent mutation rates. Also provided is the backward procedure to infer the mutation rates using maximum likelihood analysis given the observed single- and dinucleotide frequencies. Reasonable estimates of the mutation rates can be obtained very efficiently, using generic noncoding DNA sequences as input, after masking out long homonucleotide subsequences. Our method is much more convenient and versatile to use than the traditional method of deducing mutation rates by counting mutation events in carefully chosen sequences. More generally, our approach provides a more realistic but still tractable description of noncoding genomic DNA and may be used as a null model for various sequence analysis applications.

DNA↗

Physical constraints and functional characteristics of transcription factor-DNA interaction.

We study theoretical "design principles" for transcription factor (TF)-DNA interaction in bacteria, focusing particularly on the statistical interaction of the TFs with the genomic background (i.e., the genome without the target sites). We introduce and motivate the concept of programmability, i.e., the ability to set the threshold concentration for TF binding over a wide range merely by mutating the binding sequence of a target site. This functional demand, together with physical constraints arising from the thermodynamics and kinetics of TF-DNA interaction, leads us to a narrow range of "optimal" interaction parameters. We find that this parameter set agrees well with experimental data for the interaction parameters of a few exemplary prokaryotic TFs, which indicates that TF-DNA interaction is indeed programmable. We suggest further experiments to test whether this is a general feature for a large class of TFs.

Bacterial Proteins↗

On the selection and evolution of regulatory DNA motifs.

The mutation and selection of regulatory DNA sequences are presented as an ideal model system of molecular evolution where genotype, phenotype, and fitness can be explicitly and independently characterized. In this theoretical study, we construct an explicit model for the evolution of regulatory sequences, making use of the known biophysics of the binding of regulatory proteins to DNA sequences, under the assumption that fitness of a sequence depends only on its binding affinity to the regulatory protein. The model is confined to the mean field (i.e., infinite population size) limit. Using realistic values for all parameters, we determine the minimum fitness advantage needed to maintain a binding sequence, demonstrating explicitly the "error threshold" below which a binding sequence cannot survive the accumulated effect of mutation over long time. The commonly observed "fuzziness" in binding motifs arises naturally as a consequence of the balance between selection and mutation in our model. In addition, we devise a simple model for the evolution of multiple binding sequences in a given regulatory region. We find the number of evolutionarily stable binding sequences to increase in a step-like fashion with increasing fitness advantage, if multiple regulatory proteins can synergistically enhance gene transcription. We discuss possible experimental approaches to resolve open questions raised by our study.

Binding Sites↗

Hybrid alignment: high-performance with universal statistics.

The score statistics of a recently introduced 'hybrid alignment' algorithm is studied in detail numerically. An extensive survey across the 2216 models of protein domains contained in the Pfam v5.4 database (Bateman et al., Nucleic Acids Res., 28, 263-266, 2000) verifies the theoretical predictions: For the position-specific scoring functions used in the Pfam models, the score statistics of hybrid alignment obey the Gumbel distribution, with the key Gumbel parameter lambda taking on the asymptotic value 1 universally for all models. Thus, the use of hybrid alignment eliminates the time-consuming computer simulations normally needed to assign p-values to alignment scores, freeing the users to experiment with different scoring parameters and functions. The performance of the hybrid algorithm in detecting sequence homology is also studied. For protein sequences from the SCOP database (Murzin et al., J. Mol. Biol., 247, 536-540, 1995) using uniform scoring functions, the performance is found to be comparable to the best of the existing methods. Preliminary results using the PfamA database suggest that the hybrid algorithm achieves similar performance as existing methods for position-specific scoring systems as well. Hybrid alignment is thereby established as a high performance alignment algorithm with well-characterized, universal statistics.

Algorithms↗