Search PubMed⌕ Search

Biomedical subjects

Jun S Liu

Publications and source records attributed to Jun S Liu.

11 recordsLinked to original sources

The sigmaE regulon and the identification of additional sporulation genes in Bacillus subtilis.

We report the identification and characterization on a genome-wide basis of genes under the control of the developmental transcription factor sigma(E) in Bacillus subtilis. The sigma(E) factor governs gene expression in the larger of the two cellular compartments (the mother cell) created by polar division during the developmental process of sporulation. Using transcriptional profiling and bioinformatics we show that 253 genes (organized in 157 operons) appear to be controlled by sigma(E). Among these, 181 genes (organized in 121 operons) had not been previously described as members of this regulon. Promoters for many of the newly identified genes were located by transcription start site mapping. To assess the role of these genes in sporulation, we created null mutations in 98 of the newly identified genes and operons. Of the resulting mutants, 12 (in prkA, ybaN, yhbH, ykvV, ylbJ, ypjB, yqfC, yqfD, ytrH, ytrI, ytvI and yunB) exhibited defects in spore formation. In addition, subcellular localization studies were carried out using in-frame fusions of several of the genes to the coding sequence for GFP. A majority of the fusion proteins localized either to the membrane surrounding the developing spore or to specific layers of the spore coat, although some fusions showed a uniform distribution in the mother cell cytoplasm. Finally, we used comparative genomics to determine that 46 of the sigma(E)-controlled genes in B.subtilis were present in all of the Gram-positive endospore-forming bacteria whose genome has been sequenced, but absent from the genome of the closely related but not endospore-forming bacterium Listeria monocytogenes, thereby defining a core of conserved sporulation genes of probable common ancestral origin. Our findings set the stage for a comprehensive understanding of the contribution of a cell-specific transcription factor to development and morphogenesis.

Bacillus subtilis↗

Identification of co-regulated genes through Bayesian clustering of predicted regulatory binding sites.

The identification of co-regulated genes and their transcription-factor binding sites (TFBS) are key steps toward understanding transcription regulation. In addition to effective laboratory assays, various computational approaches for the detection of TFBS in promoter regions of coexpressed genes have been developed. The availability of complete genome sequences combined with the likelihood that transcription factors and their cognate sites are often conserved during evolution has led to the development of phylogenetic footprinting. The modus operandi of this technique is to search for conserved motifs upstream of orthologous genes from closely related species. The method can identify hundreds of TFBS without prior knowledge of co-regulation or coexpression. Because many of these predicted sites are likely to be bound by the same transcription factor, motifs with similar patterns can be put into clusters so as to infer the sets of co-regulated genes, that is, the regulons. This strategy utilizes only genome sequence information and is complementary to and confirmative of gene expression data generated by microarray experiments. However, the limited data available to characterize individual binding patterns, the variation in motif alignment, motif width, and base conservation, and the lack of knowledge of the number and sizes of regulons make this inference problem difficult. We have developed a Gibbs sampling-based Bayesian motif clustering (BMC) algorithm to address these challenges. Tests on simulated data sets show that BMC produces many fewer errors than hierarchical and K-means clustering methods. The application of BMC to hundreds of predicted gamma-proteobacterial motifs correctly identified many experimentally reported regulons, inferred the existence of previously unreported members of these regulons, and suggested novel regulons.

Algorithms↗

Integrating regulatory motif discovery and genome-wide expression analysis.

We propose motif regressor for discovering sequence motifs upstream of genes that undergo expression changes in a given condition. The method combines the advantages of matrix-based motif finding and oligomer motif-expression regression analysis, resulting in high sensitivity and specificity. motif regressor is particularly effective in discovering expression-mediating motifs of medium to long width with multiple degenerate positions. When applied to Saccharomyces cerevisiae, motif regressor identified the ROX1 and YAP1 motifs from Rox1p and Yap1p overexpression experiments, respectively; predicted that Gcn4p may have increased activity in YAP1 deletion mutants; reported a group of motifs (including GCN4, PHO4, MET4, STRE, USR1, RAP1, M3A, and M3B) that may mediate the transcriptional response to amino acid starvation; and found all of the known cell-cycle regulation motifs from 18 expression microarrays over two cell cycles.

Algorithms↗

Haplotype information and linkage disequilibrium mapping for single nucleotide polymorphisms.

Single nucleotide polymorphisms in the human genome have become an increasingly popular topic in that their analyses promise to be a key step toward personalized medicine. We investigate two related questions, how much the haplotype information contributes to linkage disequilibrium (LD) mapping and whether an in silico haplotype construction preceding the LD analysis can help. For disease gene mapping, using both simulated and real data sets on cystic fibrosis and the Alzheimer disease, we reached the following conclusions: (1) for simple Mendelian diseases, in which case a tractable full statistical model can be developed, the loss of haplotype information for either control or disease data do not have a great impact on LD fine mapping, and haplotype inference should be carried out jointly with LD mapping; (2) for complex diseases, inferring haplotype phases for individuals prior to LD mapping helps achieve a better accuracy. An improved version of the linkage disequilibrium mapping program, BLADE v2, is available at http://www.fas.harvard.edu/junliu/TechRept/03folder/bladev2.tgz.

Algorithms↗

Ran's C-terminal, basic patch, and nucleotide exchange mechanisms in light of a canonical structure for Rab, Rho, Ras, and Ran GTPases.

Proteins comprising the core of the eukaryotic cellular machinery are often highly conserved, presumably due to selective constraints maintaining important structural features. We have developed statistical procedures to decompose these constraints into distinct categories and to pinpoint critical structural features within each category. When applied to P-loop GTPases, this revealed within Rab, Rho, Ras, and Ran a canonical network of molecular interactions centered on bound nucleotide. This network presumably performs a crucial structural and/or mechanistic role considering that it has persisted for more than a billion years after the divergence of these families. We call these 'FY-pivot' GTPases after their most distinguishing feature, a phenylalanine or tyrosine that functions as a pivot within this network. Specific families deviate somewhat from canonical features in interesting ways, presumably reflecting their functional specialization during evolution. We illustrate this here for Ran GTPases, within which two highly conserved histidines, His30 and His139, strikingly diverge from their canonical counterparts. These, along with other residues specifically conserved in Ran, such as Tyr98, Lys99, and Phe138, appear to work in conjunction with FY-pivot canonical residues to facilitate alternative conformations in which these histidines are strategically positioned to couple Ran's basic patch and C-terminal switch to nucleotide exchange and effector binding. Other core components of the cellular machinery are likewise amenable to this approach, which we term Contrast Hierarchical Alignment and Interaction Network (CHAIN) analysis.

Amino Acid Sequence↗

An algorithm for finding protein-DNA binding sites with applications to chromatin-immunoprecipitation microarray experiments.

Chromatin immunoprecipitation followed by cDNA microarray hybridization (ChIP-array) has become a popular procedure for studying genome-wide protein-DNA interactions and transcription regulation. However, it can only map the probable protein-DNA interaction loci within 1-2 kilobases resolution. To pinpoint interaction sites down to the base-pair level, we introduce a computational method, Motif Discovery scan (MDscan), that examines the ChIP-array-selected sequences and searches for DNA sequence motifs representing the protein-DNA interaction sites. MDscan combines the advantages of two widely adopted motif search strategies, word enumeration and position-specific weight matrix updating, and incorporates the ChIP-array ranking information to accelerate searches and enhance their success rates. MDscan correctly identified all the experimentally verified motifs from published ChIP-array experiments in yeast (STE12, GAL4, RAP1, SCB, MCB, MCM1, SFF, and SWI5), and predicted two motif patterns for the differential binding of Rap1 protein in telomere regions. In our studies, the method was faster and more accurate than several established motif-finding algorithms. MDscan can be used to find DNA motifs not only in ChIP-array experiments but also in other experiments in which a subgroup of the sequences can be inferred to contain relatively abundant motif sites. The MDscan web server can be accessed at http://BioProspector.stanford.edu/MDscan/.

Algorithms↗

Methylation of histone H3 Lys 4 in coding regions of active genes.

Posttranslational modifications of histone tails regulate chromatin structure and transcription. Here we present global analyses of histone acetylation and histone H3 Lys 4 methylation patterns in yeast. We observe a significant correlation between acetylation of histones H3 and H4 in promoter regions and transcriptional activity. In contrast, we find that dimethylation of histone H3 Lys 4 in coding regions correlates with transcriptional activity. The histone methyltransferase Set1 is required to maintain expression of these active, promoter-acetylated, and coding region-methylated genes. Global comparisons reveal that genomic regions deacetylated by the yeast enzymes Rpd3 and Hda1 overlap extensively with Lys 4 hypo- but not hypermethylated regions. In the context of recent studies showing that Lys 4 methylation precludes histone deacetylase recruitment, we conclude that Set1 facilitates transcription, in part, by protecting active coding regions from deacetylation.

Acetylation↗

BALSA: Bayesian algorithm for local sequence alignment.

The Smith-Waterman algorithm yields a single alignment, which, albeit optimal, can be strongly affected by the choice of the scoring matrix and the gap penalties. Additionally, the scores obtained are dependent upon the lengths of the aligned sequences, requiring a post-analysis conversion. To overcome some of these shortcomings, we developed a Bayesian algorithm for local sequence alignment (BALSA), that takes into account the uncertainty associated with all unknown variables by incorporating in its forward sums a series of scoring matrices, gap parameters and all possible alignments. The algorithm can return both the joint and the marginal optimal alignments, samples of alignments drawn from the posterior distribution and the posterior probabilities of gap penalties and scoring matrices. Furthermore, it automatically adjusts for variations in sequence lengths. BALSA was compared with SSEARCH, to date the best performing dynamic programming algorithm in the detection of structural neighbors. Using the SCOP databases PDB40D-B and PDB90D-B, BALSA detected 19.8 and 41.3% of remote homologs whereas SSEARCH detected 18.4 and 38% at an error rate of 1% errors per query over the databases, respectively.

Algorithms↗

A Bayesian method for classification of images from electron micrographs.

Particle classification is an important component of multivariate statistical analysis methods that has been used extensively to extract information from electron micrographs of single particles. Here we describe a new Bayesian Gibbs sampling algorithm for the classification of such images. This algorithm, which is applied after dimension reduction by correspondence analysis or by principal components analysis, dynamically learns the parameters of the multivariate Gaussian distributions that characterize each class. These distributions describe tilted ellipsoidal clusters that adaptively adjust shape to capture differences in the variances of factors and the correlations of factors within classes. A novel Bayesian procedure to objectively select factors for inclusion in the classification models is a component of this procedure. A comparison of this algorithm with hierarchical ascendant classification of simulated data sets shows improved classification over a broad range of signal-to-noise ratios.

Algorithms↗

Bayesian haplotype inference for multiple linked single-nucleotide polymorphisms.

Haplotypes have gained increasing attention in the mapping of complex-disease genes, because of the abundance of single-nucleotide polymorphisms (SNPs) and the limited power of conventional single-locus analyses. It has been shown that haplotype-inference methods such as Clark's algorithm, the expectation-maximization algorithm, and a coalescence-based iterative-sampling algorithm are fairly effective and economical alternatives to molecular-haplotyping methods. To contend with some weaknesses of the existing algorithms, we propose a new Monte Carlo approach. In particular, we first partition the whole haplotype into smaller segments. Then, we use the Gibbs sampler both to construct the partial haplotypes of each segment and to assemble all the segments together. Our algorithm can accurately and rapidly infer haplotypes for a large number of linked SNPs. By using a wide variety of real and simulated data sets, we demonstrate the advantages of our Bayesian algorithm, and we show that it is robust to the violation of Hardy-Weinberg equilibrium, to the presence of missing data, and to occurrences of recombination hotspots.

Algorithms↗