Search PubMed⌕ Search

Biomedical subjects

Chengxiang Zhai

Publications and source records attributed to Chengxiang Zhai.

4 recordsLinked to original sources

Genome scan for cis-regulatory DNA motifs associated with social behavior in honey bees.

Honey bees (Apis mellifera) undergo an age-related, socially regulated transition from working in the hive to foraging, which is associated with changes in the expression of thousands of genes in the brain. To begin to study the cis-regulatory code underlying this massive social regulation of gene expression, we used the newly sequenced honey bee genome to scan the promoter regions of eight sets of behaviorally related genes differentially expressed in the brain in the context of division of labor among worker bees, for 41 cis-regulatory motifs previously characterized in Drosophila melanogaster. Binding sites for the transcription factors Hairy, GAGA, Adf1, Cf1, Snail, and Dri, known to function in nervous system development, olfactory learning, or hormone binding in Drosophila, were significantly associated with one or more gene sets. The presence of some binding sites also predicted expression patterns for as many as 71% of the genes in some gene sets. These results suggest that there is a robust relationship between cis and social regulation of brain gene expression, especially considering that we studied <15% of all known transcription factors. These results also suggest that transcriptional networks involved in the regulation of development in Drosophila are used to regulate behavioral development in adult honey bees. However, differences in gene regulation between these two processes are suggested by the finding that the promoter regions for the behaviorally related bee genes differed in both motif occurrence and G/C content relative to their Drosophila orthologs.

Animals↗

Enhancing text categorization with semantic-enriched representation and training data augmentation.

OBJECTIVE: Acquiring and representing biomedical knowledge is an increasingly important component of contemporary bioinformatics. A critical step of the process is to identify and retrieve relevant documents among the vast volume of modern biomedical literature efficiently. In the real world, many information retrieval tasks are difficult because of high data dimensionality and the lack of annotated examples to train a retrieval algorithm. Under such a scenario, the performance of information retrieval algorithms is often unsatisfactory, therefore improvements are needed. DESIGN: We studied two approaches that enhance the text categorization performance on sparse and high data dimensionality: (1) semantic-preserving dimension reduction by representing text with semantic-enriched features; and (2) augmenting training data with semi-supervised learning. A probabilistic topic model was applied to extract major semantic topics from a corpus of text of interest. The representation of documents was projected from the high-dimensional vocabulary space onto a semantic topic space with reduced dimensionality. A semi-supervised learning algorithm based on graph theory was applied to identify potential positive training cases, which were further used to augment training data. The effects of data transformation and augmentation on text categorization by support vector machine (SVM) were evaluated. RESULTS AND CONCLUSION: Semantic-enriched data transformation and the pseudo-positive-cases augmented training data enhance the efficiency and performance of text categorization by SVM.

Algorithms↗

Automatically generating gene summaries from biomedical literature.

Biologists often need to find information about genes whose function is not described in the genome databases. Currently they must try to search disparate biomedical literature to locate relevant articles, and spend considerable efforts reading the retrieved articles in order to locate the most relevant knowledge about the gene. We describe our software, the first that automatically generates gene summaries from biomedical literature. We present a two-stage summarization method, which involves first retrieving relevant articles and then extracting the most informative sentences from the retrieved articles to generate a structured gene summary. The generated summary explicitly covers multiple aspects of a gene, such as the sequence information, mutant phenotypes, and molecular interaction with other genes. We propose several heuristic approaches to improve the accuracy in both stages. The proposed methods are evaluated using 10 randomly chosen genes from FlyBase and a subset of Medline abstracts about Drosophila. The results show that the precision of the top selected sentences in the 6 aspects is typically about 50-70%, and the generated summaries are quite informative, indicating that our approaches are effective in automatically summarizing literature information about genes. The generated summaries not only are directly useful to biologists but also serve as useful entry points to enable them to quickly digest the retrieved literature articles.

Animals↗

Automatic annotation of protein motif function with Gene Ontology terms.

BACKGROUND: Conserved protein sequence motifs are short stretches of amino acid sequence patterns that potentially encode the function of proteins. Several sequence pattern searching algorithms and programs exist foridentifying candidate protein motifs at the whole genome level. However, a much needed and important task is to determine the functions of the newly identified protein motifs. The Gene Ontology (GO) project is an endeavor to annotate the function of genes or protein sequences with terms from a dynamic, controlled vocabulary and these annotations serve well as a knowledge base. RESULTS: This paper presents methods to mine the GO knowledge base and use the association between the GO terms assigned to a sequence and the motifs matched by the same sequence as evidence for predicting the functions of novel protein motifs automatically. The task of assigning GO terms to protein motifs is viewed as both a binary classification and information retrieval problem, where PROSITE motifs are used as samples for mode training and functional prediction. The mutual information of a motif and aGO term association is found to be a very useful feature. We take advantage of the known motifs to train a logistic regression classifier, which allows us to combine mutual information with other frequency-based features and obtain a probability of correct association. The trained logistic regression model has intuitively meaningful and logically plausible parameter values, and performs very well empirically according to our evaluation criteria. CONCLUSIONS: In this research, different methods for automatic annotation of protein motifs have been investigated. Empirical result demonstrated that the methods have a great potential for detecting and augmenting information about the functions of newly discovered candidate protein motifs.

Amino Acid Motifs↗