Search PubMed⌕ Search

Biomedical subjects

Runsheng Chen

Publications and source records attributed to Runsheng Chen.

At least 19 recordsLinked to original sources

Profiling the long noncoding RNA interaction network in the regulatory elements of target genes by chromatin in situ reverse transcription sequencing.

Long noncoding RNAs (lncRNAs) can regulate the activity of target genes by participating in the organization of chromatin architecture. We have devised a "chromatin-RNA in situ reverse transcription sequencing" (CRIST-seq) approach to profile the lncRNA interaction network in gene regulatory elements by combining the simplicity of RNA biotin labeling with the specificity of the CRISPR/Cas9 system. Using gene-specific gRNAs, we describe a pluripotency-specific lncRNA interacting network in the promoters of Sox2 and Pou5f1, two critical stem cell factors that are required for the maintenance of pluripotency. The promoter-interacting lncRNAs were specifically activated during reprogramming into pluripotency. Knockdown of these lncRNAs caused the stem cells to exit from pluripotency. In contrast, overexpression of the pluripotency-associated lncRNA activated the promoters of core stem cell factor genes and enhanced fibroblast reprogramming into pluripotency. These CRIST-seq data suggest that the Sox2 and Pou5f1 promoters are organized within a unique lncRNA interaction network that determines the fate of pluripotency during reprogramming. This CRIST approach may be broadly used to map lncRNA interaction networks at target loci across the genome.

Animals↗

Prediction of structured non-coding RNAs in the genomes of the nematodes Caenorhabditis elegans and Caenorhabditis briggsae.

We present a survey for non-coding RNAs and other structured RNA motifs in the genomes of Caenorhabditis elegans and Caenorhabditis briggsae using the RNAz program. This approach explicitly evaluates comparative sequence information to detect stabilizing selection acting on RNA secondary structure. We detect 3,672 structured RNA motifs, of which only 678 are known non-translated RNAs (ncRNAs) or clear homologs of known C. elegans ncRNAs. Most of these signals are located in introns or at a distance from known protein-coding genes. With an estimated false positive rate of about 50% and a sensitivity on the order of 50%, we estimate that the nematode genomes contain between 3,000 and 4,000 RNAs with evolutionary conserved secondary structures. Only a small fraction of these belongs to the known RNA classes, including tRNAs, snoRNAs, snRNAs, or microRNAs. A relatively small class of ncRNA candidates is associated with previously observed RNA-specific upstream elements.

Animals↗

Dynamic changes in subgraph preference profiles of crucial transcription factors.

Transcription factors with a large number of target genes--transcription hub(s), or THub(s)--are usually crucial components of the regulatory system of a cell, and the different patterns through which they transfer the transcriptional signal to downstream cascades are of great interest. By profiling normalized abundances (A(N)) of basic regulatory patterns of individual THubs in the yeast Saccharomyces cerevisiae transcriptional regulation network under five different cellular states and environmental conditions, we have investigated their preferences for different basic regulatory patterns. Subgraph-normalized abundances downstream of individual THubs often differ significantly from that of the network as a whole, and conversely, certain over-represented subgraphs are not preferred by any THub. The THub preferences changed substantially when the cellular or environmental conditions changed. This switching of regulatory pattern preferences suggests that a change in conditions does not only elicit a change in response by the regulatory network, but also a change in the mechanisms by which the response is mediated. The THub subgraph preference profile thus provides a novel tool for description of the structure and organization between the large-scale exponents and local regulatory patterns.

Cell Cycle↗

Phylophenetic properties of metabolic pathway topologies as revealed by global analysis.

BACKGROUND: As phenotypic features derived from heritable characters, the topologies of metabolic pathways contain both phylogenetic and phenetic components. In the post-genomic era, it is possible to measure the "phylophenetic" contents of different pathways topologies from a global perspective. RESULTS: We reconstructed phylophenetic trees for all available metabolic pathways based on topological similarities, and compared them to the corresponding 16S rRNA-based trees. Similarity values for each pair of trees ranged from 0.044 to 0.297. Using the quartet method, single pathways trees were merged into a comprehensive tree containing information from a large part of the entire metabolic networks. This tree showed considerably higher similarity (0.386) to the corresponding 16S rRNA-based tree than any tree based on a single pathway, but was, on the other hand, sufficiently distinct to preserve unique phylogenetic information not reflected by the 16S rRNA tree. CONCLUSION: We observed that the topology of different metabolic pathways provided different phylogenetic and phenetic information, depicting the compromise between phylogenetic information and varying evolutionary pressures forming metabolic pathway topologies in different organisms. The phylogenetic information content of the comprehensive tree is substantially higher than that of any tree based on a single pathway, which also gave clues to constraints working on the topology of the global metabolic networks, information that is only partly reflected by the topologies of individual metabolic pathways.

Chromosome Mapping↗

Integrated analysis of multiple data sources reveals modular structure of biological networks.

It has been a challenging task to integrate high-throughput data into investigations of the systematic and dynamic organization of biological networks. Here, we presented a simple hierarchical clustering algorithm that goes a long way to achieve this aim. Our method effectively reveals the modular structure of the yeast protein-protein interaction network and distinguishes protein complexes from functional modules by integrating high-throughput protein-protein interaction data with the added subcellular localization and expression profile data. Furthermore, we take advantage of the detected modules to provide a reliably functional context for the uncharacterized components within modules. On the other hand, the integration of various protein-protein association information makes our method robust to false-positives, especially for derived protein complexes. More importantly, this simple method can be extended naturally to other types of data fusion and provides a framework for the study of more comprehensive properties of the biological network and other forms of complex networks.

Algorithms↗

The DNA sequence, annotation and analysis of human chromosome 3.

After the completion of a draft human genome sequence, the International Human Genome Sequencing Consortium has proceeded to finish and annotate each of the 24 chromosomes comprising the human genome. Here we describe the sequencing and analysis of human chromosome 3, one of the largest human chromosomes. Chromosome 3 comprises just four contigs, one of which currently represents the longest unbroken stretch of finished DNA sequence known so far. The chromosome is remarkable in having the lowest rate of segmental duplication in the genome. It also includes a chemokine receptor gene cluster as well as numerous loci involved in multiple human cancers such as the gene encoding FHIT, which contains the most common constitutive fragile site in the genome, FRA3B. Using genomic sequence from chimpanzee and rhesus macaque, we were able to characterize the breakpoints defining a large pericentric inversion that occurred some time after the split of Homininae from Ponginae, and propose an evolutionary history of the inversion.

Animals↗

A novel scoring schema for peptide identification by searching protein sequence databases using tandem mass spectrometry data.

BACKGROUND: Tandem mass spectrometry (MS/MS) is a powerful tool for protein identification. Although great efforts have been made in scoring the correlation between tandem mass spectra and an amino acid sequence database, improvements could be made in three aspects, including characterization ofpeaks in spectra, adoption of effective scoring functions and access to thereliability of matching between peptides and spectra. RESULTS: A novel scoring function is presented, along with criteria to estimate the performance confidence of the function. Through learning the typesof product ions and the probability of generating them, a hypothetic spectrum was generated for each candidate peptide. Then relative entropy was introduced to measure the similarity between the hypothetic and the observed spectra. Based on the extreme value distribution (EVD) theory, a threshold was chosen to distinguish a true peptide assignment from a random one. Tests on a public MS/MS dataset demonstrated that this method performs better than the well-known SEQUEST. CONCLUSION: A reliable identification of proteins from the spectra promises a more efficient application of tandem mass spectrometry to proteomes with high complexity.

Algorithms↗

Association study with 33 single-nucleotide polymorphisms in 11 candidate genes for hypertension in Chinese.

Essential hypertension is considered to be a typical complex disease with multifactorial etiology, which leads to inconsistent findings in genetic studies. One possibility of failure to replicate some single-locus results is that the underlying genetics of hypertension are not only based on multiple genes with minor effects but also on gene-gene interactions. To test this hypothesis, a case-control study was constructed in Chinese subjects, detecting both single locus and multilocus effects. Eleven candidate genes were selected from biochemical pathways that have been implicated in the development and progression of hypertension, and 33 polymorphisms were evaluated in 503 hypertension patients and 490 age- and gender-matched controls. Single-locus associations, using traditional logistic regression analyses, and multilocus associations, using classification and regression trees and multivariate adaptive regression splines, were both explored in this study. Final models were selected using either Bonferroni correction or cross-validation. Three polymorphisms, TH*rs2070762, ADRB2*Q27E, and GRK4*A486V, were found to be independently associated with essential hypertension in Chinese subjects. In addition to these individual predictors, a potential interaction of CYP11B2-AGTR1 is also involved in the etiology of hypertension. These findings support the multigenic nature of the etiology of essential hypertension and propose a potential gene-gene interactive model for future studies.

Alanine↗

Identifying Hfq-binding small RNA targets in Escherichia coli.

The Hfq-binding small RNAs (sRNAs) have recently drawn much attention as regulators of translation in Escherichia coli. We attempt to identify the targets of this class of sRNAs in genome scale and gain further insight into the complexity of translational regulation induced by Hfq-binding sRNAs. Using a new alignment algorithm, most known negatively regulated targets of Hfq-binding sRNAs were identified. The results also show several interesting aspects of the regulatory function of Hfq-binding sRNAs.

Algorithms↗

Faster and more accurate global protein function assignment from protein interaction networks using the MFGO algorithm.

MOTIVATION: Predicting protein function accurately is an important issue in the post-genomic era. To achieve this goal, several approaches have been proposed deduce the function of unclassified proteins through sequence similarity, co-expression profiles, and other information. Among these methods, the global optimization method (GOM) is an interesting and powerful tool that assigns functions to unclassified proteins based on their positions in a physical interactions network [Vazquez, A., Flammini, A., Maritan, A. and Vespignani, A. (2003) Global protein function prediction from protein-protein interaction networks, Nat. Biotechnol., 21, 697-700]. To boost both the accuracy and speed of GOM, a new prediction method, MFGO (modified and faster global optimization) is presented in this paper, which employs local optimal repetition method to reduce calculation time, and takes account of topological structure information to achieve a more accurate prediction. CONCLUSION: On four proteins interaction datasets, including Vazquez dataset, YP dataset, DIP-core dataset, and SPK dataset, MFGO was tested and compared with the popular MR (majority rule) and GOM methods. Experimental results confirm MFGO's improvement on both speed and accuracy. Especially, MFGO method has a distinctive advantage in accurately predicting functions for proteins with few neighbors. Moreover, the robustness of the approach was validated both in a dataset containing a high percentage of unknown proteins and a disturbed dataset through random insertion and deletion. The analysis shows that a moderate amount of misplaced interactions do not preclude a reliable function assignment.

Algorithms↗

Plasminogen activator inhibitor-1 gene: selection of tagging single nucleotide polymorphisms and association with coronary heart disease.

OBJECTIVE: To explore the effect of plasminogen activator inhibitor-1 (PAI-1) gene variations on the risk of coronary heart disease (CHD) in Chinese Han population. METHODS AND RESULTS: We screened all exons and the promoter region of PAI-1 gene in 48 patients and identified 17 polymorphisms. Five tagging single nucleotide polymorphisms were selected and genotyped in 816 patients with CHD and 937 controls. In the total sample, no main effects of the loci or haplotypes reached statistical significance after adjusting environmental covariates. However, a strongly significant gene-smoking interaction was observed. Among nonsmokers, 2 polymorphisms located at promoter region (rs2227631 and rs1799889) showed significant association with CHD. The cases had higher frequency of rs2227631 A allele and rs1799889 4G allele than the controls (0.42 versus 0.33, P=0.001; 0.60 versus 0.52, P=0.002). Haplotype analyses confirmed the effects of the PAI-1 gene-smoking interaction on CHD risk. Compared with the most common haplotype G-5G-A-A-T (35.1%), the haplotype A-4G-A-A-C (32.7%) significantly increased the risk of CHD with adjusted odds ratio of 1.51 (95% CI, 1.12 to 2.05; P=0.008) in nonsmokers. CONCLUSIONS: This study identified a significant interaction between PAI-1 gene and smoking status. Both single locus and haplotype analyses indicated that rs2227631 A allele and rs1799889 4G allele increased the risk of CHD among nonsmokers in Chinese.

Adult↗

NPInter: the noncoding RNAs and protein related biomacromolecules interaction database.

The noncoding RNAs and protein related biomacromolecules interaction database (NPInter; http://bioinfo.ibp.ac.cn/NPInter or http://www.bioinfo.org.cn/NPInter) is a database that documents experimentally determined functional interactions between noncoding RNAs (ncRNAs) and protein related biomacromolecules (PRMs) (proteins, mRNAs or genomic DNAs). NPInter intends to provide the scientific community with a comprehensive and integrated tool for efficient browsing and extraction of information on interactions between ncRNAs and PRMs. Beyond cataloguing details of these interactions, the NPInter will be useful for understanding ncRNA function, as it adds a very important functional element, ncRNAs, to the biomolecule interaction network and sets up a bridge between the coding and the noncoding kingdoms.

Animals↗

Association of alpha1A adrenergic receptor gene variants on chromosome 8p21 with human stage 2 hypertension.

OBJECTIVE AND DESIGN: We previously reported a significant linkage between human chromosome 8p22 with essential hypertension and systolic blood pressure levels. On the basis of this, we used an efficient age, sex and area-matched case-control scheme to test the association of the polymorphisms in the human alpha1A adrenergic receptor (ADRA1A) gene, located on chromosome 8p21-p11.2, with essential hypertension in a northern Han Chinese population. METHODS: Seven polymorphisms were identified by direct sequencing of genomic DNA derived from 48 randomly recruited hypertensive and 48 healthy subjects. They were also examined for association with essential hypertension in 480 stage 2 hypertensive individuals and their individually matched controls. RESULTS: We observed significantly higher frequencies of the 347Arg allele and 2547G alleles in the cases compared with their controls (P = 0.04 and 0.007, respectively). McNemar's test revealed that carriers of 2547G alleles were at a greater risk of essential hypertension with an odds ratio of 3.00 [95% confidence interval (CI) 1.23-8.35]. We then performed a conditional logistic regression to adjust the effects of conventional risk factors, revealing an odds ratio of 2.84 for carriers of the 2547G allele (95% CI 1.15-6.99). With the haplotypic probabilities estimated using PHASE software, we performed haplotype trend regression analysis, showing a significant association between haplotype 7 and essential hypertension (P = 0.02), after adjustment for conventional risk factors. CONCLUSIONS: Our findings suggest that the genetic variations in the ADRA1A gene are significantly associated with essential hypertension, and may play an important role in the development of essential hypertension in this Chinese population.

Case-Control Studies↗

Organization of the Caenorhabditis elegans small non-coding transcriptome: genomic features, biogenesis, and expression.

Recent evidence points to considerable transcription occurring in non-protein-coding regions of eukaryote genomes. However, their lack of conservation and demonstrated function have created controversy over whether these transcripts are functional. Applying a novel cloning strategy, we have cloned 100 novel and 61 known or predicted Caenorhabditis elegans full-length ncRNAs. Studying the genomic environment and transcriptional characteristics have shown that two-thirds of all ncRNAs, including many intronic snoRNAs, are independently transcribed under the control of ncRNA-specific upstream promoter elements. Furthermore, the transcription levels of at least 60% of the ncRNAs vary with developmental stages. We identified two new classes of ncRNAs, stem-bulge RNAs (sbRNAs) and snRNA-like RNAs (snlRNAs), both featuring distinct internal motifs, secondary structures, upstream elements, and high and developmentally variable expression. Most of the novel ncRNAs are conserved in Caenorhabditis briggsae, but only one homolog was found outside the nematodes. Preliminary estimates indicate that the C. elegans transcriptome contains approximately 2700 small non-coding RNAs, potentially acting as regulatory elements in nematode development.

Animals↗

Genome-wide analysis of mammalian DNA segment fusion/fission.

As a powerful tool for gene function prediction, gene fusion has been widely studied in prokaryotes and certain groups of eukaryotes, but it has been little applied in studies of mammalian genomes. With the first fully sequenced mammalian genomes (human, mouse, rat) now available, we defined and collected a set of fusion/fission event-linked segments (FFLS) based on structured organized genomic alignment. The statistics of the sequence features highlighted the FFLSs against their random context. We found that there are three groups of FFLSs with different component pairs (i.e. gene-gene, gene-noncoding and noncoding-noncoding) in all three mammalian genomes. The proteins encoded by the components of FFLSs in the first group shown a strong tendency to interact with each other. The segmental components in the last two groups which did not contain any protein-coding genes, were found not only to be transcribed to some level, but also more conserved than the random background. Thus, these segments are possibly carrying certain biologically functional elements. We propose that FFLS may be a potential tool for prediction and analysis of function and functional interaction of genetic elements, including both genes and noncoding elements, in mammalian genomes. The full list of the FFLSs in the genomes of the three mammals is available as supporting information at doi:10.1016/j.jtbi.2005.09.016.

Animals↗

Antibody responses to individual proteins of SARS coronavirus and their neutralization activities.

A novel coronavirus, the severe acute respiratory syndrome (SARS) coronavirus (SARS-CoV), was identified as the causative agent of SARS. The profile of specific antibodies to individual proteins of the virus is critical to the development of vaccine and diagnostic tools. In this study, 13 recombinant proteins associated with four structural proteins (S, E, M and N) and five putative uncharacterized proteins (3a, 3b, 6, 7a and 9b) of the SARS-CoV were prepared and used for screening and monitoring their specific IgG antibodies in SARS patient sera by protein microarray. Antibodies to proteins S, 3a, N and 9b were detected in the sera from convalescent-phase SARS patients, whereas those to proteins E, M, 3b, 6 and 7a were undetected. In the detectable specific antibodies, anti-S and anti-N were dominant and could persist in the sera of SARS patients until week 30. Among the rabbit antisera to recombinant proteins S3, N, 3a and 9b, only anti-S3 serum showed significant neutralizing activity to the SARS-CoV infection in Vero E6 cells. The results suggest (1) that anti-S and anti-N antibodies are diagnostic markers and in particular that S3 is immunogenic and therefore is a good candidate as a subunit vaccine antigen; and (2) that, from a virus structure viewpoint, the presence in some human sera of antibodies reacting with two recombinant polypeptides, 3a and 9b, supports the hypothesis that they are synthesized during the virus cycle.

Animals↗

NONCODE: an integrated knowledge database of non-coding RNAs.

NONCODE is an integrated knowledge database dedicated to non-coding RNAs (ncRNAs), that is to say, RNAs that function without being translated into proteins. All ncRNAs in NONCODE were filtered automatically from literature and GenBank, and were later manually curated. The distinctive features of NONCODE are as follows: (i) the ncRNAs in NONCODE include almost all the types of ncRNAs, except transfer RNAs and ribosomal RNAs. (ii) All ncRNA sequences and their related information (e.g. function, cellular role, cellular location, chromosomal information, etc.) in NONCODE have been confirmed manually by consulting relevant literature: more than 80% of the entries are based on experimental data. (iii) Based on the cellular process and function, which a given ncRNA is involved in, we introduced a novel classification system, labeled process function class, to integrate existing classification systems. (iv) In addition, some 1100 ncRNAs have been grouped into nine other classes according to whether they are specific to gender or tissue or associated with tumors and diseases, etc. (v) NONCODE provides a user-friendly interface, a visualization platform and a convenient search option, allowing efficient recovery of sequence, regulatory elements in the flanking sequences, secondary structure, related publications and other information. The first release of NONCODE (v1.0) contains 5339 non-redundant sequences from 861 organisms, including eukaryotes, eubacteria, archaebacteria, virus and viroids. Access is free for all users through a web interface at http://noncode.bioinfo.org.cn.

Base Sequence↗

Expression in Escherichia coli, purification and characterization of Thermoanaerobacter tengcongensis ribosome recycling factor.

A very promising approach to understanding the mechanism of protein thermostability is to investigate the structure-function relationship of homologous proteins with different thermostabilities. Ribosome recycling factor (RRF), which is an essential factor for protein synthesis in bacteria, may be a good candidate for such study. In this report, a ribosome recycling factor from Thermoanaerobacter tengcongensis was expressed and characterized. This protein contains 184 residues, shows 51.4% identity to that of Escherichia coli RRF, and has very strong antigenic cross-reactivity with antibody to E. coli RRF. In vivo activity assay shows that weak residual activity may remain in TteRRF in E. coli cells. Circular dichroism spectral analysis shows that TteRRF has a very similar secondary structure to that of E. coli RRF, implying that they have similar tertiary structures. However, their thermostabilities are significantly different. To find which domain of RRF is mainly responsible for maintaining stability, TteDI/EcoDII and EcoDI/TteDII RRF chimeras were created. Their domain I and domain II are from E. coli and T. tengcongensis RRFs, respectively. The results of GdnHCl and heat induced denaturation of the chimeric RRFs suggest that the domain I plays a major role in maintaining the stability of the RRF molecule.

Amino Acid Sequence↗