Search PubMed⌕ Search

Biomedical subjects

Yijun Ruan

Publications and source records attributed to Yijun Ruan.

12 recordsLinked to original sources

Stereo-cell: Spatial enhanced-resolution single-cell sequencing with high-density DNA nanoball-patterned arrays.

Single-cell sequencing technologies have advanced our understanding of cellular heterogeneity and biological complexity. However, existing methods face limitations in throughput, capture uniformity, cell size flexibility, and technical extensibility. We present Stereo-cell, a spatial enhanced-resolution single-cell sequencing platform based on high-density DNA nanoball (DNB)-patterned arrays, which enables scalable and unbiased cell capture at a wide input range and supports high-fidelity transcriptome profiling. Stereo-cell further allows integration with imaging-based modalities and multiomics strategies, including immunofluorescence and epitope profiling. This platform is also compatible with profiling extracellular vesicles, microstructures, and large cells, whereas its spatial resolution facilitates in situ analysis of cell-cell interactions, cellular microenvironments, and subcellular transcript localization. Together, Stereo-cell provides a flexible framework for expanding single-cell research applications.

Animals↗

The Oct4 and Nanog transcription network regulates pluripotency in mouse embryonic stem cells.

Oct4 and Nanog are transcription factors required to maintain the pluripotency and self-renewal of embryonic stem (ES) cells. Using the chromatin immunoprecipitation paired-end ditags method, we mapped the binding sites of these factors in the mouse ES cell genome. We identified 1,083 and 3,006 high-confidence binding sites for Oct4 and Nanog, respectively. Comparative location analyses indicated that Oct4 and Nanog overlap substantially in their targets, and they are bound to genes in different configurations. Using de novo motif discovery algorithms, we defined the cis-acting elements mediating their respective binding to genomic sites. By integrating RNA interference-mediated depletion of Oct4 and Nanog with microarray expression profiling, we demonstrated that these factors can activate or suppress transcription. We further showed that common core downstream targets are important to keep ES cells from differentiating. The emerging picture is one in which Oct4 and Nanog control a cascade of pathways that are intricately connected to govern pluripotency, self-renewal, genome surveillance and cell fate determination.

Animals↗

A global map of p53 transcription-factor binding sites in the human genome.

The ability to derive a whole-genome map of transcription-factor binding sites (TFBS) is crucial for elucidating gene regulatory networks. Herein, we describe a robust approach that couples chromatin immunoprecipitation (ChIP) with the paired-end ditag (PET) sequencing strategy for unbiased and precise global localization of TFBS. We have applied this strategy to map p53 targets in the human genome. From a saturated sampling of over half a million PET sequences, we characterized 65,572 unique p53 ChIP DNA fragments and established overlapping PET clusters as a readout to define p53 binding loci with remarkable specificity. Based on this information, we refined the consensus p53 binding motif, identified at least 542 binding loci with high confidence, discovered 98 previously unidentified p53 target genes that were implicated in novel aspects of p53 functions, and showed their clinical relevance to p53-dependent tumorigenesis in primary cancer samples.

Binding Sites↗

RNA viral community in human feces: prevalence of plant pathogenic viruses.

The human gut is known to be a reservoir of a wide variety of microbes, including viruses. Many RNA viruses are known to be associated with gastroenteritis; however, the enteric RNA viral community present in healthy humans has not been described. Here, we present a comparative metagenomic analysis of the RNA viruses found in three fecal samples from two healthy human individuals. For this study, uncultured viruses were concentrated by tangential flow filtration, and viral RNA was extracted and cloned into shotgun viral cDNA libraries for sequencing analysis. The vast majority of the 36,769 viral sequences obtained were similar to plant pathogenic RNA viruses. The most abundant fecal virus in this study was pepper mild mottle virus (PMMV), which was found in high concentrations--up to 10(9) virions per gram of dry weight fecal matter. PMMV was also detected in 12 (66.7%) of 18 fecal samples collected from healthy individuals on two continents, indicating that this plant virus is prevalent in the human population. A number of pepper-based foods tested positive for PMMV, suggesting dietary origins for this virus. Intriguingly, the fecal PMMV was infectious to host plants, suggesting that humans might act as a vehicle for the dissemination of certain plant viruses.

Adult↗

Transcriptome analysis of zebrafish embryogenesis using microarrays.

Zebrafish (Danio rerio) is a well-recognized model for the study of vertebrate developmental genetics, yet at the same time little is known about the transcriptional events that underlie zebrafish embryogenesis. Here we have employed microarray analysis to study the temporal activity of developmentally regulated genes during zebrafish embryogenesis. Transcriptome analysis at 12 different embryonic time points covering five different developmental stages (maternal, blastula, gastrula, segmentation, and pharyngula) revealed a highly dynamic transcriptional profile. Hierarchical clustering, stage-specific clustering, and algorithms to detect onset and peak of gene expression revealed clearly demarcated transcript clusters with maximum gene activity at distinct developmental stages as well as co-regulated expression of gene groups involved in dedicated functions such as organogenesis. Our study also revealed a previously unidentified cohort of genes that are transcribed prior to the mid-blastula transition, a time point earlier than when the zygotic genome was traditionally thought to become active. Here we provide, for the first time to our knowledge, a comprehensive list of developmentally regulated zebrafish genes and their expression profiles during embryogenesis, including novel information on the temporal expression of several thousand previously uncharacterized genes. The expression data generated from this study are accessible to all interested scientists from our institute resource database (http://giscompute.gis.a-star.edu.sg/~govind/zebrafish/data_download.html).

Journal Article↗

SARS transmission pattern in Singapore reassessed by viral sequence variation analysis.

BACKGROUND: Epidemiological investigations of infectious disease are mainly dependent on indirect contact information and only occasionally assisted by characterization of pathogen sequence variation from clinical isolates. Direct sequence analysis of the pathogen, particularly at a population level, is generally thought to be too cumbersome, technically difficult, and expensive. We present here a novel application of mass spectrometry (MS)-based technology in characterizing viral sequence variations that overcomes these problems, and we apply it retrospectively to the severe acute respiratory syndrome (SARS) outbreak in Singapore. METHODS AND FINDINGS: The success rate of the MS-based analysis for detecting SARS coronavirus (SARS-CoV) sequence variations was determined to be 95% with 75 copies of viral RNA per reaction, which is sufficient to directly analyze both clinical and cultured samples. Analysis of 13 SARS-CoV isolates from the different stages of the Singapore outbreak identified nine sequence variations that could define the molecular relationship between them and pointed to a new, previously unidentified, primary route of introduction of SARS-CoV into the Singapore population. Our direct determination of viral sequence variation from a clinical sample also clarified an unresolved epidemiological link regarding the acquisition of SARS in a German patient. We were also able to detect heterogeneous viral sequences in primary lung tissues, suggesting a possible coevolution of quasispecies of virus within a single host. CONCLUSION: This study has further demonstrated the importance of improving clinical and epidemiological studies of pathogen transmission through the use of genetic analysis and has revealed the MS-based analysis to be a sensitive and accurate method for characterizing SARS-CoV genetic variations in clinical samples. We suggest that this approach should be used routinely during outbreaks of a wide variety of agents, in order to allow the most effective control.

DNA, Viral↗

Gene identification signature (GIS) analysis for transcriptome characterization and genome annotation.

We have developed a DNA tag sequencing and mapping strategy called gene identification signature (GIS) analysis, in which 5' and 3' signatures of full-length cDNAs are accurately extracted into paired-end ditags (PETs) that are concatenated for efficient sequencing and mapped to genome sequences to demarcate the transcription boundaries of every gene. GIS analysis is potentially 30-fold more efficient than standard cDNA sequencing approaches for transcriptome characterization. We demonstrated this approach with 116,252 PET sequences derived from mouse embryonic stem cells. Initial analysis of this dataset identified hundreds of previously uncharacterized transcripts, including alternative transcripts of known genes. We also uncovered several intergenically spliced and unusual fusion transcripts, one of which was confirmed as a trans-splicing event and was differentially expressed. The concept of paired-end ditagging described here for transcriptome analysis can also be applied to whole-genome analysis of cis-regulatory and other DNA elements and represents an important technological advance for genome annotation.

5' Flanking Region↗

Mutational dynamics of the SARS coronavirus in cell culture and human populations isolated in 2003.

BACKGROUND: The SARS coronavirus is the etiologic agent for the epidemic of the Severe Acute Respiratory Syndrome. The recent emergence of this new pathogen, the careful tracing of its transmission patterns, and the ability to propagate in culture allows the exploration of the mutational dynamics of the SARS-CoV in human populations. METHODS: We sequenced complete SARS-CoV genomes taken from primary human tissues (SIN3408, SIN3725V, SIN3765V), cultured isolates (SIN848, SIN846, SIN842, SIN845, SIN847, SIN849, SIN850, SIN852, SIN3408L), and five consecutive Vero cell passages (SIN2774_P1, SIN2774_P2, SIN2774_P3, SIN2774_P4, SIN2774_P5) arising from SIN2774 isolate. These represented individual patient samples, serial in vitro passages in cell culture, and paired human and cell culture isolates. Employing a refined mutation filtering scheme and constant mutation rate model, the mutation rates were estimated and the possible date of emergence was calculated. Phylogenetic analysis was used to uncover molecular relationships between the isolates. RESULTS: Close examination of whole genome sequence of 54 SARS-CoV isolates identified before 14th October 2003, including 22 from patients in Singapore, revealed the mutations engendered during human-to-Vero and Vero-to-human transmission as well as in multiple Vero cell passages in order to refine our analysis of human-to-human transmission. Though co-infection by different quasipecies in individual tissue samples is observed, the in vitro mutation rate of the SARS-CoV in Vero cell passage is negligible. The in vivo mutation rate, however, is consistent with estimates of other RNA viruses at approximately 5.7 x 10-6 nucleotide substitutions per site per day (0.17 mutations per genome per day), or two mutations per human passage (adjusted R-square = 0.4014). Using the immediate Hotel M contact isolates as roots, we observed that the SARS epidemic has generated four major genetic groups that are geographically associated: two Singapore isolates, one Taiwan isolate, and one North China isolate which appears most closely related to the putative SARS-CoV isolated from a palm civet. Non-synonymous mutations are centered in non-essential ORFs especially in structural and antigenic genes such as the S and M proteins, but these mutations did not distinguish the geographical groupings. However, no non-synonymous mutations were found in the 3CLpro and the polymerase genes. CONCLUSIONS: Our results show that the SARS-CoV is well adapted to growth in culture and did not appear to undergo specific selection in human populations. We further assessed that the putative origin of the SARS epidemic was in late October 2002 which is consistent with a recent estimate using cases from China. The greater sequence divergence in the structural and antigenic proteins and consistent deletions in the 3'--most portion of the viral genome suggest that certain selection pressures are interacting with the functional nature of these validated and putative ORFs.

Animals↗

5' Long serial analysis of gene expression (LongSAGE) and 3' LongSAGE for transcriptome characterization and genome annotation.

Complete genome annotation relies on precise identification of transcription units bounded by a transcription initiation site (TIS) and a polyadenylation site (PAS). To facilitate this process, we developed a set of two complementary methods, 5' Long serial analysis of gene expression (LS) and 3'LS. These analyses are based on the original SAGE and LS methods coupled with full-length cDNA cloning, and enable the high-throughput extraction of the first and the last 20 bp of each transcript. We demonstrate that the mapping of 5'LS and 3'LS tags to the genome allows the localization of TIS and PAS. By using 537 tag pairs mapping to the region of known genes, we confirmed that >90% of the tag pairs appropriately assigned to the first and last exons. Moreover, by using tag sequences as primers for RT-PCRs, we were able to recover putative full-length transcripts in 81% of the attempts. This large-scale generation of transcript terminal tags is at least 20-40 times more efficient than full-length cDNA cloning and sequencing in the identification of complete transcription units. The apparent precision and deep coverage makes 5'LS and 3'LS an advanced approach for genome annotation through whole-transcriptome characterization.

Animals↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗