Search PubMed⌕ Search

Biomedical subjects

Jingchu Luo

Publications and source records attributed to Jingchu Luo.

At least 19 recordsLinked to original sources

Efficient targeted anticoagulant with active RGD motif.

Three anticoagulants combining large peptide recombinant hirudin variants (rHV2-K47) and Arg-Gly-Asp (RGD) motif related to platelet aggregation were generated, i.e. sequences CRFPRGDADPYCE and CNPRGDFRCI were added to the C-terminus of hirudin to obtain RGD-hirudin 1 and 2, respectively, and the sequence RGDSE was inserted between residues 53-54 of hirudin to obtain RGD-hirudin 3. All products exhibited antithrombin and antiplatelet activities, especially IC50 of RGD-hirudin 1 and 2 were much lower with respect to previously reported similar peptides. Our data suggested that RGD-hirudin 1 and 2 would be promising anticoagulants in clinic. Moreover, the triangular structure of active RGD was shown by computer simulation, which might contribute to our understanding on integrin-related peptides and proteins.

Amino Acid Sequence↗

Multilocus analysis of nucleotide variation of Oryza sativa and its wild relatives: severe bottleneck during domestication of rice.

Varying degrees of reduction of genetic diversity in crops relative to their wild progenitors occurred during the process of domestication. Such information, however, has not been available for the Asian cultivated rice (Oryza sativa) despite its importance as a staple food and a model organism. To reveal levels and patterns of nucleotide diversity and to elucidate the genetic relationship and demographic history of O. sativa and its close relatives (Oryza rufipogon and Oryza nivara), we investigated nucleotide diversity data from 10 unlinked nuclear loci in species-wide samples of these species. The results indicated that O. rufipogon and O. nivara possessed comparable levels of nucleotide variation ((sil) = 0.0077 approximately 0.0095) compared with the relatives of other crops. In contrast, nucleotide diversity of O. sativa was as low as (sil) = 0.0024 and even lower ((sil) = 0.0021 for indica and 0.0011 for japonica), if we consider the 2 subspecies separately. Overall, only 20-10% of the diversity in the wild species was retained in 2 subspecies of the cultivated rice (indica and japonica), respectively. Because statistic tests did not reject the assumption of neutrality for all 10 loci, we further used coalescent to simulate bottlenecks under various lengths and population sizes to better understand the domestication process. Consistent with the dramatic reduction in nucleotide diversity, we detected a severe domestication bottleneck and demonstrated that the sequence diversity currently found in the rice genome could be explained by a founding population of 1,500 individuals if the initial domestication event occurred over a 3,000-year period. Phylogenetic analyses revealed close genetic relationships and ambiguous species boundary of O. rufipogon and O. nivara, providing additional evidence to treat them as 2 ecotypes of a single species. Lowest linkage disequilibrium (LD) was found in the perennial O. rufipogon where the r(2) value dropped to a negligible level within 400 bp, and the highest in the japonica rice where LD extended to the entirely sequenced region ( approximately 900 bp), implying that LD mapping by genome scans may not be feasible in wild rice due to the high density of markers needed.

Base Sequence↗

Statistical inference of chromosomal homology based on gene colinearity and applications to Arabidopsis and rice.

BACKGROUND: The identification of chromosomal homology will shed light on such mysteries of genome evolution as DNA duplication, rearrangement and loss. Several approaches have been developed to detect chromosomal homology based on gene synteny or colinearity. However, the previously reported implementations lack statistical inferences which are essential to reveal actual homologies. RESULTS: In this study, we present a statistical approach to detect homologous chromosomal segments based on gene colinearity. We implement this approach in a software package ColinearScan to detect putative colinear regions using a dynamic programming algorithm. Statistical models are proposed to estimate proper parameter values and evaluate the significance of putative homologous regions. Statistical inference, high computational efficiency and flexibility of input data type are three key features of our approach. CONCLUSION: We apply ColinearScan to the Arabidopsis and rice genomes to detect duplicated regions within each species and homologous fragments between these two species. We find many more homologous chromosomal segments in the rice genome than previously reported. We also find many small colinear segments between rice and Arabidopsis genomes.

Algorithms↗

KOBAS server: a web-based platform for automated annotation and pathway identification.

There is an increasing need to automatically annotate a set of genes or proteins (from genome sequencing, DNA microarray analysis or protein 2D gel experiments) using controlled vocabularies and identify the pathways involved, especially the statistically enriched pathways. We have previously demonstrated the KEGG Orthology (KO) as an effective alternative controlled vocabulary and developed a standalone KO-Based Annotation System (KOBAS). Here we report a KOBAS server with a friendly web-based user interface and enhanced functionalities. The server can support input by nucleotide or amino acid sequences or by sequence identifiers in popular databases and can annotate the input with KO terms and KEGG pathways by BLAST sequence similarity or directly ID mapping to genes with known annotations. The server can then identify both frequent and statistically enriched pathways, offering the choices of four statistical tests and the option of multiple testing correction. The server also has a 'User Space' in which frequent users may store and manage their data and results online. We demonstrate the usability of the server by finding statistically enriched pathways in a set of upregulated genes in Alzheimer's Disease (AD) hippocampal cornu ammonis 1 (CA1). KOBAS server can be accessed at http://kobas.cbi.pku.edu.cn.

Alzheimer Disease↗

DRTF: a database of rice transcription factors.

SUMMARY: DRTF contains 2025 putative transcription factors (TFs) in Oryza sativa L. ssp. indica and 2384 in ssp. japonica, distributed in 63 families, identified by computational prediction and manual curation. It includes detailed annotations of each TF including sequence features, functional domains, Gene Ontology assignment, chromosomal localization, EST and microarray expression information, as well as multiple sequence alignment of the DNA-binding domains for each TF family. The database can be browsed and searched with a user-friendly web interface. AVAILABILITY: DRTF is available at http://drtf.cbi.pku.edu.cn

Database Management Systems↗

Nucleotide substitution pattern in rice paralogues: implication for negative correlation between the synonymous substitution rate and codon usage bias.

Understanding the correlation between synonymous substitution rate and GC content is essential to decipher the gene evolution. However, it has been controversial on their relationship. We analyzed the GC content and synonymous substitution rate in 1092 paralogues produced by two large-scale duplication events in the rice genome. According to the GC content at the third codon sites (GC3), the paralogues were classified into GC3-rich and GC3-poor genes. By referring to their outgroup sequences, we inferred the last common ancestor of sister paralogues and, consequently, calculated the average synonymous substitution rate for two gene classes. The results suggest that average synonymous substitution rate is lower in GC3-rich genes than that in GC3-poor genes, indicating that the synonymous substitution rate is negatively correlated with GC content in the rice genome. Through characterizing the synonymous nucleotide substitution pattern, we found a strong synonymous nucleotide substitution frequency bias from AT to GC in GC3-rich genes. This indicates possible limitations of commonly used methods developed to estimate the synonymous substitution rate. Their estimates might produce misleading results on correlation between the synonymous substitution rate and GC content.

Base Composition↗

Systematic high-yield production of human secreted proteins in Escherichia coli.

Human secreted proteins play a very important role in signal transduction. In order to study all potential secreted proteins identified from the human genome sequence, systematic production of large amounts of biologically active secreted proteins is a prerequisite. We selected 25 novel genes as a trial case for establishing a reliable expression system to produce active human secreted proteins in Escherichia coli. Expression of proteins with or without signal peptides was examined and compared in E. coli strains. The results indicated that deletion of signal peptides, to a certain extent, can improve the expression of these proteins and their solubilities. More importantly, under expression conditions such as induction temperature, N-terminus fusion peptides need to be optimized in order to express adequate amounts of soluble proteins. These recombinant proteins were characterized as well-folded proteins. This system enables us to rapidly obtain soluble and highly purified human secreted proteins for further functional studies.

Chromosome Mapping↗

GBA server: EST-based digital gene expression profiling.

Expressed Sequence Tag-based gene expression profiling can be used to discover functionally associated genes on a large scale. Currently available web servers and tools focus on finding differentially expressed genes in different samples or tissues rather than finding co-expressed genes. To fill this gap, we have developed a web server that implements the GBA (Guilt-by-Association) co-expression algorithm, which has been successfully used in finding disease-related genes. We have also annotated UniGene clusters with links to several important databases such as GO, KEGG, OMIM, Gene, IPI and HomoloGene. The GBA server can be accessed and downloaded at http://gba.cbi.pku.edu.cn.

Algorithms↗

DATF: a database of Arabidopsis transcription factors.

UNLABELLED: We have probably developed the most comprehensive database of Arabidopsis transcription factors (DATF). The DATF contains known and predicted Arabidopsis transcription factors (1827 genes in 56 families) with the unique information of 1177 cloned sequences and many other features including 3D structure templates, EST expression information, transcription factor binding sites and nuclear location signals. AVAILABILITY: DATF is freely available at http://datf.cbi.pku.edu.cn

Amino Acid Sequence↗

SPD--a web-based secreted protein database.

With the improved secreted protein prediction approach and comprehensive data sources, including Swiss-Prot, TrEMBL, RefSeq, Ensembl and CBI-Gene, we have constructed secretomes of human, mouse and rat, with a total of 18 152 secreted proteins. All the entries are ranked according to the prediction confidence. They were further annotated via a proteome annotation pipeline that we developed. We also set up a secreted protein classification pipeline and classified our predicted secreted proteins into different functional categories. To make the dataset more convincing and comprehensive, nine reference datasets are also integrated, such as the secreted proteins from the Gene Ontology Annotation (GOA) system at the European Bioinformatics Institute, and the vertebrate secreted proteins from Swiss-Prot. All these entries were grouped via a TribeMCL based clustering pipeline. We have constructed a web-based secreted protein database, which has been publicly available at http://spd.cbi.pku.edu.cn. Users can browse the database via a GO assignment or chromosomal-location-based interface. Moreover, text query and sequence similarity search are also provided, and the sequence and annotation data can be downloaded freely from the SPD website.

Animals↗

Identification of a naturally occurring recombinant isolate of Sugarcane mosaic virus causing maize dwarf mosaic disease.

The complete nucleotide sequence of a potyvirus causing severe maize dwarf mosaic disease in Shaanxi province, northwestern China was determined (GenBank accession No. AY569692). The full genome is 9596 nucleotides in length excluding the 3 '-terminal poly (A) sequence. It contains a large open reading frame (ORF) flanked by a 149 nt 5'-untranslated region (UTR) and a 255 nt 3'-UTR. The putative polyprotein encoded by this large ORF comprises of 3063 amino acid residues. Sequence comparisons and phylogenetic analyses showed that this potyvirus is an isolate of Sugarcane mosaic virus (SCMV). The entire sequences shared identities of 89.6-97.6 % and 79.3-93.3% with 9 sequenced SCMV isolates at the nucleotide and deduced amino acid levels, respectively. But it showed much lower identities with Maize dwarf mosaic virus (MDMV), Sorghum mosaic virus (SrMV) and Johnsongrass mosaic virus (JGMV) isolates. The putative coat protein sequence is identical to that of a Chinese maize isolate SCMV-HZ. However, partition comparisons and phylogenetic profile analyses of the viral nucleotide sequences indicated that it is a recombinant isolate of SCMV. The recombination sites are located within the 6K1 and CI coding regions.

3' Untranslated Regions↗

Duplication and DNA segmental loss in the rice genome: implications for diploidization.

* Large-scale duplication events have been recently uncovered in the rice genome, but different interpretations were proposed regarding the extent of the duplications. * Through analysing the 370 Mb genome sequences assembled into 12 chromosomes of Oryza sativa subspecies indica, we detected 10 duplicated blocks on all 12 chromosomes that contained 47% of the total predicted genes. Based on the phylogenetic analysis, we inferred that this was a result of a genome duplication that occurred c. 70 million years ago, supporting the polyploidy origin of the rice genome. In addition, a segmental duplication was also identified involving chromosomes 11 and 12, which occurred c. 5 million years ago. * Following the duplications, there have been large-scale chromosomal rearrangements and deletions. About 30-65% of duplicated genes were lost shortly after the duplications, leading to a rapid diploidization. * Together with other lines of evidence, we propose that polyploidization is still an ongoing process in grasses of polyploidy origins.

Biological Evolution↗

RDfolder: a web server for prediction of RNA secondary structure.

Prediction of RNA secondary structure is important in the functional analysis of RNA molecules. The RDfolder web server described in this paper provides two methods for prediction of RNA secondary structure: random stacking of helical regions and helical regions distribution. The random stacking method predicts secondary structure by Monte Carlo simulations. The method of helical regions distribution predicts secondary structure based on the helices that appear most frequently in the set of structures, which are generated by the random stacking method. The RDfolder web server can be accessed at http://rna.cbi.pku.edu.cn.

Internet↗

PCAS--a precomputed proteome annotation database resource.

BACKGROUND: Many model proteomes or "complete" sets of proteins of given organisms are now publicly available. Much effort has been invested in computational annotation of those "draft" proteomes. Motif or domain based algorithms play a pivotal role in functional classification of proteins. Employing most available computational algorithms, mainly motif or domain recognition algorithms, we set up to develop an online proteome annotation system with integrated proteome annotation data to complement existing resources. RESULTS: We report here the development of PCAS (ProteinCentric Annotation System) as an online resource of pre-computed proteome annotation data. We applied most available motif or domain databases and their analysis methods, including hmmpfam search of HMMs in Pfam, SMART and TIGRFAM, RPS-PSIBLAST search of PSSMs in CDD, pfscan of PROSITE patterns and profiles, as well as PSI-BLAST search of SUPERFAMILY PSSMs. In addition, signal peptide and TM are predicted using SignalP and TMHMM respectively. We mapped SUPERFAMILY and COGs to InterPro, so the motif or domain databases are integrated through InterPro. PCAS displays table summaries of pre-computed data and a graphical presentation of motifs or domains relative to the protein. As of now, PCAS contains human IPI, mouse IPI, and rat IPI, A. thaliana, C. elegans, D. melanogaster, S. cerevisiae, and S. pombe proteome.PCAS is available at http://pak.cbi.pku.edu.cn/proteome/gca.php CONCLUSION: PCAS gives better annotation coverage for model proteomes by employing a wider collection of available algorithms. Besides presenting the most confident annotation data, PCAS also allows customized query so users can inspect statistically less significant boundary information as well. Therefore, besides providing general annotation information, PCAS could be used as a discovery platform. We plan to update PCAS twice a year. We will upgrade PCAS when new proteome annotation algorithms identified.

Algorithms↗

PepPat, a pattern-based oligopeptide homology search method and the identification of a novel tachykinin-like peptide.

UNLABELLED: PepPat, a hybrid method that combines pattern matching with similarity scoring, is described. We also report PepPat's application in the identification of a novel tachykinin-like peptide. PepPat takes as input a query peptide and a user-specified regular expression pattern within the peptide. It first performs a database pattern match and then ranks candidates on the basis of their similarity to the query peptide. PepPat calculates similarity over the pattern spanning region, enhancing PepPat's sensitivity for short query peptides. PepPat can also search for a user-specified number of occurrences of a repeated pattern within the target sequence. We illustrate PepPat's application in short peptide ligand mining. As a validation example, we report the identification of a novel tachykinin-like peptide, C14TKL-1, and show it is an NK1 (neuokinin receptor 1) agonist whose message is widely expressed in human periphery. AVAILABILITY: PepPat is offered online at: http://peppat.cbi.pku.edu.cn.

Algorithms↗

Secreted protein prediction system combining CJ-SPHMM, TMHMM, and PSORT.

To increase the coverage of secreted protein prediction, we describe a combination strategy. Instead of using a single method, we combine Hidden Markov Model (HMM)-based methods CJ-SPHMM and TMHMM with PSORT in secreted protein prediction. CJ-SPHMM is an HMM-based signal peptide prediction method, while TMHMM is an HMM-based transmembrane (TM) protein prediction algorithm. With CJ-SPHMM and TMHMM, proteins with predicted signal peptide and without predicted TM regions are taken as putative secreted proteins. This HMM-based approach predicts secreted protein with Ac (Accuracy) at 0.82 and Cc (Correlation coefficient) at 0.75, which are similar to PSORT with Ac at 0.82 and Cc at 0.76. When we further complement the HMM-based method, i.e., CJ-SPHMM + TMHMM with PSORT in secreted protein prediction, the Ac value is increased to 0.86 and the Cc value is increased to 0.81. Taking this combination strategy to search putative secreted proteins from the International Protein Index (IPI) maintained at the European Bioinformatics Institute (EBI), we constructed a putative human secretome with 5235 proteins. The prediction system described here can also be applied to predicting secreted proteins from other vertebrate proteomes.

Computational Biology↗

[MGAP-A microbe genome annotation platform].

A Microbe Genome Annotation Platform (MGAP) was developed and applied to the cynobacterium PCC7002 genome annotation. Various bioinformatics software tools from sequence analysis to gene identification and function prediction were implemented in MGAP. Protein sequence databases SWISSPROT and PDBseq, protein information resource InterPro and COG were also integrated in the platform. The web interface of MGAP has the functionality to display a circular map of gene distribution and GC contents throughout the genome. Detailed information such as the DNA and protein sequence, the location of genes on chromosomes can be viewed by clicking the corresponding object within the map. MGAP is based on a PC/Linux system affordable for small biological laboratories and has the advantage of using free software tools including MySQL, Apache and Perl.

Cyanobacteria↗

PGAAS: a prokaryotic genome assembly assistant system.

MOTIVATION: In order to accelerate the finishing phase of genome assembly, especially for the whole genome shotgun approach of prokaryotic species, we have developed a software package designated prokaryotic genome assembly assistant system (PGAAS). The approach upon which PGAAS is based is to confirm the order of contigs and fill gaps between contigs through peptide links obtained by searching each contig end with BLASTX against protein databases. RESULTS: We used the contig dataset of the cyanobacterium Synechococcus sp. strain PCC7002 (PCC7002), which was sequenced with six-fold coverage and assembled using the Phrap package. The subject database is the protein database of the cyanobacterium, Synechocystis sp. strain PCC6803 (PCC6803). We found more than 100 non-redundant peptide segments which can link at least 2 contigs. We tested one pair of linked contigs by sequencing and obtained satisfactory result. PGAAS provides a graphic user interface to show the bridge peptides and pier contigs. We integrated Primer3 into our package to design PCR primers at the adjacent ends of the pier contigs. AVAILABILITY: We tested PGAAS on a Linux (Redhat 6.2) PC machine. It is developed with free software (MySQL, PHP and Apache). The whole package is distributed freely and can be downloaded as UNIX compress file: ftp://ftp.cbi.pku.edu.cn/pub/software/unix/pgaas1.0.tar.gz. The package is being continually updated.

Algorithms↗