Search PubMed⌕ Search

Biomedical subjects

Pei Hao

Publications and source records attributed to Pei Hao.

14 recordsLinked to original sources

In silico discovery of human natural antisense transcripts.

BACKGROUND: Several high-throughput searches for potential natural antisense transcripts (NATs) have been performed recently, but most of the reports were focused on cis type. A thorough in silico analysis of human transcripts will help expand our knowledge of NATs. RESULTS: We have identified 568 NATs from human RefSeq RNA sequences. Among them, 403 NATs are reported for the first time, and at least 157 novel NATs are trans type. According to the pairing region of a sense and antisense RNA pair, hNATs are divided into 6 classes, of which about 87% involve 5' or 3' UTR sequences, supporting the regulatory role of UTRs. Among a total of 535 NAT pairs related with splice variants, 77.4% (414/535) have their pairing regions affected or completely eliminated by alternative splicing, suggesting significant relationship of alternative splicing and antisense-directed regulation. The extensive occurrence of splice variants in hNATs and other multiple pairing patterns results in a one-to-many relationship, allowing the formation of complex regulation networks. Based on microarray data from Stanford Microarray Database, two hNAT pairs were found to display significant inverse expression patterns before and after insulin injection. CONCLUSION: NATs might carry out more extensive and complex functions than previously thought. Combined with endogenous micro RNAs, hNATs could be regarded as a special group of transcripts contributing to the complex regulation networks.

Algorithms↗

KDE Bioscience: platform for bioinformatics analysis workflows.

Bioinformatics is a dynamic research area in which a large number of algorithms and programs have been developed rapidly and independently without much consideration so far of the need for standardization. The lack of such common standards combined with unfriendly interfaces make it difficult for biologists to learn how to use these tools and to translate the data formats from one to another. Consequently, the construction of an integrative bioinformatics platform to facilitate biologists' research is an urgent and challenging task. KDE Bioscience is a java-based software platform that collects a variety of bioinformatics tools and provides a workflow mechanism to integrate them. Nucleotide and protein sequences from local flat files, web sites, and relational databases can be entered, annotated, and aligned. Several home-made or 3rd-party viewers are built-in to provide visualization of annotations or alignments. KDE Bioscience can also be deployed in client-server mode where simultaneous execution of the same workflow is supported for multiple users. Moreover, workflows can be published as web pages that can be executed from a web browser. The power of KDE Bioscience comes from the integrated algorithms and data sources. With its generic workflow mechanism other novel calculations and simulations can be integrated to augment the current sequence analysis functions. Because of this flexible and extensible architecture, KDE Bioscience makes an ideal integrated informatics environment for future bioinformatics or systems biology research.

Biological Science Disciplines↗

Cloning and characterization of the chromosomal replication origin region of Amycolatopsis mediterranei U32.

The chromosomal replication origins (oriC) of gram positive, acid-fast actinomycetes have been investigated in streptomycetes and mycobacteria. A 1339 bp DNA fragment of the putative oriC region from the rifamycin SV producer Amycolatopsis mediterranei U32 was cloned by PCR amplification employing primers designed based on the conserved flanking genes of dnaA and dnaN. The 884 bp sequence of the intergenic region between dnaA and dnaN genes consists of 19 DnaA-boxes and two 13-mer AT-rich sequences, which is similar to the oriC structure of Streptomyces lividans. A mini-chromosome constructed by cloning the putative U32 oriC DNA fragment into an Escherichia coli plasmid was able to replicate autonomously, but was unstable, in A. mediterranei U32 with an estimated copy number of two per cell. Although efficient replication of the mini-chromosome in U32 requires the complete set of DnaA-boxes and AT-rich regions, only one of the AT-rich sequences together with part of the DnaA-boxes is sufficient, suggesting the presence of combinatorial alternatives for a functional oriC region of A. mediterranei U32. Phylogenetic analysis based on definite oriC sequences among eubacteria reflects well the relationship between these species.

Actinobacteria↗

Identification of two critical amino acid residues of the severe acute respiratory syndrome coronavirus spike protein for its variation in zoonotic tropism transition via a double substitution strategy.

Severe acute respiratory syndrome coronavirus (SARS-CoV) is a recently identified human coronavirus. The extremely high homology of the viral genomic sequences between the viruses isolated from human (huSARS-CoV) and those of palm civet origin (pcSARS-CoV) suggested possible palm civet-to-human transmission. Genetic analysis revealed that the spike (S) protein of pcSARS-CoV and huSARS-CoV was subjected to the strongest positive selection pressure during transmission, and there were six amino acid residues within the receptor-binding domain of the S protein being potentially important for SARS progression and tropism. Using the single-round infection assay, we found that a two-amino acid substitution (N479K/T487S) of a huSARS-CoV for those of pcSARS-CoV almost abolished its infection of human cells expressing the SARS-CoV receptor ACE2 but no effect upon the infection of mouse ACE2 cells. Although single substitution of these two residues had no effects on the infectivity of huSARS-CoV, these recombinant S proteins bound to human ACE2 with different levels of reduced affinity, and the two-amino acid-substituted S protein showed extremely low affinity. On the contrary, substitution of these two amino acid residues of pcSARS-CoV for those of huSRAS-CoV made pcSARS-CoV capable of infecting human ACE2-expressing cells. These results suggest that amino acid residues at position 479 and 487 of the S protein are important determinants for SARS-CoV tropism and animal-to-human transmission.

Amino Acid Sequence↗

Comparative and functional genomic analyses of the pathogenicity of phytopathogen Xanthomonas campestris pv. campestris.

Xanthomonas campestris pathovar campestris (Xcc) is the causative agent of crucifer black rot disease, which causes severe losses in agricultural yield world-wide. This bacterium is a model organism for studying plant-bacteria interactions. We sequenced the complete genome of Xcc 8004 (5,148,708 bp), which is highly conserved relative to that of Xcc ATCC 33913. Comparative genomics analysis indicated that, in addition to a significant genomic-scale rearrangement cross the replication axis between two IS1478 elements, loss and acquisition of blocks of genes, rather than point mutations, constitute the main genetic variation between the two Xcc strains. Screening of a high-density transposon insertional mutant library (16,512 clones) of Xcc 8004 against a host plant (Brassica oleraceae) identified 75 nonredundant, single-copy insertions in protein-coding sequences (CDSs) and intergenic regions. In addition to known virulence factors, full virulence was found to require several additional metabolic pathways and regulatory systems, such as fatty acid degradation, type IV secretion system, cell signaling, and amino acids and nucleotide metabolism. Among the identified pathogenicity-related genes, three of unknown function were found in Xcc 8004-specific chromosomal segments, revealing a direct correlation between genomic dynamics and Xcc virulence. The present combination of comparative and functional genomic analyses provides valuable information about the genetic basis of Xcc pathogenicity, which may offer novel insight toward the development of efficient methods for prevention of this important plant disease.

Bacterial Proteins↗

Cross-host evolution of severe acute respiratory syndrome coronavirus in palm civet and human.

The genomic sequences of severe acute respiratory syndrome coronaviruses from human and palm civet of the 2003/2004 outbreak in the city of Guangzhou, China, were nearly identical. Phylogenetic analysis suggested an independent viral invasion from animal to human in this new episode. Combining all existing data but excluding singletons, we identified 202 single-nucleotide variations. Among them, 17 are polymorphic in palm civets only. The ratio of nonsynonymous/synonymous nucleotide substitution in palm civets collected 1 yr apart from different geographic locations is very high, suggesting a rapid evolving process of viral proteins in civet as well, much like their adaptation in the human host in the early 2002-2003 epidemic. Major genetic variations in some critical genes, particularly the Spike gene, seemed essential for the transition from animal-to-human transmission to human-to-human transmission, which eventually caused the first severe acute respiratory syndrome outbreak of 2002/2003.

Amino Acid Substitution↗

MPSS: an integrated database system for surveying a set of proteins.

SUMMARY: We design and implement an integrated database system called 'multi-protein survey system' (MPSS), which provides a platform to retrieve information about many proteins at a time. This system integrates several important and widely used databases including SwissProt, TrEMBL, PDB and InterPro, plus useful references such as GO and KEGG to other databases. Users may submit a group of protein IDs, entry names, SwissProt/TrEMBL accession numbers or GenBank GIs through MPSS' web interface, and obtain protein annotation information from public databases and pre-computed molecular properties speedily. MPSS can also supply comprehensive information about query proteins, including 3D structures, domains, pathway, gene ontology and visual presentation of mapping to the GO tree and KEGG pathway, to provide an up-to-date view of available knowledge with regard to the structures and molecular functions of proteins under study. AVAILABILITY: MPSS is freely accessible at http://www.scbit.org/mpss/

Database Management Systems↗

A molecular docking model of SARS-CoV S1 protein in complex with its receptor, human ACE2.

The exact residues within severe acute respiratory syndrome coronavirus (SARS-CoV) S1 protein and its receptor, human ACE2, involved in their interaction still remain largely undetermined. Identification of exact amino acid residues that are crucial for the interaction of S1 with ACE2 could provide working hypotheses for experimental studies and might be helpful for the development of antiviral inhibitor. In this paper, a molecular docking model of SARS-CoV S1 protein in complex with human ACE2 was constructed. The interacting residue pairs within this complex model and their contact types were also identified. Our model, supported by significant biochemical evidence, suggested receptor-binding residues were concentrated in two segments of S1 protein. In contrast, the interfacial residues in ACE2, though close to each other in tertiary structure, were found to be widely scattered in the primary sequence. In particular, the S1 residue ARG453 and ACE2 residue LYS341 might be the key residues in the complex formation.

Amino Acid Sequence↗

Prediction of quaternary assembly of SARS coronavirus peplomer.

The tertiary structures of the S1 and S2 domains of the spike protein of the coronavirus which is responsible of the severe acute respiratory syndrome (SARS) have been recently predicted. Here a molecular assembly of SARS coronavirus peplomer which accounts for the available functional data is suggested. The interaction between S1 and S2 appears to be stabilised by a large hydrophobic network of aromatic side chains present in both domains. This feature results to be common to all coronaviruses, suggesting potential targeting for drugs preventing coronavirus-related infections.

Amino Acid Substitution↗

A high-throughput approach for subcellular proteome: identification of rat liver proteins using subcellular fractionation coupled with two-dimensional liquid chromatography tandem mass spectrometry and bioinformatic analysis.

Four fractions from rat liver (a crude mitochondria (CM) and cytosol (C) fraction obtained with differential centrifugation, a purified mitochondrial (PM) fraction obtained with nycodenz density gradient centrifugation, and a total liver (TL) fraction) were analyzed with two-dimensional liquid chromatography tandem mass spectrometry analysis. A total of 564 rat proteins were identified and were bioinformatically annotated according to their physicochemical characteristics and functions. While most extreme alkaline ribosomal proteins were identified in the TL fraction, the C fraction mainly included neutral enzymes and the PM fraction enriched alkaline proteins and proteins with electron transfer activity or oxygen binding activity. Such characteristics were more apparent in proteins identified only in the TL, C, or PM fraction. The Swiss-Prot annotation and the bioinformatic prediction results proved that the C and PM fractions had enriched cytoplasmic or mitochondrial proteins, respectively. Combination usage of subcellular fractionation with two-dimensional liquid chromatography tandem mass spectrometry was proved to be a high-throughput, sensitive, and effective analytical approach for subcellular proteomics research. Using such a strategy, we have constructed the largest proteome database to date for rat liver (564 rat proteins) and its cytosol (222 rat proteins) and mitochondrial fractions (227 rat proteins). Moreover, the 352 proteins with Swiss-Prot subcellular location annotation in the 564 identified proteins were used as an actual subcellular proteome dataset to evaluate the widely used bioinformatics tools such as PSORT, TargetP, TMHMM, and GRAVY.

Animals↗

Putative hAPN receptor binding sites in SARS_CoV spike protein.

AIM: To obtain the information of ligand-receptor binding between the S protein of SARS-CoV and CD13, identify the possible interacting domains or motifs related to binding sites, and provide clues for studying the functions of SARS proteins and designing anti-SARS drugs and vaccines. METHODS: On the basis of comparative genomics, the homology search, phylogenetic analyses, and multi-sequence alignment were used to predict CD13 related interacting domains and binding sites in the S protein of SARS-CoV. Molecular modeling and docking simulation methods were employed to address the interaction feature between CD13 and S protein of SARS-CoV in validating the bioinformatics predictions. RESULTS: Possible binding sites in the SARS-CoV S protein to CD13 have been mapped out by using bioinformatics analysis tools. The binding for one protein-protein interaction pair (D757-R761 motif of the SARS-CoV S protein to P585-A653 domain of CD13) has been simulated by molecular modeling and docking simulation methods. CONCLUSION: CD13 may be a possible receptor of the SARS-CoV S protein, which may be associated with the SARS infection. This study also provides a possible strategy for mapping the possible binding receptors of the proteins in a genome.

Amino Acid Sequence↗

Identification of probable genomic packaging signal sequence from SARS-CoV genome by bioinformatics analysis.

AIM: To predict the probable genomic packaging signal of SARS-CoV by bioinformatics analysis. The derived packaging signal may be used to design antisense RNA and RNA interfere (RNAi) drugs treating SARS. METHODS: Based on the studies about the genomic packaging signals of MHV and BCoV, especially the information about primary and secondary structures, the putative genomic packaging signal of SARS-CoV were analyzed by using bioinformatic tools. Multi-alignment for the genomic sequences was performed among SARS-CoV, MHV, BCoV, PEDV and HCoV 229E. Secondary structures of RNA sequences were also predicted for the identification of the possible genomic packaging signals. Meanwhile, the N and M proteins of all five viruses were analyzed to study the evolutionary relationship with genomic packaging signals. RESULTS: The putative genomic packaging signal of SARS-CoV locates at the 3' end of ORF1b near that of MHV and BCoV, where is the most variable region of this gene. The RNA secondary structure of SARS-CoV genomic packaging signal is very similar to that of MHV and BCoV. The same result was also obtained in studying the genomic packaging signals of PEDV and HCoV 229E. Further more, the genomic sequence multi-alignment indicated that the locations of packaging signals of SARS-CoV, PEDV, and HCoV overlaped each other. It seems that the mutation rate of packaging signal sequences is much higher than the N protein, while only subtle variations for the M protein. CONCLUSIONS: The probable genomic packaging signal of SARS-CoV is analogous to that of MHV and BCoV, with the corresponding secondary RNA structure locating at the similar region of ORF1b. The positions where genomic packaging signals exist have suffered rounds of mutations, which may influence the primary structures of the N and M proteins consequently.

Amino Acid Sequence↗

Sequence and analysis of rice chromosome 4.

Rice is the principal food for over half of the population of the world. With its genome size of 430 megabase pairs (Mb), the cultivated rice species Oryza sativa is a model plant for genome research. Here we report the sequence analysis of chromosome 4 of O. sativa, one of the first two rice chromosomes to be sequenced completely. The finished sequence spans 34.6 Mb and represents 97.3% of the chromosome. In addition, we report the longest known sequence for a plant centromere, a completely sequenced contig of 1.16 Mb corresponding to the centromeric region of chromosome 4. We predict 4,658 protein coding genes and 70 transfer RNA genes. A total of 1,681 predicted genes match available unique rice expressed sequence tags. Transposable elements have a pronounced bias towards the euchromatic regions, indicating a close correlation of their distributions to genes along the chromosome. Comparative genome analysis between cultivated rice subspecies shows that there is an overall syntenic relationship between the chromosomes and divergence at the level of single-nucleotide polymorphisms and insertions and deletions. By contrast, there is little conservation in gene order between rice and Arabidopsis.

Arabidopsis↗

A fine physical map of the rice chromosome 4.

As part of an international effort to completely sequence the rice genome, we have produced a fine bacterial artificial chromosome (BAC)-based physical map of the Oryza sativa japonica Nipponbare chromosome 4 through an integration of 114 sequenced BAC clones from a taxonomically related subspecies O. sativa indica Guangluai 4 and 182 RFLP and 407 expressed sequence tag (EST) markers with the fingerprinted data of the Nipponbare genome. The map consists of 11 contigs with a total length of 34.5 Mb covering 94% of the estimated chromosome size (36.8 Mb). BAC clones corresponding to telomeres, as well as to the centromere position, were determined by BAC-pachytene chromosome fluorescence in situ hybridization (FISH). This gave rise to an estimated length ratio of 5.13 for the long arm and 2.9 for the short arm (on the basis of the physical map), which indicates that the short arm is a highly condensed one. The FISH analysis and physical mapping also showed that the short arm and the pericentromeric region of the long arm are rich in heterochromatin, which occupied 45% of the chromosome, indicating that this chromosome is likely very difficult to sequence. To our knowledge, this map provides the first example of a rapid and reliable physical mapping on the basis of the integration of the data from two taxonomically related subspecies.

Chromosomes↗