Search PubMed⌕ Search

Biomedical subjects

Jong Bhak

Publications and source records attributed to Jong Bhak.

6 recordsLinked to original sources

Impact of transcriptional properties on essentiality and evolutionary rate.

We characterized general transcriptional activity and variability of eukaryotic genes from global expression profiles of human, mouse, rat, fly, plants, and yeast. The variability shows a higher degree of divergence between distant species, implying that it is more closely related to phenotypic evolution, than the activity. More specifically, we show that transcriptional variability should be a true indicator of evolutionary rate. If we rule out the effect of translational selection, which seems to operate only in yeast, the apparent slow evolution of highly expressed genes should be attributed to their low variability. Meanwhile, rapidly evolving genes may acquire a high level of transcriptional variability and contribute to phenotypic variations. Essentiality also seems to be correlated with the variability, not the activity. We show that indispensable or highly interactive proteins tend to be present in high abundance to maintain a low variability. Our results challenge the current theory that highly expressed genes are essential and evolve slowly. Transcriptional variability, rather than transcriptional activity, might be a common indicator of essentiality and evolutionary rate, contributing to the correlation between the two variables.

Animals↗

Localizome: a server for identifying transmembrane topologies and TM helices of eukaryotic proteins utilizing domain information.

The Localizome server predicts the transmembrane (TM) helix number and TM topology of a user-supplied eukaryotic protein and presents the result as an intuitive graphic representation. It utilizes hmmpfam to detect the presence of Pfam domains and a prediction algorithm, Phobius, to predict the TM helices. The results are combined and checked against the TM topology rules stored in a protein domain database called LocaloDom. LocaloDom is a curated database that contains TM topologies and TM helix numbers of known protein domains. It was constructed from Pfam domains combined with Swiss-Prot annotations and Phobius predictions. The Localizome server corrects the combined results of the user sequence to conform to the rules stored in LocaloDom. Compared with other programs, this server showed the highest accuracy for TM topology prediction: for soluble proteins, the accuracy and coverage were 99 and 75%, respectively, while for TM protein domain regions, they were 96 and 68%, respectively. With a graphical representation of TM topology and TM helix positions with the domain units, the Localizome server is a highly accurate and comprehensive information source for subcellular localization for soluble proteins as well as membrane proteins. The Localizome server can be found at http://localizome.org/.

Cell Compartmentation↗

Functional annotation and analysis of Korean patented biological sequences using bioinformatics.

A recent report of the Korean Intellectual Property Office (KIPO) showed that the number of biological sequence-based patents is rapidly increasing in Korea. We present biological features of Korean patented sequences though bioinformatic analysis. The analysis is divided into two steps. The first is an annotation step in which the patented sequences were annotated with the Reference Sequence (RefSeq) database. The second is an association step in which the patented sequences were linked to genes, diseases, pathway, and biological functions. We used Entrez Gene, Online Mendelian Inheritance in Man (OMIM), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Gene Ontology (GO) databases. Through the association analysis, we found that nearly 2.6% of human genes were associated with Korean patenting, compared to 20% of human genes in the U.S. patent. The association between the biological functions and the patented sequences indicated that genes whose products act as hormones on defense responses in the extra-cellular environments were the most highly targeted for patenting. The analysis data are available at http://www.patome.net.

Base Sequence↗

Sequence-level analysis of the diploidization process in the triplicated FLOWERING LOCUS C region of Brassica rapa.

Strong evidence exists for polyploidy having occurred during the evolution of the tribe Brassiceae. We show evidence for the dynamic and ongoing diploidization process by comparative analysis of the sequences of four paralogous Brassica rapa BAC clones and the homologous 124-kb segment of Arabidopsis thaliana chromosome 5. We estimated the times since divergence of the paralogous and homologous lineages. The three paralogous subgenomes of B. rapa triplicated 13 to 17 million years ago (MYA), very soon after the Arabidopsis and Brassica divergence occurred at 17 to 18 MYA. In addition, a pair of BACs represents a more recent segmental duplication, which occurred approximately 0.8 MYA, and provides an exception to the general expectation of three paralogous segments within the B. rapa genome. The Brassica genome segments show extensive interspersed gene loss relative to the inferred structure of the ancestral genome, whereas the Arabidopsis genome segment appears little changed. Representatives of all 32 genes in the Arabidopsis genome segment are represented in Brassica, but the hexaploid complement of 96 has been reduced to 54 in the three subgenomes, with compression of the genomic region lengths they occupy to between 52 and 110 kb. The gene content of the recently duplicated B. rapa genome segments is identical, but intergenic sequences differ.

Brassica rapa↗

A protein domain interaction interface database: InterPare.

BACKGROUND: Most proteins function by interacting with other molecules. Their interaction interfaces are highly conserved throughout evolution to avoid undesirable interactions that lead to fatal disorders in cells. Rational drug discovery includes computational methods to identify the interaction sites of lead compounds to the target molecules. Identifying and classifying protein interaction interfaces on a large scale can help researchers discover drug targets more efficiently. DESCRIPTION: We introduce a large-scale protein domain interaction interface database called InterPare http://interpare.net. It contains both inter-chain (between chains) interfaces and intra-chain (within chain) interfaces. InterPare uses three methods to detect interfaces: 1) the geometric distance method for checking the distance between atoms that belong to different domains, 2) Accessible Surface Area (ASA), a method for detecting the buried region of a protein that is detached from a solvent when forming multimers or complexes, and 3) the Voronoi diagram, a computational geometry method that uses a mathematical definition of interface regions. InterPare includes visualization tools to display protein interior, surface, and interaction interfaces. It also provides statistics such as the amino acid propensities of queried protein according to its interior, surface, and interface region. The atom coordinates that belong to interface, surface, and interior regions can be downloaded from the website. CONCLUSION: InterPare is an open and public database server for protein interaction interface information. It contains the large-scale interface data for proteins whose 3D-structures are known. As of November 2004, there were 10,583 (Geometric distance), 10,431 (ASA), and 11,010 (Voronoi diagram) entries in the Protein Data Bank (PDB) containing interfaces, according to the above three methods. In the case of the geometric distance method, there are 31,620 inter-chain domain-domain interaction interfaces and 12,758 intra-chain domain-domain interfaces.

Computers, Molecular↗

Comparative interactomics analysis of protein family interaction networks using PSIMAP (protein structural interactome map).

MOTIVATION: Many genomes have been completely sequenced. However, detecting and analyzing their protein-protein interactions by experimental methods such as co-immunoprecipitation, tandem affinity purification and Y2H is not as fast as genome sequencing. Therefore, a computational prediction method based on the known protein structural interactions will be useful to analyze large-scale protein-protein interaction rules within and among complete genomes. RESULTS: We confirmed that all the predicted protein family interactomes (the full set of protein family interactions within a proteome) of 146 species are scale-free networks, and they share a small core network comprising 36 protein families related to indispensable cellular functions. We found two fundamental differences among prokaryotic and eukaryotic interactomes: (1) eukarya had significantly more hub families than archaea and bacteria and (2) certain special hub families determined the topology of the eukaryotic interactomes. Our comparative analysis suggests that a very small number of expansive protein families led to the evolution of interactomes and seemed to have played a key role in species diversification. SUPPLEMENTARY INFORMATION: http://interactomics.org.

Algorithms↗