Search PubMed⌕ Search

Biomedical subjects

In-Geol Choi

Publications and source records attributed to In-Geol Choi.

6 recordsLinked to original sources

Protein conformational space in higher order phi-Psi maps.

We have mapped protein conformational space from two to seven residue lengths by employing multidimensional scaling on a data matrix composed of pair-wise angular distances for multiple phi-Psi values collected from high-resolution protein structures. The resulting global maps show clustering of peptide conformations that reveals a dramatic reduction of conformational space as sampled by experimentally observed peptides. Each map can be viewed as a higher order phi-Psi plot defining regions of space that are conformationally allowed.

Cluster Analysis↗

Structure of OsmC from Escherichia coli: a salt-shock-induced protein.

The crystal structure of an osmotically inducible protein (OsmC) from Escherichia coli has been determined at 2.4 A resolution. OsmC is a representative protein of the OsmC sequence family, which is composed of three sequence subfamilies. The structure of OsmC provides a view of a salt-shock-induced protein. Two identical monomers form a cylindrically shaped dimer in which six helices are located on the inside and two six-stranded beta-sheets wrap around these helices. Structural comparison suggests that the OsmC sequence family has a peroxiredoxin function and has a unique structure compared with other peroxiredoxin families. A detailed analysis of structures and sequence comparisons in the OsmC sequence family revealed that each subfamily has unique motifs. In addition, the molecular function of the OsmC sequence family is discussed based on structural comparisons among the subfamily members.

Amino Acid Sequence↗

Local feature frequency profile: a method to measure structural similarity in proteins.

Measures of structural similarity between known protein structures provide an objective basis for classifying protein folds and for revealing a global view of the protein structure universe. Here, we describe a rapid method to measure structural similarity based on the profiles of representative local features of C(alpha) distance matrices of compared protein structures. We first extract a finite number of representative local feature (LF) patterns from the distance matrices of all protein fold families by medoid analysis. Then, each C(alpha) distance matrix of a protein structure is encoded by labeling all its submatrices by the index of the nearest representative LF patterns. Finally, the structure is represented by the frequency distribution of these indices, which we call the LF frequency (LFF) profile of the protein. The LFF profile allows one to calculate structural similarity scores among a large number of protein structures quickly, and also to construct and update the "map" of the protein structure universe easily. The LFF profile method efficiently maps complex protein structures into a common Euclidean space without prior assignment of secondary structure information or structural alignment.

Proteins↗

Structure-based functional inference in structural genomics.

The dramatically increasing number of new protein sequences arising from genomics and proteomics requires the need for methods to rapidly and reliably infer the molecular and cellular functions of these proteins. One such approach, structural genomics, aims to delineate the total repertoire of protein folds in nature, thereby providing three-dimensional folding patterns for all proteins and to infer molecular functions of the proteins based on the combined information of structures and sequences. The goal of obtaining protein structures on a genomic scale has motivated the development of high throughput technologies and protocols for macromolecular structure determination that have begun to produce structures at a greater rate than previously possible. These new structures have revealed many unexpected functional inferences and evolutionary relationships that were hidden at the sequence level. Here, we present samples of structures determined at Berkeley Structural Genomics Center and collaborators' laboratories to illustrate how structural information provides and complements sequence information to deduce the functional inferences of proteins with unknown molecular functions. Two of the major premises of structural genomics are to discover a complete repertoire of protein folds in nature and to find molecular functions of the proteins whose functions are not predicted from sequence comparison alone. To achieve these objectives on a genomic scale, new methods, protocols, and technologies need to be developed by multi-institutional collaborations worldwide. As part of this effort, the Protein Structure Initiative has been launched in the United States (PSI; www.nigms.nih.gov/funding/psi.html). Although infrastructure building and technology development are still the main focus of structural genomics programs, a considerable number of protein structures have already been produced, some of them coming directly out of semiautomated structure determination pipelines. The Berkeley Structural Genomics Center (BSGC) has focused on the proteins of Mycoplasma or their homologues from other organisms as its structural genomics targets because of the minimal genome size of the Mycoplasmas as well as their relevance to human and animal pathogenicity (http://www.strgen.org). Here we present several protein examples encompassing a spectrum of functional inferences obtainable from their three-dimensional structures in five situations, where the inferences are new and testable, and are not predictable from protein sequence information alone.

Adenosine Triphosphate↗

Target selection for structural genomics: a single genome approach.

We describe our strategy for selecting targets for protein structure determination in context of structural genomics of a single genome. In the course of target selection, we have studied two of the smallest microbial genomes, Mycoplasma genitalium and Mycoplasma pneumoniae. To our surprise, we found that only 71 Mycoplasma genes or their orthologues can be considered as easy targets for high-throughput structural studies--far fewer than expected. We discuss the methods and criteria used for target selection and the reasons explaining rarity of easy targets. First, despite the common opinion that protein folds can be predicted for only 30-50% of genes, the number of "truly unknown" structures is less than one-third. Second, due to the different codon usage, two thirds of Mycoplasma proteins cannot be directly expressed in E. coli in high-throughput manner and require substitution by their homologues from other organisms. Third, membrane or large multi-domain proteins are difficult targets because of solubility and size issues and often require identification and structure determination of protein domains. Finally, we propose different approaches to address the difficult targets.

Cell Membrane↗