Search PubMed⌕ Search

Biomedical subjects

Jimin Pei

Publications and source records attributed to Jimin Pei.

15 recordsLinked to original sources

MUMMALS: multiple sequence alignment improved by using hidden Markov models with local structural information.

We have developed MUMMALS, a program to construct multiple protein sequence alignment using probabilistic consistency. MUMMALS improves alignment quality by using pairwise alignment hidden Markov models (HMMs) with multiple match states that describe local structural information without exploiting explicit structure predictions. Parameters for such models have been estimated from a large library of structure-based alignments. We show that (i) on remote homologs, MUMMALS achieves statistically best accuracy among several leading aligners, such as ProbCons, MAFFT and MUSCLE, albeit the average improvement is small, in the order of several percent; (ii) a large collection (>10 000) of automatically computed pairwise structure alignments of divergent protein domains is superior to smaller but carefully curated datasets for estimation of alignment parameters and performance tests; (iii) reference-independent evaluation of alignment quality using sequence alignment-dependent structure superpositions correlates well with reference-dependent evaluation that compares sequence-based alignments to structure-based reference alignments.

Markov Chains↗

Substrate and functional diversity of lysine acetylation revealed by a proteomics survey.

Acetylation of proteins on lysine residues is a dynamic posttranslational modification that is known to play a key role in regulating transcription and other DNA-dependent nuclear processes. However, the extent of this modification in diverse cellular proteins remains largely unknown, presenting a major bottleneck for lysine-acetylation biology. Here we report the first proteomic survey of this modification, identifying 388 acetylation sites in 195 proteins among proteins derived from HeLa cells and mouse liver mitochondria. In addition to regulators of chromatin-based cellular processes, nonnuclear localized proteins with diverse functions were identified. Most strikingly, acetyllysine was found in more than 20% of mitochondrial proteins, including many longevity regulators and metabolism enzymes. Our study reveals previously unappreciated roles for lysine acetylation in the regulation of diverse cellular pathways outside of the nucleus. The combined data sets offer a rich source for further characterization of the contribution of this modification to cellular physiology and human diseases.

Acetylation↗

Prediction of functional specificity determinants from protein sequences using log-likelihood ratios.

MOTIVATION: A number of methods have been developed to predict functional specificity determinants in protein families based on sequence information. Most of these methods rely on pre-defined functional subgroups. Manual subgroup definition is difficult because of the limited number of experimentally characterized subfamilies with differing specificity, while automatic subgroup partitioning using computational tools is a non-trivial task and does not always yield ideal results. RESULTS: We propose a new approach SPEL (specificity positions by evolutionary likelihood) to detect positions that are likely to be functional specificity determinants. SPEL, which does not require subgroup definition, takes a multiple sequence alignment of a protein family as the only input, and assigns a P-value to every position in the alignment. Positions with low P-values are likely to be important for functional specificity. An evolutionary tree is reconstructed during the calculation, and P-value estimation is based on a random model that involves evolutionary simulations. Evolutionary log-likelihood is chosen as a measure of amino acid distribution at a position. To illustrate the performance of the method, we carried out a detailed analysis of two protein families (LacI/PurR and G protein alpha subunit), and compared our method with two existing methods (evolutionary trace and mutual information based). All three methods were also compared on a set of protein families with known ligand-bound structures. AVAILABILITY: SPEL is freely available for non-commercial use. Its pre-compiled versions for several platforms and alignments used in this work are available at ftp://iole.swmed.edu/pub/SPEL/

Algorithms↗

COG3926 and COG5526: a tale of two new lysozyme-like protein families.

We have identified two new lysozyme-like protein families by using a combination of sequence similarity searches, domain architecture analysis, and structural predictions. First, the P5 protein from bacteriophage phi8, which belongs to COG3926 and Pfam family DUF847, is predicted to have a new lysozyme-like domain. This assignment is consistent with the lytic function of P5 proteins observed in several related double-stranded RNA bacteriophages. Domain architecture analysis reveals two lysozyme-associated transmembrane modules (LATM1 and LATM2) in a few COG3926/DUF847 members. LATM2 is also present in two proteins containing a peptidoglycan binding domain (PGB) and an N-terminal region that corresponds to COG5526 with uncharacterized function. Second, structure prediction and sequence analysis suggest that COG5526 represents another new lysozyme-like family. Our analysis offers fold and active-site assignments for COG3926/DUF847 and COG5526. The predicted enzymatic activity is consistent with an experimental study on the zliS gene product from Zymomonas mobilis, suggesting that bacterial COG3926/DUF847 members might be activators of macromolecular secretion.

Amino Acid Sequence↗

The P5 protein from bacteriophage phi-6 is a distant homolog of lytic transglycosylases.

Peptidases are classical objects of enzymology and structural studies. However, a few protein families with experimentally characterized proteolytic activity, but unknown catalytic mechanism and three-dimensional structures, still exist. Using comparative sequence analysis, we deduce spatial structure for one of such families, namely, U40, which contains just one P5 protein from bacteriophage phi-6. We show that this singleton sequence possesses conserved sequence motifs characteristic of lysozymes and is a distant homolog of lytic transglycosylases that cleave bacterial peptidoglycan. The structure of the P5 protein is therefore predicted to adopt the lysozyme-like fold shared by T4, lambda, C-type, G-type lysozymes, and lytic transglycosylases. Since previous biochemical experiments with P5 of phi-6 have indicated that the purified enzyme possesses endopeptidase activity and not glycosidase activity, our results point to the possibility of a newly evolved molecular function and call for further experimental characterization of this unusual P5 protein.

Amino Acid Sequence↗

Three-dimensional structure of the rSly1 N-terminal domain reveals a conformational change induced by binding to syntaxin 5.

Sec1/Mun18-like (SM) proteins and soluble N-ethylmaleimide-sensitive factor attachment protein receptors (SNAREs) play central roles in intracellular membrane fusion. Diverse modes of interaction between SM proteins and SNAREs from the syntaxin family have been described. However, the observation that the N-terminal domains of Sly1 and Vps45, the SM proteins involved in traffic at the endoplasmic reticulum, the Golgi, the trans-Golgi network and the endosomes, bind to similar N-terminal sequences of their cognate syntaxins suggested a unifying theme for SM protein/SNARE interactions in most internal membrane compartments. To further understand this mechanism of SM protein/SNARE coupling, we have elucidated the structure in solution of the isolated N-terminal domain of rat Sly1 (rSly1N) and analyzed its complex with an N-terminal peptide of rat syntaxin 5 by NMR spectroscopy. Comparison with the crystal structure of a complex between Sly1p and Sed5p, their yeast homologues, shows that syntaxin 5 binding requires a striking conformational change involving a two-residue shift in the register of the C-terminal beta-strand of rSly1N. This conformational change is likely to induce a significant alteration in the overall shape of full-length rSly1 and may be critical for its function. Sequence analyses indicate that this conformational change is conserved in the Sly1 family but not in other SM proteins, and that the four families represented by the four SM proteins found in yeast (Sec1p, Sly1p, Vps45p and Vps33p) diverged early in evolution. These results suggest that there are marked distinctions between the mechanisms of action of each of the four families of SM proteins, which may have arisen from different regulatory requirements of traffic in their corresponding membrane compartments.

Animals↗

Reconstruction of ancestral protein sequences and its applications.

BACKGROUND: Modern-day proteins were selected during long evolutionary history as descendants of ancient life forms. In silico reconstruction of such ancestral protein sequences facilitates our understanding of evolutionary processes, protein classification and biological function. Additionally, reconstructed ancestral protein sequences could serve to fill in sequence space thus aiding remote homology inference. RESULTS: We developed ANCESCON, a package for distance-based phylogenetic inference and reconstruction of ancestral protein sequences that takes into account the observed variation of evolutionary rates between positions that more precisely describes the evolution of protein families. To improve the accuracy of evolutionary distance estimation and ancestral sequence reconstruction, two approaches are proposed to estimate position-specific evolutionary rates. Comparisons show that at large evolutionary distances our method gives more accurate ancestral sequence reconstruction than PAML, PHYLIP and PAUP*. We apply the reconstructed ancestral sequences to homology inference and functional site prediction. We show that the usage of hypothetical ancestors together with the present day sequences improves profile-based sequence similarity searches; and that ancestral sequence reconstruction methods can be used to predict positions with functional specificity. CONCLUSIONS: As a computational tool to reconstruct ancestral protein sequences from a given multiple sequence alignment, ANCESCON shows high accuracy in tests and helps detection of remote homologs and prediction of functional sites. ANCESCON is freely available for non-commercial use. Pre-compiled versions for several platforms can be downloaded from ftp://iole.swmed.edu/pub/ANCESCON/.

Amino Acid Sequence↗

Combining evolutionary and structural information for local protein structure prediction.

We study the effects of various factors in representing and combining evolutionary and structural information for local protein structural prediction based on fragment selection. We prepare databases of fragments from a set of non-redundant protein domains. For each fragment, evolutionary information is derived from homologous sequences and represented as estimated effective counts and frequencies of amino acids (evolutionary frequencies) at each position. Position-specific amino acid preferences called structural frequencies are derived from statistical analysis of discrete local structural environments in database structures. Our method for local structure prediction is based on ranking and selecting database fragments that are most similar to a target fragment. Using secondary structure type as a local structural property, we test our method in a number of settings. The major findings are: (1) the COMPASS-type scoring function for fragment similarity comparison gives better prediction accuracy than three other tested scoring functions for profile-profile comparison. We show that the COMPASS-type scoring function can be derived both in the probabilistic framework and in the framework of statistical potentials. (2) Using the evolutionary frequencies of database fragments gives better prediction accuracy than using structural frequencies. (3) Finer definition of local environments, such as including more side-chain solvent accessibility classes and considering the backbone conformations of neighboring residues, gives increasingly better prediction accuracy using structural frequencies. (4) Combining evolutionary and structural frequencies of database fragments, either in a linear fashion or using a pseudocount mixture formula, results in improvement of prediction accuracy. Combination at the log-odds score level is not as effective as combination at the frequency level. This suggests that there might be better ways of combining sequence and structural information than the commonly used linear combination of log-odds scores. Our method of fragment selection and frequency combination gives reasonable results of secondary structure prediction tested on 56 CASP5 targets (average SOV score 0.77), suggesting that it is a valid method for local protein structure prediction. Mixture of predicted structural frequencies and evolutionary frequencies improve the quality of local profile-to-profile alignment by COMPASS.

Algorithms↗

Double-stranded DNA bacteriophage prohead protease is homologous to herpesvirus protease.

Double-stranded DNA bacteriophages and herpesviruses assemble their heads in a similar fashion; a pre-formed precursor called a prohead or procapsid undergoes a conformational transition to give rise to a mature head or capsid. A virus-encoded prohead or procapsid protease is often required in this maturation process. Through computational analysis, we infer homology between bacteriophage prohead proteases (MEROPS families U9 and U35) and herpesvirus protease (MEROPS family S21), and unify them into a procapsid protease superfamily. We also extend this superfamily to include an uncharacterized cluster of orthologs (COG3566) and many other phage or bacteria-encoded hypothetical proteins. On the basis of this homology and the herpesvirus protease structure and catalytic mechanism, we predict that bacteriophage prohead proteases adopt the herpesvirus protease fold and exploit a conserved Ser and His residue pair in catalysis. Our study provides further support for the proposed evolutionary link between dsDNA bacteriophages and herpesviruses.

Amino Acid Sequence↗

Using protein design for homology detection and active site searches.

We describe a method of designing artificial sequences that resemble naturally occurring sequences in terms of their compatibility with a template structure and its functional constraints. The design procedure is a Monte Carlo simulation of amino acid substitution process. The selective fixation of substitutions is dictated by a simple scoring function derived from the template structure and a multiple alignment of its homologs. Designed sequences represent an enlargement of sequence space around native sequences. We show that the use of designed sequences improves the performance of profile-based homology detection. The difference in position-specific conservation between designed sequences and native sequences is helpful for prediction of functionally important residues. Our sequence selection criteria in evolutionary simulations introduce amino acid substitution rate variation among sites in a natural way, providing a better model to test phylogenetic methods.

Binding Sites↗

PCMA: fast and accurate multiple sequence alignment based on profile consistency.

UNLABELLED: PCMA (profile consistency multiple sequence alignment) is a progressive multiple sequence alignment program that combines two different alignment strategies. Highly similar sequences are aligned in a fast way as in ClustalW, forming pre-aligned groups. The T-Coffee strategy is applied to align the relatively divergent groups based on profile-profile comparison and consistency. The scoring function for local alignments of pre-aligned groups is based on a novel profile-profile comparison method that is a generalization of the PSI-BLAST approach to profile-sequence comparison. PCMA balances speed and accuracy in a flexible way and is suitable for aligning large numbers of sequences. AVAILABILITY: PCMA is freely available for non-commercial use. Pre-compiled versions for several platforms can be downloaded from ftp://iole.swmed.edu/pub/PCMA/.

Algorithms↗

CASP5 assessment of fold recognition target predictions.

We present an overview of the fifth round of Critical Assessment of Protein Structure Prediction (CASP5) fold recognition category. Prediction models were evaluated by using six different structural measures and four different alignment measures, and these scores were compared to those assigned manually over a diverse subset of target domains. Scores were combined to compare overall performance of participating groups and to estimate rank significance. The methods used by a few groups outperformed all other methods in terms of the evaluated criteria and could be considered state-of-the-art in structure prediction. We discuss a few examples of difficult fold recognition targets to highlight the progress of ab initio-type methods on difficult structure analogs and the difficulties of predicting multidomain targets and selecting prediction models. We also compared the results of manual groups to those of automatic servers evaluated in parallel by CAFASP, showing that the top performing automated server structure predictions approached those of the best manual predictors.

Algorithms↗

Peptidase family U34 belongs to the superfamily of N-terminal nucleophile hydrolases.

Peptidase family U34 consists of enzymes with unclear catalytic mechanism, for instance, dipeptidase A from Lactobacillus helveticus. Using extensive sequence similarity searches, we infer that U34 family members are homologous to penicillin V acylases (PVA) and thus potentially adopt the N-terminal nucleophile (Ntn) hydrolase fold. Comparative sequence and structural analysis reveals a cysteine as the catalytic nucleophile as well as other conserved residues important for catalysis. The PVA/U34 family is variable in sequence and exhibits great diversity in substrate specificity, to include enzymes such as choloyglycine hydrolases, acid ceramidases, isopenicillin N acyltransferases, and a subgroup of eukaryotic proteins with unclear function.

Amidohydrolases↗

C-terminal domain of gyrase A is predicted to have a beta-propeller structure.

Two different type II topoisomerases are known in bacteria. DNA gyrase (Gyr) introduces negative supercoils into DNA. Topoisomerase IV (Par) relaxes DNA supercoils. GyrA and ParC subunits of bacterial type II topoisomerases are involved in breakage and reunion of DNA. The spatial structure of the C-terminal fragment in GyrA/ParC is not available. We infer homology between the C-terminal domain of GyrA/ParC and a regulator of chromosome condensation (RCC1), a eukaryotic protein that functions as a guanine-nucleotide-exchange factor for the nuclear G protein Ran. This homology, complemented by detection of 6 sequence repeats with 4 predicted beta-strands each in GyrA/ParC sequences, allows us to predict that the GyrA/ParC C-terminal domain folds into a 6-bladed beta-propeller. The prediction rationalizes available experimental data and sheds light on the spatial properties of the largest topoisomerase domain that lacks structural information.

Amino Acid Sequence↗

Breaking the singleton of germination protease.

Germination protease (GPR) plays an important role in the germination of spores of Bacillus and Clostridium species. A few very similar GPRs form a singleton group without significant sequence similarities to any other proteins. Their active site locations and catalytic mechanisms are unclear, despite the recent 3-D structure determination of Bacillus megaterium GPR. Using structural comparison and sequence analysis, we show that GPR is homologous to bacterial hydrogenase maturation protease (HybD). HybD's activity relies on the recognition and binding of metal ions in Ni-Fe hydrogenase, its substrate. Two highly conserved motifs are shared among GPRs, hydrogenase maturation proteases, and another group of hypothetical proteins. Conservation of two acidic residues in all these homologs indicates that metal binding is important for their function. Our analysis helps localize the active site of GPRs and provides insight into the catalytic mechanisms of a superfamily of putative metal-regulated proteases.

Amino Acid Sequence↗