Search PubMed⌕ Search

Biomedical subjects

S Batalov

Publications and source records attributed to S Batalov.

4 recordsLinked to original sources

Functional annotation of a full-length mouse cDNA collection.

The RIKEN Mouse Gene Encyclopaedia Project, a systematic approach to determining the full coding potential of the mouse genome, involves collection and sequencing of full-length complementary DNAs and physical mapping of the corresponding genes to the mouse genome. We organized an international functional annotation meeting (FANTOM) to annotate the first 21,076 cDNAs to be analysed in this project. Here we describe the first RIKEN clone collection, which is one of the largest described for any organism. Analysis of these cDNAs extends known gene families and identifies new ones.

Animals↗

Estimating local backbone structural deviation in homology models.

After the atomic coordinates themselves, the most important data in a homology model are the spatial reliability estimates associated with each of the atoms (atom annotation). Recent blind homology modeling predictions have demonstrated that principally correct sequence-structure alignments are achievable to sequence identities as low as 25% [Martin, A.C., MacArthur, M.W., Thornton, J.M., 1997. Assessment of comparative modeling in CASP2. Proteins Suppl(1), 14-28]. The locations and extent of spatial deviations in the backbone between correctly aligned homologous protein structures remained very poorly estimated however, and these errors were the cause of errant loop predictions [Abagyan, R., Batalov, S., Cardozo, T., Totrov, M., Webber, J., Zhou, Y., 1997. Homology modeling with internal coordinate mechanics: deformation zone mapping and improvements of models via conformational search. Proteins Suppl(1), 29-37]. In order to derive accurate measures for local backbone deviations, we made a systematic study of static local backbone deviations between homologous pairs of protein structures. We found that 'through space' proximity to gaps and chain termini, local three-dimensional 'density', three-dimensional environment conservation, and B-factor of the template contribute to local deviations in the backbone in addition to local sequence identity. Based on these finding, we have identified the meaningful ranges of values within which each of these parameters correlates with static local backbone deviation and produced a combined scoring function to greatly improve the estimation of local backbone deviations. The optimized function has more than twice the accuracy of local sequence identity or B-factor alone and was validated in a recent blind structure prediction experiment. This method may be used to evaluate the utility of a preliminary homology model for a particular biological investigation (e.g. drug design) or to provide an improved starting point for molecular mechanics loop prediction methods.

Computer Simulation↗

Do aligned sequences share the same fold?

Sequence comparison remains a powerful tool to assess the structural relatedness of two proteins. To develop a sensitive sequence-based procedure for fold recognition, we performed an exhaustive global alignment (with zero end gap penalties) between sequences of protein domains with known three-dimensional folds. The subset of 1.3 million alignments between sequences of structurally unrelated domains was used to derive a set of analytical functions that represent the probability of structural significance for any sequence alignment at a given sequence identity, sequence similarity and alignment score. Analysis of overlap between structurally significant and insignificant alignments shows that sequence identity and sequence similarity measures are poor indicators of structural relatedness in the "twilight zone", while the alignment score allows much better discrimination between alignments of structurally related and unrelated sequences for a wide variety of alignment settings. A fold recognition benchmark was used to compare eight different substitution matrices with eight sets of gap penalties. The best performing matrices were Gonnet and Blosum50 with normalized gap penalties of 2.4/0.15 and 2.0/0.15, respectively, while the positive matrices were the worst performers. The derived functions and parameters can be used for fold recognition via a multilink chain of probability weighted pairwise sequence alignments.

Databases as Topic↗

Homology modeling with internal coordinate mechanics: deformation zone mapping and improvements of models via conformational search.

Five models by homology containing insertions and deletions and ranging from 33% to 48% sequence identity to the known homologue, and one high sequence identity (85%) model were built for the CASP2 meeting. For all five low identity targets: (i) our starting models were improved by the Internal Coordinate Mechanics (ICM) energy optimization, (ii) the refined models were consistently better than those built with the automatic SWISS-MODEL program, and (iii) the refined models differed by less than 2% from the best model submitted, as judged by the residue contact area difference (CAD) measure [Abagyan, R.A., Totrov, M.J. Mol. Biol. 268:678-685, 1997]. The CAD measure is proposed for ranking models built by homology instead of global root-mean-square deviation, which is frequently dominated by insignificant yet large contributions from incorrectly predicted fragments or side chains. We demonstrate that the precise identification of regions of local backbone deviation is an independent and crucial step in the homology modeling procedure after alignment, since aligned fragments can strongly deviate from the template at various distances from the alignment gap or even in the ungapped parts of the alignment. We show that a local alignment score can be used as an indicator of such local deviation. While four short loops of the meeting targets were predicted by database search, the best loop 1 target T0028, for which the correct database fragment was not found, was predicted by Internal Coordinate Mechanics global energy optimization at 1.2 A accuracy. A classification scheme for errors in homology modeling is proposed.

Models, Molecular↗