Search PubMedSearch

Biomedical subjects

N N Alexandrov

Publications and source records attributed to N N Alexandrov.

10 recordsLinked to original sources

Analysis of topological and nontopological structural similarities in the PDB: new examples with old structures.

We have developed a new method and program, SARF2, for fast comparison of protein structures, which can detect topological as well as nontopological similarities. The method searches for large ensembles of secondary structure elements, which are mutually compatible in two proteins. These ensembles consist of small fragments of C alpha-trace, similarly arranged in three-dimensional space in two proteins, but not necessarily equally-ordered along the polypeptide chains. The program SARF2 is available for everyone through the World-Wide Web (WWW). We have performed an exhaustive pairwise comparison of all the entries from a recent issue of the Protein Data Bank (PDB) and report here on the results of an automated hierarchical cluster analysis. In addition, we report on several new cases of significant structural resemblance between proteins. To this end, a new definition of the significance of structural similarity is introduced, which effectively distinguishes the biologically meaningful equivalences from those occurring by chance. Analyzing the distribution of sequence similarity in significant structural matches, we show that sequence similarity as low as 20% in structurally-prealigned proteins can be a strong indication for the biological relevance of structural similarity.

Algorithms

SARFing the PDB.

Fast growth of the number of the solved protein structures is increasing the role of their comparative analysis. In this paper I describe a new program, SARF2, for protein structure comparison and discuss new examples of the non-topological structural resemblance. SARF2 is designed to detect ensembles of secondary structure elements, which form similar spatial arrangements with possible different topological connections. The program is available to everyone through the World Wide Web (URL http:@www-lmmb.ncifcrf.gov/approximately nicka/sarf2.html). The performance of the program is demonstrated by previously unnoticed cases of the significant similarities. One similarity discussed in this paper, between heme-binding proteins (cytochrome P450 and globin), consists of six alpha-helices, arranged into a globin fold. Another pair of structures (pectate lyase and snowdrop lectin) achieve similar beta-prism architecture through different topologies. The significance of these similarities is validated by (i) the distribution of a similarity score, (ii) the comparison of the aligned contact maps and/ or (iii) the location of the active site. The observation of recurrent non-topological structural motifs implies their energetic stability and opens new possibilities for sequence-structure alignment (threading) methods.

Computer Communication Networks

DNASUN: a package of computer programs for the biotechnology laboratory.

The paper describes a new software package DNASUN developed for supporting gene engineering laboratories. The package provides a user-friendly interface for experimental researches and supports the traditional nucleotide/protein sequence analysis as well as physical mapping, sequencing, plasmid manipulations, optimal oligonucleotide probe selection and other common molecular biology procedures.

Biotechnology

Biological meaning, statistical significance, and classification of local spatial similarities in nonhomologous proteins.

We have completed an exhaustive search for the common spatial arrangements of backbone fragments (SARFs) in nonhomologous proteins. This type of local structural similarity, incorporating short fragments of backbone atoms, arranged not necessarily in the same order along the polypeptide chain, appears to be important for protein function and stability. To estimate the statistical significance of the similarities, we have introduced a similarity score. We present several locally similar structures, with a large similarity score, which have not yet been reported. On the basis of the results of pairwise comparison, we have performed hierarchical cluster analysis of protein structures. Our analysis is not limited by comparison of single chains but also includes complex molecules consisting of several subunits. The SARFs with backbone fragments from different polypeptide chains provide a stable interaction between subunits in protein molecules. In many cases the active site of enzyme is located at the same position relative to the common SARFs, implying a function of the certain SARFs as a universal interface of the protein-substrate interaction.

Binding Sites

Common spatial arrangements of backbone fragments in homologous and non-homologous proteins.

We have developed a new method of detecting common spatial arrangements of backbone fragments in proteins. This method allows corresponding fragments to occur in a different order in respective amino acid sequences. We applied this method to detect structural similarities between an acid protease, endothiapepsin, and all other proteins in the protein data bank. Significant similarities were found not only with other acid proteases but also with virus proteases and with proteins having different functions. The possible biological meaning of these similarities is discussed.

Aspartic Acid Endopeptidases

Local multiple alignment by consensus matrix.

A new algorithm for aligning several sequences based on the calculation of a consensus matrix and the comparison of all the sequences using this consensus matrix is described. This consensus matrix contains the preference scores of each nucleotide/amino acid and gaps in every position of the alignment. Two modifications of the algorithm corresponding to the evolutionary and functional meanings of the alignment were developed. The first one solves the best-fitting problem without any penalty for end gaps and with an internal gap penalty function independent on the gap length. This algorithm should be used when comparing evolutionary-related proteins for identifying the most conservative residues. The other modification of the algorithm finds the most similar segments in the given sequences. It can be used for finding those parts of the sequences that are responsible for the same biological function. In this case the gap penalty function was chosen to be proportional to the gap length. The result of aligning amino acid sequences of neutral proteases and a compilation of 65 allosteric effectors and substrates of PEP carboxylase are presented.

Algorithms

Application of a new method of pattern recognition in DNA sequence analysis: a study of E. coli promoters.

An algorithm from the pattern recognition theory 'generalized portrait' was used to find a distinguishing vector (scoring matrix) for E. coli promoters. We have attempted to solve three closely linked problems: (i) the selection of significant features of the signal; (ii) subsequent multiple alignment and (iii) calculation of the vector coordinates. Promoters with known strength have been successfully ranked in the correct order using this vector. We demonstrate the use of this method in predicting the location of promoters. A revised consensus promoter sequence is also presented.

Algorithms

Statistical method for rapid homology search.

A new method for homology search of DNA sequences is suggested. This method may be used to find extensive and not strong homologies with point mutations and deletions. The running program time for comparing sequences is less then the dynamic program algorithms at least at two orders of magnitude. It makes possible to use the method for homology searching throughover the nucleotide bank by personal computers.

Algorithms