Search PubMed⌕ Search

Biomedical subjects

F L Stahura

Publications and source records attributed to F L Stahura.

7 recordsLinked to original sources

Computational techniques for diversity analysis and compound classification.

Molecular similarity and diversity analysis has played a significant role in computer-aided drug discovery for more than a decade. Compound classification methods have also become increasingly important for the design and organization of compound databases and in silico screening. Here we review these related methodologies and discuss selected applications.

Cluster Analysis↗

Searching for molecules with similar biological activity: analysis by fingerprint profiling.

We have recently developed a mini-fingerprint (MFP) representation for small molecules that performs well in database searches for compounds with similar biological activity. The MFP consists of only 54 bit positions that account for numerical ranges of three two-dimensional (2D) descriptors or the presence or absence of defined structural fragments. Here we present an analysis method, termed fingerprint profiling, to systematically compare bit patterns of compounds belonging to different biological activity classes. Some but not all bit positions were variably occupied in seven different activity classes and responsible for the detection of structure-activity differences. The analysis has made it possible to rank bit positions and encoded molecular descriptors according to their importance for our similarity search calculations. Fingerprint profiling can be applied to any keyed bit string representation and should be helpful, for example, to analyze descriptor distributions in large compound databases.

Computer-Aided Design↗

Molecular scaffold-based design and comparison of combinatorial libraries focused on the ATP-binding site of protein kinases.

Compound libraries were designed to target specifically the ATP cofactor-binding site in protein kinases by combining knowledge- and diversity-based design elements. A key aspect of the approach is the identification of molecular building blocks or scaffolds that are compatible with the binding site and therefore capture some aspects of target specificity. Scaffolds were selected on the basis of docking calculations and analysis of known inhibitors. We have generated 75 molecular scaffolds and applied different strategies to compute diverse compounds from scaffolds or, alternatively, to screen compound databases for molecules containing these scaffolds. The resulting libraries had a similar degree of molecular diversity, with at most 12% of the compounds being identical. However, their scaffold distributions differed significantly and a small number of scaffolds dominated the majority of compounds in each library.

Adenosine Triphosphate↗

Site-specific recognition by an isolated DNA-binding domain of the sine oculis protein.

The sine oculis (so) gene is required for the development of the Drosophila visual system. The 416 amino acid SO protein contains a 40 amino acid region homologous to the helix-turn-helix (HtH) region of the homeodomain. Three HtH-containing peptides ranging in size from 63 to 93 amino acids (SO(218-279), SO(204-279), and SO(188-279)) were expressed in Escherichia coli and characterized in vitro. These fragments show circular dichroism spectra characteristic of helical proteins and cooperative unfolding transitions. Derivatization of these three peptides with the chemical nuclease 1,10-phenanthroline:copper (OP-Cu) allowed the identification of specific DNA-binding sites within the 3.1 kb pUC119 plasmid. Similar cleavage patterns with similar relative affinities were obtained for all three peptides. Nucleotide resolution mapping of the predominant cleavage area identified two primary cleavage sites with a similar core sequence. The DNA cleavage sites were confirmed by DNase I footprinting with both native and OP-Cu-conjugated SO HtH peptides. This study identifies a 63 amino acid peptide as sufficient for specific DNA binding.

Copper↗

Mini-fingerprints detect similar activity of receptor ligands previously recognized only by three-dimensional pharmacophore-based methods.

Mini-fingerprints (MFPs) are short binary bit string representations of molecular structure and properties, composed of few selected two-dimensional (2D) descriptors and a number of structural keys. MFPs were specifically designed to recognize compounds with similar activity. Here we report that MFPs are capable of detecting similar activities of some druglike molecules, including endothelin A antagonists and alpha(1)-adrenergic receptor ligands, the recognition of which was previously thought to depend on the use of multiple point three-dimensional (3D) pharmacophore methods. Thus, in these cases, MFPs and pharmacophore fingerprints produce similar results, although they define, in terms of their complexity, opposite ends of the spectrum of methods currently used to study molecular similarity or diversity. For each of the studied compound classes, comparison of MFP bit settings identified a consensus or signature pattern. Scaling factors can be applied to these bits in order to increase the probability of finding compounds with similar activity by virtual screening.

Angiotensin II↗

Fingerprint scaling increases the probability of identifying molecules with similar activity in virtual screening calculations.

Results of systematic virtual screening calculations using a structural key-type fingerprint are reported for compounds belonging to 14 activity classes added to randomly selected synthetic molecules. For each class, a fingerprint profile was calculated to monitor the relative occupancy of fingerprint bit positions. Consensus bit patterns were determined consisting of all bits that were always set on in compounds belonging to a specific activity class. In virtual screening calculations, scale factors were applied to each consensus bit position in fingerprints of query molecules. This technique, called "fingerprint scaling", effectively increases the weight of consensus bit positions in fingerprint comparisons. Although overall prediction accuracy was satisfactory using unscaled calculations, scaling significantly increased the number of correct predictions but only slightly increased the rate of false positives. These observations suggest that fingerprint scaling is an attractive approach to increase the probability of identifying molecules with similar activity by virtual screening. It requires the availability of a series of related compounds and can be easily applied to any keyed fingerprint representation that associates bit positions with specific molecular features.

Algorithms↗

Distinguishing between natural products and synthetic molecules by descriptor Shannon entropy analysis and binary QSAR calculations.

Molecular descriptors were identified by Shannon entropy analysis that correctly distinguished, in binary QSAR calculations, between naturally occurring molecules and synthetic compounds. The Shannon entropy concept was first used in digital communication theory and has only very recently been applied to descriptor analysis. Binary QSAR methodology was originally developed to correlate structural features and properties of compounds with a binary formulation of biological activity (i.e., active or inactive) and has here been adapted to correlate molecular features with chemical source (i.e., natural or synthetic). We have identified a number of molecular descriptors with significantly different Shannon entropy and/or "entropic separation" in natural and synthetic compound databases. Different combinations of such descriptors and variably distributed structural keys were applied to learning sets consisting of natural and synthetic molecules and used to derive predictive binary QSAR models. These models were then applied to predict the source of compounds in different test sets consisting of randomly collected natural and synthetic molecules, or, alternatively, sets of natural and synthetic molecules with specific biological activities. On average, greater than 80% prediction accuracy was achieved with our best models. For the test case consisting of molecules with specific activities, greater than 90% accuracy was achieved. From our analysis, some chemical features were identified that systematically differ in many naturally occurring versus synthetic molecules.

Algorithms↗