CD28/CTLA-4 receptor structure, binding stoichiometry and aggregation during T-cell activation.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to J Bajorath.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Molecular modeling was used to build a three-dimensional model of the variable regions of the tumor-reactive monoclonal antibody BR96. An immunoconjugate of this antibody with the anticancer drug doxorubicin is currently in a phase I clinical trial for the treatment of solid tumors. A model structure of the BR96 variable fragment was generated to guide site-specific mutagenesis experiments and further improve the affinity of the antibody. The model displayed a distinct groove-type binding site which contained a significant number of aromatic residues. The dimensions and nature of the proposed binding site were consistent with the binding of the Ley tetrasaccharide which was found to bind to BR96. On the basis of the model, some BR96 residues are proposed to be crucial for antigen binding. BR96 and its complex with the Ley determinant have recently been crystallized, and structure determination is currently underway. Therefore, the detailed prediction of the BR96 combining site will soon be assessed, as a "blind test", based on crystallographic data.
E- and P-selectin are cell adhesion molecules implicated in the early events of inflammation. Three-dimensional models of the lectin domains have been reported by us and others prior to the availability of X-ray structural information. The models have been used to outline the ligand binding site in the selectins and to identify residues critical for function. Recently, the crystal structure of E-selectin has been reported, and thus, comparison of our E-selectin model with the X-ray data is now possible. The comparison shows that the assumptions on which the modeling was based were generally correct and provides an instructive example for the opportunities and the limitations of comparative modeling.
Explore the source record for details and available documents.
Mini-fingerprints (MFPs) are short binary bit string representations of molecular structure and properties, composed of few selected two-dimensional (2D) descriptors and a number of structural keys. MFPs were specifically designed to recognize compounds with similar activity. Here we report that MFPs are capable of detecting similar activities of some druglike molecules, including endothelin A antagonists and alpha(1)-adrenergic receptor ligands, the recognition of which was previously thought to depend on the use of multiple point three-dimensional (3D) pharmacophore methods. Thus, in these cases, MFPs and pharmacophore fingerprints produce similar results, although they define, in terms of their complexity, opposite ends of the spectrum of methods currently used to study molecular similarity or diversity. For each of the studied compound classes, comparison of MFP bit settings identified a consensus or signature pattern. Scaling factors can be applied to these bits in order to increase the probability of finding compounds with similar activity by virtual screening.
Results of systematic virtual screening calculations using a structural key-type fingerprint are reported for compounds belonging to 14 activity classes added to randomly selected synthetic molecules. For each class, a fingerprint profile was calculated to monitor the relative occupancy of fingerprint bit positions. Consensus bit patterns were determined consisting of all bits that were always set on in compounds belonging to a specific activity class. In virtual screening calculations, scale factors were applied to each consensus bit position in fingerprints of query molecules. This technique, called "fingerprint scaling", effectively increases the weight of consensus bit positions in fingerprint comparisons. Although overall prediction accuracy was satisfactory using unscaled calculations, scaling significantly increased the number of correct predictions but only slightly increased the rate of false positives. These observations suggest that fingerprint scaling is an attractive approach to increase the probability of identifying molecules with similar activity by virtual screening. It requires the availability of a series of related compounds and can be easily applied to any keyed fingerprint representation that associates bit positions with specific molecular features.
We have evaluated combinations of 111 descriptors that were calculated from two-dimensional representations of molecules to classify 455 compounds belonging to seven biological activity classes using a method based on principal component analysis. The analysis was facilitated by application of a genetic algorithm. Using scoring functions that related the number of compounds in pure classes (i.e., compounds with the same biological activity), singletons, and mixed classes, effective descriptor sets were identified. A combination of only four molecular descriptors accounting for aromatic character, hydrogen bond acceptors, estimated polar van der Waals surface area, and a single structural key gave overall best results. At this performance level, approximately 91% of the compounds occurred in pure classes and mixed classes were absent. The results indicate that combinations of only a few critical descriptors are preferred to partition compounds according to their biological activity, at least in the test cases studied here.
Combinations of 65 preferred 1D/2D molecular descriptors and 143 single structural keys were evaluated for their performance in compound classification focused on biological activity. The analysis was based on principal component analysis of descriptor combinations and facilitated by use of a genetic algorithm and different scoring functions. In these calculations, several descriptor combinations with greater than 95% prediction accuracy were identified. A set of 40 preferred structural keys was incorporated into a small binary fingerprint designed to search databases for compounds with biological activity similar to query molecules. The performance of mini-fingerprints was tested by systematic similarity search calculations in a database consisting of compounds belonging to seven biological activity classes, which had not been used to select effective descriptors. In these blind test calculations, mini-fingerprints correctly identified approximately 54% of compounds sharing similar biological activity and with 1% false positives. Thus, although the design of mini-fingerprints is conceptually simple, they perform well in activity-oriented similarity searching.
Molecular descriptors were identified by Shannon entropy analysis that correctly distinguished, in binary QSAR calculations, between naturally occurring molecules and synthetic compounds. The Shannon entropy concept was first used in digital communication theory and has only very recently been applied to descriptor analysis. Binary QSAR methodology was originally developed to correlate structural features and properties of compounds with a binary formulation of biological activity (i.e., active or inactive) and has here been adapted to correlate molecular features with chemical source (i.e., natural or synthetic). We have identified a number of molecular descriptors with significantly different Shannon entropy and/or "entropic separation" in natural and synthetic compound databases. Different combinations of such descriptors and variably distributed structural keys were applied to learning sets consisting of natural and synthetic molecules and used to derive predictive binary QSAR models. These models were then applied to predict the source of compounds in different test sets consisting of randomly collected natural and synthetic molecules, or, alternatively, sets of natural and synthetic molecules with specific biological activities. On average, greater than 80% prediction accuracy was achieved with our best models. For the test case consisting of molecules with specific activities, greater than 90% accuracy was achieved. From our analysis, some chemical features were identified that systematically differ in many naturally occurring versus synthetic molecules.
A method termed Differential Shannon Entropy (DSE) is introduced to compare differences in information content and variance of molecular descriptors between compound databases. The analysis is based on histograms recording the individual and grouped distributions of molecular descriptors and calculation of Shannon entropy (SE), a formalism originally applied to digital communication. We have recently shown that SE values reflect the nonparametric variability of descriptor settings. Now the analysis has been advanced to assess differences in information content of 143 molecular descriptors in databases containing synthetic compounds, natural products, or drug-like molecules. The DSE metric captures the degree to which descriptor distributions complement or duplicate information contained in molecular databases. In our analysis, we observe significant differences for a number of descriptors and rank them according to their associated DSE values. Using DSE calculations, relative information content of different types of descriptors can be quantified, even if differences are subtle.
The use of high throughput screening (HTS) to identify lead compounds has greatly challenged conventional quantitative structure-activity relationship (QSAR) techniques that typically correlate structural variations in similar compounds with continuous changes in biological activity. A new QSAR-like methodology that can correlate less quantitative assay data (i.e., "active" versus "inactive"), as initially generated by HTS, has been introduced. In the present study, we have, for the first time, applied this approach to a drug discovery problem; that is, the study of the estrogen receptor ligands. The binding affinities of 463 estrogen analogues were transformed into a binary data format, and a predictive binary QSAR model was derived using 410 estrogen analogues as a training set. The model was applied to predict the activity of 53 estrogen analogues not included in the training set. An overall accuracy of 94% was obtained.
In an effort to identify biologically active molecules in compound databases, we have investigated similarity searching using short binary bit strings with a maximum of 54 bit positions. These "minifingerprints" (MFPs) were designed to account for the presence or absence of structural fragments and/or aromatic character, flexibility, and hydrogen-bonding capacity of molecules. MFP design was based on an analysis of distributions of molecular descriptors and structural fragments in two large compound collections. The performance of different MFPs and a reference fingerprint was tested by systematic "one-against-all" similarity searches of molecules in a database containing 364 compounds with different biological activities. For each fingerprint, the most effective similarity cutoff value was determined. An MFP accounting for only 32 structural fragments showed less than 2% false positive similarity matches and correctly assigned on average approximately 40% of the compounds with the same biological activity to a query molecule. Inclusion of three numerical two-dimensional (2D) molecular descriptors increased the performance by 15%. This MFP performed better than a complex 2D fingerprint. At a similarity cutoff value of 0.85, the 2D fingerprint totally eliminated false positives but recognized less than 10% of the compounds within the same activity class.
Binary and conventional 2D QSAR have been derived for a set of carbonic anhydrase II (CA II) inhibitors. An overall predictive accuracy of 94% was obtained by binary QSAR and of 84% by 2D QSAR model. For both models, preferred molecular descriptor sets were identified, which were overlapping but not identical. Both binary and 2D QSAR captured important molecular features of CA II inhibitors, notably the presence of a sulfonamido group, which is critical for binding, but also hydrophobicity. Promising results were obtained when the derived QSAR models were used to test a set of CA II inhibitors not included in the training set. In binary QSAR, previously unobserved boundary effects were detected both in the analysis of known inhibitors and when screening a large combinatorial library for putative inhibitors. The complementary use of binary and conventional 2D QSAR is thought to increase the accuracy of the lead discovery process by QSAR techniques.
Explore the source record for details and available documents.