Search PubMed⌕ Search

Biomedical subjects

J Bajorath

Publications and source records attributed to J Bajorath.

At least 19 recordsLinked to original sources

Computational techniques for diversity analysis and compound classification.

Molecular similarity and diversity analysis has played a significant role in computer-aided drug discovery for more than a decade. Compound classification methods have also become increasingly important for the design and organization of compound databases and in silico screening. Here we review these related methodologies and discuss selected applications.

Cluster Analysis↗

Mini-fingerprints for virtual screening: design principles and generation of novel prototypes based on information theory.

Binary fingerprint representations of molecular structure and properties are convenient computational tools for similarity searching in compound databases and virtual screening (VS). We are investigating the design of relatively simple fingerprints for the identification of molecules having similar biological activity and recognition of remote similarity relationships. Since our designs are considerably shorter than other fingerprints used in VS, we have previously termed them "mini-fingerprints" (MFPs). A key aspect of the design strategy is the identification of suitable molecular descriptors. Whereas our initial fingerprint designs have relied on descriptor combinations that performed well in compound classification according to biological activity, second generation MFPs encode combinations of descriptors with high information content in large compound databases and high frequency of occurrence in drug-like molecules. Thus, the design of these new fingerprints does not depend on the analysis of specific classes of bioactive compounds, but rather on descriptor information content in large compound databases. Systematic evaluation of fingerprint performance in VS test calculations demonstrates that these new prototypes perform better than previously generated MFPs. The analysis described herein provides an example for the development of search tools for VS.

Environmental Pollutants↗

Rational drug discovery revisited: interfacing experimental programs with bio- and chemo-informatics.

Over the past few years, bio- and chemo-informatics have rapidly evolved as related yet distinct disciplines. In drug discovery, it is increasingly recognized that combining and integrating these approaches is crucial for their successful application. In addition, the use of complementary experimental and informatics techniques increases the chances of success in many stages of the discovery process, from the identification of novel targets and elucidation of their functions to the discovery and development of lead compounds with desired properties. This review highlights recent trends that emphasize the role of integrated bio- and chemo-informatics research in drug discovery and discusses representative concepts and methodologies.

Journal Article↗

Cell surface receptors and their ligands: in vitro analysis of CD6-CD166 interactions.

CD6 is a cell surface receptor belonging to the scavenger receptor cysteine-rich (SRCR) protein superfamily (SRCRSF). It specifically binds activated leukocyte cell adhesion molecule (ALCAM, CD166), a member of the immunoglobulin (Ig) superfamily (IgSF). CD166 was among the first molecules identified as a ligand for an SRCRSF receptor, and the CD6-CD166 interaction was the first interaction characterized involving SRCRSF and IgSF proteins. We focus here on what has been learned about the specifics of the CD6-CD166 interaction from in vitro analysis. The studies are thought to provide an instructive example for the analysis of interactions between single-path transmembrane cell surface proteins. Using soluble recombinant forms, the extracellular binding domains of receptor and ligand have been identified and characterized in a variety of assay systems. Both CD6 and CD166 have been subjected to intense mutagenesis and monoclonal antibody (mAb) binding studies and residues critical for their interaction have been identified. The availability of structural prototypes of both superfamilies has made it possible to map the binding site in CD166 and, more recently, in CD6 and compare these regions to epitopes of mAbs that block, or do not block, the interaction. In addition, the molecular basis of observed cross-species receptor-ligand interactions could be rationalized. These studies illustrate the value of structural templates for the interpretation of sequence and mutagenesis analyses. Proteins 2000;40:420-428.

Activated-Leukocyte Cell Adhesion Molecule↗

Molecular organization, structural features, and ligand binding characteristics of CD44, a highly variable cell surface glycoprotein with multiple functions.

CD44 is a type I transmembrane protein and member of the cartilage link protein family. It is involved in cell-cell and cell-matrix interactions and signal transduction. Several CD44 ligands have been identified. CD44 is a major cell surface receptor for hyaluronan, a component of the extracellular matrix. It is implicated in diseases such as cancer and inflammation and therefore intensely studied. A characteristic feature of CD44 is the occurrence of many isoforms that are expressed in a cell-specific manner and differentially glycosylated. Although a number of CD44 isoforms have been characterized, the structural diversity of CD44 makes it often challenging to study (isoform-specific) CD44-ligand interactions at the molecular level of detail. The structural organization and ligand binding characteristics of CD44 are focal points of this review. On the basis of recent structural and mutagenesis studies, details of the CD44-hyaluronan interaction are beginning to be understood. Proteins 2000;39:103-111.

Humans↗

Variability of molecular descriptors in compound databases revealed by Shannon entropy calculations

A method is introduced to calculate and compare the variability of molecular descriptors in compound databases. Descriptor variability analysis is based on histograms recording the distribution of molecular descriptors and calculation of Shannon entropy (SE), a metric originally applied in digital communication. SE values reflect the variability of descriptor settings. We have calculated a total of 92 molecular descriptors in the ACD and NCI databases and ranked them according to their variability. Significant differences in entropy are observed for a number of descriptors. However, the most variable descriptors are similar in the ACD and NCI databases. Such high-entropy descriptors are preferred tools to discriminate between compounds or account for the diversity of chemical libraries.

Journal Article↗

Combinatorial preferences affect molecular similarity/diversity calculations using binary fingerprints and Tanimoto coefficients

A combinatorial method was developed to calculate complete distributions of the Tanimoto coefficient (Tc) for binary fingerprint (FP) representations of specified length, regardless of the chemical parameters they reflect. Theoretical Tc distributions were calculated for FPs consisting of up to 67 bit positions which revealed significant statistical preferences of certain Tc values. Calculation of Tc distributions in a large compound database using different FPs mirrored the effects identified by our general analysis. On the basis of these findings, an average Tc is biased by statistically preferred values.

Journal Article↗

Searching for molecules with similar biological activity: analysis by fingerprint profiling.

We have recently developed a mini-fingerprint (MFP) representation for small molecules that performs well in database searches for compounds with similar biological activity. The MFP consists of only 54 bit positions that account for numerical ranges of three two-dimensional (2D) descriptors or the presence or absence of defined structural fragments. Here we present an analysis method, termed fingerprint profiling, to systematically compare bit patterns of compounds belonging to different biological activity classes. Some but not all bit positions were variably occupied in seven different activity classes and responsible for the detection of structure-activity differences. The analysis has made it possible to rank bit positions and encoded molecular descriptors according to their importance for our similarity search calculations. Fingerprint profiling can be applied to any keyed bit string representation and should be helpful, for example, to analyze descriptor distributions in large compound databases.

Computer-Aided Design↗

Molecular descriptors in chemoinformatics, computational combinatorial chemistry, and virtual screening.

Many contemporary applications in computer-aided drug discovery and chemoinformatics depend on representations of molecules by descriptors that capture their structural characteristics and properties. Such applications include, among others, diversity analysis, library design, and virtual screening. Hundreds of molecular descriptors have been reported in the literature, ranging from simple bulk properties to elaborate three-dimensional formulations and complex molecular fingerprints, which sometimes consist of thousands of bit positions. Knowledge-based selection of descriptors that are suitable for specific applications is an important task in chemoinformatics research. If descriptors are to be selected on rational grounds, rather than guesses or chemical intuition, detailed evaluation of their performance is required. A number of studies have been reported that investigate the performance of molecular descriptors in specific applications and/or introduce novel types of descriptors. Progress made in this area is reviewed here in the context of other computational developments in combinatorial chemistry and compound screening.

Algorithms↗

Identification of the ligand binding site in Fas (CD95) and analysis of Fas-ligand interactions.

Fas (CD95), a member of the tumor necrosis factor receptor superfamily, and its ligand (FasL), a tumor necrosis factor-like protein, are intensely studied because their interaction on the cell surface is critical for the induction of programmed cell death (apoptosis) and the regulation of immune responses. The structure and specificity of the extracellular binding domains of Fas and its ligand were studied, in different laboratories, by combining molecular modeling, mutagenesis, and a variety of binding and functional experiments. Residues critical for the receptor-ligand interaction were identified and, in the absence of experimentally determined structures, binding sites and details of the Fas-ligand interactions were predicted. These studies provide an instructive example for the close combination of prediction and experiment and illustrate how insights into the structure and binding characteristics of Fas and its ligand were gradually refined. Discussed methodological aspects are representative of structure-function studies on extracellular domains of other single-path transmembrane proteins.

Amino Acid Sequence↗

Molecular scaffold-based design and comparison of combinatorial libraries focused on the ATP-binding site of protein kinases.

Compound libraries were designed to target specifically the ATP cofactor-binding site in protein kinases by combining knowledge- and diversity-based design elements. A key aspect of the approach is the identification of molecular building blocks or scaffolds that are compatible with the binding site and therefore capture some aspects of target specificity. Scaffolds were selected on the basis of docking calculations and analysis of known inhibitors. We have generated 75 molecular scaffolds and applied different strategies to compute diverse compounds from scaffolds or, alternatively, to screen compound databases for molecules containing these scaffolds. The resulting libraries had a similar degree of molecular diversity, with at most 12% of the compounds being identical. However, their scaffold distributions differed significantly and a small number of scaffolds dominated the majority of compounds in each library.

Adenosine Triphosphate↗

Analysis of Fas-ligand interactions using a molecular model of the receptor-ligand interface.

A molecular model of the complex between Fas and its ligand was generated to better understand the location and putative effects of site-specific mutations, analyze interactions at the Fas-FasL interface, and identify contact residues. The modeling study was conservative in the sense that regions in Fas and its ligand which could not be predicted with confidence were omitted from the model to ensure accuracy of the analysis. Using the model, it was possible to map four of five N-linked glycosylation sites in Fas and FasL and to study 10 of 11 residues previously identified by mutagenesis as important for binding. Interactions involving six of these residues could be analyzed in detail and their importance for binding was rationalized based on the model. The predicted structure of the Fas-FasL interface was consistent with the experimentally established importance of these residues for binding. In addition, five previously not targeted residues were identified and predicted to contribute to binding via electrostatic interactions. Despite its limitations, the study provided a much improved basis to understand the role of Fas and FasL residues for binding compared to previous residue mapping studies using only a molecular model of Fas.

Amino Acid Sequence↗

Escherichia coli and Porphyromonas gingivalis lipopolysaccharide interactions with CD14: implications for myeloid and nonmyeloid cell activation.

Porphyromonas gingivalis, a gram-negative bacterium, is an etiologic agent for adult periodontitis. Lipopolysaccharide (LPS) released from this bacterium can react with numerous host cell types. P. gingivalis LPS stimulates tumor necrosis factor alpha and interleukin-1beta secretion from monocytes (myeloid) but does not elicit E-selectin expression from human endothelial cells (nonmyeloid). In contrast, Escherichia coli LPS facilitates expression of these inflammatory mediators through CD14-dependent pathways on both myeloid and nonmyeloid cells. LPS binding studies have revealed that although P. gingivalis and E. coli LPSs bind to CD14 differently, this fact does not adequately explain the lack of endothelial cell activation by P. gingivalis LPS. Rather, LPS binding site and blocking monoclonal antibody epitope mapping studies have suggested that CD14 presents a charged surface that captures different microbial ligands by electrostatic interactions. We propose that human endothelial cells do not respond to P. gingivalis LPS because of their inability to "recognize" CD14-P. gingivalis LPS complexes.

Amino Acid Sequence↗

Detailed comparison of two molecular models of the human CD40 ligand with an x-ray structure and critical assessment of model-based mutagenesis and residue mapping studies.

The interactions between the B cell receptor CD40 and its ligand on T cells are critical for the integrity of immune responses. The human CD40 ligand gp39, a tumor necrosis factor-like protein, has been the subject of intense efforts to identify the receptor-binding site and to analyze naturally occurring mutations that compromise gp39 function in vivo. These investigations relied heavily on molecular models of gp39, built in the presence of only approximately 25% sequence identity to tumor necrosis factor. The x-ray structure of gp39 has made it possible to assess modeling accuracy and to evaluate the results of model-based mutagenesis analyses. Although the models display local errors, their accuracy was sufficient to predict the CD40-binding site, to map natural mutations, and to rationalize their effects. One of five gp39 residues critical for CD40 binding was displaced in the models, and 1 of 21 point mutants was incorrectly classified. Factors most important for the reliability of the molecular models and their successful applications were valid sequence alignments and the focus of experimental studies on regions of high prediction confidence. Analysis of mutagenesis experiments correlated with anti-gp39 monoclonal antibody binding studies to assess the conformational integrity of mutant proteins.

Amino Acid Sequence↗