Search PubMed⌕ Search

Biomedical subjects

Chihae Yang

Publications and source records attributed to Chihae Yang.

8 recordsLinked to original sources

Comparison of methods for sequential screening of large compound sets.

Sequential screening is an iterative procedure that can greatly increase hit rates over random screening or non-iterative procedures. We studied the effects of three factors on enrichment rates: the method used to rank compounds, the molecular descriptor set and the selection of initial training set. The primary factor influencing recovery rates was the method of selecting the initial training set. Rates for recovering active compounds were substantially lower with the diverse training sets than they were with training sets selected by other methods. Because structure-activity information is incrementally enhanced in intermediate training sets, sequential screening provides significant improvement in the average rate of recovery of active compounds when compared with non-iterative selection procedures.

Chemistry, Pharmaceutical↗

Landscape of current toxicity databases and database standards.

Having readily available historical information for modeling toxicity has become important throughout the various stages of research and development. The high cost of late-phase attrition and recent international regulatory legislations have made even more acute the need to be able to mine the fragmented data and information available across diverse databases. In addition, the general trend to accelerate regulatory processes globally makes the effective use of existing data an imperative. To achieve efficient screening, develop profiles and gain the ability to cross reference, databases must be interoperated to allow data exchange and integration. Several database standards and controlled vocabulary initiatives have been used in the development of toxicity data models to transform the current landscape. This review describes the major databases of toxicological information now available, and provides a simple example of standardization that demonstrates the benefits of a toxicity database containing such qualified data.

Animals↗

Chemical effects in biological systems--data dictionary (CEBS-DD): a compendium of terms for the capture and integration of biological study design description, conventional phenotypes, and 'omics data.

A critical component in the design of the Chemical Effects in Biological Systems (CEBS) Knowledgebase is a strategy to capture toxicogenomics study protocols and the toxicity endpoint data (clinical pathology and histopathology). A Study is generally an experiment carried out during a period of time for the purpose of obtaining data, and the Study Design Description captures the methods, timing, and organization of the Study. The CEBS Data Dictionary (CEBS-DD) has been designed to define and organize terms in an attempt to standardize nomenclature needed to describe a toxicogenomics Study in a structured yet intuitive format and provide a flexible means to describe a Study as conceptualized by the investigator. The CEBS-DD will organize and annotate information from a variety of sources, thereby facilitating the capture and display of toxicogenomics data in biological context in CEBS, i.e., associating molecular events detected in highly-parallel data with the toxicology/pathology phenotype as observed in the individual Study Subjects and linked to the experimental treatments. The CEBS-DD has been developed with a focus on acute toxicity studies, but with a design that will permit it to be extended to other areas of toxicology and biology with the addition of domain-specific terms. To illustrate the utility of the CEBS-DD, we present an example of integrating data from two proteomics and transcriptomics studies of the response to acute acetaminophen toxicity (A. N. Heinloth et al., 2004, Toxicol. Sci. 80, 193-202).

Acetaminophen↗

Building predictive models for protein tyrosine phosphatase 1B inhibitors based on discriminating structural features by reassembling medicinal chemistry building blocks.

A new approach to predicting the biological activity of small molecule pharmaceutics is demonstrated. Structural features of medicinal chemistry building blocks are used as 2-D molecular descriptors. These descriptors include predefined structural features and macrostructures obtained from a supervised process in which features in the core library are reassembled to provide larger features that strongly differentiate the desired biological response variable. Chemical features derived in this manner can serve as predictor variables for diverse modeling algorithms, and application using partial least squares techniques is demonstrated here. Models are presented for inhibition by benzofuran and benzothiophene biphenyl analogues of protein tyrosine phosphatase 1B (PTP1B), a target for insulin-resistant disease states. Results are compared to models for PTP1B inhibitors available in the literature based on CoMFA-related techniques and 3-D molecular descriptors.

Models, Molecular↗

Systematic analysis of large screening sets in drug discovery.

Each year large pharmaceutical companies produce massive amounts of primary screening data for lead discovery. To make better use of the vast amount of information in pharmaceutical databases, companies have begun to scrutinize the lead generation stage to ensure that more and better qualified lead series enter the downstream optimization and development stages. This article describes computational techniques for end to end analysis of large drug discovery screening sets. The analysis proceeds in three stages: In stage 1 the initial screening set is filtered to remove compounds that are unsuitable as lead compounds. In stage 2 local structural neighborhoods around active compound classes are identified, including similar but inactive compounds. In stage 3 the structure-activity relationships within local structural neighborhoods are analyzed. These processes are illustrated by analyzing two large, publicly available databases.

Algorithms↗

Finding discriminating structural features by reassembling common building blocks.

We present a new method for constructing discriminating substructures by reassembling common medicinal chemistry building blocks. The algorithm can be parametrized to meet differing objectives: (1) to build features that discriminate for biological activity in a local structural neighborhood, (2) to build scaffolds for R-group analysis, (3) to construct cluster signatures that discriminate for membership in the cluster and provide a graphical representation for its members, and (4) to identify substructures that characterize major classes in a heterogeneous compound set. We illustrated the results of the algorithm on a literature dataset is of 118 compounds with in vitro inhibition data against recombinant human protein tyrosine phosphatase 1B (PTP-1B).

Algorithms↗

Multiscale and Bayesian approaches to data analysis in genomics high-throughput screening.

Tremendous amounts of data are produced by high-throughput screening methods currently employed in drug discovery and product development. A typical cDNA microarray or oligonucleotide-based gene chip experiment easily generates over 10,000 data points for each array or chip. The challenge of inferring meaningful information is formidable given the size and number of these datasets. This paper reviews the current status of statistical tools available for gene expression analysis, with emphasis on Bayesian approaches and multiscale wavelet filtering. Fundamental concepts of Bayesian and multiscale modeling are discussed from the perspective of their potential to address important issues related to the analysis of gene expression data, such as the fact that genomic data often have non-Gaussian distributions and feature localization and multiple scales in both frequency and measurement dimension. Recent publications in these areas are reviewed. Wavelet filtering and the advantages of multiscale methods are demonstrated by application to publicly available gene expression data from the National Cancer Institute (NCI). Multiscale methods, including multiscale principal component analysis (MSPCA), are applied to extract gene subsets and to visualize data in multidimensions for comparisons. Similarity in cell lines and gene selection are effectively visualized and quantitatively compared.

Animals↗