Search PubMed⌕ Search

Biomedical subjects

S Stanley Young

Publications and source records attributed to S Stanley Young.

6 recordsLinked to original sources

Testing association of statistically inferred haplotypes with discrete and continuous traits in samples of unrelated individuals.

There have been increasing efforts to relate drug efficacy and disease predisposition with genetic polymorphisms. We present statistical tests for association of haplotype frequencies with discrete and continuous traits in samples of unrelated individuals. Haplotype frequencies are estimated through the expectation-maximization algorithm, and each individual in the sample is expanded into all possible haplotype configurations with corresponding probabilities, conditional on their genotype. A regression-based approach is then used to relate inferred haplotype probabilities to the response. The relationship of this technique to commonly used approaches developed for case-control data is discussed. We confirm the proper size of the test under H(0) and find an increase in power under the alternative by comparing test results using inferred haplotypes with single-marker tests using simulated data. More importantly, analysis of real data comprised of a dense map of single nucleotide polymorphisms spaced along a 12-cM chromosomal region allows us to confirm the utility of the haplotype approach as well as the validity and usefulness of the proposed statistical technique. The method appears to be successful in relating data from multiple, correlated markers to response.

Algorithms↗

Optimization of focused chemical libraries using recursive partitioning.

A number of methods currently exist for designing chemical libraries. General or universal libraries use a measurement of chemical diversity in their design and seek to cover as much of chemical space as possible in order to maximize the likelihood of discovering a novel lead class of active compounds. Focused chemical libraries are then synthesized to expand on this particular class and thoroughly explore the space about it. Rarely, however, is relevant biological data tightly incorporated in the design of focused libraries. Recursive partitioning is a statistical technique that is used to quickly build SAR models from high-throughput screening data sets and associated chemical descriptors. Using these models in a virtual screening mode significantly increases the probability of finding other active compounds. The predicted activity can be also be used as the fitness function for a genetic algorithm that is designed to select monomer subsets having a higher probability of being active. This dramatically reduces the number of compounds that need to be synthesized in focused libraries thus saving considerable time, effort and expense. This paper describes how recursive partitioning models are used to optimize the design of focused chemical libraries.

Chemistry, Pharmaceutical↗

Initial compound selection for sequential screening.

Initial leads for drug development often originate from high-throughput screening (HTS), where hundreds of thousands of compounds are tested for biological activity. As the number of both targets for screening and compounds available for screening increase, there is a need to consider methods for making this process more efficient. One approach is to screen sequentially, whereby a relatively small set of compounds is assayed and the results are statistically analyzed to produce a mathematical model. The model is used to predict activity and select additional compounds for screening. The new compound bioassay results are added to the results for the initial set and a new model is determined. The process iterates. The focus of this review is on how to select the initial screening set (ISS). It is presumed that the size and quality of the initial set will affect the subsequent model building and, hence, the efficiency of finding active compounds.

Algorithms↗

The construction and assessment of a statistical model for the prediction of protein assay data.

The focus of this work is the development of a statistical model for a bioinformatics database whose distinctive structure makes model assessment an interesting and challenging problem. The key components of the statistical methodology, including a fast approximation to the singular value decomposition and the use of adaptive spline modeling and tree-based methods, are described, and preliminary results are presented. These results are shown to compare favorably to selected results achieved using comparitive methods. An attempt to determine the predictive ability of the model through the use of cross-validation experiments is discussed. In conclusion a synopsis of the results of these experiments and their implications for the analysis of bioinformatic databases in general is presented.

Computational Biology↗

A factorial design to optimize cell-based drug discovery analysis.

Drug discovery is dependent on finding a very small number of biologically active or potent compounds among millions of compounds stored in chemical collections. Quantitative structure-activity relationships suggest that potency of a compound is highly related to that compound's chemical makeup or structure. To improve the efficiency of cell-based analysis methods for high throughput screening, where information of a compound's structure is used to predict potency, we consider a number of potentially influential factors in the cell-based approach. A fractional factorial design is implemented to evaluate the effects of these factors, and lift chart results show that the design scheme is able to find conditions that enhance hit rates.

Anti-HIV Agents↗