Search PubMed⌕ Search

Biomedical subjects

Leming Shi

Publications and source records attributed to Leming Shi.

36 records · Page 2Linked to original sources

Construction of a virtual combinatorial library using SMILES strings to discover potential structure-diverse PPAR modulators.

Based on the structural characters of PPAR modulators, a virtual combinatorial library containing 1226,625 compounds was constructed using SMILES strings. Selected ADME filters were employed to compel compounds having poor drug-like properties from this library. This library was converted to sdf and mol2 files by CONCORD 4.0, and was then docked to PPARgamma by DOCK 4.0 to identify new chemical entities that may be potential drug leads against type 2 diabetes and other metabolic diseases. The method to construct virtual combinatorial library using SMILES strings was further visualized by Visual Basic.net that can facilitate the needs of generating other type virtual combinatorial libraries.

Combinatorial Chemistry Techniques↗

Multi-class cancer classification by total principal component regression (TPCR) using microarray gene expression data.

DNA microarray technology provides a promising approach to the diagnosis and prognosis of tumors on a genome-wide scale by monitoring the expression levels of thousands of genes simultaneously. One problem arising from the use of microarray data is the difficulty to analyze the high-dimensional gene expression data, typically with thousands of variables (genes) and much fewer observations (samples), in which severe collinearity is often observed. This makes it difficult to apply directly the classical statistical methods to investigate microarray data. In this paper, total principal component regression (TPCR) was proposed to classify human tumors by extracting the latent variable structure underlying microarray data from the augmented subspace of both independent variables and dependent variables. One of the salient features of our method is that it takes into account not only the latent variable structure but also the errors in the microarray gene expression profiles (independent variables). The prediction performance of TPCR was evaluated by both leave-one-out and leave-half-out cross-validation using four well-known microarray datasets. The stabilities and reliabilities of the classification models were further assessed by re-randomization and permutation studies. A fast kernel algorithm was applied to decrease the computation time dramatically. (MATLAB source code is available upon request.).

Acute Disease↗

The External RNA Controls Consortium: a progress report.

Standard controls and best practice guidelines advance acceptance of data from research, preclinical and clinical laboratories by providing a means for evaluating data quality. The External RNA Controls Consortium (ERCC) is developing commonly agreed-upon and tested controls for use in expression assays, a true industry-wide standard control.

Animals↗

Design, synthesis, and evaluation of a new class of noncyclic 1,3-dicarbonyl compounds as PPARalpha selective activators.

Lipid accumulation in nonadipose tissues is increasingly linked to the development of type 2 diabetes in obese individuals. We report here the design, synthesis, and evaluation of a series of novel PPARalpha selective activators containing 1,3-dicarbonyl moieties. Structure-activity relationship studies led to the identification of PPARalpha selective activators (compounds 10, 14, 17, 18, and 21) with stronger potency and efficacy to activate PPARalpha over PPARgamma and PPARdelta. Experiments in vivo showed that compounds 10, 14, and 17 had blood glucose lowering effect in diabetic db/db mouse model after two weeks oral dosing. The data strongly support further testing of these lead compounds in other relevant disease animal models to evaluate their potential therapeutic benefits.

Adipose Tissue↗

Development of public toxicogenomics software for microarray data management and analysis.

A robust bioinformatics capability is widely acknowledged as central to realizing the promises of toxicogenomics. Successful application of toxicogenomic approaches, such as DNA microarray, inextricably relies on appropriate data management, the ability to extract knowledge from massive amounts of data and the availability of functional information for data interpretation. At the FDA's National Center for Toxicological Research (NCTR), we are developing a public microarray data management and analysis software, called ArrayTrack. ArrayTrack is Minimum Information About a Microarray Experiment (MIAME) supportive for storing both microarray data and experiment parameters associated with a toxicogenomics study. A quality control mechanism is implemented to assure the fidelity of entered expression data. ArrayTrack also provides a rich collection of functional information about genes, proteins and pathways drawn from various public biological databases for facilitating data interpretation. In addition, several data analysis and visualization tools are available with ArrayTrack, and more tools will be available in the next released version. Importantly, gene expression data, functional information and analysis methods are fully integrated so that the data analysis and interpretation process is simplified and enhanced. ArrayTrack is publicly available online and the prospective user can also request a local installation version by contacting the authors.

Databases, Genetic↗

3D QSAR studies on peroxisome proliferator-activated receptor gamma agonists using CoMFA and CoMSIA.

The peroxisome proliferator-activated receptors (PPARs) have increasingly become attractive targets for developing novel anti-type 2 diabetic drugs. We employed comparative molecular field analysis (CoMFA) and comparative molecular similarity indices analysis (CoMSIA) to study three-dimensional quantitative structure-activity relationship (3D QSAR) based on existing agonists of PPARgamma (including five thiazolidinediones and 74 tyrosine-based compounds). Predictive 3D QSAR models with conventional r2 and cross-validated coefficient (q2) values up to 0.974 and 0.642 for CoMFA and 0.979 and 0.686 for COMSIA were established using the SYBYL package. These models were validated by a test set containing 18 compounds. The CoMFA and CoMSIA field distributions are in general agreement with the structural characteristics of the binding pockets of PPARgamma, which demonstrates that the 3D QSAR models built here are very useful in predicting activities of novel compounds for activating PPARgamma.

Computer Simulation↗

Multi-class tumor classification by discriminant partial least squares using microarray gene expression data and assessment of classification models.

High-throughput DNA microarray provides an effective approach to the monitoring of expression levels of thousands of genes in a sample simultaneously. One promising application of this technology is the molecular diagnostics of cancer, e.g. to distinguish normal tissue from tumor or to classify tumors into different types or subtypes. One problem arising from the use of microarray data is how to analyze the high-dimensional gene expression data, typically with thousands of variables (genes) and much fewer observations (samples). There is a need to develop reliable classification methods to make full use of microarray data and to evaluate accurately the predictive ability and reliability of such derived models. In this paper, discriminant partial least squares was used to classify the different types of human tumors using four microarray datasets and showed good prediction performance. Four different cross-validation procedures (leave-one-out versus leave-half-out; incomplete versus full) were used to evaluate the classification model. Our results indicate that discriminant partial least squares using leave-half-out cross-validation provides a more realistic estimate of the predictive ability of a classification model, which may be overestimated by some of the cross-validation procedures, and the information obtained from different cross-validation procedures can be used to evaluate the reliability of the classification model.

Algorithms↗

Classification of cDNA array genes that have a highly significant discriminative power due to their unique distribution in four brain regions.

Novel statistical methods were used to distinguish functionally distinct brain regions using their cDNA array gene expression profiles, and it was found that one of four specific factors is often associated with the most regionally discriminative genes. The gene expression profiles for the substantia nigra (SN), striatum (STR), parietal cortex (PC), and posterolateral cortical amygdaloid nucleus (PLCo) brain regions were determined from each brain region. An F-test identified 339 genes of the 1185 array genes as having a P < or = 0.01 and applied a gene ranking and selection method based on Soft Independent Modeling of Class Analogy (SIMCA) to obtain 59 of the most discriminative genes. Their discriminative power was validated in three steps. The most convincing step showed their ability to correctly predict the brain regional classifications for 18 "test" gene expression sets obtained from the four regions. A two-way Hierarchical Cluster Analysis organized the 59 genes in six clusters according to their expression differences in the brain regions. Expression patterns in the SN and STR regions greatly differed from each other and the PC and PLCo. The closer similarity in the gene expression patterns of the PC and PLCo was probably due to their functional similarity. The important factors in determining differences in the regional gene expression profiles in six clusters were (1) regional myelin/oligodendrocyte levels, (2) resident neuron types, (3) neurotransmitter innervation profiles, and (4) Ca++-dependent signaling and second messenger systems.

Animals↗

Multiclass Decision Forest--a novel pattern recognition method for multiclass classification in microarray data analysis.

The wealth of knowledge imbedded in gene expression data from DNA microarrays portends rapid advances in both research and clinic. Turning the prodigious and noisy data into knowledge is a challenge to the field of bioinformatics, and development of classifiers using supervised learning techniques is the primary methodological approach for clinical application using gene expression data. In this paper, we present a novel classification method, multiclass Decision Forest (DF), that is the direct extension of the two-class DF previously developed in our lab. Central to DF is the synergistic combining of multiple heterogenic but comparable decision trees to reach a more accurate and robust classification model. The computationally inexpensive multiclass DF algorithm integrates gene selection and model development, and thus eliminates the bias of gene preselection in crossvalidation. Importantly, the method provides several statistical means for assessment of prediction accuracy, prediction confidence, and diagnostic capability. We demonstrate the method by application to gene expression data for 83 small round blue-cell tumors (SRBCTs) samples belonging to one of four different classes. Based on 500 runs of 10-fold crossvalidation, tumor prediction accuracy was approximately 97%, sensitivity was approximately 95%, diagnostic sensitivity was approximately 91%, and diagnostic accuracy was approximately 99.5%. Among 25 genes selected to distinguish tumor class, 12 have functional information in the literature implicating their involvement in cancer. The four types of SRBCTs samples are also distinguishable in a clustering analysis based on the expression profiles of these 25 genes. The results demonstrated that the multiclass DF is an effective classification method for analysis of gene expression data for the purpose of molecular diagnostics.

Carcinoma, Small Cell↗

Using decision forest to classify prostate cancer samples on the basis of SELDI-TOF MS data: assessing chance correlation and prediction confidence.

Class prediction using "omics" data is playing an increasing role in toxicogenomics, diagnosis/prognosis, and risk assessment. These data are usually noisy and represented by relatively few samples and a very large number of predictor variables (e.g., genes of DNA microarray data or m/z peaks of mass spectrometry data). These characteristics manifest the importance of assessing potential random correlation and overfitting of noise for a classification model based on omics data. We present a novel classification method, decision forest (DF), for class prediction using omics data. DF combines the results of multiple heterogeneous but comparable decision tree (DT) models to produce a consensus prediction. The method is less prone to overfitting of noise and chance correlation. A DF model was developed to predict presence of prostate cancer using a proteomic data set generated from surface-enhanced laser deposition/ionization time-of-flight mass spectrometry (SELDI-TOF MS). The degree of chance correlation and prediction confidence of the model was rigorously assessed by extensive cross-validation and randomization testing. Comparison of model prediction with imposed random correlation demonstrated biologic relevance of the model and the reduction of overfitting in DF. Furthermore, two confidence levels (high and low confidences) were assigned to each prediction, where most misclassifications were associated with the low-confidence region. For the high-confidence prediction, the model achieved 99.2% sensitivity and 98.2% specificity. The model also identified a list of significant peaks that could be useful for biomarker identification. DF should be equally applicable to other omics data such as gene expression data or metabolomic data. The DF algorithm is available upon request.

Decision Support Techniques↗

Assessment of prediction confidence and domain extrapolation of two structure-activity relationship models for predicting estrogen receptor binding activity.

Quantitative structure-activity relationship (QSAR) methods have been widely applied in drug discovery, lead optimization, toxicity prediction, and regulatory decisions. Despite major advances in algorithms and software, QSAR models have inherent limitations associated with a size and chemical-structure diversity of the training set, experimental error, and many characteristics of structure representation and correlation algorithms. Whereas excellent fit to the training data may be readily attainable, often models fail to predict accurately chemicals that are outside their domain of applicability. A QSAR's utility and, in the case of regulatory decisions, justification for usage increasingly depend on the ability to quantify a model's potential for predicting unknown chemicals with some known degree of certainty. It is never possible to predict an unknown chemical with absolute certainty. Here we report on two QSAR models based on different data sets for classification of chemicals according to their ability to bind to the estrogen receptor. The models were developed by using a novel QSAR method, Decision Forest, which combines the results of multiple heterogeneous but comparable Decision Tree models to produce a consensus prediction. We used an extensive cross-validation process to define an applicability domain for model predictions based on two quantitative measures: prediction confidence and domain extrapolation. Together, these measures quantify the accuracy of each prediction within and outside of the training domain. Despite being based on large and diverse training sets, both QSAR models had poor accuracy for chemicals within the domain of low confidence, whereas good accuracy was obtained for those within the domain of high confidence. For prediction in the high confidence domain, accuracy was inversely proportional to the degree of domain extrapolation. The model with a larger training set of 1,092, compared with 232 for the other, was more accurate in predicting chemicals at larger domain extrapolation, and could be particularly useful for rapidly prioritizing potential endocrine disruptors from large chemical universe.

Animals↗

QA/QC: challenges and pitfalls facing the microarray community and regulatory agencies.

The scientific community has been enthusiastic about DNA microarray technology for pharmacogenomic and toxicogenomic studies in the hope of advancing personalized medicine and drug development. The US Food and Drug Administration has been proactive in promoting the use of pharmacogenomic data in drug development and has issued a draft guidance for the pharmaceutical industry on data submissions. However, many challenges and pitfalls are facing the microarray community and regulatory agencies before microarray data can be reliably applied to support regulatory decision making. Four types of factors (i.e., technical, instrumental, computational and interpretative) affect the outcome of a microarray study, and a major concern about microarray studies has been the lack of reproducibility and accuracy. Intralaboratory data consistency is the foundation of reliable knowledge extraction and meaningful crosslaboratory or crossplatform comparisons; unfortunately, it has not been seriously evaluated and demonstrated in every study. Profound problems in data quality have been observed from analyzing published data sets, and many laboratories have been struggling with technical troubleshooting rather than generating reliable data of scientific significance. The microarray community and regulatory agencies must work together to establish a set of consensus quality assurance and quality control criteria for assessing and ensuring data quality, to identify critical factors affecting data quality, and to optimize and standardize microarray procedures so that biologic interpretation and decision-making are not based on unreliable data. These fundamental issues must be adequately addressed before microarray technology can be transformed from a research tool to clinical practices.

Drug Approval↗

Quantitative structure-activity relationship study of histone deacetylase inhibitors.

Histone deacetylases (HDACs) play a critical role in gene transcription and have become a novel target for the discovery of drugs against cancer and other diseases. During the past several years there have been extensive efforts in the identification and optimization of histone deacetylase inhibitors (HDACIs) as novel anticancer drugs. Here we report a comprehensive quantitative structure-activity relationship (QSAR) study of HDACIs in the hope of identifying the structural determinants for anticancer activity. We have identified, collected, and verified the structural and biological activity data for 124 compounds from various literature sources and performed an extensive QSAR study on this comprehensive data set by using various QSAR and classification methods. A highly predictive QSAR model with R(2) of 0.76 and leave-one-out cross-validated R(2) of 0.73 was obtained. The overall rate of cross-validated correct prediction of the classification model is around 92%. The QSAR and classification models provided direct guidance to our internal programs of identifying and optimizing HDAC inhibitors. Limitations of the models were also discussed.

Animals↗

ArrayTrack--supporting toxicogenomic research at the U.S. Food and Drug Administration National Center for Toxicological Research.

The mapping of the human genome and the determination of corresponding gene functions, pathways, and biological mechanisms are driving the emergence of the new research fields of toxicogenomics and systems toxicology. Many technological advances such as microarrays are enabling this paradigm shift that indicates an unprecedented advancement in the methods of understanding the expression of toxicity at the molecular level. At the National Center for Toxicological Research (NCTR) of the U.S. Food and Drug Administration, core facilities for genomic, proteomic, and metabonomic technologies have been established that use standardized experimental procedures to support centerwide toxicogenomic research. Collectively, these facilities are continuously producing an unprecedented volume of data. NCTR plans to develop a toxicoinformatics integrated system (TIS) for the purpose of fully integrating genomic, proteomic, and metabonomic data with the data in public repositories as well as conventional (Italic)in vitro(/Italic) and (Italic)in vivo(/Italic) toxicology data. The TIS will enable data curation in accordance with standard ontology and provide or interface a rich collection of tools for data analysis and knowledge mining. In this article the design, practical issues, and functions of the TIS are discussed through presenting its prototype version, ArrayTrack, for the management and analysis of DNA microarray data. ArrayTrack is logically constructed of three linked components: a) a library (LIB) that mirrors critical data in public databases; b) a database (MicroarrayDB) that stores microarray experiment information that is Minimal Information About a Microarray Experiment (MIAME) compliant; and c) tools (TOOL) that operate on experimental and public data for knowledge discovery. Using ArrayTrack, we can select an analysis method from the TOOL and apply the method to selected microarray data stored in the MicroarrayDB; the analysis results can be linked directly to gene information in the LIB.

Databases, Factual↗

Structure-activity relationship approaches and applications.

New techniques and software have enabled ubiquitous use of structure-activity relationships (SARs) in the pharmaceutical industry and toxicological sciences. We review the status of SAR technology by using examples to underscore the advances as well as the unique technical challenges. Applying SAR involves two steps: Characterization of the chemicals under investigation, and application of chemometric approaches to explore data patterns or to establish the relationships between structure and activity. We describe generally but not exhaustively the SAR methodologies popular use in toxicology, including representation of chemical structure, and chemometric techniques where models are both unsupervised and supervised. The utility of SAR technology is most evident when supervised methods are used to predict toxicity of untested chemicals based only on chemical structure. Such models can predict on both an ordinal scale (e.g., active vs inactive) or a continuouis scale (e.g., median lethal dose [LD50] dose). The reader is also referred to a companion paper in this issue that discusses quantitative structure-activity relationship (QSAR) methods that have advanced markedly over the past decade.

Forecasting↗

Phytoestrogens and mycoestrogens bind to the rat uterine estrogen receptor.

Consumption of phytoestrogens and mycoestrogens in food products or as dietary supplements is of interest because of both the potential beneficial and adverse effects of these compounds in estrogen-responsive target tissues. Although the hazards of exposure to potent estrogens such as diethylstilbestrol in developing male and female reproductive tracts are well characterized, less is known about the effects of weaker estrogens including phytoestrogens. With some exceptions, ligand binding to the estrogen receptor (ER) predicts uterotrophic activity. Using a well-established and rigorously validated ER-ligand binding assay, we assessed the relative binding affinity (RBA) for 46 chemicals from several chemical structure classes of potential phytoestrogens and mycoestrogens. Although none of the test compounds bound to ER with the affinity of the standard, 17beta-estradiol (E(2)), ER binding was found among all classes of chemical structures (flavones, isoflavones, flavanones, coumarins, chalcones and mycoestrogens). Estrogen receptor relative binding affinities were distributed across a wide range (from approximately 43 to 0.00008; E(2) = 100). These data can be utilized before animal testing to rank order estimates of the potential for in vivo estrogenic activity of a wide range of untested plant chemicals (as well as other chemicals) based on ER binding.

Animals↗

Prediction of estrogen receptor binding for 58,000 chemicals using an integrated system of a tree-based model with structural alerts.

A number of environmental chemicals, by mimicking natural hormones, can disrupt endocrine function in experimental animals, wildlife, and humans. These chemicals, called "endocrine-disrupting chemicals" (EDCs), are such a scientific and public concern that screening and testing 58,000 chemicals for EDC activities is now statutorily mandated. Computational chemistry tools are important to biologists because they identify chemicals most important for in vitro and in vivo studies. Here we used a computational approach with integration of two rejection filters, a tree-based model, and three structural alerts to predict and prioritize estrogen receptor (ER) ligands. The models were developed using data for 232 structurally diverse chemicals (training set) with a 10(6) range of relative binding affinities (RBAs); we then validated the models by predicting ER RBAs for 463 chemicals that had ER activity data (testing set). The integrated model gave a lower false negative rate than any single component for both training and testing sets. When the integrated model was applied to approximately 58,000 potential EDCs, 80% (approximately 46,000 chemicals) were predicted to have negligible potential (log RBA < -4.5, with log RBA = 2.0 for estradiol) to bind ER. The ability to process large numbers of chemicals to predict inactivity for ER binding and to categorically prioritize the remainder provides one biologic measure to prioritize chemicals for entry into more expensive assays (most chemicals have no biologic data of any kind). The general approach for predicting ER binding reported here may be applied to other receptors and/or reversible binding mechanisms involved in endocrine disruption.

Animals↗

Eigenvalue analysis of peroxisome proliferator-activated receptor gamma agonists.

Eigenvalue analysis (EVA) was conducted on a series of potent agonists of peroxisome proliferator-activated receptor gamma (PPARgamma). Predictive EVA quantitative structure-activity relationship (QSAR) models were established using the SYBYL package, which had conventional r2 and cross-validated coefficient (q2) values up to 0.920 and 0.587 for the AM1 method and 0.863 and 0.586 for the PM3 method, respectively. These models were validated by a test set containing 18 compounds. The capability to predict by these two models for PPARgamma agonists, with the best predictive r2pred value of 0.614 for AM1 and 0.822 for PM3 methods, set a successful example for applying a similar approach in building QSAR models for PPARalpha and -delta that could potentially offer a new opportunity in the design of novel PPAR modulators.

Models, Chemical↗