Search PubMed⌕ Search

Biomedical subjects

Weida Tong

Publications and source records attributed to Weida Tong.

At least 37 records · Page 2Linked to original sources

Quality control and quality assessment of data from surface-enhanced laser desorption/ionization (SELDI) time-of flight (TOF) mass spectrometry (MS).

BACKGROUND: Proteomic profiling of complex biological mixtures by the ProteinChip technology of surface-enhanced laser desorption/ionization time-of-flight (SELDI-TOF) mass spectrometry (MS) is one of the most promising approaches in toxicological, biological, and clinic research. The reliable identification of protein expression patterns and associated protein biomarkers that differentiate disease from health or that distinguish different stages of a disease depends on developing methods for assessing the quality of SELDI-TOF mass spectra. The use of SELDI data for biomarker identification requires application of rigorous procedures to detect and discard low quality spectra prior to data analysis. RESULTS: The systematic variability from plates, chips, and spot positions in SELDI experiments was evaluated using biological and technical replicates. Systematic biases on plates, chips, and spots were not found. The reproducibility of SELDI experiments was demonstrated by examining the resulting low coefficient of variances of five peaks presented in all 144 spectra from quality control samples that were loaded randomly on different spots in the chips of six bioprocessor plates. We developed a method to detect and discard low quality spectra prior to proteomic profiling data analysis, which uses a correlation matrix to measure the similarities among SELDI mass spectra obtained from similar biological samples. Application of the correlation matrix to our SELDI data for liver cancer and liver toxicity study and myeloma-associated lytic bone disease study confirmed this approach as an efficient and reliable method for detecting low quality spectra. CONCLUSION: This report provides evidence that systematic variability between plates, chips, and spots on which the samples were assayed using SELDI based proteomic procedures did not exist. The reproducibility of experiments in our studies was demonstrated to be acceptable and the profiling data for subsequent data analysis are reliable. Correlation matrix was developed as a quality control tool to detect and discard low quality spectra prior to data analysis. It proved to be a reliable method to measure the similarities among SELDI mass spectra and can be used for quality control to decrease noise in proteomic profiling data prior to data analysis.

Female↗

Multi-class cancer classification by total principal component regression (TPCR) using microarray gene expression data.

DNA microarray technology provides a promising approach to the diagnosis and prognosis of tumors on a genome-wide scale by monitoring the expression levels of thousands of genes simultaneously. One problem arising from the use of microarray data is the difficulty to analyze the high-dimensional gene expression data, typically with thousands of variables (genes) and much fewer observations (samples), in which severe collinearity is often observed. This makes it difficult to apply directly the classical statistical methods to investigate microarray data. In this paper, total principal component regression (TPCR) was proposed to classify human tumors by extracting the latent variable structure underlying microarray data from the augmented subspace of both independent variables and dependent variables. One of the salient features of our method is that it takes into account not only the latent variable structure but also the errors in the microarray gene expression profiles (independent variables). The prediction performance of TPCR was evaluated by both leave-one-out and leave-half-out cross-validation using four well-known microarray datasets. The stabilities and reliabilities of the classification models were further assessed by re-randomization and permutation studies. A fast kernel algorithm was applied to decrease the computation time dramatically. (MATLAB source code is available upon request.).

Acute Disease↗

Building an organ-specific carcinogenic database for SAR analyses.

FDA reviewers need a means to rapidly predict organ-specific carcinogenicity to aid in evaluating new chemicals submitted for approval. This research addressed the building of a database to use in developing a predictive model for such an application based on structure-activity relationships (SAR). The Internet availability of the Carcinogenic Potency Database (CPDB) provided a solid foundation on which to base such a model. The addition of molecular structures to the CPDB provided the extra ingredient necessary for SAR analyses. However, the CPDB had to be compressed from a multirecord to a single record per chemical database; multiple records representing each gender, species, route of administration, and organ-specific toxicity had to be summarized into a single record for each study. Multiple studies on a single chemical had to be further reduced based on a hierarchical scheme. Structural cleanup involved removal of all chemicals that would impede the accurate generation of SAR type descriptors from commercial software programs; that is, inorganic chemicals, mixtures, and organometallics were removed. Counterions such as Na, K, sulfates, hydrates, and salts were also removed for structural consistency. Structural modification sometimes resulted in duplicate records that also had to be reduced to a single record based on the hierarchical scheme. The modified database containing 999 chemicals was evaluated for liver-specific carcinogenicity using a variety of analysis techniques. These preliminary analyses all yielded approximately the same results with an overall predictability of about 63%, which was comprised of a sensitivity of about 30% and a specificity of about 77%.

Animals↗

Development of public toxicogenomics software for microarray data management and analysis.

A robust bioinformatics capability is widely acknowledged as central to realizing the promises of toxicogenomics. Successful application of toxicogenomic approaches, such as DNA microarray, inextricably relies on appropriate data management, the ability to extract knowledge from massive amounts of data and the availability of functional information for data interpretation. At the FDA's National Center for Toxicological Research (NCTR), we are developing a public microarray data management and analysis software, called ArrayTrack. ArrayTrack is Minimum Information About a Microarray Experiment (MIAME) supportive for storing both microarray data and experiment parameters associated with a toxicogenomics study. A quality control mechanism is implemented to assure the fidelity of entered expression data. ArrayTrack also provides a rich collection of functional information about genes, proteins and pathways drawn from various public biological databases for facilitating data interpretation. In addition, several data analysis and visualization tools are available with ArrayTrack, and more tools will be available in the next released version. Importantly, gene expression data, functional information and analysis methods are fully integrated so that the data analysis and interpretation process is simplified and enhanced. ArrayTrack is publicly available online and the prospective user can also request a local installation version by contacting the authors.

Databases, Genetic↗

Multi-class tumor classification by discriminant partial least squares using microarray gene expression data and assessment of classification models.

High-throughput DNA microarray provides an effective approach to the monitoring of expression levels of thousands of genes in a sample simultaneously. One promising application of this technology is the molecular diagnostics of cancer, e.g. to distinguish normal tissue from tumor or to classify tumors into different types or subtypes. One problem arising from the use of microarray data is how to analyze the high-dimensional gene expression data, typically with thousands of variables (genes) and much fewer observations (samples). There is a need to develop reliable classification methods to make full use of microarray data and to evaluate accurately the predictive ability and reliability of such derived models. In this paper, discriminant partial least squares was used to classify the different types of human tumors using four microarray datasets and showed good prediction performance. Four different cross-validation procedures (leave-one-out versus leave-half-out; incomplete versus full) were used to evaluate the classification model. Our results indicate that discriminant partial least squares using leave-half-out cross-validation provides a more realistic estimate of the predictive ability of a classification model, which may be overestimated by some of the cross-validation procedures, and the information obtained from different cross-validation procedures can be used to evaluate the reliability of the classification model.

Algorithms↗

Classification of cDNA array genes that have a highly significant discriminative power due to their unique distribution in four brain regions.

Novel statistical methods were used to distinguish functionally distinct brain regions using their cDNA array gene expression profiles, and it was found that one of four specific factors is often associated with the most regionally discriminative genes. The gene expression profiles for the substantia nigra (SN), striatum (STR), parietal cortex (PC), and posterolateral cortical amygdaloid nucleus (PLCo) brain regions were determined from each brain region. An F-test identified 339 genes of the 1185 array genes as having a P < or = 0.01 and applied a gene ranking and selection method based on Soft Independent Modeling of Class Analogy (SIMCA) to obtain 59 of the most discriminative genes. Their discriminative power was validated in three steps. The most convincing step showed their ability to correctly predict the brain regional classifications for 18 "test" gene expression sets obtained from the four regions. A two-way Hierarchical Cluster Analysis organized the 59 genes in six clusters according to their expression differences in the brain regions. Expression patterns in the SN and STR regions greatly differed from each other and the PC and PLCo. The closer similarity in the gene expression patterns of the PC and PLCo was probably due to their functional similarity. The important factors in determining differences in the regional gene expression profiles in six clusters were (1) regional myelin/oligodendrocyte levels, (2) resident neuron types, (3) neurotransmitter innervation profiles, and (4) Ca++-dependent signaling and second messenger systems.

Animals↗

Multiclass Decision Forest--a novel pattern recognition method for multiclass classification in microarray data analysis.

The wealth of knowledge imbedded in gene expression data from DNA microarrays portends rapid advances in both research and clinic. Turning the prodigious and noisy data into knowledge is a challenge to the field of bioinformatics, and development of classifiers using supervised learning techniques is the primary methodological approach for clinical application using gene expression data. In this paper, we present a novel classification method, multiclass Decision Forest (DF), that is the direct extension of the two-class DF previously developed in our lab. Central to DF is the synergistic combining of multiple heterogenic but comparable decision trees to reach a more accurate and robust classification model. The computationally inexpensive multiclass DF algorithm integrates gene selection and model development, and thus eliminates the bias of gene preselection in crossvalidation. Importantly, the method provides several statistical means for assessment of prediction accuracy, prediction confidence, and diagnostic capability. We demonstrate the method by application to gene expression data for 83 small round blue-cell tumors (SRBCTs) samples belonging to one of four different classes. Based on 500 runs of 10-fold crossvalidation, tumor prediction accuracy was approximately 97%, sensitivity was approximately 95%, diagnostic sensitivity was approximately 91%, and diagnostic accuracy was approximately 99.5%. Among 25 genes selected to distinguish tumor class, 12 have functional information in the literature implicating their involvement in cancer. The four types of SRBCTs samples are also distinguishable in a clustering analysis based on the expression profiles of these 25 genes. The results demonstrated that the multiclass DF is an effective classification method for analysis of gene expression data for the purpose of molecular diagnostics.

Carcinoma, Small Cell↗

Three new consensus QSAR models for the prediction of Ames genotoxicity.

Three QSAR methods, artificial neural net (ANN), k-nearest neighbors (kNN), and Decision Forest (DF), were applied to 3363 diverse compounds tested for their Ames genotoxicity. The ratio of mutagens to non-mutagens was 60/40 for this dataset. This group of compounds includes >300 therapeutic drugs. All models were developed using the same initial set of 148 topological indices: molecular connectivity chi indices and electrotopological state indices (atom-type, bond-type and group-type E-state), as well as binary indicators. While previous studies have found logP to be a determining factor in genotoxicity, it was not found to be important by any modeling method employed in this study. The three models yielded an average training/test concordance value of 88%, with a low percentage of false positives and false negatives. External validation testing on 400 compounds not used for QSAR model development gave an average concordance of 82%. This value increased to 92% upon removal of less reliable outcomes, as determined by a reliability criterion used within each model. The ANN model showed the best performance in predicting drug compounds, yielding 97% concordance (34/35 drugs) after the removal of less reliable predictions. The appreciable commonality found among the top 10 ranked descriptors from each model is of particular interest because of the diversity in the learning algorithms and descriptor selection techniques employed in this study. Forty percent of the most important descriptors in any one model are found in one or two other models. Fourteen of the most important descriptors relate directly to known toxicophores involved in potent genotoxic responses in Salmonella typhimurium. A comparison of the validation results with those of MULTICASE and DEREK indicated that the new models presented in this work perform substantially better than the former models in predicting genotoxicity of therapeutic drugs. Substantially higher specificity was achieved with these new models as compared with MULTICASE or DEREK with comparable sensitivities among all models.

Algorithms↗

Using decision forest to classify prostate cancer samples on the basis of SELDI-TOF MS data: assessing chance correlation and prediction confidence.

Class prediction using "omics" data is playing an increasing role in toxicogenomics, diagnosis/prognosis, and risk assessment. These data are usually noisy and represented by relatively few samples and a very large number of predictor variables (e.g., genes of DNA microarray data or m/z peaks of mass spectrometry data). These characteristics manifest the importance of assessing potential random correlation and overfitting of noise for a classification model based on omics data. We present a novel classification method, decision forest (DF), for class prediction using omics data. DF combines the results of multiple heterogeneous but comparable decision tree (DT) models to produce a consensus prediction. The method is less prone to overfitting of noise and chance correlation. A DF model was developed to predict presence of prostate cancer using a proteomic data set generated from surface-enhanced laser deposition/ionization time-of-flight mass spectrometry (SELDI-TOF MS). The degree of chance correlation and prediction confidence of the model was rigorously assessed by extensive cross-validation and randomization testing. Comparison of model prediction with imposed random correlation demonstrated biologic relevance of the model and the reduction of overfitting in DF. Furthermore, two confidence levels (high and low confidences) were assigned to each prediction, where most misclassifications were associated with the low-confidence region. For the high-confidence prediction, the model achieved 99.2% sensitivity and 98.2% specificity. The model also identified a list of significant peaks that could be useful for biomarker identification. DF should be equally applicable to other omics data such as gene expression data or metabolomic data. The DF algorithm is available upon request.

Decision Support Techniques↗

Assessment of prediction confidence and domain extrapolation of two structure-activity relationship models for predicting estrogen receptor binding activity.

Quantitative structure-activity relationship (QSAR) methods have been widely applied in drug discovery, lead optimization, toxicity prediction, and regulatory decisions. Despite major advances in algorithms and software, QSAR models have inherent limitations associated with a size and chemical-structure diversity of the training set, experimental error, and many characteristics of structure representation and correlation algorithms. Whereas excellent fit to the training data may be readily attainable, often models fail to predict accurately chemicals that are outside their domain of applicability. A QSAR's utility and, in the case of regulatory decisions, justification for usage increasingly depend on the ability to quantify a model's potential for predicting unknown chemicals with some known degree of certainty. It is never possible to predict an unknown chemical with absolute certainty. Here we report on two QSAR models based on different data sets for classification of chemicals according to their ability to bind to the estrogen receptor. The models were developed by using a novel QSAR method, Decision Forest, which combines the results of multiple heterogeneous but comparable Decision Tree models to produce a consensus prediction. We used an extensive cross-validation process to define an applicability domain for model predictions based on two quantitative measures: prediction confidence and domain extrapolation. Together, these measures quantify the accuracy of each prediction within and outside of the training domain. Despite being based on large and diverse training sets, both QSAR models had poor accuracy for chemicals within the domain of low confidence, whereas good accuracy was obtained for those within the domain of high confidence. For prediction in the high confidence domain, accuracy was inversely proportional to the degree of domain extrapolation. The model with a larger training set of 1,092, compared with 232 for the other, was more accurate in predicting chemicals at larger domain extrapolation, and could be particularly useful for rapidly prioritizing potential endocrine disruptors from large chemical universe.

Animals↗

QA/QC: challenges and pitfalls facing the microarray community and regulatory agencies.

The scientific community has been enthusiastic about DNA microarray technology for pharmacogenomic and toxicogenomic studies in the hope of advancing personalized medicine and drug development. The US Food and Drug Administration has been proactive in promoting the use of pharmacogenomic data in drug development and has issued a draft guidance for the pharmaceutical industry on data submissions. However, many challenges and pitfalls are facing the microarray community and regulatory agencies before microarray data can be reliably applied to support regulatory decision making. Four types of factors (i.e., technical, instrumental, computational and interpretative) affect the outcome of a microarray study, and a major concern about microarray studies has been the lack of reproducibility and accuracy. Intralaboratory data consistency is the foundation of reliable knowledge extraction and meaningful crosslaboratory or crossplatform comparisons; unfortunately, it has not been seriously evaluated and demonstrated in every study. Profound problems in data quality have been observed from analyzing published data sets, and many laboratories have been struggling with technical troubleshooting rather than generating reliable data of scientific significance. The microarray community and regulatory agencies must work together to establish a set of consensus quality assurance and quality control criteria for assessing and ensuring data quality, to identify critical factors affecting data quality, and to optimize and standardize microarray procedures so that biologic interpretation and decision-making are not based on unreliable data. These fundamental issues must be adequately addressed before microarray technology can be transformed from a research tool to clinical practices.

Drug Approval↗

Study of 202 natural, synthetic, and environmental chemicals for binding to the androgen receptor.

A number of environmental and industrial chemicals are reported to possess androgenic or antiandrogenic activities. These androgenic endocrine disrupting chemicals may disrupt the endocrine system of humans and wildlife by mimicking or antagonizing the functions of natural hormones. The present study developed a low cost recombinant androgen receptor (AR) competitive binding assay that uses no animals. We validated the assay by comparing the protocols and results from other similar assays, such as the binding assay using prostate cytosol. We tested 202 natural, synthetic, and environmental chemicals that encompass a broad range of structural classes, including steroids, diethylstilbestrol and related chemicals, antiestrogens, flutamide derivatives, bisphenol A derivatives, alkylphenols, parabens, alkyloxyphenols, phthalates, siloxanes, phytoestrogens, DDTs, PCBs, pesticides, organophosphate insecticides, and other chemicals. Some of these chemicals are environmentally persistent and/or commercially important, but their AR binding affinities have not been previously reported. To the best of our knowledge, these results represent the largest and most diverse data set publicly available for chemical binding to the AR. Through a careful structure-activity relationship (SAR) examination of the data set in conjunction with knowledge of the recently reported ligand-AR crystal structures, we are able to define the general structural requirements for chemical binding to AR. Hydrophobic interactions are important for AR binding. The interaction between ligand and AR at the 3- and 17-positions of testosterone and R1881 found in other chemical classes are discussed in depth. The SAR studies of ligand binding characteristics for AR are compared to our previously reported results for estrogen receptor binding.

Androgen Receptor Antagonists↗

ArrayTrack--supporting toxicogenomic research at the U.S. Food and Drug Administration National Center for Toxicological Research.

The mapping of the human genome and the determination of corresponding gene functions, pathways, and biological mechanisms are driving the emergence of the new research fields of toxicogenomics and systems toxicology. Many technological advances such as microarrays are enabling this paradigm shift that indicates an unprecedented advancement in the methods of understanding the expression of toxicity at the molecular level. At the National Center for Toxicological Research (NCTR) of the U.S. Food and Drug Administration, core facilities for genomic, proteomic, and metabonomic technologies have been established that use standardized experimental procedures to support centerwide toxicogenomic research. Collectively, these facilities are continuously producing an unprecedented volume of data. NCTR plans to develop a toxicoinformatics integrated system (TIS) for the purpose of fully integrating genomic, proteomic, and metabonomic data with the data in public repositories as well as conventional (Italic)in vitro(/Italic) and (Italic)in vivo(/Italic) toxicology data. The TIS will enable data curation in accordance with standard ontology and provide or interface a rich collection of tools for data analysis and knowledge mining. In this article the design, practical issues, and functions of the TIS are discussed through presenting its prototype version, ArrayTrack, for the management and analysis of DNA microarray data. ArrayTrack is logically constructed of three linked components: a) a library (LIB) that mirrors critical data in public databases; b) a database (MicroarrayDB) that stores microarray experiment information that is Minimal Information About a Microarray Experiment (MIAME) compliant; and c) tools (TOOL) that operate on experimental and public data for knowledge discovery. Using ArrayTrack, we can select an analysis method from the TOOL and apply the method to selected microarray data stored in the MicroarrayDB; the analysis results can be linked directly to gene information in the LIB.

Databases, Factual↗

Quantitative structure-activity relationship methods: perspectives on drug discovery and toxicology.

Quantitative structure-activity relationships (QSARs) attempt to correlate chemical structure with activity using statistical approaches. The QSAR models are useful for various purposes including the prediction of activities of untested chemicals. Quantitative structure-activity relationships and other related approaches have attracted broad scientific interest, particularly in the pharmaceutical industry for drug discovery and in toxicology and environmental science for risk assessment. An assortment of new QSAR methods have been developed during the past decade, most of them focused on drug discovery. Besides advancing our fundamental knowledge of QSARs, these scientific efforts have stimulated their application in a wider range of disciplines, such as toxicology, where QSARs have not yet gained full appreciation. In this review, we attempt to summarize the status of QSAR with emphasis on illuminating the utility and limitations of QSAR technology. We will first review two-dimensional (2D) QSAR with a discussion of the availability and appropriate selection of molecular descriptors. We will then proceed to describe three-dimensional (3D) QSAR and key issues associated with this technology, then compare the relative suitability of 2D and 3D QSAR for different applications. Given the recent technological advances in biological research for rapid identification of drug targets, we mention several examples in which QSAR approaches are employed in conjunction with improved knowledge of the structure and function of the target receptor. The review will conclude by discussing statistical validation of QSAR models, a topic that has received sparse attention in recent years despite its critical importance.

Drug Design↗

Structure-activity relationship approaches and applications.

New techniques and software have enabled ubiquitous use of structure-activity relationships (SARs) in the pharmaceutical industry and toxicological sciences. We review the status of SAR technology by using examples to underscore the advances as well as the unique technical challenges. Applying SAR involves two steps: Characterization of the chemicals under investigation, and application of chemometric approaches to explore data patterns or to establish the relationships between structure and activity. We describe generally but not exhaustively the SAR methodologies popular use in toxicology, including representation of chemical structure, and chemometric techniques where models are both unsupervised and supervised. The utility of SAR technology is most evident when supervised methods are used to predict toxicity of untested chemicals based only on chemical structure. Such models can predict on both an ordinal scale (e.g., active vs inactive) or a continuouis scale (e.g., median lethal dose [LD50] dose). The reader is also referred to a companion paper in this issue that discusses quantitative structure-activity relationship (QSAR) methods that have advanced markedly over the past decade.

Forecasting↗

Influence of the structural diversity of data sets on the statistical quality of three-dimensional quantitative structure-activity relationship (3D-QSAR) models: predicting the estrogenic activity of xenoestrogens.

Federal legislation has resulted in the two-tiered in vitro and in vivo screening of some 80 000 structurally diverse chemicals for possible endocrine disrupting effects. To maximize efficiency and minimize expense, prioritization of these chemicals with respect to their estrogenic disrupting potential prior to this time-consuming and labor-intensive screening process is essential. Computer-based quantitative structure-activity relationship (QSAR) models, such as those obtained using comparative molecular field analysis (CoMFA), have been demonstrated as useful for risk assessment in this application. In general, however, CoMFA models to predict estrogenicity have been developed from data sets with limited structural diversity. In this study, we constructed CoMFA models based on biological data for a structurally diverse set of compounds spanning eight chemical families. We also compared two standard alignment schemes employed in CoMFA, namely, atom-fit and flexible field-fit, with respect to the predictive capabilities of their respective models for structurally diverse data sets. The present analysis indicates that flexible field-fit alignment fares better than atom-fit alignment as the structural diversity of the data set increases. Values of log(RP), where RP = relative potency, predicted by the final flexible field-fit CoMFA models are in good agreement with the corresponding experimental values. These models should be effective for predicting the endocrine disrupting potential of existing chemicals as well as prospective and newly prepared chemicals before they enter the environment.

Animals↗

Phytoestrogens and mycoestrogens bind to the rat uterine estrogen receptor.

Consumption of phytoestrogens and mycoestrogens in food products or as dietary supplements is of interest because of both the potential beneficial and adverse effects of these compounds in estrogen-responsive target tissues. Although the hazards of exposure to potent estrogens such as diethylstilbestrol in developing male and female reproductive tracts are well characterized, less is known about the effects of weaker estrogens including phytoestrogens. With some exceptions, ligand binding to the estrogen receptor (ER) predicts uterotrophic activity. Using a well-established and rigorously validated ER-ligand binding assay, we assessed the relative binding affinity (RBA) for 46 chemicals from several chemical structure classes of potential phytoestrogens and mycoestrogens. Although none of the test compounds bound to ER with the affinity of the standard, 17beta-estradiol (E(2)), ER binding was found among all classes of chemical structures (flavones, isoflavones, flavanones, coumarins, chalcones and mycoestrogens). Estrogen receptor relative binding affinities were distributed across a wide range (from approximately 43 to 0.00008; E(2) = 100). These data can be utilized before animal testing to rank order estimates of the potential for in vivo estrogenic activity of a wide range of untested plant chemicals (as well as other chemicals) based on ER binding.

Animals↗