Search PubMed⌕ Search

Biomedical subjects

Keith A Baggerly

Publications and source records attributed to Keith A Baggerly.

At least 19 recordsLinked to original sources

Surface PD-L1, E-cadherin, CD24, and VEGFR2 as markers of epithelial cancer stem cells associated with rapid tumorigenesis.

Cancer cells require both migratory and tumorigenic property to establish metastatic tumors outside the primary microenvironment. Identifying the characteristic features of migratory cancer stem cells with tumorigenic property is important to predict patient prognosis and combat metastasis. Here we established one epithelial and two mesenchymal cell lines from ascites of a bladder cancer patient (i.e. cells already migrated outside primary tumor). Analyses of these cell lines demonstrated that the epithelial cells with surface expression of PD-L1, E-cadherin, CD24, and VEGFR2 rapidly formed tumors outside the primary tumor microenvironment in nude mice, exhibited signatures of immune evasion, increased stemness, increased calcium signaling, transformation, and novel E-cadherin-RalBP1 interaction. The mesenchymal cells on the other hand, exhibited constitutive TGF-β signaling and were less tumorigenic. Hence, targeting epithelial cancer stem cells with rapid tumorigenesis signatures in future might help to combat metastasis.

B7-H1 Antigen↗

PrepMS: TOF MS data graphical preprocessing tool.

UNLABELLED: We introduce a simple-to-use graphical tool that enables researchers to easily prepare time-of-flight mass spectrometry data for analysis. For ease of use, the graphical executable provides default parameter settings, experimentally determined to work well in most situations. These values, if desired, can be changed by the user. PrepMS is a stand-alone application made freely available (open source), and is under the General Public License (GPL). Its graphical user interface, default parameter settings, and display plots allow PrepMS to be used effectively for data preprocessing, peak detection and visual data quality assessment. AVAILABILITY: Stand-alone executable files and Matlab toolbox are available for download at: http://sourceforge.net/projects/prepms

Algorithms↗

Synchronous selection of homing peptides for multiple tissues by in vivo phage display.

In vivo phage display is a technology used to reveal organ-specific vascular ligand-receptor systems in animal models and, recently, in patients, and to validate them as potential therapy targets. Here, we devised an efficient approach to simultaneously screen phage display libraries for peptides homing to any number of tissues without the need for an individual subject for each target tissue. We tested this approach in mice by selecting homing peptides for six different organs in a single screen and prioritizing them by using software compiled for statistical validation of peptide biodistribution specificity. We identified a number of motif-containing biological candidates for ligands binding to organ-selective receptors based on similarity of the selected peptide motifs to mouse proteins. To demonstrate that this methodology can lead to targetable ligand-receptor systems, we validated one of the pancreas-homing peptides as a mimic peptide of natural prolactin receptor ligands. This new comprehensive strategy for screening phage libraries in vivo provides an advantage over the conventional approach because multiple organs internally control for organ selectivity of each other in the successive rounds of selection. It may prove particularly relevant for patient studies, allowing efficient high-throughput selection of targeting ligands for multiple organs in a single screen.

Amino Acid Motifs↗

Patterns of gene expression in different histotypes of epithelial ovarian cancer correlate with those in normal fallopian tube, endometrium, and colon.

PURPOSE: Epithelial ovarian cancers are thought to arise from flattened epithelial cells that cover the ovarian surface or that line inclusion cysts. During malignant transformation, different histotypes arise that resemble epithelial cells from normal fallopian tube, endometrium, and intestine. This study compares gene expression in serous, endometrioid, clear cell, and mucinous ovarian cancers with that in the normal tissues that they resemble. EXPERIMENTAL DESIGN: Expression of 63,000 probe sets was measured in 50 ovarian cancers, in 5 pools of normal ovarian epithelial brushings, and in mucosal scrapings from 4 normal fallopian tube, 5 endometrium, and 4 colon specimens. Using rank-sum analysis, genes whose expressions best differentiated the ovarian cancer histotypes and normal ovarian epithelium were used to determine whether a correlation based on gene expression existed between ovarian cancer histotypes and the normal tissues they resemble. RESULTS: When compared with normal ovarian epithelial brushings, alterations in serous tumors correlated with those in normal fallopian tube (P = 0.0042) but not in other normal tissues. Similarly, mucinous cancers correlated with those in normal colonic mucosa (P = 0.0003), and both endometrioid and clear cell histotypes correlated with changes in normal endometrium (P = 0.0172 and 0.0002, respectively). Mucinous cancers displayed the greatest number of alterations in gene expression when compared with normal ovarian epithelial cells. CONCLUSION: Studies at a molecular level show distinct expression profiles of different histologies of ovarian cancer and support the long-held belief that histotypes of ovarian cancers come to resemble normal fallopian tube, endometrial, and colonic epithelium. Several potential molecular markers for mucinous ovarian cancers have been identified.

Adenocarcinoma, Clear Cell↗

Signal in noise: evaluating reported reproducibility of serum proteomic tests for ovarian cancer.

Proteomic profiling of serum initially appeared to be dramatically effective for diagnosis of early-stage ovarian cancer, but these results have proven difficult to reproduce. A recent publication reported good classification in one dataset using results from training on a much earlier dataset, but the authors have since reported that they did not perform the analysis as described. We examined the reproducibility of the proteomic patterns across datasets in more detail. Our analysis reveals that the pattern that enabled successful classification is biologically implausible and that the method, properly applied, does not classify the data accurately. We show that the method used in previously published studies does not establish reproducibility and performs no better than chance for classifying the second dataset, in part because the second dataset is easy to classify correctly. We conclude that the reproducibility of the proteomic profiling approach has yet to be established.

Female↗

Feature extraction and quantification for mass spectrometry in biomedical applications using the mean spectrum.

MOTIVATION: Mass spectrometry yields complex functional data for which the features of scientific interest are peaks. A common two-step approach to analyzing these data involves first extracting and quantifying the peaks, then analyzing the resulting matrix of peak quantifications. Feature extraction and quantification involves a number of interrelated steps. It is important to perform these steps well, since subsequent analyses condition on these determinations. Also, it is difficult to compare the performance of competing methods for analyzing mass spectrometry data since the true expression levels of the proteins in the population are generally not known. RESULTS: In this paper, we introduce a new method for feature extraction in mass spectrometry data that uses translation-invariant wavelet transforms and performs peak detection using the mean spectrum. We examine the method's performance through examples and simulation, and demonstrate the advantages of using the mean spectrum to detect peaks. We also describe a new physics-based computer model of mass spectrometry and demonstrate how one may design simulation studies based on this tool to systematically compare competing methods. AVAILABILITY: MATLAB scripts to implement the methods described in this paper and R code for the virtual mass spectrometer are available at http://bioinformatics.mdanderson.org/software.html SUPPLEMENTARY INFORMATION: http://bioinformatics.mdanderson.org/supplements.html.

Algorithms↗

Improved peak detection and quantification of mass spectrometry data acquired from surface-enhanced laser desorption and ionization by denoising spectra with the undecimated discrete wavelet transform.

Mass spectrometry is being used to find disease-related patterns in mixtures of proteins derived from biological fluids. Questions have been raised about the reproducibility and reliability of peak quantifications using this technology. We collected nipple aspirate fluid from breast cancer patients and healthy women, pooled them into a quality control sample, and produced 24 replicate SELDI spectra. We developed a novel algorithm to process the spectra, denoising with the undecimated discrete wavelet transform (UDWT), and evaluated it for consistency and reproducibility. UDWT efficiently decomposes spectra into noise and signal. The noise is consistent and uncorrelated. Baseline correction produces isolated peak clusters separated by flat regions. Our method reproducibly detects more peaks than the method implemented in Ciphergen software. After normalization and log transformation, the mean coefficient of variation of peak heights is 10.6%. Our method to process spectra provides improvements over existing methods. Denoising using the UDWT appears to be an important step toward obtaining results that are more accurate. It improves the reproducibility of quantifications and supplies tools for investigation of the variations in the technology more carefully. Further study will be required, because we do not have a gold standard providing an objective assessment of which peaks are present in the samples.

Algorithms↗

Diagnostic protein discovery using liquid chromatography/mass spectrometry for proteolytic peptide targeting.

A peptide targeting method has been developed for diagnostic protein discovery, which combines proteolytic digestion of fractionated plasma proteins and liquid chromatography coupled to electrospray time-of-flight mass spectrometry (LC/ESI-TOFMS) profiling. Proteolysis prior to profiling overcomes molecular weight limitations and compensates for the poor sensitivity of matrix-assisted laser desorption/ionization (MALDI) protein profiling. LC/MS increases the peak capacity compared to crude fractionation techniques or single sample MALDI analysis. Differentially expressed peptides are targeted in the mass chromatograms using bioinformatic techniques and subsequently sequenced with MALDI tandem MS. In a model study comparing pancreatic cancer patients to controls, 74% of the peptide targets were successfully sequenced. This profiling method was superior to previous experiments using single sample MALDI analysis for protein profiling or proteolytic peptide profiling, because more potential protein markers were identified.

Adult↗

The importance of experimental design in proteomic mass spectrometry experiments: some cautionary tales.

Proteomic expression patterns derived from mass spectrometry have been put forward as potential biomarkers for the early diagnosis of cancer and other diseases. This approach has generated much excitement and has led to a large number of new experiments and vast amounts of new data. The data, derived at great expense, can have very little value if careful attention is not paid to the experimental design and analysis. Using examples from surface-enhanced laser desorption/ionisation time-of-flight (SELDI-TOF) and matrix-assisted laser desorption-ionisation/time-of-flight (MALDI-TOF) experiments, we describe several experimental design issues that can corrupt a dataset. Fortunately, the problems we identify can be avoided if attention is paid to potential sources of bias before the experiment is run. With an appropriate experimental design, proteomics technology can be a useful tool for discovering important information relating protein expression to disease.

Biomarkers, Tumor↗

Overdispersed logistic regression for SAGE: modelling multiple groups and covariates.

BACKGROUND: Two major identifiable sources of variation in data derived from the Serial Analysis of Gene Expression (SAGE) are within-library sampling variability and between-library heterogeneity within a group. Most published methods for identifying differential expression focus on just the sampling variability. In recent work, the problem of assessing differential expression between two groups of SAGE libraries has been addressed by introducing a beta-binomial hierarchical model that explicitly deals with both of the above sources of variation. This model leads to a test statistic analogous to a weighted two-sample t-test. When the number of groups involved is more than two, however, a more general approach is needed. RESULTS: We describe how logistic regression with overdispersion supplies this generalization, carrying with it the framework for incorporating other covariates into the model as a byproduct. This approach has the advantage that logistic regression routines are available in several common statistical packages. CONCLUSIONS: The described method provides an easily implemented tool for analyzing SAGE data that correctly handles multiple types of variation and allows for more flexible modelling.

Analysis of Variance↗

Selection of potential markers for epithelial ovarian cancer with gene expression arrays and recursive descent partition analysis.

PURPOSE: Advanced-stage epithelial ovarian cancer has a poor prognosis with long-term survival in less than 30% of patients. When the disease is detected in stage I, more than 90% of patients can be cured by conventional therapy. Screening for early-stage disease with individual serum tumor markers, such as CA125, is limited by the fact that no single marker is up-regulated and shed in adequate amounts by all ovarian cancers. Consequently, use of multiple markers in combination might detect a larger fraction of early-stage ovarian cancers. EXPERIMENTAL DESIGN: To identify potential candidates for novel markers, we have used Affymetrix human genome arrays (U95 series) to analyze differences in gene expression of 41,441 known genes and expressed sequence tags between five pools of normal ovarian surface epithelial cells (OSE) and 42 epithelial ovarian cancers of different stages, grades, and histotypes. Recursive descent partition analysis (RDPA) was performed with 102 probe sets representing 86 genes that were up-regulated at least 3-fold in epithelial ovarian cancers when compared with normal OSE. In addition, a panel of 11 genes known to encode potential tumor markers [mucin 1, transmembrane (MUC1), mucin 16 (CA125), mesothelin, WAP four-disulfide core domain 2 (HE4), kallikrein 6, kallikrein 10, matrix metalloproteinase 2, prostasin, osteopontin, tetranectin, and inhibin] were similarly analyzed. RESULTS: The 3-fold up-regulated genes were examined and four genes [Notch homologue 3 (NOTCH3), E2F transcription factor 3 (E2F3), GTPase activating protein (RACGAP1), and hematological and neurological expressed 1 (HN1)] distinguished all tumor samples from normal OSE. The 3-fold up-regulated genes were analyzed using RDPA, and the combination of elevated claudin 3 (CLDN3) and elevated vascular endothelial growth factor (VEGF) distinguished the cancers from normal OSE. The 11 known markers were analyzed using RDPA, and a combination of HE4, CA125, and MUC1 expression could distinguish tumor from normal specimens. Expression at the mRNA level in the candidate markers was examined via semiquantitative reverse transcription-PCR and was found to correlate well with the array data. Immunohistochemistry was performed to identify expression of the genes at the protein level in 158 ovarian cancers of different histotypes. A combination of CLDN3, CA125, and MUC1 stained 157 (99.4%) of 158 cancers, and all of the tumors were detected with a combination of CLDN3, CA125, MUC1, and VEGF. CONCLUSIONS: Our data are consistent with the possibility that a limited number of markers in combination might identify >99% of epithelial ovarian cancers despite the heterogeneity of the disease.

Biomarkers, Tumor↗

Pharmacoproteomic analysis of prechemotherapy and postchemotherapy plasma samples from patients receiving neoadjuvant or adjuvant chemotherapy for breast carcinoma.

BACKGROUND: In this study, proteomic changes were examined in response to paclitaxel chemotherapy or 5-fluorouracil, doxorubicin, and cyclophosphamide (FAC) chemotherapy in plasma from patients with Stage I-III breast carcinoma. The authors also compared the plasma profiles of patients with cancer with the plasma profiles of healthy women to identify breast carcinoma-associated protein markers. METHODS: Sixty-nine patients and 15 healthy volunteers participated in the study. Plasma was sampled on Day 0 before chemotherapy and on Day 3 posttreatment in the 69 patients or 3 days apart in the 15 healthy women. Twenty-nine patients received preoperative chemotherapy, and 40 received postoperative chemotherapy. Surface-enhanced laser desorption/ionization mass spectrometry was used to generate protein mass profiles. RESULTS: Few changes were observed in plasma during treatment. Only 1 protein peak was identified (mass/charge ratio [m/z], 2790) that was induced by paclitaxel and, to a lesser extent, by FAC chemotherapy. This proteomic response was detectable in 80% of patients who were treated preoperatively but also was present with lesser intensity in approximately 40% of patients treated postoperatively. There was no clear correlation between induction of m/z 2790 during a single course of treatment and final tumor response to preoperative chemotherapy. Five other peaks also were identified that discriminated between plasma from patients with breast carcinoma and plasma from normal women. These same peaks also were detectable in a subset of patients who already had undergone surgery to remove their tumors. CONCLUSIONS: A single chemotherapy-inducible SELDI-MS peak and five other peaks that distinguished plasma obtained from patients with breast carcinoma from plasma obtained from normal, healthy women were identified. The (as yet unsequenced) proteins represented by these peaks are candidate markers of micrometastatic disease after surgery.

Adult↗

Reproducibility of SELDI-TOF protein patterns in serum: comparing datasets from different experiments.

MOTIVATION: There has been much interest in using patterns derived from surface-enhanced laser desorption and ionization (SELDI) protein mass spectra from serum to differentiate samples from patients both with and without disease. Such patterns have been used without identification of the underlying proteins responsible. However, there are questions as to the stability of this procedure over multiple experiments. RESULTS: We compared SELDI proteomic spectra from serum from three experiments by the same group on separating ovarian cancer from normal tissue. These spectra are available on the web at http://clinicalproteomics.steem.com. In general, the results were not reproducible across experiments. Baseline correction prevents reproduction of the results for two of the experiments. In one experiment, there is evidence of a major shift in protocol mid-experiment which could bias the results. In another, structure in the noise regions of the spectra allows us to distinguish normal from cancer, suggesting that the normals and cancers were processed differently. Sets of features found to discriminate well in one experiment do not generalize to other experiments. Finally, the mass calibration in all three experiments appears suspect. Taken together, these and other concerns suggest that much of the structure uncovered in these experiments could be due to artifacts of sample processing, not to the underlying biology of cancer. We provide some guidelines for design and analysis in experiments like these to ensure better reproducible, biologically meaningfully results. AVAILABILITY: The MATLAB and Perl code used in our analyses is available at http://bioinformatics.mdanderson.org

Algorithms↗

Differential expression in SAGE: accounting for normal between-library variation.

MOTIVATION: In contrasting levels of gene expression between groups of SAGE libraries, the libraries within each group are often combined and the counts for the tag of interest summed, and inference is made on the basis of these larger 'pseudolibraries'. While this captures the sampling variability inherent in the procedure, it fails to allow for normal variation in levels of the gene between individuals within the same group, and can consequently overstate the significance of the results. The effect is not slight: between-library variation can be hundreds of times the within-library variation. RESULTS: We introduce a beta-binomial sampling model that correctly incorporates both sources of variation. We show how to fit the parameters of this model, and introduce a test statistic for differential expression similar to a two-sample t-test.

Algorithms↗

A comprehensive approach to the analysis of matrix-assisted laser desorption/ionization-time of flight proteomics spectra from serum samples.

For our analysis of the data from the First Annual Proteomics Data Mining Conference, we attempted to discriminate between 24 disease spectra (group A) and 17 normal spectra (group B). First, we processed the raw spectra by (i) correcting for additive sinusoidal noise (periodic on the time scale) affecting most spectra, (ii) correcting for the overall baseline level, (iii) normalizing, (iv) recombining fractions, and (v) using variable-width windows for data reduction. Also, we identified a set of polymeric peaks (at multiples of 180.6 Da) that is present in several normal spectra (B1-B8). After data processing, we found the intensities at the following mass to charge (m/z) values to be useful discriminators: 3077, 12 886 and 74 263. Using these values, we were able to achieve an overall classification accuracy of 38/41 (92.6%). Perfect classification could be achieved by adding two additional peaks, at 2476 and 6955. We identified these values by applying a genetic algorithm to a filtered list of m/z values using Mahalanobis distance between the group means as a fitness function.

Blood Proteins↗

Bayesian shrinkage estimation of the relative abundance of mRNA transcripts using SAGE.

Serial analysis of gene expression (SAGE) is a technology for quantifying gene expression in biological tissue that yields count data that can be modeled by a multinomial distribution with two characteristics: skewness in the relative frequencies and small sample size relative to the dimension. As a result of these characteristics, a given SAGE sample may fail to capture a large number of expressed mRNA species present in the tissue. Empirical estimators of mRNA species' relative abundance effectively ignore these missing species, and as a result tend to overestimate the abundance of the scarce observed species comprising a vast majority of the total. We have developed a new Bayesian estimation procedure that quantifies our prior information about these characteristics, yielding a nonlinear shrinkage estimator with efficiency advantages over the MLE. Our prior is mixture of Dirichlets, whereby species are stochastically partitioned into abundant and scarce classes, each with its own multivariate prior. Simulation studies reveal our estimator has lower integrated mean squared error (IMSE) than the MLE for the SAGE scenarios simulated, and yields relative abundance profiles closer in Euclidean distance to the truth for all samples simulated. We apply our method to a SAGE library of normal colon tissue, and discuss its implications for assessing differential expression.

Bayes Theorem↗

Quality control and peak finding for proteomics data collected from nipple aspirate fluid by surface-enhanced laser desorption and ionization.

BACKGROUND: Recently, researchers have been using mass spectroscopy to study cancer. For use of proteomics spectra in a clinical setting, stringent quality-control procedures will be needed. METHODS: We pooled samples of nipple aspirate fluid from healthy breasts and breasts with cancer to prepare a control sample. Aliquots of the control sample were used on two spots on each of three IMAC ProteinChip arrays (Ciphergen Biosystems, Inc.) on 4 successive days to generate 24 SELDI spectra. In 36 subsequent experiments, the control sample was applied to two spots of each ProteinChip array, and the resulting spectra were analyzed to determine how closely they agreed with the original 24 spectra. RESULTS: We describe novel algorithms that (a) locate peaks in unprocessed proteomics spectra and (b) iteratively combine peak detection with baseline correction. These algorithms detected approximately 200 peaks per spectrum, 68 of which are detected in all 24 original spectra. The peaks were highly correlated across samples. Moreover, we could explain 80% of the variance, using only six principal components. Using a criterion that rejects a chip if the Mahalanobis distance from both control spectra to the center of the six-dimensional principal component space exceeds the 95% confidence limit threshold, we rejected 5 of the 36 chips. CONCLUSIONS: Mahalanobis distance in principal component space provides a method for assessing the reproducibility of proteomics spectra that is robust, effective, easily computed, and statistically sound.

Algorithms↗