Search PubMed⌕ Search

Biomedical subjects

Giovanni Parmigiani

Publications and source records attributed to Giovanni Parmigiani.

At least 19 recordsLinked to original sources

When should one subtract background fluorescence in 2-color microarrays?

Two-color microarrays are a powerful tool for genomic analysis, but have noise components that make inferences regarding gene expression inefficient and potentially misleading. Background fluorescence, whether attributable to nonspecific binding or other sources, is an important component of noise. The decision to subtract fluorescence surrounding spots of hybridization from spot fluorescence has been controversial, with no clear criteria for determining circumstances that may favor, or disfavor, background subtraction. While it is generally accepted that subtracting background reduces bias but increases variance in the estimates of the ratios of interest, no formal analysis of the bias-variance trade off of background subtraction has been undertaken. In this paper, we use simulation to systematically examine the bias-variance trade off under a variety of possible experimental conditions. Our simulation is based on data obtained from 2 self versus self microarray experiments and is free of distributional assumptions. Our results identify factors that are important for determining whether to background subtract, including the correlation of foreground to background intensity ratios. Using these results, we develop recommendations for diagnostic visualizations that can help decisions about background subtraction.

Algorithms↗

The genome and transcriptomes of the anti-tumor agent Clostridium novyi-NT.

Bacteriolytic anti-cancer therapies employ attenuated bacterial strains that selectively proliferate within tumors. Clostridium novyi-NT spores represent one of the most promising of these agents, as they generate potent anti-tumor effects in experimental animals. We have determined the 2.55-Mb genomic sequence of C. novyi-NT, identifying a new type of transposition and 139 genes that do not have homologs in other bacteria. The genomic sequence was used to facilitate the detection of transcripts expressed at various stages of the life cycle of this bacterium in vitro as well as in infections of tumors in vivo. Through this analysis, we found that C. novyi-NT spores contained mRNA and that the spore transcripts were distinct from those in vegetative forms of the bacterium.

Animals↗

Gene expression profiling reveals reproducible human lung adenocarcinoma subtypes in multiple independent patient cohorts.

PURPOSE: Published reports suggest that DNA microarrays identify clinically meaningful subtypes of lung adenocarcinomas not recognizable by other routine tests. This report is an investigation of the reproducibility of the reported tumor subtypes. METHODS: Three independent cohorts of patients with lung cancer were evaluated using a variety of DNA microarray assays. Using the integrative correlations method, a subset of genes was selected, the reliability of which was acceptable across the different DNA microarray platforms. Tumor subtypes were selected using consensus clustering and genes distinguishing subtypes were identified using the weighted difference statistic. Gene lists were compared across cohorts using centroids and gene set enrichment analysis. RESULTS: Cohorts of 31, 72, and 128 adenocarcinomas were generated for a total of 231 microarrays, each with 2,553 reliable genes. Three adenocarcinoma subtypes were identified in each cohort. These were named bronchioid, squamoid, and magnoid according to their respective correlations with gene expression patterns from histologically defined bronchioalveolar carcinoma, squamous cell carcinoma, and large-cell carcinoma. Tumor subtypes were distinguishable by many hundreds of genes, and lists generated in one cohort were predictive of tumor subtypes in the two other cohorts. Tumor subtypes correlated with clinically relevant covariates, including stage-specific survival and metastatic pattern. Most notably, bronchioid tumors were correlated with improved survival in early-stage disease, whereas squamoid tumors were associated with better survival in advanced disease. CONCLUSION: DNA microarray analysis of lung adenocarcinomas identified reproducible tumor subtypes which differ significantly in clinically important behaviors such as stage-specific survival.

Adenocarcinoma↗

Prediction of germline mutations and cancer risk in the Lynch syndrome.

CONTEXT: Identifying families at high risk for the Lynch syndrome (ie, hereditary nonpolyposis colorectal cancer) is critical for both genetic counseling and cancer prevention. Current clinical guidelines are effective but limited by applicability and cost. OBJECTIVE: To develop and validate a genetic counseling and risk prediction tool that estimates the probability of carrying a deleterious mutation in mismatch repair genes MLH1, MSH2, or MSH6 and the probability of developing colorectal or endometrial cancer. DESIGN, SETTING, AND PATIENTS: External validation of the MMRpro model was conducted on 279 individuals from 226 clinic-based families in the United States, Canada, and Australia (referred between 1993-2005) by comparing model predictions with results of highly sensitive germline mutation detection techniques. MMRpro models the autosomal dominant inheritance of mismatch repair mutations, with parameters based on meta-analyses of the penetrance and prevalence of mutations and of the predictive values of tumor characteristics. The model's prediction is tailored to each individual's detailed family history information on colorectal and endometrial cancer and to tumor characteristics including microsatellite instability. MAIN OUTCOME MEASURE: Ability of MMRpro to correctly predict mutation carrier status, as measured by operating characteristics, calibration, and overall accuracy. RESULTS: In the independent validation, MMRpro provided a concordance index of 0.83 (95% confidence interval, 0.78-0.88) and a ratio of observed to predicted cases of 0.94 (95% confidence interval, 0.84-1.05). This results in higher accuracy than existing alternatives and current clinical guidelines. CONCLUSIONS: MMRpro is a broadly applicable, accurate prediction model that can contribute to current screening and genetic counseling practices in a high-risk population. It is more sensitive and more specific than existing clinical guidelines for identifying individuals who may benefit from MMR germline testing. It is applicable to individuals for whom tumor samples are not available and to individuals in whom germline testing finds no mutation.

Adaptor Proteins, Signal Transducing↗

The consensus coding sequences of human breast and colorectal cancers.

The elucidation of the human genome sequence has made it possible to identify genetic alterations in cancers in unprecedented detail. To begin a systematic analysis of such alterations, we determined the sequence of well-annotated human protein-coding genes in two common tumor types. Analysis of 13,023 genes in 11 breast and 11 colorectal cancers revealed that individual tumors accumulate an average of approximately 90 mutant genes but that only a subset of these contribute to the neoplastic process. Using stringent criteria to delineate this subset, we identified 189 genes (average of 11 per tumor) that were mutated at significant frequency. The vast majority of these genes were not known to be genetically altered in tumors and are predicted to affect a wide range of cellular functions, including transcription, adhesion, and invasion. These data define the genetic landscape of two human cancer types, provide new targets for diagnostic and therapeutic intervention, and open fertile avenues for basic research in tumor biology.

Amino Acid Substitution↗

Three allele combinations associated with multiple sclerosis.

BACKGROUND: Multiple sclerosis (MS) is an immune-mediated disease of polygenic etiology. Dissection of its genetic background is a complex problem, because of the combinatorial possibilities of gene-gene interactions. As genotyping methods improve throughput, approaches that can explore multigene interactions appropriately should lead to improved understanding of MS. METHODS: 286 unrelated patients with definite MS and 362 unrelated healthy controls of Russian descent were genotyped at polymorphic loci (including SNPs, repeat polymorphisms, and an insertion/deletion) of the DRB1, TNF, LT, TGFbeta1, CCR5 and CTLA4 genes and TNFa and TNFb microsatellites. Each allele carriership in patients and controls was compared by Fisher's exact test, and disease-associated combinations of alleles in the data set were sought using a Bayesian Markov chain Monte Carlo-based method recently developed by our group. RESULTS: We identified two previously unknown MS-associated tri-allelic combinations:-509TGFbeta1*C, DRB1*18(3), CTLA4*G and -238TNF*B1,-308TNF*A2, CTLA4*G, which perfectly separate MS cases from controls, at least in the present sample. The previously described DRB1*15(2) allele, the microsatellite TNFa9 allele and the biallelic combination CCR5Delta32, DRB1*04 were also reidentified as MS-associated. CONCLUSION: These results represent an independent validation of MS association with DRB1*15(2) and TNFa9 in Russians and are the first to find the interplay of three loci in conferring susceptibility to MS. They demonstrate the efficacy of our approach for the identification of complex-disease-associated combinations of alleles.

Adult↗

Gene expression patterns in dendritic cells infected with measles virus compared with other pathogens.

Gene expression patterns supply insight into complex biological networks that provide the organization in which viruses and host cells interact. Measles virus (MV) is an important human pathogen that induces transient immunosuppression followed by life-long immunity in infected individuals. Dendritic cells (DCs) are potent antigen-presenting cells that initiate the immune response to pathogens and are postulated to play a role in MV-induced immunosuppression. To better understand the interaction of MV with DCs, we examined the gene expression changes that occur over the first 24 h after infection and compared these changes to those induced by other viral, bacterial, and fungal pathogens. There were 1,553 significantly regulated genes with nearly 60% of them down-regulated. MV-infected DCs up-regulated a core of genes associated with maturation of antigen-presenting function and migration to lymph nodes but also included genes for IFN-regulatory factors 1 and 7, 2'5' oligoadenylate synthetase, Mx, and TNF superfamily proteins 2, 7, 9, and 10 (TNF-related apoptosis-inducing ligand). MV induced genes for IFNs, ILs, chemokines, antiviral proteins, histones, and metallothioneins, many of which were also induced by influenza virus, whereas genes for protein synthesis and oxidative phosphorylation were down-regulated. Unique to MV were the induction of genes for a broad array of IFN-alphas and the failure to up-regulate dsRNA-dependent protein kinase. These results provide a modular view of common and unique DC responses after infection and suggest mechanisms by which MV may modulate the immune response.

Animals↗

Characterization of BRCA1 and BRCA2 mutations in a large United States sample.

PURPOSE: An accurate evaluation of the penetrance of BRCA1 and BRCA2 mutations is essential to the identification and clinical management of families at high risk of breast and ovarian cancer. Existing studies have focused on Ashkenazi Jews (AJ) or on families from outside the United States. In this article, we consider the US population using the largest US-based cohort to date of both AJ and non-AJ families. METHODS: We collected 676 AJ families and 1,272 families of other ethnicities through the Cancer Genetics Network. Two hundred eighty-two AJ families were population based, whereas the remainder was collected through counseling clinics. We used a retrospective likelihood approach to correct for bias induced by oversampling of participants with a positive family history. Our approach takes full advantage of detailed family history information and the Mendelian transmission of mutated alleles in the family. RESULTS: In the US population, the estimated cumulative breast cancer risk at age 70 years was 0.46 (95% CI, 0.39 to 0.54) in BRCA1 carriers and 0.43 (95% CI, 0.36 to 0.51) in BRCA2 carriers, whereas ovarian cancer risk was 0.39 (95% CI, 0.30 to 0.50) in BRCA1 carriers and 0.22 (95% CI, 0.14 to 0.32) in BRCA2 carriers. We also reported the prospective risks of developing cancer for cancer-free carriers in 10-year age intervals. We noted a rapid decrease in the relative risk of breast cancer with age and derived its implication for genetic counseling. CONCLUSION: The penetrance of BRCA mutations in the United States is largely consistent with previous studies on Western populations given the large CIs on existing estimates. However, the absolute cumulative risks are on the lower end of the spectrum.

Adult↗

Analysis of the human protein interactome and comparison with yeast, worm and fly interaction datasets.

We present the first analysis of the human proteome with regard to interactions between proteins. We also compare the human interactome with the available interaction datasets from yeast (Saccharomyces cerevisiae), worm (Caenorhabditis elegans) and fly (Drosophila melanogaster). Of >70,000 binary interactions, only 42 were common to human, worm and fly, and only 16 were common to all four datasets. An additional 36 interactions were common to fly and worm but were not observed in humans, although a coimmunoprecipitation assay showed that 9 of the interactions do occur in humans. A re-examination of the connectivity of essential genes in yeast and humans indicated that the available data do not support the presumption that the number of interaction partners can accurately predict whether a gene is essential. Finally, we found that proteins encoded by genes mutated in inherited genetic disorders are likely to interact with proteins known to cause similar disorders, suggesting the existence of disease subnetworks. The human interaction map constructed from our analysis should facilitate an integrative systems biology approach to elucidating the cellular networks that contribute to health and disease states.

Animals↗

Amplification of a chromatin remodeling gene, Rsf-1/HBXAP, in ovarian carcinoma.

A genomewide technology, digital karyotyping, was used to identify subchromosomal alterations in ovarian cancer. Amplification at 11q13.5 was found in three of seven ovarian carcinomas, and amplicon mapping delineated a 1.8-Mb core of amplification that contained 13 genes. FISH analysis demonstrated amplification of this region in 13.2% of high-grade ovarian carcinomas but not in any of low-grade carcinomas or benign ovarian tumors. Combined genetic and transcriptome analyses showed that Rsf-1 (HBXAPalpha) was the only gene that demonstrated consistent overexpression in all of the tumors harboring the 11q13.5 amplification. Patients with Rsf-1 amplification or overexpression had a significantly shorter overall survival than those without. Overexpression of Rsf-1 gene stimulated cell proliferation and transform nonneoplastic cells by conferring serum-independent and anchorage-independent growth. Furthermore, Rsf-1 gene knockdown inhibited cell growth in OVCAR3 cells, which harbor Rsf-1 amplification. Taken together, these findings indicate an important role of Rsf-1 amplification in ovarian cancer.

Carcinoma↗

Searching for differentially expressed gene combinations.

We propose 'CorScor', a novel approach for identifying gene pairs with joint differential expression. This is defined as a situation with good phenotype discrimination in the bivariate, but not in the two marginal distributions. CorScor can be used to detect phenotype-related dependencies and interactions among genes. Our easily interpretable approach is scalable to current microarray dimensions and yields promising results on several cancer-gene-expression datasets.

Gene Expression Profiling↗

A Markov chain Monte Carlo technique for identification of combinations of allelic variants underlying complex diseases in humans.

In recent years, the number of studies focusing on the genetic basis of common disorders with a complex mode of inheritance, in which multiple genes of small effect are involved, has been steadily increasing. An improved methodology to identify the cumulative contribution of several polymorphous genes would accelerate our understanding of their importance in disease susceptibility and our ability to develop new treatments. A critical bottleneck is the inability of standard statistical approaches, developed for relatively modest predictor sets, to achieve power in the face of the enormous growth in our knowledge of genomics. The inability is due to the combinatorial complexity arising in searches for multiple interacting genes. Similar "curse of dimensionality" problems have arisen in other fields, and Bayesian statistical approaches coupled to Markov chain Monte Carlo (MCMC) techniques have led to significant improvements in understanding. We present here an algorithm, APSampler, for the exploration of potential combinations of allelic variations positively or negatively associated with a disease or with a phenotype. The algorithm relies on the rank comparison of phenotype for individuals with and without specific patterns (i.e., combinations of allelic variants) isolated in genetic backgrounds matched for the remaining significant patterns. It constructs a Markov chain to sample only potentially significant variants, minimizing the potential of large data sets to overwhelm the search. We tested APSampler on a simulated data set and on a case-control MS (multiple sclerosis) study for ethnic Russians. For the simulated data, the algorithm identified all the phenotype-associated allele combinations coded into the data and, for the MS data, it replicated the previously known findings.

Algorithms↗

Assessing reproducibility of a protein dynamics study using in vivo labeling and liquid chromatography tandem mass spectrometry.

Measuring dynamics of proteins abundance in cells in response to stimuli such as growth factors or drugs requires analysis of more than one time point. Proteomic approaches have traditionally been used to measure only one state at a time because quantitation is difficult, especially when mass spectrometry is used as a readout. Isotopically labeled reagents have recently been introduced that allow comparison of two or three different states by mass spectrometry. Here, we evaluate the reproducibility of an experiment that measures three states simultaneously through stable isotope labeling of cells with amino acids in cell culture (SILAC) using light, medium, and heavy versions of amino acids. The major goal of this study was to assess the reproducibility of such experiments in combination with liquid chromatography tandem mass spectrometry (LC-MS/MS). Our results show that it is possible to obtain reproducible quantitative data to study protein dynamics based on our analysis of more than 220 peptide sets derived from 20 proteins from 3 different LC-MS/MS runs.

Arginine↗

Accuracy of MSI testing in predicting germline mutations of MSH2 and MLH1: a case study in Bayesian meta-analysis of diagnostic tests without a gold standard.

Microsatellite instability (MSI) testing is a common screening procedure used to identify families that may harbor mutations of a mismatch repair (MMR) gene and therefore may be at high risk for hereditary colorectal cancer. A reliable estimate of sensitivity and specificity of MSI for detecting germline mutations of MMR genes is critical in genetic counseling and colorectal cancer prevention. Several studies published results of both MSI and mutation analysis on the same subjects. In this article we perform a meta-analysis of these studies and obtain estimates that can be directly used in counseling and screening. In particular, we estimate the sensitivity of MSI for detecting mutations of MSH2 and MLH1 to be 0.81 (0.73-0.89). Statistically, challenges arise from the following: (a) traditional mutation analysis methods used in these studies cannot be considered a gold standard for the identification of mutations; (b) studies are heterogeneous in both the design and the populations considered; and (c) studies may include different patterns of missing data resulting from partial testing of the populations sampled. We address these challenges in the context of a Bayesian meta-analytic implementation of the Hui-Walter design, tailored to account for various forms of incomplete data. Posterior inference is handled via a Gibbs sampler.

Adaptor Proteins, Signal Transducing↗

Impact of the Cancer Risk Intake System on patient-clinician discussions of tamoxifen, genetic counseling, and colonoscopy.

The Cancer Risk Intake System (CRIS), a computerized program that "matches" objective cancer risks to appropriate risk management recommendations, was designed to facilitate patient-clinician discussion. We evaluated CRIS in primary care settings via a single-group, self-report, pretest-posttest design. Participants completed baseline telephone surveys, used CRIS during clinic visits, and completed follow-up surveys 1 to 2 months postvisit. Compared with proportions reporting having had discussions at baseline, significantly greater proportions of participants reported having discussed tamoxifen, genetic counseling, and colonoscopy, as appropriate, after using CRIS. Most (79%) reported CRIS had "caused" their discussion. CRIS is an easily used, disseminable program that showed promising results in primary care settings.

Adult↗

A model-based comparison of breast cancer screening strategies: mammograms and clinical breast examinations.

In screening for secondary prevention of breast cancer, clinical breast examination (CBE) combined with mammography may improve overall screening sensitivity compared with mammography alone. A systematic evaluation of the relative expenses and projected benefit of combining these two screening modalities is not presently available. We addressed this issue using a microsimulation model incorporating age-specific preclinical duration of the disease, age-specific sensitivities of the two modalities, age-specific incidence of the disease, screening strategy, and competing causes of mortality. We examined a total of 48 screening strategies, depending on the age range, the examination interval, and whether mammography or CBE is given at every one or two exam. Our results indicate that a biennial mammography can be cost-effective if coupled with annual CBE. For each screening interval and starting age, giving mammography every two exams and CBE at every exam has the lowest marginal cost per year of quality-adjusted life saved, whereas giving both at every exam has the highest. Comparing annual mammography and CBE to biennial mammography and annual CBE from 50 to 79, the total cost was reduced by 35%, whereas the marginal quality-adjusted life years only decreased by 12%. Similar reductions are observed for other starting ages. It is cost-effective to have a biennial mammography if coupled with an annual CBE. Annual mammography combined with CBE every 6 months will lead to a 41% increase in the quality-adjusted life years compared with annual mammography and CBE from 50 to 79, whereas the total cost increases by 30%.

Adult↗

Identification of a gene expression profile that differentiates between ischemic and nonischemic cardiomyopathy.

BACKGROUND: Gene expression profiling refines diagnostic and prognostic assessment in oncology but has not yet been applied to myocardial diseases. We hypothesized that gene expression differentiates ischemic and nonischemic cardiomyopathy, demonstrating that gene expression profiling by clinical parameters is feasible in cardiology. METHODS AND RESULTS: Affymetrix U133A microarrays of 48 myocardial samples from Johns Hopkins Hospital (JHH) and the University of Minnesota (UM) obtained (1) at transplantation or left ventricular assist device (LVAD) placement (end-stage; n=25), (2) after LVAD support (post-LVAD; n=16), and (3) from newly diagnosed patients (biopsy; n=7) were analyzed with prediction analysis of microarrays. A training set was used to develop the profile and test sets to validate the accuracy of the profile. An etiology prediction profile developed in end-stage JHH samples was tested in independent samples from both JHH and UM with 100% sensitivity and 100% specificity in end-stage samples and 33% sensitivity and 100% specificity in both post-LVAD and biopsy samples. The overall sensitivity was 89% (95% CI 75% to 100%), and specificity was 89% (95% CI 60% to 100%) over 210 random partitions of end-stage samples into training and test sets. Age, gender, and hemodynamic differences did not affect the profile's accuracy in stratified analyses. Select gene expression was confirmed with quantitative polymerase chain reaction. CONCLUSIONS: Gene expression profiling accurately predicts cardiomyopathy etiology, is generalizable to samples from separate institutions, is specific to disease stage, and is unaffected by differences in clinical characteristics. This strongly supports ongoing efforts to incorporate expression profiling-based biomarkers in determining prognosis and response to therapy in heart failure.

Adult↗

MergeMaid: R tools for merging and cross-study validation of gene expression data.

Cross-study validation of gene expression investigations is critical in genomic analysis. We developed an R package and associated object definitions to merge and visualize multiple gene expression datasets. Our merging functions use arbitrary character IDs and generate objects that can efficiently support a variety of joint analyses. Visualization tools support exploration and cross-study validation of the data, without requiring normalization across platforms. Tools include "integrative correlation'' plots that is, scatterplots of all pairwise correlations in one study against the corresponding pairwise correlations of another, both for individual genes and all genes combined. Gene-specific plots can be used to identify genes whose changes are reliably measured across studies. Visualizations also include scatterplots of gene-specific statistics quantifying relationships between expression and phenotypes of interest, using linear, logistic and Cox regression.

Journal Article↗