Search PubMed⌕ Search

Biomedical subjects

Roger E Bumgarner

Publications and source records attributed to Roger E Bumgarner.

At least 19 recordsLinked to original sources

On the persistence of supplementary resources in biomedical publications.

BACKGROUND: Providing for long-term and consistent public access to scientific data is a growing concern in biomedical research. One aspect of this problem can be demonstrated by evaluating the persistence of supplementary data associated with published biomedical papers. METHODS: We manually evaluated 655 supplementary data links extracted from PubMed abstracts published 1998-2005 (Method 1) as well as a further focused subset of 162 full-text manuscripts published within three representative high-impact biomedical journals between September and December 2004 (Method 2). RESULTS: For Method 1 we found that since 2001, only 71 - 92% of supplementary data were still accessible via the links provided, with 93% of these inaccessible links occurring where supplementary data was not stored with the publishing journal. Of the manuscripts evaluated in Method 2, we found that only 83% of these links were available approximately a year after publication, with 55% of these inaccessible links were at locations outside the journal of publication. CONCLUSION: We conclude that if supplemental data is required to support the publication, journals policies must take-on the responsibility to accept and store such data or require that it be maintained with a credible independent institution or under the terms of a strategic data storage plan specified by the authors. We further recommend that publishers provide automated systems to ensure that supplementary links remain persistent, and that granting bodies such as the NIH develop policies and funding mechanisms to maintain long-term persistent access to these data.

Abstracting and Indexing↗

Development of the Minimum Information Specification for In Situ Hybridization and Immunohistochemistry Experiments (MISFISHIE).

We describe the creation process of the Minimum Information Specification for In Situ Hybridization and Immunohistochemistry Experiments (MISFISHIE). Modeled after the existing minimum information specification for microarray data, we created a new specification for gene expression localization experiments, initially to facilitate data sharing within a consortium. After successful use within the consortium, the specification was circulated to members of the wider biomedical research community for comment and refinement. After a period of acquiring many new suggested requirements, it was necessary to enter a final phase of excluding those requirements that were deemed inappropriate as a minimum requirement for all experiments. The full specification will soon be published as a version 1.0 proposal to the community, upon which a more full discussion must take place so that the final specification may be achieved with the involvement of the whole community.

Computational Biology↗

Bayesian robust inference for differential gene expression in microarrays with multiple samples.

We consider the problem of identifying differentially expressed genes under different conditions using gene expression microarrays. Because of the many steps involved in the experimental process, from hybridization to image analysis, cDNA microarray data often contain outliers. For example, an outlying data value could occur because of scratches or dust on the surface, imperfections in the glass, or imperfections in the array production. We develop a robust Bayesian hierarchical model for testing for differential expression. Errors are modeled explicitly using a t-distribution, which accounts for outliers. The model includes an exchangeable prior for the variances, which allows different variances for the genes but still shrinks extreme empirical variances. Our model can be used for testing for differentially expressed genes among multiple samples, and it can distinguish between the different possible patterns of differential expression when there are three or more samples. Parameter estimation is carried out using a novel version of Markov chain Monte Carlo that is appropriate when the model puts mass on subspaces of the full parameter space. The method is illustrated using two publicly available gene expression data sets. We compare our method to six other baseline and commonly used techniques, namely the t-test, the Bonferroni-adjusted t-test, significance analysis of microarrays (SAM), Efron's empirical Bayes, and EBarrays in both its lognormal-normal and gamma-gamma forms. In an experiment with HIV data, our method performed better than these alternatives, on the basis of between-replicate agreement and disagreement.

Bayes Theorem↗

Human rhinovirus attenuates the type I interferon response by disrupting activation of interferon regulatory factor 3.

The type I interferon (IFN) response requires the coordinated activation of the latent transcription factors NF-kappaB, interferon regulatory factor 3 (IRF-3), and ATF-2, which in turn activate transcription from the IFN-beta promoter. Synthesis and subsequent secretion of IFN-beta activate the Jak/STAT signaling pathway, resulting in the transcriptional induction of the full spectrum of antiviral gene products. We utilized high-density microarrays to examine the transcriptional response to rhinovirus type 14 (RV14) infection in HeLa cells, with particular emphasis on the type I interferon response and production of IFN-beta. We found that, although RV14 infection results in altered levels of a wide variety of host mRNAs, induction of IFN-beta mRNA or activation of the Jak/STAT pathway is not seen. Prior work has shown, and our results have confirmed, that NF-kappaB and ATF-2 are activated following infection. Since many viruses are known to target IRF-3 to inhibit the induction of IFN-beta mRNA, we analyzed the status of IRF-3 in infected cells. IRF-3 was translocated to the nucleus and phosphorylated in RV14-infected cells. Despite this apparent activation, very little homodimerization of IRF-3 was evident following infection. Similar results in A549 lung alveolar epithelial cells demonstrated the biological relevance of these findings to RV14 pathogenesis. In addition, prior infection of cells with RV14 prevented the induction of IFN-beta mRNA following treatment with double-stranded RNA, indicating that RV14 encodes an activity that specifically inhibits this innate host defense pathway. Collectively, these results indicate that RV14 infection inhibits the host type I interferon response by interfering with IRF-3 activation.

Activating Transcription Factor 2↗

Chlamydia trachomatis variant with nonfusing inclusions: growth dynamic and host-cell transcriptional response.

We compared growth rate and host-cell transcriptional responses of a Chlamydia trachomatis variant strain and a prototype strain. Growth dynamics were estimated by 16S rRNA level and by inclusion-forming units (IFUs) at different times after infection in HeLa cells. When inoculated at the same multiplicity of infection and observed 24-48 h after infection, the variant 16S rRNA transcriptional level was 3%-4% that of the prototype, and the IFUs of the variant strain were 0.1%-1% those of the prototype. Specific host-cell transcriptional responses to the variant were identified in a global-expression microarray in which variant strain-infected cells were compared with mock-infected and prototype strain-infected cells. In variant strain-infected cells, 47% (16/34) of specifically induced host genes were related to immunity and 32% (8/25) of specifically suppressed genes were related to lipid metabolism. The variant strain grew significantly more slowly and induced a modified host-cell transcriptional response, compared with the prototype strain.

Chlamydia trachomatis↗

Identification of high and low responders to lipopolysaccharide in normal subjects: an unbiased approach to identify modulators of innate immunity.

LPS stimulates a vigorous inflammatory response from circulating leukocytes that varies greatly from individual to individual. The goal of this study was to use an unbiased approach to identify differences in gene expression that may account for the high degree of interindividual variability in inflammatory responses to LPS in the normal human population. We measured LPS-induced cytokine production ex vivo in whole blood from 102 healthy human subjects and identified individuals who consistently showed either very high or very low responses to LPS (denoted lps(high) and lps(low), respectively). Comparison of gene expression profiles between the lps(high) and lps(low) individuals revealed 80 genes that were differentially expressed in the presence of LPS and 21 genes that were differentially expressed in the absence of LPS (p < 0.005, ANOVA). Expression of a subset of these genes was confirmed using real-time RT-PCR. Functional relevance for one gene confirmed to be expressed at a higher level in lps(high), adipophilin, was inferred when reduction in adipophilin mRNA by small interfering RNA in the human monocyte-like cell line THP-1 resulted in a modest but significant reduction in LPS-induced MCP-1 mRNA expression. These data illustrate a novel approach to the identification of factors that determine interindividual variability in innate immune inflammatory responses and identify adipophilin as a novel potential regulator of LPS-induced MCP-1 production in human monocytes.

Adolescent↗

Standardizing global gene expression analysis between laboratories and across platforms.

To facilitate collaborative research efforts between multi-investigator teams using DNA microarrays, we identified sources of error and data variability between laboratories and across microarray platforms, and methods to accommodate this variability. RNA expression data were generated in seven laboratories, which compared two standard RNA samples using 12 microarray platforms. At least two standard microarray types (one spotted, one commercial) were used by all laboratories. Reproducibility for most platforms within any laboratory was typically good, but reproducibility between platforms and across laboratories was generally poor. Reproducibility between laboratories increased markedly when standardized protocols were implemented for RNA labeling, hybridization, microarray processing, data acquisition and data normalization. Reproducibility was highest when analysis was based on biological themes defined by enriched Gene Ontology (GO) categories. These findings indicate that microarray results can be comparable across multiple laboratories, especially when a common platform and set of procedures are used.

Gene Expression Profiling↗

Donuts, scratches and blanks: robust model-based segmentation of microarray images.

MOTIVATION: Inner holes, artifacts and blank spots are common in microarray images, but current image analysis methods do not pay them enough attention. We propose a new robust model-based method for processing microarray images so as to estimate foreground and background intensities. The method starts with a very simple but effective automatic gridding method, and then proceeds in two steps. The first step applies model-based clustering to the distribution of pixel intensities, using the Bayesian Information Criterion (BIC) to choose the number of groups up to a maximum of three. The second step is spatial, finding the large spatially connected components in each cluster of pixels. The method thus combines the strengths of the histogram-based and spatial approaches. It deals effectively with inner holes in spots and with artifacts. It also provides a formal inferential basis for deciding when the spot is blank, namely when the BIC favors one group over two or three. RESULTS: We apply our methods for gridding and segmentation to cDNA microarray images from an HIV infection experiment. In these experiments, our method had better stability across replicates than a fixed-circle segmentation method or the seeded region growing method in the SPOT software, without introducing noticeable bias when estimating the intensities of differentially expressed genes. AVAILABILITY: spotSegmentation, an R language package implementing both the gridding and segmentation methods is available through the Bioconductor project (http://www.bioconductor.org). The segmentation method requires the contributed R package MCLUST for model-based clustering (http://cran.us.r-project.org). CONTACT: fraley@stat.washington.edu.

Algorithms↗

Bayesian model averaging: development of an improved multi-class, gene selection and classification tool for microarray data.

MOTIVATION: Selecting a small number of relevant genes for accurate classification of samples is essential for the development of diagnostic tests. We present the Bayesian model averaging (BMA) method for gene selection and classification of microarray data. Typical gene selection and classification procedures ignore model uncertainty and use a single set of relevant genes (model) to predict the class. BMA accounts for the uncertainty about the best set to choose by averaging over multiple models (sets of potentially overlapping relevant genes). RESULTS: We have shown that BMA selects smaller numbers of relevant genes (compared with other methods) and achieves a high prediction accuracy on three microarray datasets. Our BMA algorithm is applicable to microarray datasets with any number of classes, and outputs posterior probabilities for the selected genes and models. Our selected models typically consist of only a few genes. The combination of high accuracy, small numbers of genes and posterior probabilities for the predictions should make BMA a powerful tool for developing diagnostics from expression data. AVAILABILITY: The source codes and datasets used are available from our Supplementary website.

Algorithms↗

Multifunctionality of PAI-1 in fibrogenesis: evidence from obstructive nephropathy in PAI-1-overexpressing mice.

BACKGROUND: Plasminogen activator inhibitor-1 (PAI-1) has been implicated in the pathogenesis of chronic kidney disease based on its up-regulated expression and on the beneficial effects of PAI-1 inhibition or depletion in experimental models. PAI-1 is a multifunctional protein and the mechanisms that account for its profibrotic effects have not been fully elucidated. METHODS: The present study was designed to investigate PAI-1-dependent fibrogenic pathways by comparing the unilateral ureteral obstruction model (UUO) (days 3, 7, and 14) in PAI-1-overexpressing mice (PAI-1 tg) to wild-type mice, both on a C57BL6 background. RESULTS: Following UUO, total kidney PAI-1 mRNA and/or protein levels were significantly higher in the PAI-1 tg mice (N= 6 to 8/group) and fibrosis severity was significantly worse (days 3, 7, and 14), measured both as Sirius red-positive interstitial area (e.g., 10 +/- 3.2% vs. 4.5 +/- 1.0%) (day 14) and total kidney collagen (e.g., 11.1 +/- 1.7 vs. 6.2 +/- 1.3 microg/mg) (day 14). By day 14, the expression of two normal tubular proteins, E-cadherin and Ksp-cadherin, were significantly lower in the PAI-1 tg mice (3.2 +/- 0.5% vs. 11.7 +/- 5.9% and 2.6 +/- 1.6) vs. 6.2 +/- 0.8%, respectively), implying more extensive tubular damage. At least four fibrogenic pathways were differentially expressed in the PAI-1 tg mice. First, interstitial macrophage recruitment was more intense (P < 0.05 days 3 and 14). Second, interstitial myofibroblast density was greater (P < 0.05 days 3 and 7) despite similar numbers of proliferating tubulointerstitial cells. Third, transforming growth factor-beta1 (TGF-beta1) and collagen I mRNA were significantly higher. Finally, urokinase activity was significantly lower (P < 0.05 days 7 and 14) despite similar mRNA levels. Gene microarray studies documented that that the deletion of this single profibrotic gene had far-reaching consequences on renal cellular responses to chronic injury. CONCLUSION: These data provide further evidence that PAI-1 is directly involved in interstitial fibrosis and tubular damage via two primary overlapping mechanisms: early effects on interstitial cell recruitment and late effects associated with decreased urokinase activity.

Animals↗

Sample size for detecting differentially expressed genes in microarray experiments.

BACKGROUND: Microarray experiments are often performed with a small number of biological replicates, resulting in low statistical power for detecting differentially expressed genes and concomitant high false positive rates. While increasing sample size can increase statistical power and decrease error rates, with too many samples, valuable resources are not used efficiently. The issue of how many replicates are required in a typical experimental system needs to be addressed. Of particular interest is the difference in required sample sizes for similar experiments in inbred vs. outbred populations (e.g. mouse and rat vs. human). RESULTS: We hypothesize that if all other factors (assay protocol, microarray platform, data pre-processing) were equal, fewer individuals would be needed for the same statistical power using inbred animals as opposed to unrelated human subjects, as genetic effects on gene expression will be removed in the inbred populations. We apply the same normalization algorithm and estimate the variance of gene expression for a variety of cDNA data sets (humans, inbred mice and rats) comparing two conditions. Using one sample, paired sample or two independent sample t-tests, we calculate the sample sizes required to detect a 1.5-, 2-, and 4-fold changes in expression level as a function of false positive rate, power and percentage of genes that have a standard deviation below a given percentile. CONCLUSIONS: Factors that affect power and sample size calculations include variability of the population, the desired detectable differences, the power to detect the differences, and an acceptable error rate. In addition, experimental design, technical variability and data pre-processing play a role in the power of the statistical tests in microarrays. We show that the number of samples required for detecting a 2-fold change with 90% probability and a p-value of 0.01 in humans is much larger than the number of samples commonly used in present day studies, and that far fewer individuals are needed for the same statistical power when using inbred animals rather than unrelated human subjects.

Animals↗

From co-expression to co-regulation: how many microarray experiments do we need?

BACKGROUND: Cluster analysis is often used to infer regulatory modules or biological function by associating unknown genes with other genes that have similar expression patterns and known regulatory elements or functions. However, clustering results may not have any biological relevance. RESULTS: We applied various clustering algorithms to microarray datasets with different sizes, and we evaluated the clustering results by determining the fraction of gene pairs from the same clusters that share at least one known common transcription factor. We used both yeast transcription factor databases (SCPD, YPD) and chromatin immunoprecipitation (ChIP) data to evaluate our clustering results. We showed that the ability to identify co-regulated genes from clustering results is strongly dependent on the number of microarray experiments used in cluster analysis and the accuracy of these associations plateaus at between 50 and 100 experiments on yeast data. Moreover, the model-based clustering algorithm MCLUST consistently outperforms more traditional methods in accurately assigning co-regulated genes to the same clusters on standardized data. CONCLUSIONS: Our results are consistent with respect to independent evaluation criteria that strengthen our confidence in our results. However, when one compares ChIP data to YPD, the false-negative rate is approximately 80% using the recommended p-value of 0.001. In addition, we showed that even with large numbers of experiments, the false-positive rate may exceed the true-positive rate. In particular, even when all experiments are included, the best results produce clusters with only a 28% true-positive rate using known gene transcription factor interactions.

Algorithms↗

Multiclass classification of microarray data with repeated measurements: application to cancer.

Prediction of the diagnostic category of a tissue sample from its gene-expression profile and selection of relevant genes for class prediction have important applications in cancer research. We have developed the uncorrelated shrunken centroid (USC) and error-weighted, uncorrelated shrunken centroid (EWUSC) algorithms that are applicable to microarray data with any number of classes. We show that removing highly correlated genes typically improves classification results using a small set of genes.

Algorithms↗

Clustering gene-expression data with repeated measurements.

Clustering is a common methodology for the analysis of array data, and many research laboratories are generating array data with repeated measurements. We evaluated several clustering algorithms that incorporate repeated measurements, and show that algorithms that take advantage of repeated measurements yield more accurate and more stable clusters. In particular, we show that the infinite mixture model-based approach with a built-in error model produces superior results.

Algorithms↗

Identification of novel tumor markers in hepatitis C virus-associated hepatocellular carcinoma.

Hepatocellular carcinoma (HCC) is a common primary cancer associated frequently with hepatitis C virus (HCV). To gain insight into the molecular mechanisms of hepatocarcinogenesis, and to identify potential HCC markers, we performed cDNA microarray analysis on surgical liver samples from 20 HCV-infected patients. RNA from individual tumors was compared with RNA isolated from adjacent nontumor tissue that was cirrhotic in all of the cases. Gene expression changes related to cirrhosis were filtered out using experiments in which pooled RNA from HCV-infected cirrhotic liver without tumors was compared with pooled RNA from normal liver. Expression of approximately 13,600 genes was analyzed using the advanced analysis tools of the Rosetta Resolver System. This analysis revealed a set of 50 potential HCC marker genes, which were up-regulated in the majority of the tumors analyzed, much more widely than common clinical markers such as cell proliferation-related genes. This HCC marker set contained several cancer-related genes, including serine/threonine kinase 15 (STK15), which has been implicated in chromosome segregation abnormalities but which has not been linked previously with liver cancer. In addition, a set of genes encoding secreted or plasma proteins was identified, including plasma glutamate carboxypeptidase (PGCP) and two secreted phospholipases A2 (PLA2G13 and PLA2G7). These genes may provide potential HCC serological markers because of their strong up-regulation in more than half of the tumors analyzed. Thus, high throughput methods coupled with high-order statistical analyses may result in the development of new diagnostic tools for liver malignancies.

Aged↗

Chlamydia trachomatis infection alters host cell transcription in diverse cellular pathways.

To study the responses of the host cell to chlamydial infection, differentially transcribed genes of the host cells were examined. Complementary DNA (cDNA) probes were made from messenger RNAs of HeLa cells infected with Chlamydia trachomatis and were hybridized to a high-density human DNA microarray of 15,000 genes and expressed sequence tags. C. trachomatis alters host cell transcription at both the early and middle phases of its developmental cycle. At 2 h after infection, 13 host genes showed mean expression ratios >/=2-fold. At 16 h after infection, 130 genes were differentially transcribed. These genes encoded factors inhibiting apoptosis and factors regulating cell differentiation, components of the cytoskeleton, transcription factors, and proinflammatory cytokines. This indicates that chlamydial infection, despite its intravacuolar location, alters the transcription of a broad range of host genes in diverse cellular pathways and provides a framework for future studies.

Chlamydia Infections↗

Gene expression profiling of the cellular transcriptional network regulated by alpha/beta interferon and its partial attenuation by the hepatitis C virus nonstructural 5A protein.

Alpha/beta interferons (IFN-alpha/beta) induce potent antiviral and antiproliferative responses and are used to treat a wide range of human diseases, including chronic hepatitis C virus (HCV) infection. However, for reasons that remain poorly understood, many HCV isolates are resistant to IFN therapy. To better understand the nature of the cellular IFN response, we examined the effects of IFN treatment on global gene expression by using several types of human cells, including HeLa cells, liver cell lines, and primary fetal hepatocytes. In response to IFN, 50 of the approximately 4,600 genes examined were consistently induced in each of these cell types and another 60 were induced in a cell type-specific manner. A search for IFN-stimulated response elements (ISREs) in genomic DNA located upstream of IFN-stimulated genes revealed both previously identified and novel putative ISREs. To determine whether HCV can alter IFN-regulated gene expression, we performed microarray analyses on IFN-treated HeLa cells expressing the HCV nonstructural 5A (NS5A) protein and on IFN-treated Huh7 cells containing an HCV subgenomic replicon. NS5A partially blocked the IFN-mediated induction of 14 IFN-stimulated genes, an effect that may play a role in HCV resistance to IFN. This block may occur through repression of ISRE-mediated transcription, since NS5A also inhibited the IFN-mediated induction of a reporter gene driven from an ISRE-containing promoter. In contrast, the HCV replicon had very little effect on IFN-regulated gene expression. These differences highlight the importance of comparing results from multiple model systems when investigating complex phenomena such as the cellular response to IFN and viral mechanisms of IFN resistance.

Cell Line↗

Cellular gene expression upon human immunodeficiency virus type 1 infection of CD4(+)-T-cell lines.

The expression levels of approximately 4,600 cellular RNA transcripts were assessed in CD4(+)-T-cell lines at different times after infection with human immunodeficiency virus type 1 strain BRU (HIV-1(BRU)) using DNA microarrays. We found that several classes of genes were inhibited by HIV-1(BRU) infection, consistent with the G(2) arrest of HIV-1-infected cells induced by Vpr. These included genes involved in cell division and transcription, a family of DEAD-box proteins (RNA helicases), and all genes involved in translation and splicing. However, the overall level of cell activation and signaling was increased in infected cells, consistent with strong virus production. These included a subgroup of transcription factors, including EGR1 and JUN, suggesting they play a specific role in the HIV-1 life cycle. Some regulatory changes were cell line specific; however, the majority, including enzymes involved in cholesterol biosynthesis, of changes were regulated in most infected cell lines. Compendium analysis comparing gene expression profiles of our HIV-1 infection experiments to those of cells exposed to heat shock, interferon, or influenza A virus indicated that HIV-1 infection largely induced specific changes rather than simply activating stress response or cytokine response pathways. Thus, microarray analysis confirmed several known HIV-1 host cell interactions and permitted identification of specific cellular pathways not previously implicated in HIV-1 infection. Continuing analyses are expected to suggest strategies for impacting HIV-1 replication in vivo by targeting these pathways.

Base Sequence↗