Search PubMed⌕ Search

Biomedical subjects

Wei Pan

Publications and source records attributed to Wei Pan.

At least 19 recordsLinked to original sources

Statistical significance analysis of longitudinal gene expression data.

MOTIVATION: Time-course microarray experiments are designed to study biological processes in a temporal fashion. Longitudinal gene expression data arise when biological samples taken from the same subject at different time points are used to measure the gene expression levels. It has been observed that the gene expression patterns of samples of a given tumor measured at different time points are likely to be much more similar to each other than are the expression patterns of tumor samples of the same type taken from different subjects. In statistics, this phenomenon is called the within-subject correlation of repeated measurements on the same subject, and the resulting data are called longitudinal data. It is well known in other applications that valid statistical analyses have to appropriately take account of the possible within-subject correlation in longitudinal data. RESULTS: We apply estimating equation techniques to construct a robust statistic, which is a variant of the robust Wald statistic and accounts for the potential within-subject correlation of longitudinal gene expression data, to detect genes with temporal changes in expression. We associate significance levels to the proposed statistic by either incorporating the idea of the significance analysis of microarrays method or using the mixture model method to identify significant genes. The utility of the statistic is demonstrated by applying it to an important study of osteoblast lineage-specific differentiation. Using simulated data, we also show pitfalls in drawing statistical inference when the within-subject correlation in longitudinal gene expression data is ignored.

Adaptation, Physiological↗

P19ARF inhibits the functions of the HPV16 E7 oncoprotein.

The E7 oncoprotein encoded by high-risk types of human papillomavirus (HPV) plays a significant role in the development of HPV-related cancers. E7 is a potent stimulator of S phase and host DNA replication. These functions of E7 are linked to the deregulation of the Rb family of proteins. For example, E7 binds and induces proteolysis of Rb through the ubiquitin-proteasome pathway. Despite advances in our understanding of E7, reagents that inhibit E7 with promise in therapy have not been developed or identified. Here, we provide evidence that the tumor suppressor ARF can inhibit E7. We show that the expression of ARF causes a relocalization of E7 from the nucleoplasm to the nucleolus. Two distinct regions in ARF overlapping with the MDM2-binding sites are necessary for the relocalization of E7. Furthermore, we show that ARF blocks the proteolysis of Rb induced by E7. In addition, ARF expression inhibits DNA replication induced by E7. Although it is not known whether the endogenous ARF, which is expressed at a low level, interferes with E7, our results suggest that ARF is an effective inhibitor of E7. We speculate that ARF or an ARF-derived molecule might have a significant impact in therapy against HPV-related tumors.

Cell Nucleolus↗

On the use of permutation in and the performance of a class of nonparametric methods to detect differential gene expression.

MOTIVATION: Recently a class of nonparametric statistical methods, including the empirical Bayes (EB) method, the significance analysis of microarray (SAM) method and the mixture model method (MMM), have been proposed to detect differential gene expression for replicated microarray experiments conducted under two conditions. All the methods depend on constructing a test statistic Z and a so-called null statistic z. The null statistic z is used to provide some reference distribution for Z such that statistical inference can be accomplished. A common way of constructing z is to apply Z to randomly permuted data. Here we point our that the distribution of z may not approximate the null distribution of Z well, leading to possibly too conservative inference. This observation may apply to other permutation-based nonparametric methods. We propose a new method of constructing a null statistic that aims to estimate the null distribution of a test statistic directly. RESULTS: Using simulated data and real data, we assess and compare the performance of the existing method and our new method when applied in EB, SAM and MMM. Some interesting findings on operating characteristics of EB, SAM and MMM are also reported. Finally, by combining the idea of SAM and MMM, we outline a simple nonparametric method based on the direct use of a test statistic and a null statistic.

Algorithms↗

A mixture model approach to detecting differentially expressed genes with microarray data.

An exciting biological advancement over the past few years is the use of microarray technologies to measure simultaneously the expression levels of thousands of genes. The bottleneck now is how to extract useful information from the resulting large amounts of data. An important and common task in analyzing microarray data is to identify genes with altered expression under two experimental conditions. We propose a nonparametric statistical approach, called the mixture model method (MMM), to handle the problem when there are a small number of replicates under each experimental condition. Specifically, we propose estimating the distributions of a t -type test statistic and its null statistic using finite normal mixture models. A comparison of these two distributions by means of a likelihood ratio test, or simply using the tail distribution of the null statistic, can identify genes with significantly changed expression. Several methods are proposed to effectively control the false positives. The methodology is applied to a data set containing expression levels of 1,176 genes of rats with and without pneumococcal middle ear infection.

Animals↗

Modified nonparametric approaches to detecting differentially expressed genes in replicated microarray experiments.

MOTIVATION: An important goal in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Various parametric tests, such as the two-sample t-test, have been used, but their possibly too strong parametric assumptions or large sample justifications may not hold in practice. As alternatives, a class of three nonparametric statistical methods, including the empirical Bayes method of Efron et al. (2001), the significance analysis of microarray (SAM) method of Tusher et al. (2001) and the mixture model method (MMM) of Pan et al. (2001), have been proposed. All the three methods depend on constructing a test statistic and a so-called null statistic such that the null statistic's distribution can be used to approximate the null distribution of the test statistic. However, relatively little effort has been directed toward assessment of the performance or the underlying assumptions of the methods in constructing such test and null statistics. RESULTS: We point out a problem of a current method to construct the test and null statistics, which may lead to largely inflated Type I errors (i.e. false positives). We also propose two modifications that overcome the problem. In the context of MMM, the improved performance of the modified methods is demonstrated using simulated data. In addition, our numerical results also provide evidence to support the utility and effectiveness of MMM.

Algorithms↗

Identification of genes responsible for osteoblast differentiation from human mesodermal progenitor cells.

Single human bone marrow-derived mesodermal progenitor cells (MPCs) differentiate into osteoblasts, chondrocytes, adipocytes, myocytes, and endothelial cells. To identify genes involved in the commitment of MPCs to osteoblasts we examined the expressed gene profile of undifferentiated MPCs and MPCs induced to the osteoblast lineage for 1-7 days by cDNA microarray analysis. As expected, growth factor, hormone, and signaling pathway genes known to be involved in osteogenesis were activated during differentiation. In addition, 41 transcription factors (TFs) were differentially expressed over time, including TFs with known roles in osteoblast differentiation and TFs not known to be involved in osteoblast differentiation. As the latter group of TFs coclustered with osteogenesis-specific TFs, they may play a role in osteoblast differentiation. When we compared the gene expression profile of MPCs induced to differentiate to chondroblasts and osteoblasts, significant differences in the nature and/or timing of gene activation were seen. These studies indicate that in vitro differentiation cultures in which MPCs are induced to one of multiple cell fates should be very useful for defining signals important for lineage-specific differentiation.

Cell Differentiation↗

An examination of the process of relapse prevention therapy designed to aid smoking cessation.

The process of relapse prevention (RP) therapy is examined. Patients' responses were recorded primarily during telephonic, RP counseling designed to facilitate smoking cessation. A computer program that prompted counselor initiatives and provided a framework for the recording of patient responses guided counselor interaction with patients. A total of 437 patients took part in 1650 counseling sessions and reported 2882 urge/lapse situations. The 2531 situations, for which complete data were available, and 4879 coping responses were analyzed. The main findings are (1) the descriptions of urge/lapse situations provided by patients in treatment are similar to those derived by research that aimed to discover the determinants of relapse without specific treatment, (2) number of coping responses rather than number of situations is related to treatment outcome, and (3) the more coping responses discussed during treatment, the better the treatment outcome.

Adaptation, Psychological↗

Identification of gene expression profiles in rat ears with cDNA microarrays.

The physiological processes of hearing implicate thousands of molecules acting in harmony; however, their identities are only partially understood. We used cDNA microarrays containing 1,176 genes to identify >150 genes expressed in rat middle and inner ear tissue. Expressed genes covered several gene families and biological pathways, many of which have previously not been described. Transcription factor genes that were expressed included inhibitors of DNA binding protein (Id). These were localized to the spiral ganglion, organ of Corti and stria vascularis, and they are possibly involved in neurogenesis and angiogenesis. Transcriptional factors that were highly expressed included Gax (homeobox) and I-kappaB, which inhibit cellular proliferation. Their presence suggests that inhibitory programs for cell proliferation are enforced in the ear. Ion channel genes that were expressed included voltage-dependent L-type calcium channels (LTCC) and proton-gated cation channels (PGCC). Genes involved in neurotransmitter production and release included glutamic acid decarboxylase (GAD1). Genes involved in postsynaptic inhibition included neuropeptide Y5 receptors (NPY5) and GAD1. Due to the existence of receptors and/or enzymes involved in their biochemical synthesis, neurotransmitters associated with these might include serotonin, glutamide, acetylcholine, gamma-aminobutyric acid (GABA), neurotensin, and dopamine.

Animals↗

[Effect of N-terminal deletion on biological activity of vascular endothelial cell growth inhibitor].

Vascular endothelial cell growth inhibitor (VEGI) is a novel cytokine which belongs to the TNF superfamily. It can inhibit the proliferation of endothelial cells and neovascularization. However, little is known about the structure-function relationship of VEGI. In order to study the effect of the N-terminus of VEGI on biological properties, the sequence alignment among VEGI and TNF superfamily members based on structure knowledge was done, and then two truncated forms of VEGI were constructed, in which 43 and 51 amino acids from N-terminus were deleted and named VEGI(131) and VEGI(123). Recombinant proteins were generated from E. coli. The expression rates were 25.2% (VEGI(131)) and 27.8% (VEGI(123)) of total bacterial proteins. After purification the purity reached 92.5% (VEGI(131)) and 91.6% (VEGI(123)). VEGI(131) showed significant inhibitory effect on growth of human umbilical vein endothelial cells (HUVEC), IC(50) of VEGI(131) being 35 mg/L. Under the same conditions, IC(50) of VEGI(151) (the wild type of VEGI) was 27 mg/L, but VEGI(123) showed no inhibitory effect. On chick choriallantic membrane (CAM) assay, VEGI(151) markedly reduced the number of main vessels, and VEGI131 decreased capillary number, while the effect of VEGI(123) was almost the same as control. These results suggest that the first 43 amino acids from N-terminus of VEGI have no significant effect on biological activity, but the amino acids 44-51 at N-terminus are required for full biological activity.

Allantois↗

Analysis by cDNA microarrays of altered gene expression in middle ears of rats following pneumococcal infection.

OBJECTIVE: Streptococcus pneumoniae is the most common pathogen in otitis media. Infection of the middle ear with S. pneumoniae potentiates development of thick effusion in the middle ear which frequently causes hearing loss and communication disorders in children. What has changed immediately in the middle ear cleft following pneumococcal infection is extensively studied and characterized but what has changed ever after remains elusive. The purpose of this study is to explore the cellular and molecular basis that remains on a longer time after acute pneucmococcal middle ear infection and potentiates development of thick effusion in the middle ear. METHODS: 12 rats were intrabullarly inoculated with pneumococcus at 2.5x10(6) CFU/ear and profiles of gene expression in the middle ear were examined by cDNA microarrays in combination with reverse transcription-polymer chain reaction (RT-PCR) 6 weeks after infection while the morphologic changes in middle ear were simultaneously characterized by histopathologic techniques. Twelve rats receiving phosphate-buffered saline (PBS) served as controls. RESULTS: it demonstrated that pneumococcus infected ears had the expression of the following genes at a high level compared to the controls: mitogenic signaling proteins (mitogen-activated protein kinase [MEK1 and MEK2], helix-loop-helix transcriptional regulators (Id3 and Id1), ion channels (sodium channel beta 1 and sodium channel 2), and mucin glycoproteins (Muc2 and Muc5). The morphology demonstrated a thickened mucosa and submucosa with increased expression of macroglycoconjugates compared to the controls. CONCLUSION: the expression of several genes remains high even after the acute episode of pneumococcal otitis media has been resolved. The up-regulated expression of these genes may serve as the basis for the development of thick effusion and mucous cell metaplasia/hyperplasia once it is complicated with other factors such as dysfunction of the Eustachian tube.

Animals↗

ApoA-I structure on discs and spheres. Variable helix registry and conformational states.

Apolipoprotein A-I (apoA-I) readily forms discoidal high density lipoprotein (HDL) particles with phospholipids serving as an ideal transporter of plasma cholesterol. In the lipid-bound conformation, apoA-I activates the enzyme lecithin:cholesterol acyltransferase stimulating the formation of cholesterol esters from free cholesterol. As esterification proceeds cholesterol esters accumulate within the hydrophobic core of the discoidal phospholipid bilayer transforming it into a spherical HDL particle. To investigate the change in apoA-I conformation as it adapts to a spherical surface, fluorescence resonance energy transfer studies were performed. Discoidal rHDL particles containing two lipid-bound apoA-I molecules were prepared with acceptor and donor fluorescent probes attached to cysteine residues located at specific positions. Fluorescence quenching was measured for probe combinations located within repeats 5 and 5 (residue 132), repeats 5 and 6 (residues 132 and 154), and repeats 6 and 6 (residue 154). Results from these experiments indicated that each of the 2 molecules of discoidal bound apoA-I exists in multiple conformations and support the concept of a "variable registry" rather than a "fixed helix-helix registry." Additionally, discoidal rHDL were transformed in vitro to core-containing particles by incubation with lecithin:cholesterol acyltransferase. Compositional analysis showed that core-containing particles contained 11% less phospholipid and 633% more cholesterol ester and a total of 3 apoA-I molecules per particle. Spherical particles showed a lowering of acceptor to donor probe quenching when compared with starting rHDL. Therefore, we conclude that as lipid-bound apoA-I adjusts from a discoidal to a spherical surface its intermolecular interactions are significantly reduced presumably to cover the increased surface area of the particle.

Apolipoprotein A-I↗

Comparing three methods for variance estimation with duplicated high density oligonucleotide arrays.

Microarray experiments are being increasingly used in molecular biology. A common task is to detect genes with differential expression across two experimental conditions, such as two different tissues or the same tissue at two time points of biological development. To take proper account of statistical variability, some statistical approaches based on the t-statistic have been proposed. In constructing the t-statistic, one needs to estimate the variance of gene expression levels. With a small number of replicated array experiments, the variance estimation can be challenging. For instance, although the sample variance is unbiased, it may have large variability, leading to a large mean squared error. For duplicated array experiments, a new approach based on simple averaging has recently been proposed in the literature. Here we consider two more general approaches based on nonparametric smoothing. Our goal is to assess the performance of each method empirically. The three methods are applied to a colon cancer data set containing 2,000 genes. Using two arrays, we compare the variance estimates obtained from the three methods. We also consider their impact on the t-statistics. Our results indicate that the three methods give variance estimates close to each other. Due to its simplicity and generality, we recommend the use of the smoothed sample variance for data with a small number of replicates.

Algorithms↗

Small-sample adjustments in using the sandwich variance estimator in generalized estimating equations.

The generalized estimating equation (GEE) approach is widely used in regression analyses with correlated response data. Under mild conditions, the resulting regression coefficient estimator is consistent and asymptotically normal with its variance being consistently estimated by the so-called sandwich estimator. Statistical inference is thus accomplished by using the asymptotic Wald chi-squared test. However, it has been noted in the literature that for small samples the sandwich estimator may not perform well and may lead to much inflated type I errors for the Wald chi-squared test. Here we propose using an approximate t- or F-test that takes account of the variability of the sandwich estimator. The level of type I error of the proposed t- or F-test is guaranteed to be no larger than that of the Wald chi-squared test. The satisfactory performance of the proposed new tests is confirmed in a simulation study. Our proposal also has some advantages when compared with other new approaches based on direct modifications of the sandwich estimator, including the one that corrects the downward bias of the sandwich estimator. In addition to hypothesis testing, our result has a clear implication on constructing Wald-type confidence intervals or regions.

Adult↗

How many replicates of arrays are required to detect gene expression changes in microarray experiments? A mixture model approach.

BACKGROUND: It has been recognized that replicates of arrays (or spots) may be necessary for reliably detecting differentially expressed genes in microarray experiments. However, the often-asked question of how many replicates are required has barely been addressed in the literature. In general, the answer depends on several factors: a given magnitude of expression change, a desired statistical power (that is, probability) to detect it, a specified Type I error rate, and the statistical method being used to detect the change. Here, we discuss how to calculate the number of replicates in the context of applying a nonparametric statistical method, the normal mixture model approach, to detect changes in gene expression. RESULTS: The methodology is applied to a data set containing expression levels of 1,176 genes in rats with and without pneumococcal middle-ear infection. We illustrate how to calculate the power functions for 2, 4, 6 and 8 replicates. CONCLUSIONS: The proposed method is potentially useful in designing microarray experiments to discover differentially expressed genes. The same idea can be applied to other statistical methods.

Animals↗

Model-based cluster analysis of microarray gene-expression data.

BACKGROUND: Microarray technologies are emerging as a promising tool for genomic studies. The challenge now is how to analyze the resulting large amounts of data. Clustering techniques have been widely applied in analyzing microarray gene-expression data. However, normal mixture model-based cluster analysis has not been widely used for such data, although it has a solid probabilistic foundation. Here, we introduce and illustrate its use in detecting differentially expressed genes. In particular, we do not cluster gene-expression patterns but a summary statistic, the t-statistic. RESULTS: The method is applied to a data set containing expression levels of 1,176 genes of rats with and without pneumococcal middle-ear infection. Three clusters were found, two of which contain more than 95% genes with almost no altered gene-expression levels, whereas the third one has 30 genes with more or less differential gene-expression levels. CONCLUSIONS: Our results indicate that model-based clustering of t-statistics (and possibly other summary statistics) can be a useful statistical tool to exploit differential gene expression for microarray data.

Animals↗

Persistence of the effect of the Lung Health Study (LHS) smoking intervention over eleven years.

BACKGROUND: Research on the long-term persistence of effects of interventions aimed at smoking cessation is limited. This paper examined the quitting behavior of individuals who were randomized to a smoking cessation intervention (SI) or to usual care (UC), at a point approximately 11 years later. METHODS: The initial sample consisted of 5,887 adult smokers in 10 clinics who had evidence of airways obstruction. Two-thirds of the original participants were offered an intensive 12-week smoking cessation intervention. Of these, 4,517 were enrolled in the long-term follow-up study. RESULTS: Randomized group assignment was a strong predictor of smoking behavior after 11 years, in that 21.9% of SI participants and only 6.0% of UC participants maintained abstinence throughout the interval. Logistic regressions identified covariates associated with abstinence. A higher proportion of abstinence was observed in participants that had been assigned to SI (OR = 4.45), were older (OR = 1.11, increment 5 years), had more years of education (OR = 1.05), and fewer cigarettes/day at baseline (OR = 0.90, increment 10 cigarettes). CONCLUSIONS: Smokers exposed to an aggressive smoking intervention program and who sustain abstinence for a five-year period are very likely to still be abstinent after 11 years.

Adult↗

A comparative review of statistical methods for discovering differentially expressed genes in replicated microarray experiments.

MOTIVATION: A common task in analyzing microarray data is to determine which genes are differentially expressed across two kinds of tissue samples or samples obtained under two experimental conditions. Recently several statistical methods have been proposed to accomplish this goal when there are replicated samples under each condition. However, it may not be clear how these methods compare with each other. Our main goal here is to compare three methods, the t-test, a regression modeling approach (Thomas et al., Genome Res., 11, 1227-1236, 2001) and a mixture model approach (Pan et al., http://www.biostat.umn.edu/cgi-bin/rrs?print+2001,2001a,b) with particular attention to their different modeling assumptions. RESULTS: It is pointed out that all the three methods are based on using the two-sample t-statistic or its minor variation, but they differ in how to associate a statistical significance level to the corresponding statistic, leading to possibly large difference in the resulting significance levels and the numbers of genes detected. In particular, we give an explicit formula for the test statistic used in the regression approach. Using the leukemia data of Golub et al. (Science, 285, 531-537, 1999), we illustrate these points. We also briefly compare the results with those of several other methods, including the empirical Bayesian method of Efron et al. (J. Am. Stat. Assoc., to appear, 2001) and the Significance Analysis of Microarray (SAM) method of Tusher et al. (PROC: Natl Acad. Sci. USA, 98, 5116-5121, 2001).

Acute Disease↗

Application of conditional moment tests to model checking for generalized linear models.

Generalized linear models (GLMs) are increasingly being used in daily data analysis. However, model checking for GLMs with correlated discrete response data remains difficult. In this paper, through a case study on marginal logistic regression using a real data set, we illustrate the flexibility and effectiveness of using conditional moment tests (CMTs), along with other graphical methods, to do model checking for generalized estimation equation (GEE) analyses. Although CMTs provide an array of powerful diagnostic tests for model checking, they were originally proposed in the econometrics literature and, to our knowledge, have never been applied to GEE analyses. CMTs cover many existing tests, including the (generalized) score test for an omitted covariate, as special cases. In summary, we believe that CMTs provide a class of useful model checking tools.

Journal Article↗