Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Royal Statistical Society”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A proportional hazards model for arbitrarily censored and truncated data.

Turnbull (1976, Journal of Royal Statistical Society, Series B 38, 290-295) proposed a method for nonparametric estimation of the distribution function when the data are incomplete because of censoring and truncation. However, as noted by Frydman (1994, Journal of Royal Statistical society, Series B 56, 71-74), Turnbull's method has to be modified to accommodate both truncation and censoring. This paper presents a detailed correction of Turnbull's method and an extension to the regression analysis: a method of fitting the proportional hazards model for arbitrarily censored and truncated data is developed. The method allows partial testing for zero regression coefficients. The test can be performed using the likelihood ratio test or the Wald test. The methodology is applied to estimate the distribution of the induction time of patients diagnosed with transfusion-associated AIDS and to estimate the distribution of time from diabetes onset to development of diabetic nephropathy for insulin-dependent diabetics.

Acquired Immunodeficiency Syndrome↗

A Monte Carlo method for Bayesian inference in frailty models.

Many analyses in epidemiological and prognostic studies and in studies of event history data require methods that allow for unobserved covariates or "frailties." Clayton and Cuzick (1985, Journal of the Royal Statistical Society, Series A 148, 82-117) proposed a generalization of the proportional hazards model that implemented such random effects, but the proof of the asymptotic properties of the method remains elusive, and practical experience suggests that the likelihoods may be markedly nonquadratic. This paper sets out a Bayesian representation of the model in the spirit of Kalbfleisch (1978, Journal of the Royal Statistical Society, Series B 40, 214-221) and discusses inference using Monte Carlo methods.

Algorithms↗

On the apparent clustering of clonal albumin production and enzyme activity levels.

This communication is a critique of a novel, but inappropriate, use of the correlation coefficient to demonstrate the clustering of biological activity levels about a purported geometric progression. The data are re-examined by using a Fourier analysis approach to test for periodicity on a logarithmic scale; this approach follows from the methods of Kendall (1974, Philosophical Transactions of the Royal Society, Series A 276, 231-266) and Fisher (1929, Proceedings of the Royal Statistical Society, Series A 125, 54-59).

Albumins↗

Relative vs. absolute statistical analysis of compositions: a comparative study of surface waters of a Mediterranean river.

Most hydrogeological research includes some sort of statistical study, which is generally conducted on the raw measures of chemical variables, though there are several theoretical and practical studies warning against this practice. Arguments refer mainly to the positive character of this type of data, and to the fact that they carry only information about the relative abundance of each component on the whole, what makes techniques based on correlation, like the widely used Principal Component Analysis (PCA), loose their meaning. The solution proposed by Aitchison (1982, Journal of the Royal Statistical Society, Series B 44(2), 139-177)-based on working with log-ratios of observations-is equivalent to define a new distance between compositions and to adapt usual statistical techniques to it. To illustrate its effect, our study compares the performance of the biplot-a PCA graphical technique-according to the usual Euclidean and to the Aitchison distance. The study is conducted on a set of 14 molarities measured monthly through the years 1997-1999 at 30 different stations along the Llobregat River and its tributaries (Barcelona, NE Spain). Ordinary analysis, implicitly based on an Euclidean distance, presents some deficiencies, mainly because it only captures major ion variations and the inferred relationship between them actually depends on other non-relevant variables, such as water mass. An analysis based on compositional distances captures variations of all the ions; it is robust against the inclusion of non-relevant variables in the analysis; and it offers a way to build factors expressed as equilibrium equations. In our case, two promising factors are extracted, showing the different anthropogenic and geological pollution sources of the rivers.

Environmental Monitoring↗

A model of traffic crashes in New Zealand.

The aim of this study was to examine the changes in the trend and seasonal patterns in fatal crashes in New Zealand in relation to changes in economic conditions between 1970 and 1994. The Harvey and Durbin (Journal of the Royal Statistical Society 149 (3) (1986) 187-227) structural time series model (STSM), an 'unobserved components' class of model, was used to estimate models for quarterly fatal traffic crashes. The dependent variable was modelled as the number of crashes and three variants of the crash rate (crashes per 10,000 km travelled, crashes per 1,000 vehicles, and crashes per 1000 population). Independent variables included in the models were unemployment rate (UER), real gross domestic product per capita, the proportion of motorcycles, the proportion of young males in the population, alcohol consumption per capita, the open road speed limit, and dummy variables for the 1973 and 1979 oil crises and seat belt wearing laws. UERs, real GDP per capita, and alcohol consumption were all significant and important factors in explaining the short-run dynamics of the models. In the long-run, real GDP per capita was directly related to the number of crashes but after controlling for distance travelled was not significant. This suggests increases in income are associated with a short-run reduction in risk but increases in exposure to a crash (i.e. distance travelled) in the long-run. A 1% increase in the open road speed limit was associated with a long-run 0.5% increase in fatal crashes. Substantial reductions in fatal crashes were associated with the 1979 oil crisis and seat belt wearing laws. The 1984 universal seat belt wearing law was associated with a sustained 15.6% reduction in fatal crashes. These road policy factors appeared to have a greater influence on crashes than the role of demographic and economic factors.

Accidents, Traffic↗

Frailty modeling for spatially correlated survival data, with application to infant mortality in Minnesota.

The use of survival models involving a random effect or 'frailty' term is becoming more common. Usually the random effects are assumed to represent different clusters, and clusters are assumed to be independent. In this paper, we consider random effects corresponding to clusters that are spatially arranged, such as clinical sites or geographical regions. That is, we might suspect that random effects corresponding to strata in closer proximity to each other might also be similar in magnitude. Such spatial arrangement of the strata can be modeled in several ways, but we group these ways into two general settings: geostatistical approaches, where we use the exact geographic locations (e.g. latitude and longitude) of the strata, and lattice approaches, where we use only the positions of the strata relative to each other (e.g. which counties neighbor which others). We compare our approaches in the context of a dataset on infant mortality in Minnesota counties between 1992 and 1996. Our main substantive goal here is to explain the pattern of infant mortality using important covariates (sex, race, birth weight, age of mother, etc.) while accounting for possible (spatially correlated) differences in hazard among the counties. We use the GIS ArcView to map resulting fitted hazard rates, to help search for possible lingering spatial correlation. The DIC criterion (Spiegelhalter et al., Journal of the Royal Statistical Society, Series B 2002, to appear) is used to choose among various competing models. We investigate the quality of fit of our chosen model, and compare its results when used to investigate neonatal versus post-neonatal mortality. We also compare use of our time-to-event outcome survival model with the simpler dichotomous outcome logistic model. Finally, we summarize our findings and suggest directions for future research.

Adult↗

A local influence approach applied to binary data from a psychiatric study.

Recently, a lot of concern has been raised about assumptions needed in order to fit statistical models to incomplete multivariate and longitudinal data. In response, research efforts are being devoted to the development of tools that assess the sensitivity of such models to often strong but always, at least in part, unverifiable assumptions. Many efforts have been devoted to longitudinal data, primarily in the selection model context, although some researchers have expressed interest in the pattern-mixture setting as well. A promising tool, proposed by Verbeke et al. (2001, Biometrics 57, 43-50), is based on local influence (Cook, 1986, Journal of the Royal Statistical Society, Series B 48, 133-169). These authors considered the Diggle and Kenward (1994, Applied Statistics 43, 49-93) model, which is based on a selection model, integrating a linear mixed model for continuous outcomes with logistic regression for dropout. In this article, we show that a similar idea can be developed for multivariate and longitudinal binary data, subject to nonmonotone missingness. We focus on the model proposed by Baker, Rosenberger, and DerSimonian (1992, Statistics in Medicine 11, 643-657). The original model is first extended to allow for (possibly continuous) covariates, whereafter a local influence strategy is developed to support the model-building process. The model is able to deal with nonmonotone missingness but has some limitations as well, stemming from the conditional nature of the model parameters. Some analytical insight is provided into the behavior of the local influence graphs.

Antidepressive Agents, Tricyclic↗

The evidential value in the DNA database search controversy and the two-stain problem.

Does the evidential strength of a DNA match depend on whether the suspect was identified through database search or through other evidence ("probable cause")? In Balding and Donnelly (1995, Journal of the Royal Statistical Society, Series A 158, 21-53) and elsewhere, it has been argued that the evidential strength is slightly larger in a database search case than in a probable cause case, while Stockmarr (1999, Biometrics 55, 671-677) reached the opposite conclusion. Both these approaches use likelihood ratios. By making an excursion to a similar problem, the two-stain problem, we argue in this article that there are certain fundamental difficulties with the use of a likelihood ratio, which can be avoided by concentrating on the posterior odds. This approach helps resolving the above-mentioned conflict.

Biometry↗

Semiparametric methods for multiple exposure mismeasurement and a bivariate outcome in HIV vaccine trials.

Exposure to infection information is important for estimating vaccine efficacy, but it is difficult to collect and prone to missingness and mismeasurement. We discuss study designs that collect detailed exposure information from only a small subset of participants while collecting crude exposure information from all participants and treat estimation of vaccine efficacy in the missing data/measurement error framework. We extend the discordant partner design for HIV vaccine trials of Golm, Halloran, and Longini (1998, Statistics in Medicine, 17, 2335-2352.) to the more complex augmented trial design of Longini, Datta, and Halloran (1996, Journal of Acquired Immune Deficiency Syndromes and Human Retrovirology 13, 440-447) and Datta, Halloran, and Longini (1998, Statistics in Medicine 17, 185-200). The model for this design includes three exposure covariates and both univariate and bivariate outcomes. We adapt recently developed semiparametric missing data methods of Reilly and Pepe (1995, Biometrika 82, 299 314), Carroll and Wand (1991, Journal of the Royal Statistical Society, Series B 53, 573-585), and Pepe and Fleming (1991, Journal of the American Statistical Association 86, 108-113) to the augmented vaccine trial design. We demonstrate with simulated HIV vaccine trial data the improvements in bias and efficiency when combining the different levels of exposure information to estimate vaccine efficacy for reducing both susceptibility and infectiousness. We show that the semiparametric methods estimate both efficacy parameters without bias when the good exposure information is either missing completely at random or missing at random. The pseudolikelihood method of Carroll and Wand (1991) and Pepe and Fleming (1991) was the more efficient of the two semiparametric methods.

AIDS Vaccines↗

A general maximum likelihood analysis of variance components in generalized linear models.

This paper describes an EM algorithm for nonparametric maximum likelihood (ML) estimation in generalized linear models with variance component structure. The algorithm provides an alternative analysis to approximate MQL and PQL analyses (McGilchrist and Aisbett, 1991, Biometrical Journal 33, 131-141; Breslow and Clayton, 1993; Journal of the American Statistical Association 88, 9-25; McGilchrist, 1994, Journal of the Royal Statistical Society, Series B 56, 61-69; Goldstein, 1995, Multilevel Statistical Models) and to GEE analyses (Liang and Zeger, 1986, Biometrika 73, 13-22). The algorithm, first given by Hinde and Wood (1987, in Longitudinal Data Analysis, 110-126), is a generalization of that for random effect models for overdispersion in generalized linear models, described in Aitkin (1996, Statistics and Computing 6, 251-262). The algorithm is initially derived as a form of Gaussian quadrature assuming a normal mixing distribution, but with only slight variation it can be used for a completely unknown mixing distribution, giving a straightforward method for the fully nonparametric ML estimation of this distribution. This is of value because the ML estimates of the GLM parameters can be sensitive to the specification of a parametric form for the mixing distribution. The nonparametric analysis can be extended straightforwardly to general random parameter models, with full NPML estimation of the joint distribution of the random parameters. This can produce substantial computational saving compared with full numerical integration over a specified parametric distribution for the random parameters. A simple method is described for obtaining correct standard errors for parameter estimates when using the EM algorithm. Several examples are discussed involving simple variance component and longitudinal models, and small-area estimation.

Adrenergic beta-Antagonists↗

Maximum likelihood analysis for heteroscedastic one-way random effects ANOVA in interlaboratory studies.

This article presents results for the maximum likelihood analysis of several groups of measurements made on the same quantity. Following Cochran (1937, Journal of the Royal Statistical Society 4(Supple), 102-118; 1954, Biometrics 10, 101-129; 1980, in Proceedings of the 25th Conference on the Design of Experiments in Army Research, Development and Testing, 21-33) and others, this problem is formulated as a one-way unbalanced random-effects ANOVA with unequal within-group variances. A reparametrization of the likelihood leads to simplified computations, easier identification and interpretation of multimodality of the likelihood, and (through a non-informative-prior Bayesian approach) approximate confidence regions for the mean and between-group variance.

Analysis of Variance↗

Adaptive regression splines in the Cox model.

We develop a method for constructing adaptive regression spline models for the exploration of survival data. The method combines Cox's (1972, Journal of the Royal Statistical Society, Series B 34, 187-200) regression model with a weighted least-squares version of the multivariate adaptive regressi on spline (MARS) technique of Friedman (1991, Annals of Statistics 19, 1-141) to adaptively select the knots and covariates. The new technique can automatically fit models with terms that represent nonlinear effects and interactions among covariates. Applications based on simulated data and data from a clinical trial for myeloma are presented. Results from the myeloma application identified several important prognostic variables, including a possible nonmonotone relationship with survival in one laboratory variable. Results are compared to those from the adaptive hazard regression (HARE) method of Kooperberg, Stone, and Truong (1995, Journal of the American Statistical Association 90, 78-94).

Biometry↗

Robustness of the latent variable model for correlated binary data.

The marginal regression model offers a useful alternative to conditional approaches to analyzing binary data (Liang, Zeger, and Qaqish, 1992, Journal of the Royal Statistical Society, Series B 54, 3-40). Instead of modelling the binary data directly as do Liang and Zeger (1986, Biometrika 73, 13-22), the parametric marginal regression model developed by Qu et al. (1992, Biometrics 48, 1095-1102) assumes that there is an underlying multivariate normal vector that gives rise to the observed correlated binary outcomes. Although this parametric approach provides a flexible way to model different within-cluster correlation structures and does not restrict the parameter space, it is of interest to know how robust the parameter estimates are with respect to choices of the latent distribution. We first extend the latent modelling to include multivariate t-distributed latent vectors and assess the robustness in this class of distributions. Then we show through a simulation that the parameter estimates are robust with respect to the latent distribution even if latent distribution is skewed. In addtion to this empirical evidence for robustness, we show through the iterative algorithm that the robustness of the regression coefficents with respect to misspecifications of covariance structure in Liang and Zeger's model in fact indicates robustness with respect to underlying distributional assumptions of the latent vector in the latent variable model.

Algorithms↗

Optimum experimental designs for multinomial logistic models.

Multinomial responses frequently occur in dose level experiments. For example, in a study of the influence of gamma radiation on the emergence of house flies (Musca domestica L., 1758), three disjoint outcomes occurred: death before the pupae opened, death during emergence, and life after emergence. Although the flies are easy to breed, this sort of bioassay is, in general, very expensive since it requires the use of a gamma radiation source. Experiments therefore need to be designed to involve the minimum number of different doses. Here the theory of optimum experimental design is applied to provide efficient experiments to estimate the parameters of those multinomial logistic models that are a special case of the multivariate logistic models of Glonek and McCullagh (1995, Journal of the Royal Statistical Society, Series B 57, 533-546). The purpose is to reduce the overall experimental cost. The general equivalence theorem (Fedorov, 1972, Theory of Optimal Experiments) is adapted to this class of models, providing an effective method of generating and checking the optimality of designs. One example on flies demonstrates the method, which can be easily implemented.

Animals↗

Discrete-time nonparametric estimation for semi-Markov models of chain-of-events data subject to interval censoring and truncation.

Chain-of-events data are longitudinal observations on a succession of events that can only occur in a prescribed order. One goal in an analysis of this type of data is to determine the distribution of times between the successive events. This is difficult when individuals are observed periodically rather than continuously because the event times are then interval censored. Chain-of-events data may also be subject to truncation when individuals can only be observed if a certain event in the chain (e.g., the final event) has occurred. We provide a nonparametric approach to estimate the distributions of times between successive events in discrete time for data such as these under the semi-Markov assumption that the times between events are independent. This method uses a self-consistency algorithm that extends Turnbull's algorithm (1976, Journal of the Royal Statistical Society, Series B 38, 290-295). The quantities required to carry out the algorithm can be calculated recursively for improved computational efficiency. Two examples using data from studies involving HIV disease are used to illustrate our methods.

Biometry↗

Semiparametric regression analysis of interval-censored data.

We propose a semiparametric approach to the proportional hazards regression analysis of interval-censored data. An EM algorithm based on an approximate likelihood leads to an M-step that involves maximizing a standard Cox partial likelihood to estimate regression coefficients and then using the Breslow estimator for the unknown baseline hazards. The E-step takes a particularly simple form because all incomplete data appear as linear terms in the complete-data log likelihood. The algorithm of Turnbull (1976, Journal of the Royal Statistical Society, Series B 38, 290-295) is used to determine times at which the hazard can take positive mass. We found multiple imputation to yield an easily computed variance estimate that appears to be more reliable than asymptotic methods with small to moderately sized data sets. In the right-censored survival setting, the approach reduces to the standard Cox proportional hazards analysis, while the algorithm reduces to the one suggested by Clayton and Cuzick (1985, Applied Statistics 34, 148-156). The method is illustrated on data from the breast cancer cosmetics trial, previously analyzed by Finkelstein (1986, Biometrics 42, 845-854) and several subsequent authors.

Algorithms↗

Survival analysis in clinical trials: past developments and future directions.

The field of survival analysis emerged in the 20th century and experienced tremendous growth during the latter half of the century. The developments in this field that have had the most profound impact on clinical trials are the Kaplan-Meier (1958, Journal of the American Statistical Association 53, 457-481) method for estimating the survival function, the log-rank statistic (Mantel, 1966, Cancer Chemotherapy Report 50, 163-170) for comparing two survival distributions, and the Cox (1972, Journal of the Royal Statistical Society, Series B 34, 187-220) proportional hazards model for quantifying the effects of covariates on the survival time. The counting-process martingale theory pioneered by Aalen (1975, Statistical inference for a family of counting processes, Ph.D. dissertation, University of California, Berkeley) provides a unified framework for studying the small- and large-sample properties of survival analysis statistics. Significant progress has been achieved and further developments are expected in many other areas, including the accelerated failure time model, multivariate failure time data, interval-censored data, dependent censoring, dynamic treatment regimes and causal inference, joint modeling of failure time and longitudinal data, and Baysian methods.

Biometry↗

Sensitivity analysis for nonrandom dropout: a local influence approach.

Diggle and Kenward (1994, Applied Statistics 43, 49-93) proposed a selection model for continuous longitudinal data subject to nonrandom dropout. It has provoked a large debate about the role for such models. The original enthusiasm was followed by skepticism about the strong but untestable assumptions on which this type of model invariably rests. Since then, the view has emerged that these models should ideally be made part of a sensitivity analysis. This paper presents a formal and flexible approach to such a sensitivity assessment based on local influence (Cook, 1986, Journal of the Royal Statistical Society, Series B 48, 133-169). The influence of perturbing a missing-at-random dropout model in the direction of nonrandom dropout is explored. The method is applied to data from a randomized experiment on the inhibition of testosterone production in rats.

Animals↗