Search PubMedSearch

Biomedical subjects

G M Fitzmaurice

Publications and source records attributed to G M Fitzmaurice.

10 recordsLinked to original sources

Goodness-of-fit for GEE: an example with mental health service utilization.

Suppose we use generalized estimating equations to estimate a marginal regression model for repeated binary observations. There are no established summary statistics available for assessing the adequacy of the fitted model. In this paper we propose a goodness-of-fit test statistic which has an approximate chi-squared distribution when we have specified the model correctly. The proposed statistic can be viewed as an extension of the Hosmer and Lemeshow goodness-of-fit statistic for ordinary logistic regression to marginal regression models for repeated binary responses. We illustrate the methods using data from a study of mental health service utilization by children. The repeated responses are a set of binary measures of service use. We fit a marginal logistic regression model to the data using generalized estimating equations, and we apply the proposed goodness-of-fit statistic to assess the adequacy of the fitted model.

Age Factors

Regression models for mixed discrete and continuous responses with potentially missing values.

In this paper a likelihood-based method for analyzing mixed discrete and continuous regression models is proposed. We focus on marginal regression models, that is, models in which the marginal expectation of the response vector is related to covariates by known link functions. The proposed model is based on an extension of the general location model of Olkin and Tate (1961, Annals of Mathematical Statistics 32, 448-465), and can accommodate missing responses. When there are no missing data, our particular choice of parameterization yields maximum likelihood estimates of the marginal mean parameters that are robust to misspecification of the association between the responses. This robustness property does not, in general, hold for the case of incomplete data. There are a number of potential benefits of a multivariate approach over separate analyses of the distinct responses. First, a multivariate analysis can exploit the correlation structure of the response vector to address intrinsically multivariate questions. Second, multivariate test statistics allow for control over the inflation of the type I error that results when separate analyses of the distinct responses are performed without accounting for multiple comparisons. Third, it is generally possible to obtain more precise parameter estimates by accounting for the association between the responses. Finally, separate analyses of the distinct responses may be difficult to interpret when there is nonresponse because different sets of individuals contribute to each analysis. Furthermore, separate analyses can introduce bias when the missing responses are missing at random (MAR). A multivariate analysis can circumvent both of these problems. The proposed methods are applied to two biomedical datasets.

Air Pollution

The score test for independence in R x C contingency tables with missing data.

In this paper, the score test statistic for testing independence in R x C contingency tables with missing data is proposed. Under the null hypothesis of independence, the statistic has an approximate chi-squared distribution with (R - 1)(C - 1) degrees of freedom. The proposed test statistic is quite similar to the Pearson chi-squared statistic with complete data and, unlike the likelihood ratio statistic for testing independence, its computation is simple and noniterative. In addition, a score test statistic is proposed for testing independence when the rows and columns of the R x C table are ordinal. Finally, extensions of the score statistics to test for conditional independence in a set of (R x C) contingency tables with missing data are described. This yields score test statistics that are natural extensions of the Mantel-Haenszel statistic. An example, using a subset of data from the Six Cities Study, is presented to illustrate the methods.

Air Pollution

Estimating equations for measures of association between repeated binary responses.

Moment-based methods for analyzing repeated binary responses using the marginal odds ratio as a measure of association have been proposed by a number of authors. Carey, Zeger, and Diggle (1993, Biometrika 80, 517-526) have recently described how the marginal odds ratio can be estimated using generalized estimating equations (GEE) based on conditional residuals (deviations about conditional expectations). In this paper, we show that other measures of association between pairs of binary responses, e.g., the correlation, can also be estimated using conditional residuals. We demonstrate that the estimator of the correlation based on conditional residuals is nearly efficient when compared with maximum likelihood or second order estimating equations (GEE2) except when the correlation is large. This estimator also yields more efficient estimates of the correlation than the usual GEE estimator that is based on unconditional residuals. Furthermore, the gains in efficiency can be quite considerable when some of the responses are missing or incomplete, or, alternatively, when cluster sizes are unequal (in the clustered data setting).

Air Pollution

Bivariate logistic regression analysis of childhood psychopathology ratings using multiple informants.

A central issue in studies of risk factors for childhood psychopathology is utilization of the information obtained about the child's mental health status from multiple informants. In this paper, the authors propose a new approach to the analysis of risk factor data when the outcomes are binary ratings (presence/absence of symptoms). This new approach has several attractive features in this setting. The strategy taken is to perform a single analysis using multivariate modeling, in which simultaneous logistic regressions are conducted for the outcomes given by each of several informants. The advantages of this approach include the following: 1) it retains the complete information about case status for each informant; 2) it permits assessment of informant-risk factor interactions as well as "overall" risk factor effects; 3) it provides measures of association between the multiple informants and adjusts for the association between responses in the analysis; and 4) missing data on a subset of respondents can be incorporated in a straightforward way, permitting all subjects with at least one informant to be used in the analysis. To illustrate the methods, the authors present findings on risk factors for measures of "Internalizing" and "Externalizing" behaviors from two surveys using parent and teacher ratings of 6- to 11-year-old children in Connecticut between 1986 and 1989.

Child

Estimation methods for the join distribution of repeated binary observations.

The joint distribution of repeated binary observations is multinomial, and can be specified using a representation first suggested by Bahadur (1961, in Studies in Item Analysis and Prediction,158-168. Stanford, California: Stanford University Press), and later by Cox (1972, Applied Statistics 21, 13-120). Using the Bahadur representation, the marginal probabilities of success can be related to a set of covariates using the logistic link function, or any other suitable link function. Besides the parameters of the marginal regression model, we may also have interest in the probability of success on any of the repeated measures. For example, in the Six Cities study, a longitudinal study of the health effects of air pollution, we have interest in both the marginal probability of a child wheezing at age t (t = 10, 11, 12), and the union probability of wheezing at any of the three ages. This "union" probability can be specified in terms of the joint probabilities and the second higher-order correlations. We discuss several methods of estimating the parameters of the Bahadur model.

Adult

A caveat concerning independence estimating equations with multivariate binary data.

Clustered binary data occur commonly in both the biomedical and health sciences. In this paper, we consider logistic regression models for multivariate binary responses, where the association between the responses is largely regarded as a nuisance characteristic of the data. In particular, we consider the estimator based on independence estimating equations (IEE), which assumes that the responses are independent. This estimator has been shown to be nearly efficient when compared with maximum likelihood (ML) and generalized estimating equations (GEE) in a variety of settings. The purpose of this paper is to highlight a circumstance where assuming independence can lead to quite substantial losses of efficiency. In particular, when the covariate design includes within-cluster covariates, assuming independence can lead to a considerable loss of efficiency in estimating the regression parameters associated with those covariates.

Biometry

Sample size for repeated measures studies with binary responses.

We consider the sample size required for repeated measures studies when the response variable is binary. We propose the use of weighted least squares (WLS) for calculating the minimum sample size required to detect some minimum clinically important treatment effect. We provide tabulated values of the estimated sample sizes for a simple example and we discuss some practical considerations in determination of sample size with repeated binary responses.

Bias

Analysing incomplete longitudinal binary responses: a likelihood-based approach.

In this paper, we describe a likelihood-based method for analysing balanced but incomplete longitudinal binary responses that are assumed to be missing at random. Following the approach outlined in Zhao and Prentice (1990, Biometrika 77, 642-648), we focus on "marginal models" in which the marginal expectation of the response variable is related to a set of covariates. The association between binary responses is modelled in terms of conditional log odds-ratios. We describe a set of scoring equations for jointly estimating both the marginal parameters and the conditional association parameters. An outline of the EM algorithm used to obtain the maximum likelihood estimates is presented. This approach yields valid and efficient estimates when the responses are missing at random, but not necessarily missing completely at random. An example, using data from the Muscatine Coronary Risk Factor Study, is presented to illustrate this methodology.

Age Factors

Performance of generalized estimating equations in practical situations.

Moment methods for analyzing repeated binary responses have been proposed by Liang and Zeger (1986, Biometrika 73, 13-22), and extended by Prentice (1988, Biometrics 44, 1033-1048). In their generalized estimating equations (GEE), both Liang and Zeger (1986) and Prentice (1988) estimate the parameters associated with the expected value of an individual's vector of binary responses as well as the correlations between pairs of binary responses. In this paper, we discuss one-step estimators, i.e., estimators obtained from one step of the generalized estimating equations, and compare their performance to that of the fully iterated estimators in small samples. In simulations, we find the performance of the one-step estimator to be qualitatively similar to that of the fully iterated estimator. When the sample size is small and the association between binary responses is high, we recommend using the one-step estimator to circumvent convergence problems associated with the fully iterated GEE algorithm. Furthermore, we find the GEE methods to be more efficient than ordinary logistic regression with variance correction for estimating the effect of a time-varying covariate.

Air Pollution