Search PubMedSearch

PubMed · 8962448

Regression analysis with missing covariate data using estimating equations.

Abstract

In regression analysis, missing covariate data has been among the most common problems. Frequently, practitioners adopt the so-called complete-case analysis, i.e., performing the analysis on only a complete dataset after excluding records with missing covariates. Performing a complete-case analysis is convenient with existing statistical packages, but it may be inefficient since the observed outcomes and covariates on those records with missing covariates are not used. It can even give misleading statistical inference if missing is not completely at random. This paper introduces a joint estimating equation (JEE) for regression analysis in the presence of missing observations on one covariate, which may be thought of as a method in a general framework for the missing covariate data problem proposed by Robins, Rotnitzky, and Zhao (1994, Journal of the American Statistical Association 89, 846-866). A generalization of JEE to more than one such covariate is discussed. The JEE is generally applicable to estimating regression coefficients from a regression model, including linear and logistic regression. Provided that the missing covariate data is either missing completely at random or missing at random (in addition to mild regularity conditions), estimates of regression coefficients from the JEE are consistent and have an asymptotic normal distribution. Simulation results show that the asymptotic distribution of estimated coefficients performs well in finite samples. Also shown through the simulation study is that the validity of JEE estimates depends on the correct specification of the probability function that characterizes the missing mechanism, suggesting a need for further research on how to robustify the estimation from making this nuisance assumption. Finally, the JEE is illustrated with an application from a case-control study of diet and thyroid cancer.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

L P Zhao, S Lipsitz, D Lew. 1996. Regression analysis with missing covariate data using estimating equations.. https://pubmed.ncbi.nlm.nih.gov/8962448/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A multivariate approach for assessing severity of acute graft-versus-host disease in bone marrow transplantation.

Patients undergoing bone marrow transplantation are at high risk of developing acute graft-versus-host disease (GVHD) which is a primary limiting factor for this procedure inasmuch as it is responsible for high morbidity rates and is associated with poor survival outcome. To provide improved treatment assessment and better interpretation of clinical outcomes, we need a precise and objective assessment of GVHD. Severity of GVHD is commonly assessed using an imprecise categorical grading system that incorporates skin, gut and liver grades, as well as subjective assessment of clinical performance. These organ grades are based on arbitrary cutpoints of skin rash, diarrhoea volume and bilirubin level. The International Bone Marrow Transplant Registry proposed an alternative grading system based on different combinations of organ involvement and provided estimates of relative risk of treatment failure. On the basis of that work, we developed an empirical mathematical model that quantifies GVHD severity, and that uses continuous, rather than categorical, daily measurements for each organ system. We use model-predicted values as an index of severity for any combination of values. The proposed index allows a more precise comparison of GVHD profiles across different treatment protocols and also permits more refined analyses to address relationships between GVHD and clinical outcomes.

Biometry

Using missing data methods in genetic studies with missing mutation status.

Because of current techniques of determining gene mutation, investigators are now interested in estimating the odds ratio between genetic status (mutation, no mutation) and an outcome variable such as disease cell type (A, B). In this paper we consider the mutation of the RAS genetic family. To determine if the genes have mutated, investigators look at five specific locations on the RAS gene. RAS mutated is a mutation in at least one of the five gene locations and RAS non-mutated is no mutation in any of the five locations. Owing to limited time and financial resources, one cannot obtain a complete genetic evaluation of all five locations on the gene for all patients. We propose the use of maximum likelihood (ML) with a 2(6) multinomial distribution formed by cross-classifying the binary mutation status at five locations by binary disease cell type. This ML method includes all patients regardless of completeness of data, treats the locations not evaluated as missing data, and uses the EM algorithm to estimate the odds ratio between genetic mutation status and the disease type. We compare the ML method to complete case estimates, and a method used by clinical investigators, which excludes patients with data on less than five locations who have no mutations on these sites.

Biometry

Construction of uniform-balanced cross-over designs for any odd number of treatments.

Cross-over designs balanced for simple carry-over effects are commonly applied in clinical studies for the comparison of treatments for chronic conditions such as hypertension or asthma. Uniform-balanced cross-over designs have the desirable property that the treatment sequences are arranged so that, in the full design, each treatment is followed by every other treatment equally often. Such designs for an even number of treatments and the same number of sequences and periods are readily constructed using suitable cyclic Latin squares. For an odd number of treatments, pairs of squares may be combined to give uniform-balanced designs. Recently, computer search techniques have been used to find nearly-balanced Latin squares which may be combined in pairs or in sets of three to produce designs with the overall properties of uniformity and balance. In this paper, simple generating formulae are described which will give, for any odd number of treatments t > 3, uniform-balanced cross-over designs with p = t periods and n = kt treatment sequences for any k > or = 2. Tables of cross-over designs obtained from these simple formulae are presented for t < or = 15.

Biometry