Search PubMed⌕ Search

PubMed · 14709437

Advanced statistics: linear regression, part II: multiple linear regression.

Abstract

The applications of simple linear regression in medical research are limited, because in most situations, there are multiple relevant predictor variables. Univariate statistical techniques such as simple linear regression use a single predictor variable, and they often may be mathematically correct but clinically misleading. Multiple linear regression is a mathematical technique used to model the relationship between multiple independent predictor variables and a single dependent outcome variable. It is used in medical research to model observational data, as well as in diagnostic and therapeutic studies in which the outcome is dependent on more than one factor. Although the technique generally is limited to data that can be expressed with a linear function, it benefits from a well-developed mathematical framework that yields unique solutions and exact confidence intervals for regression coefficients. Building on Part I of this series, this article acquaints the reader with some of the important concepts in multiple regression analysis. These include multicollinearity, interaction effects, and an expansion of the discussion of inference testing, leverage, and variable transformations to multivariate models. Examples from the first article in this series are expanded on using a primarily graphic, rather than mathematical, approach. The importance of the relationships among the predictor variables and the dependence of the multivariate model coefficients on the choice of these variables are stressed. Finally, concepts in regression model building are discussed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Keith A Marill. 2004. Advanced statistics: linear regression, part II: multiple linear regression.. https://doi.org/10.1197/j.aem.2003.09.006

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Assessment of blinding in pharmacotherapy and noninvasive neuromodulation randomized controlled trials for neuropathic pain in adults.

In randomized controlled trials (RCTs), study participants and research personnel are often blinded to minimize biases related to knowing treatment allocation. To determine if blinding was effective, participants may be asked which treatment they believe they received ("treatment guess"). This descriptive review characterized blinding assessment (BA) reporting in pharmacotherapy and neuromodulation neuropathic pain RCTs. Of 288 papers, 36 (12.5%) reported a BA. One paper reported the results of 2 studies, so in total 37 studies with a BA were assessed. Of these, 19 were crossover, 17 parallel, and 1 partial crossover in design. All 37 studies assessed participant blinding, and 10 also assessed investigator blinding. Approximately 27% included an "unsure" answer option for treatment guess, and 38% asked the reason for the guess. There were no clear patterns in BA reporting across time nor based on treatment type. Seventeen trials provided sufficient data to calculate Bang Blinding Index (BI) to determine blinding success. Participants remained blinded (BI = 0 &#xb1; 0.2) in 10/17 placebo and 10/17 treatment arms, 6 placebo and 5 treatment arms had a BI > 0.2 suggesting possible unblinding, whereas 1 placebo and 2 treatment arms had a BI < -0.2 suggesting misinformed guessing. Overall, we found that BAs are done in a minority of published neuropathic pain trials and with variable methodology. Given the importance of minimizing risk of bias because of treatment unblinding, future studies should consider including BAs, and further consensus building is necessary to determine if and how BAs should be conducted and interpreted in analgesic clinical trials.

Bias↗

A residuals-based transition model for longitudinal analysis with estimation in the presence of missing data.

We propose a transition model for analysing data from complex longitudinal studies. Because missing values are practically unavoidable in large longitudinal studies, we also present a two-stage imputation method for handling general patterns of missing values on both the outcome and the covariates by combining multiple imputation with stochastic regression imputation. Our model is a time-varying auto-regression on the past innovations (residuals), and it can be used in cases where general dynamics must be taken into account, and where the model selection is important. The entire estimation process was carried out using available procedures in statistical packages such as SAS and S-PLUS. To illustrate the viability of the proposed model and the two-stage imputation method, we analyse data collected in an epidemiological study that focused on various factors relating to childhood growth. Finally, we present a simulation study to investigate the behaviour of our two-stage imputation procedure.

Bias↗

HIV viral dynamic models with dropouts and missing covariates.

In recent years HIV viral dynamic models have received great attention in AIDS studies. Often, subjects in these studies may drop out for various reasons such as drug intolerance or drug resistance, and covariates may also contain missing data. Statistical analyses ignoring informative dropouts and missing covariates may lead to misleading results. We consider appropriate methods for HIV viral dynamic models with informative dropouts and missing covariates and evaluate these methods via simulations. A real data set is analysed, and the results show that the initial viral decay rate, which may reflect the efficacy of the anti-HIV treatment, may be over-estimated if dropout patients are ignored. We also find that the current or immediate previous viral load values may be most predictive for patients' dropout. These results may be important for HIV/AIDS studies.

Bias↗