Patient-based assessment: tools for monitoring and improving healthcare outcomes.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to J E Ware.
Explore the source record for details and available documents.
This article addresses an aspect of outcomes assessment that includes outcomes from the perspective of the patient; namely, functional health status and patient satisfaction. For rehabilitation care they are the only outcomes that can be measured. Measurement tools exist and are being used to quantify the qualitative perspectives of patient care to produce valuable information about the quality and the efficiency of rehab services. The biggest challenge is lack of standardization to allow comparability of data. Different users bring different perspectives to the issue. The researcher, the software developer, the health care facility, and the payer all have an interest in using functional status and patient satisfaction to evaluate outcomes.
BACKGROUND: The Breast Cancer Prevention Trial (BCPT) is a large, multicenter chemoprevention trial testing the efficacy of the antiestrogen drug tamoxifen for prevention of breast cancer and coronary heart disease in healthy women at high risk of breast cancer. The BCPT evolved from a series of prior studies in early stage breast cancer demonstrating the efficacy of tamoxifen in the prevention of systemic breast cancer recurrence and in the reduction of contralateral breast cancers. PURPOSE: The purpose of this article is to describe the methodologic considerations in the collection of health-related quality-of-life (HRQL) data in the BCPT and to present base-line HRQL data on the first 9749 participants. METHODS: An HRQL questionnaire that included the Center for Epidemiologic Studies-Depression Scale, a symptom checklist, the Medical Outcomes Study 36-item short form (MOS-SF-36), and the MOS sexual problems questions was completed by participants in the BCPT at base line (prior to random assignment). Medical and demographic information, as well as projected risk of breast cancer, were collected as part of study eligibility. Descriptive and correlational data were examined for these study participants. RESULTS: BCPT participants report high levels of functioning compared with U.S. general population norms but still report an average of 8.9 distinct symptoms during the past 4 weeks. Depression is less prevalent among the participants than in community samples, which reflects the exclusion of clinically depressed individuals. Sixty-five percent reported being sexually active in the past 6 months, with an age-related decline in sexual activity. Younger women reported fewer sexual problems than older women. There is a strong correlation between the two mental health measures, moderate to weak correlations between HRQL scales and levels of self-reported symptoms, and only weak correlations between measures of breast cancer risk and HRQL scales. The MOS-SF-36 scores were examined for three consecutive recruitment samples (0-6 months, 7-12 months, and 13-20 months), and the base-line scores were slightly better for the earliest group of participants. CONCLUSIONS: This article demonstrates the feasibility of collecting HRQL data in a large, multicenter, chemoprevention trial for women at high risk of breast cancer. The successful integration of HRQL data collection into this clinical trial attests to its value as a safety-monitoring end point and as an explicit and measurable outcome for the entire trial. IMPLICATIONS: HRQL data are important for studies in which healthy populations are involved and in which the potential for decrements in quality of life are real or perceived.
We studied 31 previously validated and newly developed generic and epilepsy-specific scales to evaluate their usefulness for assessing the impact of epilepsy and anti-epileptic drug (AED) therapy on health-related quality of life (HRQOL). Included were the MOS SF-36 Health Survey, additional measures of mental health, cognition, epilepsy-specific perception of control, behavioural problems, distress, worries and experiences, the Liverpool Epilepsy Impact and Seizure Severity scales, and a patient-completed symptom checklist. Questionnaires were completed twice by 136 patients on AED therapy in a multicentre study in the UK. Validity was assessed in relation to disease severity, defined as time since last seizure, and to patient-reported symptoms. Statistical analyses to estimate the contribution of HRQOL information of each scale relative to that of others were conducted. The 171-item questionnaire could be completed by out-patients with epilepsy with good data quality. With few exceptions, generic and epilepsy-specific measures satisfied psychometric tests of hypothesized item groupings and scale score reliability (internal consistency and test-retest reliability) and differentiated well between groups of patients differing in time since last seizure and in symptom impact, regardless of time since last seizure. However, scales differed widely in their validity in discriminating between groups of patients known to differ clinically. The SF-36 Role Physical scale best discriminated among groups differing in disease severity. The epilepsy-specific Mastery, Impact, Experience, Worry, Distress, and Agitation scales were among the 10 best measures in discriminating among groups differing in disease severity. Generic measures, especially measures of social and role functioning and mental health, were best at differentiating groups of patients differing in symptom impact. Recommendations are offered for concepts and specific scales most likely to be useful in future studies of the HRQOL burden of epilepsy and the HRQOL benefits of AED therapy.
We document the applicability of the SF-36 Health Survey, which was translated into Swedish using methods later adopted by the International Quality of Life Assessment (IQOLA) Project procedures. To test its appropriateness for use in Sweden, it was administered through mail-out/mail-back questionnaires in seven general population studies with an average response rate of 68%. The 8930 respondents varied by gender (48.2% men), age (range 15-93 years, mean age 42.7), marital status, education, socio-economic status, and geographical area. Psychometric methods used in the evaluation of the SF-36 in the U.S. were replicated. Over 90% of respondents had complete items for each of the eight SF-36 scales, although more missing data were observed for subjects 75 years and over. Scale scores could be computed for the vast majority of respondents (95% and over); slightly fewer in the oldest subgroup. Item-internal consistency was consistently high across socio-demographic subgroups and the eight scales. Most reliability estimates exceeded the 0.80 level. The highest reliability was observed for the Bodily Pain Scale where all subgroups met the 0.90 level recommended for individual comparisons; coefficients at or above 0.90 were also observed in most subgroups for the Physical Functioning Scale. Tests of scaling assumptions including hypothesized item groupings, which reflect the construct validity of scales, were consistently favorable across subgroups, although lower rates were noted in the oldest age group. In conclusion, these studies have yielded empirical evidence supporting the feasibility of a non-English language reproduction of the SF-36 Health Survey. The Swedish SF-36 is ready for further evaluation.
There is growing demand for translations of health status questionnaires for use in multinational drug therapy studies and for population comparisons of health statistics. The International Quality of Life Assessment (IQOLA) Project is conducting a three-stage research program to determine the feasibility of translating the SF-36 Health Survey, widely used in English-speaking countries, into other languages. In stage 1, the conceptual equivalence and acceptability of translated questionnaires are evaluated and improved using qualitative and quantitative methods. In stage 2, assumptions underlying the construction and scoring of questionnaire scales are tested empirically. In stage 3, the equivalence of the interpretation of questionnaire scores across countries is tested using methods that closely approximate their intended use, and empirical results are compared. Data analyses from Sweden and the United Kingdom, as well as other research cited, support the feasibility of cross-cultural health measurement using the SF-36.
In a randomized 6-week trial comparing fluoxetine with placebo, the Medical Outcomes Study 36-Item Short-Form Health Status Survey (SF-36) scales were used to measure the effects of treatment on functional health and well-being among elderly (age > or = 60 years) outpatients with major depression. In the fluoxetine and placebo groups, 261 and 271 patients, respectively, completed the SF-36 before treatment and at Weeks 3 and 6. Compared with national norms for individuals over age 60, study patients before treatment exhibited baseline decrements on the following SF-36 scales: mental health, role limitations due to emotional problems, social functioning, vitality, role limitations due to physical problems, and bodily pain. Analyses of SF-36 changed scores from baseline to Week 6 revealed that the fluoxetine group improved more than the placebo group across all scales. Differences in changes of scores between groups were significant (p < .05), favoring the fluoxetine group for the scales of mental health, role limitations due to emotional problems, physical functioning, and bodily pain. Improvements observed in the fluoxetine group were both clinically and socially significant.
Alternate-form health measures are useful for clinical trials or health services research requiring repeated administrations over a short interval of time. Further, by using alternate-form methodology, they can be utilized to estimate score reliability. Data from the Medical Outcomes Study were used to evaluate five alternate forms of the Short-Form 36-Item Health Survey (SF-36) general mental health scale (MHI-5). Well-established psychometric criteria were used to select the best alternate form and to estimate the reliability of the MHI-5 using the alternate-form methodology. Although a considerable degree of comparability across the five alternate forms was observed for criteria pertaining to estimates of item-internal consistency and reliability, distributional characteristics of scales, tests of empirical validity, and score equivalence at the individual level, we recommend one alternate form that satisfied all evaluation criteria and did so better than any other alternate form. Using the alternate-form methodology of estimating reliability, results suggest that the internal-consistency method underestimates the reliability of the MHI-5 by 3%. The methodology presented here should prove useful to others interested in constructing and evaluating alternate forms, and the alternate form recommended here (MHI-5AF) should prove useful across many health status assessment applications.
This article identifies the characteristics of patients and office visits associated with decreased mutual decision-making between physicians and patients. In the baseline cross-sectional survey of the Medical Outcomes Study we measured specific patient characteristics hypothesized to influence participatory decision-making (PDM) styles of physicians. We related these characteristics to the PDM style scores for their physicians. The study was conducted in solo practices, multi-specialty groups, and health maintenance organizations in Boston, Chicago, and Los Angeles. Over a 9-day period in 1986, 8,316 patients were sampled from the practices of 344 participating Medical Outcome Study physicians, representing general internal medicine, family practice, cardiology and endocrinology. Physicians' PDM style was measured using a 3-item scale included on the baseline questionnaire completed by patients after office visits to their Medical Outcome Study physicians. We found that the elderly (age 75 and older) and young adult (younger than age 30) patients, patients with high school education or less, minority patients, and male patients had the least participatory visits with their physicians. We also found that male patients seeing male physicians had the least participatory visits compared with male patients seeing female physicians, and compared with female patients seeing physicians of either gender. Our data indicated that PDM style increased as duration or tenure of the physician-patient relationship increased. Participatory decision-making style also increased with increasing length of office visits. The role of effective interpersonal care in optimizing patients' health outcomes may be underappreciated. We have identified seven patient and visit characteristics that maximize or compromise the effectiveness of interpersonal care. Recognizing those at risk for suboptimal interpersonal care may be a first step in improving the management of chronic disease. Key words: participatory decision-making style; interpersonal care; doctor-patient communication.
General health status and a broader concept of quality of life are discussed and methods of widely used surveys are reviewed. A consensus regarding the inclusion of measures of physical, mental, social, and role functioning and general health perceptions is noted for comprehensive assessments of health. A schematic of relationships among condition-specific and generic measures is presented along with results expected for objective and subjective measures of physical and mental dimensions of health. Suggestions are offered for the labeling of disease-specific and generic measures and ways to avoid confounding of content. Applications of health surveys in general population monitoring, health policy evaluation, clinical trials of alternative treatments, monitoring and improving of health care outcomes, and in everyday clinical practice are exemplified and discussed. A unified measurement strategy is proposed and arguments in favor of standardizing the content of health surveys across applications are offered.
Physical component summary (PCS) and mental component summary (MCS) measures make it possible to reduce the number of statistical comparisons and thereby the role of chance in testing hypotheses about health outcomes. To test their usefulness relative to a profile of eight scores, results were compared across 16 tests involving patients (N = 1,440) participating in the Medical Outcomes Study. Comparisons were made between groups known to differ at a point in time or to change over time in terms of age, diagnosis, severity of disease, comorbid conditions, acute symptoms, self-reported changes in health, and recovery from clinical depression. The relative validity (RV) of each measure was estimated by a comparison of statistical results with those for the best scales in the same tests. Differences in RV among scales from the Medical Outcomes Study 36-Item Short-Form Health Survey (SF-36) were consistent with those in previous studies. One or both of the summary measures were significant for 14 of 15 differences detected in multivariate analyses of profiles and detected differences missed by the profile in one test. Relative validity coefficients ranged from .20 to .94 (median, .79) for PCS in tests involving physical criteria and from .93 to 1.45 (median, 1.02) for MCS in tests involving mental criteria. The MCS was superior to the best SF-36 scale in three of four tests involving mental health. Results suggest that the two summary measures may be useful in most studies and that their empiric validity, relative to the best SF-36 scale, will depend on the application. Surveys offering the option of analyzing both a profile and psychometrically based summary measures have an advantage over those that do not.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Indexes developed to measure physical functioning as an essential component of general health status are often based on sets of hierarchically-structured items intended to represent a broad underlying concept. Rasch Item Response Theory (IRT) provides a methodology to examine the hierarchical structure, unidimensionality, and reproducibility of item positions (calibrations) along a scale. Data gathered on the 10-item Physical Functioning Scale (PF-10) from a large sample of Medical Outcomes Study patients (N = 3445) were used to examine the hierarchical order, unidimensionality, and reproducibility of item calibrations. Rasch-IRT analyses generated an empirical item hierarchy, confirmed the unidimensionality of the PF-10 for most patients, and established the reproducibility of item calibrations across patient populations and repeated tests. These findings support the content validity of the PF-10 as a measure of physical functioning and suggest that valid Rasch-IRT summary scores could be generated as an alternative to the current Likert summative scores. Unidimensionality and reproducibility of the item scale are essential prerequisites for the development of Rasch-based person measures of physical functioning that can be used across populations and over repeated tests.
The widespread use of standardized health surveys is predicated on the largely untested assumption that scales constructed from those surveys will satisfy minimum psychometric requirements across diverse population groups. Data from the Medical Outcomes Study (MOS) were used to evaluate data completeness and quality, test scaling assumptions, and estimate internal-consistency reliability for the eight scales constructed from the MOS SF-36 Health Survey. Analyses were conducted among 3,445 patients and were replicated across 24 subgroups differing in sociodemographic characteristics, diagnosis, and disease severity. For each scale, item-completion rates were high across all groups (88% to 95%), but tended to be somewhat lower among the elderly, those with less than a high school education, and those in poverty. On average, surveys were complete enough to compute scales scores for more than 96% of the sample. Across patient groups, all scales passed tests for item-internal consistency (97% passed) and item-discriminant validity (92% passed). Reliability coefficients ranged from a low of 0.65 to a high of 0.94 across scales (median = 0.85) and varied somewhat across patient subgroups. Floor effects were negligible except for the two role disability scales. Noteworthy ceiling effects were observed for both role disability scales and the social functioning scale. These findings support the use of the SF-36 survey across the diverse populations studied and identify population groups in which use of standardized health status measures may or may not be problematic.
Many health status surveys have been designed for mail, telephone, or in-person administration. However, with rare exception, investigators have not studied the effect the survey mode of administration has on the way respondents assess their health and other important parameters (such as response rates, nonresponse bias, and data quality), which can affect the generalizability of results. Using a national sampling frame of noninstitutionalized adults from the General Social Survey, we randomly assigned adults to a mail survey (80%) or a computer-assisted telephone survey (20%). The surveys were designed to provide national norms for the SF-36 Health Survey. Total data collection costs per case for the telephone survey ($47.86) were 77% higher than that for the mail survey ($27.07). A significantly higher response rate was achieved among respondents randomly assigned to the mail (79.2%) than telephone survey (68.9%). Nonresponse bias was evident in both modes but, with the exception of age, was not differential between modes. The rate of missing responses was higher for mail than telephone respondents (1.59 vs. 0.49 missing items). Health ratings based on the SF-36 scales were less favorable, and reports of chronic conditions were more frequent, for mail than telephone respondents. Results are discussed in light of the trade-offs involved in choosing a survey methodology for health status assessment applications. Norms for mail and telephone versions of the SF-36 survey are provided for use in interpreting individual and group scores.
OBJECTIVE: Compare adult migraineurs' health related quality of life to adults in the general U.S. population reporting no chronic conditions, and to samples of patients with other chronic conditions. METHODS: Subjects (n = 845) were surveyed 2-6 months after participation in a placebo-controlled clinical trial and asked to complete a questionnaire including the SF-36 Health Survey, a migraine severity measurement scale and demographics. Results were adjusted for severity of illness and comorbidities. Scores were compared with responses to the same survey by the U.S. sample and by patients with other chronic conditions. RESULTS: Response rate was 67%. After adjustment for comorbid conditions, SF-36 scale scores were significantly (P 0.001) lower in migraineurs, relative to age and sex-adjusted norms for the U.S. sample with no chronic conditions. Some health dimensions were more affected by migraine than other chronic conditions, while other dimensions were less affected by migraine. Measures of bodily pain, role disability due to physical health and social functioning discriminated best between migraineurs, the U.S. sample, and patients with other chronic conditions. Patients reporting moderate, severe and very severe migraines scored significantly (P < or = 0.001) lower on five of the eight SF-36 scales than the U.S. sample. CONCLUSIONS: Migraine has a unique, significant quality of life burden.
The patient's opinion is central to the monitoring and improvement of health outcomes. The goal of treatment should be the preservation of function and well-being of the patient. Despite this, standardized assessments of patients' experiences of disease and treatment are not routinely collected in clinical research and medical practice. In an era of cost containment, it is essential to monitor health outcomes. A prototype for collection of relevant data can be found in the Medical Outcomes Study which tests methods for monitoring the results of medical care among patients with hypertension and other conditions. Results from this study have shown that there is good reason to be optimistic about the feasibility of standardized, self-administered questionnaires as a primary means of collection for patient outcome data. Furthermore, it is possible to create an enhanced data base and add to it routinely on a large scale across diverse health care settings. Details of the health outcome measures used in the Medical Outcomes Study are presented in this paper.