Search PubMed⌕ Search

Biomedical subjects

J E Ware

Publications and source records attributed to J E Ware.

At least 37 records · Page 2Linked to original sources

The SF-36 Arthritis-Specific Health Index (ASHI): II. Tests of validity in four clinical trials.

OBJECTIVE: The SF-36 Arthritis-Specific Health Index (ASHI) was constructed to improve the responsiveness of the SF-36 Health Survey to changes in the severity of arthritis through the use of arthritis-specific scoring algorithms. This study compared the responsiveness of the ASHI and other generic scales and summary measures scored from the SF-36 in clinical trials of health outcomes for patients with arthritis. METHODS: Longitudinal data for patients (n = 835) participating in four placebo-controlled trials were analyzed. Study participants had at least a 6-month history of moderate to severe osteoarthritis or rheumatoid arthritis of the knee or hip. All had undergone a washout period of 3 to 14 days before baseline assessment to bring about a flare state in osteoarthritis or rheumatoid arthritis symptoms. Their average age was 60 years, and 72% were female. Responders and nonresponders were classified on the basis of physician assessments of changes in arthritis severity, with blinding as to treatment group; treated and untreated (placebo) groups were also compared. For the SF-36 ASHI, generic physical (PCS) and mental (MCS) component summary measures and each of eight subscales scored from the SF-36 (acute version) change scores were computed by subtracting scores before treatment from scores at 2-week follow-up. To evaluate empirical validity, analyses of variance were performed. For each measure, an F-ratio was computed for the comparison between clinically defined groups of responders and nonresponders and between groups of patients assigned to placebo versus drug therapy. Relative validity (RV) coefficients were computed for the ASHI in comparison with PCS, MCS, and the best SF-36 scale to determine which was more responsive. RESULTS: In analyses of each of the four trials and all trials combined, RV coefficients for the ASHI were higher than those for both of the generic SF-36 summary measures and for the most valid SF-36 scale (Bodily Pain), with only one exception. Across 40 tests of validity in distinguishing treated from untreated patients, the ASHI was 5% to 19% more valid than the best SF-36 scale (RV = 1.05-1.19; RV = 1.10 in all trials combined). The generic summary measures (PCS and MCS) were much less valid in these tests (RV = 0.67 and 0.27, respectively). In analyses of responders and nonresponders, RV coefficients for the ASHI ranged from 0.70 to 1.22 (RV = 1.04 in all trials combined), in comparison with the best SF-36 subscale, which was always Bodily Pain. RV coefficients were lower for PCS (RV = 0.75) and much lower than the MCS (RV = 0.18) in comparisons of treatment outcomes based on all trials combined. CONCLUSION: The ASHI appears to be more valid than the eight SF-36 scales and PCS and MCS summary measures for purposes of distinguishing between treated and untreated patients and between clinical responders and nonresponders. This study demonstrates the feasibility of improving the validity of the SF-36 through the use of arthritis-specific scoring while retaining the option of generic scoring, which makes it possible to also compare results across diseases and treatments.

Arthritis, Rheumatoid↗

Reliability and validity of French, German, Italian, Dutch, and UK English translations of the Medical Outcomes Study HIV Health Survey.

OBJECTIVES: Test the reliability and validity of 5 translations of the 34-item version of the MOS HIV for use in multinational clinical trials. RESEARCH DESIGN: Investigators in five countries followed a standardized protocol and recruited HIV+ patients stratified by disease stage: asymptomatic; symptomatic; and AIDS. During routine clinic visits, patients completed the MOS HIV and a checklist of HIV-related symptoms. Clinicians reported patients' demographics, most recent CD4+ count and disease stage. SUBJECTS: Three hundred and sixty three HIV+ outpatients attending AIDS clinics in The Netherlands, France, Germany, Italy, and England. MEASURES: Dutch, French, German, Italian, and UK English translations of the MOS HIV CD4+ cell count and the SCL-57. RESULTS: All translations recruited roughly equal proportions of each disease stage, although the number of patients recruited differed by translation (n: German = 92, French = 86; Italian = 88; UK English = 72; and Dutch = 25). Internal consistency reliability was similar across translations and adequate (alpha >.70) for all scales except for Mental Health in the French sample. Multi-trait analyses supported structural validity of the MOS HIV scales in each translation. Principal component analysis of scale scores identified 2 dimensions for all translations except German. For all translations, scores were significantly correlated with symptom severity scores but were uncorrelated with CD4+ cell counts. CONCLUSIONS: In general, the 5 translations of the MOS HIV had similar psychometric properties to those reported in the validation study for the original US English version of the MOS HIV. With some revision, these translations promise to provide useful quality of life data from HIV+ subjects in clinical trials.

Activities of Daily Living↗

Sustained virologic response is associated with improved health-related quality of life in relapsed chronic hepatitis C patients.

Although evidence of virologic elimination, normalization of serum alanine aminotransferase levels, and reduction in liver inflammation are the principal therapeutic outcome goals in chronic hepatitis C patients, improvement in health-related quality of life (HQL) is also an important aspect of therapeutic outcome. In a recent report of chronic hepatitis C patients treated for 24 weeks with interferon, sustained virologic response (24 weeks post-treatment) was associated with improvement in HQL compared with nonresponse. We report on the relationship between sustained virologic response and Hepatitis Quality-of-Life Questionnaire (HQLQ) survey results of patients who relapsed after a previous course of interferon alfa who were subsequently treated with recombinant interferon alfa-2b (rIFN-alpha 2b) either alone or in combination with ribavirin. The HQLQ was administered at baseline, at treatment Weeks 12 and 24, and at follow-up Weeks 12 and 24. All patients received rIFN-alpha 2b 3 million International Units by subcutaneous injection three times weekly plus either oral ribavirin (1,000 or 1,200 mg) or placebo daily for 24 weeks. At baseline, patients scored lower than adjusted population norms in HQL. Relative to patients treated with rIFN-alpha 2b monotherapy, patients receiving combination therapy showed better HQL in 6 of 13 domains. Furthermore, sustained virologic response in either treatment group was associated with improvement in the scores of both generic and hepatitis-specific HQL survey domains. These results indicate that successful therapeutic resolution of hepatitis C infection improves HQL as assessed by generic and hepatitis C-specific measures of functional health and well-being. Furthermore, improvements in HQL outcome measures may predict reduced demand for health care resources and greater productivity in the workplace.

Adult↗

Overview of the SF-36 Health Survey and the International Quality of Life Assessment (IQOLA) Project.

This article presents information about the development and evaluation of the SF-36 Health Survey, a 36-item generic measure of health status. It summarizes studies of reliability and validity and provides administrative and interpretation guidelines for the SF-36. A brief history of the International Quality of Life Assessment (IQOLA) Project is also included.

Activities of Daily Living↗

Translating health status questionnaires and evaluating their quality: the IQOLA Project approach. International Quality of Life Assessment.

This article describes the methods adopted by the International Quality of Life Assessment (IQOLA) project to translate the SF-36 Health Survey. Translation methods included the production of forward and backward translations, use of difficulty and quality ratings, pilot testing, and cross-cultural comparison of the translation work. Experience to date suggests that the SF-36 can be adapted for use in other countries with relatively minor changes to the content of the form, providing support for the use of these translations in multinational clinical trials and other studies. The most difficult items to translate were physical functioning items, which used examples of activities and distances that are not common outside of the United States; items that used colloquial expressions such as pep or blue; and the social functioning items. Quality ratings were uniformly high across countries. While the IQOLA approach to translation and validation was developed for use with the SF-36, it is applicable to other translation efforts.

Cross-Cultural Comparison↗

Cross-cultural comparisons of the content of SF-36 translations across 10 countries: results from the IQOLA Project. International Quality of Life Assessment.

Increasingly, translated and culturally adapted health-related quality of life measures are being used in cross-cultural research. To assess comparability of results, researchers need to know the comparability of the content of the questionnaires used in different countries. Based on an item-by-item discussion among International Quality of Life Assessment (IQOLA) investigators of the content of the translated versions of the SF-36 in 10 countries, we discuss the difficulties that arose in translating the SF-36. We also review the solutions identified by IQOLA investigators to translate items and response choices so that they are appropriate within each country as well as comparable across countries. We relate problems and solutions to ratings of difficulty and conceptual equivalence for each item. The most difficult items to translate were physical functioning items that refer to activities not common outside the United States and items that use colloquial expressions in the source version. Identifying the origin of the source items, their meaning to American English-speaking respondents and American English synonyms, in response to country-specific translation issues, greatly helped the translation process. This comparison of the content of translated SF-36 items suggests that the translations are culturally appropriate and comparable in their content.

Cross-Cultural Comparison↗

Testing the equivalence of translations of widely used response choice labels: results from the IQOLA Project. International Quality of Life Assessment.

The similarity in meaning assigned to response choice labels from the SF-36 Health Survey (SF-36) was evaluated across countries. Convenience samples of judges (range, 10 to 117; median = 48) from 13 countries rated translations of response choice labels, using a variation of the Thurstone method of equal appearing intervals. Judges marked a point on a 10-cm line-representing the magnitude of a response choice label (e.g., "good" relative to the anchors of "poor" and "excellent"). Ratings were evaluated to determine the ordinal consistency of response choice labels within a response scale; the degree to which differences between adjacent response choice labels were equal interval; and the amount of variance due to response choice label, country, judge, and interaction between response choice label and country. Results confirmed the hypothesized ordering of response choice labels; the percentage of ordinal pairs ranged from 88.7% to 100% (median = 98.2%) across countries and response scales. Examination of the average magnitudes of response choice labels supported the "quasi-interval" nature of the scales. Analysis of variance (ANOVA) results supported the generalizability of response choice magnitudes across countries; labels explained 64% to 77% of the variance in ratings, and country explained 1% to 3%. These results support the equivalence of SF-36 response choice labels across countries. Departures from the assumption of equal intervals, when observed, were similar across countries and were greatest for the two response scales that are recalibrated under standard SF-36 scoring. Results provide justification for scoring translations of individual items using standard SF-36 scoring; whether these items form the same scales in other countries as they do in the United States is evaluated with tests of scaling assumptions.

Analysis of Variance↗

Methods for testing data quality, scaling assumptions, and reliability: the IQOLA Project approach. International Quality of Life Assessment.

Following the translation development stage, the second research stage of the IQOLA Project tests the assumptions underlying item scoring and scale construction. This article provides detailed information on the research methods used by the IQOLA Project to evaluate data quality, scaling and scoring assumptions, and the reliability of the SF-36 scales. Tests include evaluation of item and scale-level descriptive statistics; examination of the equality of item-scale correlations, item internal consistency and item discriminant validity; and estimation of scale score reliability using internal consistency and test-retest methods. Results from these tests are used to determine if standard algorithms for the construction and scoring of the eight SF-36 scales can be used in each country and to provide information that can be used in translation improvement.

Activities of Daily Living↗

Psychometric and clinical tests of validity of the Japanese SF-36 Health Survey.

Cross-sectional data from a representative sample of the general population in Japan were analyzed to test the validity of Japanese SF-36 Health Survey scales as measures of physical and mental health. Results from psychometric and clinical tests of validity were compared. Principal components analyses were used to test for the hypothesized physical and mental dimensions of health and the pattern of scale correlations with those components. To test the clinical validity of SF-36 scale scores, self-reports of chronic medical conditions and the Zung Self-Rating Depression Scale were used to create mutually exclusive groups differing in the severity of physical and mental conditions. The pattern of correlations between the SF-36 scales and the two empirically derived components generally confirmed hypotheses for most scales. Results of psychometric and clinical tests of validity were in agreement for the Physical Functioning, Role-Physical, Vitality, Social Functioning, and Mental Health scales. Relatively less agreement between psychometric and clinical tests of validity was observed for the Bodily Pain, General Health, and Role-Emotional scales, and the physical and mental health factor content of those scales was not consistent with hypotheses. In clinical tests of validity, the General Health, Bodily Pain, and Physical Functioning scales were the most valid scales in discriminating between groups with and without a severe physical condition. Scales that correlated highest with mental health in the components analysis (Mental Health and Vitality) also were most valid in discriminating between groups with and without depression. The results of this study provide preliminary interpretation guidelines for all SF-36 scales, although caution is recommended in the interpretation of the Role-Emotional, Bodily Pain, and General Health scales pending further studies in Japan.

Adult↗

Tests of data quality, scaling assumptions, and reliability of the SF-36 in eleven countries: results from the IQOLA Project. International Quality of Life Assessment.

Data from general population samples in 11 countries (n = 1483 to 9151) were used to assess data quality and test the assumptions underlying the construction and scoring of multi-item scales from the SF-36 Health Survey. Across all countries, the rate of item-level missing data generally was low, although slightly higher for items printed in the grid format. In each country, item means generally were clustered as hypothesized within scales. Correlations between items and hypothesized scales were greater than 0.40 with one exception, supporting item internal consistency. Items generally correlated significantly higher with their own scale than with competing scales, supporting item discriminant validity. Scales could be constructed for 93-100% of respondents. Internal consistency reliability of the eight SF-36 scales was above 0.70 for all scales, with two exceptions. Floor effects were low for all except the two role functioning scales; ceiling effects were high for both role functioning scales and also were noteworthy for the Physical Functioning, Bodily Pain, and Social Functioning scales in some countries. These results support the construction and scoring of the SF-36 translations in these 11 countries using the method of summated ratings.

Cross-Cultural Comparison↗

The factor structure of the SF-36 Health Survey in 10 countries: results from the IQOLA Project. International Quality of Life Assessment.

Studies of the factor structure of the SF-36 Health Survey are an important step in its construct validation. Its structure is also the psychometric basis for scoring physical and mental health summary scales, which are proving useful in simplifying and interpreting statistical analyses. To test the generalizability of the SF-36 factor structure, product-moment correlations among the eight SF-36 Health Survey scales were estimated for representative samples of general populations in each of 10 countries. Matrices were independently factor analyzed using identical methods to test for hypothesized physical and mental health components, and results were compared with those published for the United States. Following simple orthogonal rotation of two principal components, they were easily interpreted as dimensions of physical and mental health in all countries. These components accounted for 76% to 85% of the reliable variance in scale scores across nine European countries, in comparison with 82% in the United States. Similar patterns of correlations between the eight scales and the components were observed across all countries and across age and gender subgroups within each country. Correlations with the physical component were highest (0.64 to 0.86) for the Physical Functioning, Role Physical, and Bodily Pain scales, whereas the Mental Health, Role Emotional, and Social Functioning scales correlated highest (0.62 to 0.91) with the mental component. Secondary correlations for both clusters of scales were much lower. Scales measuring General Health and Vitality correlated moderately with both physical and mental health components. These results support the construct validity of the SF-36 translations and the scoring of physical and mental health components in all countries studied.

Cross-Cultural Comparison↗

The equivalence of SF-36 summary health scores estimated using standard and country-specific algorithms in 10 countries: results from the IQOLA Project. International Quality of Life Assessment.

Data from general population surveys (n = 1771 to 9151) in nine European countries (Denmark, France, Germany, Italy, the Netherlands, Norway, Spain, Sweden, and the United Kingdom) were analyzed to test the algorithms used to score physical and mental component summary measures (PCS-36/MCS-36) based on the SF-36 Health Survey. Scoring coefficients for principal components were estimated independently in each country using identical methods of factor extraction and orthogonal rotation. PCS-36 and MCS-36 scores were also estimated using standard (U.S.-derived) scoring algorithms, and results were compared. Product-moment correlations between scores estimated from standard and country-specific scoring coefficients were very high (0.98 to 1.00) for both physical and mental health components in all countries. As hypothesized for orthogonal components, correlations between physical and mental components within each country were very low (0.00 to 0.12) for both estimation methods. Mean scores for PCS-36 differed by as much as 3.0 points across countries using standard scoring, and mean scores for MCS-36 differed across countries by as much as 6.4 points. In view of the high degree of equivalence observed within each country, using standard and country-specific algorithms, we recommend use of standard scoring algorithms for purposes of multinational studies involving these 10 countries.

Algorithms↗

Cross-validation of item selection and scoring for the SF-12 Health Survey in nine countries: results from the IQOLA Project. International Quality of Life Assessment.

Data from general population surveys (n = 1483 to 9151) in nine European countries (Denmark, France, Germany, Italy, the Netherlands, Norway, Spain, Sweden, and the United Kingdom) were analyzed to cross-validate the selection of questionnaire items for the SF-12 Health Survey and scoring algorithms for 12-item physical and mental component summary measures. In each country, multiple regression methods were used to select 12 SF-36 items that best reproduced the physical and mental health summary scores for the SF-36 Health Survey. Summary scores then were estimated with 12 items in three ways: using standard (U.S.-derived) SF-12 items and scoring algorithms; standard items and country-specific scoring; and country-specific sets of 12 items and scoring. Replication of the 36-item summary measures by the 12-item summary measures was then evaluated through comparison of mean scores and the strength of product-moment correlations. Product-moment correlations between SF-36 summary measures and SF-12 summary measures (standard and country-specific) were very high, ranging from 0.94-0.96 and 0.94-0.97 for the physical and mental summary measures, respectively. Mean 36-item summary measures and comparable 12-item summary measures were within 0.0 to 1.5 points (median = 0.5 points) in each country and were comparable across age groups. Because of the high degree of correspondence between summary physical and mental health measures estimated using the SF-12 and SF-36, it appears that the SF-12 will prove to be a practical alternative to the SF-36 in these countries, for purposes of large group comparisons in which the focus is on overall physical and mental health outcomes.

Cross-Cultural Comparison↗

Use of structural equation modeling to test the construct validity of the SF-36 Health Survey in ten countries: results from the IQOLA Project. International Quality of Life Assessment.

A crucial prerequisite to the use of the SF-36 Health Survey in multinational studies is the reproduction of the conceptual model underlying its scoring and interpretation. Structural equation modeling (SEM) was used to test these aspects of the construct validity of the SF-36 in ten IQOLA countries: Denmark, France, Germany, Italy, the Netherlands, Norway, Spain, Sweden, the United Kingdom, and the United States. Data came from general population surveys fielded to gather normative data. Measurement and structural models developed in the United States were cross-validated in random halves of the sample in each country. SEM analyses supported the eight first-order factor model of health that underlies the scoring of SF-36 scales and two second-order factors that are the basis for summary physical and mental health measures. A single third-order factor was also observed in support of the hypothesis that all responses to the SF-36 are generated by a single, underlying construct--health. In addition, a third second-order factors, interpreted as general well-being, was shown to improve the fit of the model. This model (including eight first-order factors, three second-order factors, and one third-order factor) was cross-validated using a holdout sample within the United States and in each of the nine other countries. These results confirm the hypothesized relationships between SF-36 items and scales and justify their scoring in each country using standard algorithms. Results also suggest that SF-36 scales and summary physical and mental health measures will have similar interpretations across countries. The practical implications of a third second-order SF-36 factor (general well-being) warrant further study.

Cross-Cultural Comparison↗

Differential item functioning in the Danish translation of the SF-36.

Statistical analyses of Differential Item Functioning (DIF) can be used for rigorous translation evaluations. DIF techniques test whether each item functions in the same way, irrespective of the country, language, or culture of the respondents. For a given level of health, the score on any item should be independent of nationality. This requirement can be tested through contingency-table methods, which are efficient for analyzing all types of items. We investigated DIF in the Danish translation of the SF-36 Health Survey, using two general population samples (USA, n = 1,506; Denmark, n = 3,950). DIF was identified for 12 out of 35 items. These results agreed with independent ratings of translation quality, but the statistical techniques were more sensitive. When included in scales, the items exhibiting DIF had only a little impact on conclusions about cross-national differences in health in the general population. However, if used as single items, the DIF items could seriously bias results from cross-national comparisons. Also, the DIF items might have larger impact on cross-national comparison of groups with poorer health status. We conclude that analysis of DIF is useful for evaluating questionnaire translations.

Adolescent↗

Comparison of Rasch and summated rating scales constructed from SF-36 physical functioning items in seven countries: results from the IQOLA Project. International Quality of Life Assessment.

Rasch models for polytomous items were used to assess the scaling assumptions and compare item response patterns in the 10-item SF-36 physical functioning scale (PF-10) for general population respondents in Denmark, Germany, Italy, the Netherlands, Sweden, the United Kingdom, and the United States. The Rasch model of physical functioning developed in the United States was compared to models for other countries, and each country was compared to a multinational composite. Strong scale congruence across the seven countries was demonstrated; items that varied between countries and from the composite may reflect unique cultural response patterns or differences in translation. Scoring algorithms based on the Rasch model for each country were superior to the current Likert scoring in tests of relative validity (RV) in discriminating among age groups in all countries. In relation to the Likert PF-10 scoring (RV = 1.00), scores estimated using the Rasch rating scale model achieve a median RV of 1.31 (range: 1.01-1.59), while the Rasch partial credit model attained a median RV of 1.44 (range: 1.01-2.23). Rasch models hold good potential for improving health status measures, estimating individual scores when responses to scale items are missing, and equating scores across countries.

Aged↗

Canadian-French, German and UK versions of the Child Health Questionnaire: methodology and preliminary item scaling results.

Using emerging international guidelines, stringent procedures were used to develop and evaluate Canadian-French, German and UK translations/adaptions of the 50 item, parent-completed Child Health Questionnaire (CHQ-PF50). Multitrait analysis was used to evaluate the convergent and discriminant validity of the hypothesized item sets across countries relative to the results obtained for a representative sample of children in the US. Cronbach's alpha coefficient was used to estimate the internal consistency reliability for each of the health scales. Floor and ceiling effects were also examined. Seventy-nine percent of all the item-scale correlations achieved acceptable internal consistency (0.40 or higher). The tests of the item convergent and discriminant validity were successful at least 87% of the time across all scales and countries. Equal item variance was observed 90% of the time across all countries. The reliability coefficients ranged from a low of 0.43 (parental time impact, Canadian English) to a high of 0.97 (physical functioning index, Canadian French) across all scales (median 0.80). Negligible floor effects were observed across countries. Noteworthy ceiling effects were observed, as expected, for the hypothesized physical scales (mean effect 73%). Conversely, fewer ceiling effects were observed for the psychosocial scales (range 3-17% behaviour-parental emotional impact). The item-scaling results obtained in these pilot studies support the psychometric properties of the American-English CHQ-PF50 and its respective translations.

Adolescent↗