Search PubMed⌕ Search

Biomedical subjects

Neil Aaronson

Publications and source records attributed to Neil Aaronson.

8 recordsLinked to original sources

Multidimensional computerized adaptive testing of the EORTC QLQ-C30: basic developments and evaluations.

OBJECTIVE: Self-report questionnaires are widely used to measure health-related quality of life (HRQOL). Ideally, such questionnaires should be adapted to the individual patient and at the same time scores should be directly comparable across patients. This may be achieved using computerized adaptive testing (CAT). Usually, CAT is carried out for a single domain at a time. However, many HRQOL domains are highly correlated. Multidimensional CAT may utilize these correlations to improve measurement efficiency. We investigated the possible advantages and difficulties of multidimensional CAT. STUDY DESIGN AND SETTING: We evaluated multidimensional CAT of three scales from the EORTC QLQ-C30: the physical functioning, emotional functioning, and fatigue scales. Analyses utilised a database with 2958 European cancer patients. RESULTS: It was possible to obtain scores for the three domains with five to seven items administered using multidimensional CAT that were very close to the scores obtained using all 12 items and with no or little loss of measurement precision. CONCLUSION: The findings suggest that multidimensional CAT may significantly improve measurement precision and efficiency and encourage further research into multidimensional CAT. Particularly, the estimation of the model underlying the multidimensional CAT and the conceptual aspects need further investigations.

Adult↗

Item response theory was used to shorten EORTC QLQ-C30 scales for use in palliative care.

BACKGROUND AND OBJECTIVE: The goal was to develop a shortened version of the EORTC QLQ-C30 for use in palliative care. We wanted to keep as few items as possible in each scale while still being able to compare results with studies using the original scales. We examined the possibilities of shortening the physical functioning, cognitive functioning, fatigue, and nausea and vomiting scales. STUDY DESIGN AND SETTING: The shortening was based on 2,366 (physical functioning) and 10,815 (three other scales) observations, respectively. We used item response theory to construct scoring algorithms for predicting scores on the original scales. RESULTS: Evaluations showed that a three-item physical scale, a two-item fatigue scale, and a one-item nausea or vomiting scale predicted the scores on the original scales with excellent agreement and had measurement abilities similar to the original scales with no loss or only a little loss in power to detect group differences. The results of the cognitive functioning scale indicated problems when predicting scores from a shortened version. CONCLUSION: Given the favorable results for the physical functioning, fatigue, and nausea or vomiting scales we expect that the shortened versions of these scales will be included in the abbreviated version of the EORTC QLQ-C30 for palliative care.

Cognition↗

Estimating clinically significant differences in quality of life outcomes.

OBJECTIVE: This report extracts important considerations for determining and applying clinically significant differences in quality of life (QOL) measures from six published articles written by 30 international experts, in the field of QOL assessment and evaluation. The original six articles were presented at the Symposium on Clinical Significance of Quality of Life Measures in Cancer Patients at the Mayo Clinic in April 2002 and subsequently were published in Mayo Clinic Proceedings. PRINCIPAL FINDINGS: Specific examples and formulas are given for anchor-based methods, as well as distribution-based methods that correspond to known or relevant anchors to determine important differences in QOL measures. Important prerequisites for clinical significance associated with instrument selection, responsiveness, and the reporting of QOL trial results are provided. We also discuss estimating the number needed to treat (NNT) relative to clinically significant thresholds. Finally, we provide a rationale for applying group-derived standards to individual assessments. CONCLUSIONS: While no single method for determining clinical significance is unilaterally endorsed, the investigation and full reporting of multiple methods for establishing clinically significant change levels for a QOL measure, and greater direct involvement of clinicians in clinical significance studies are strongly encouraged.

Cross-Sectional Studies↗

Scoring based on item response theory did not alter the measurement ability of EORTC QLQ-C30 scales.

BACKGROUND AND OBJECTIVES: Most health-related quality-of-life questionnaires include multi-item scales. Scale scores are usually estimated as simple sums of the item scores. However, scoring procedures utilizing more information from the items might improve measurement abilities, and thereby reduce the needed sample sizes. We investigated whether item response theory (IRT)-based scoring improved the measurement abilities of the EORTC QLQ-C30 physical functioning, emotional functioning, and fatigue scales. METHODS: Using a database of 13,010 subjects we estimated the relative validities of IRT scoring compared to sum scoring of the scales. RESULTS: The mean relative validities were 1.04 (physical), 1.03 (emotional), and 0.97 (fatigue). None of these were significantly larger than 1. Thus, no gain in measurement abilities using IRT scoring was found for these scales. Possible explanations include that the items in the scales are not constructed for IRT scoring and that the scales are relatively short. CONCLUSION: IRT scoring of the three longest EORTC QLQ-C30 scales did not improve measurement abilities compared to the traditional sum scoring of the scales.

Adult↗

Use of differential item functioning analysis to assess the equivalence of translations of a questionnaire.

In cross-national comparisons based on questionnaires, accurate translations are necessary to obtain valid results. Differential item functioning (DIF) analysis can be used to test whether translations of items in multi-item scales are equivalent to the original. In data from 10,815 respondents representing 10 European languages we tested for DIF in the nine translations of the EORTC QLQ-C30 emotional function scale when compared to the original English version. We tested for DIF using two different methods in parallel, a contingency table method and logistic regression. The DIF results obtained with the two methods were similar. We found indications of DIF in seven of the nine translations. At least two of the DIF findings seem to reflect linguistic problems in the translation. 'Imperfect' translations can affect conclusions drawn from cross-national comparisons. Given that translations can never be identical to the original we discuss how findings of DIF can be interpreted and discuss the difference between linguistic DIF and DIF caused by confounding, cross-cultural differences, or DIF in other items in the scale. We conclude that testing for DIF is a useful way to validate questionnaire translations.

Adult↗

Assessing health status and quality-of-life instruments: attributes and review criteria.

The field of health status and quality of life (QoL) measurement - as a formal discipline with a cohesive theoretical framework, accepted methods, and diverse applications--has been evolving for the better part of 30 years. To identify health status and QoL instruments and review them against rigorous criteria as a precursor to creating an instrument library for later dissemination, the Medical Outcomes Trust in 1994 created an independently functioning Scientific Advisory Committee (SAC). In the mid-1990s, the SAC defined a set of attributes and criteria to carry out instrument assessments; 5 years later, it updated and revised these materials to take account of the expanding theories and technologies upon which such instruments were being developed. This paper offers the SAC's current conceptualization of eight key attributes of health status and QoL instruments (i.e., conceptual and measurement model; reliability; validity; responsiveness; interpretability; respondent and administrative burden; alternate forms; and cultural and language adaptations) and the criteria by which instruments would be reviewed on each of those attributes. These are suggested guidelines for the field to consider and debate; as measurement techniques become both more familiar and more sophisticated, we expect that experts will wish to update and refine these criteria accordingly.

Health Services Research↗

Assessing the clinical significance of single items relative to summated scores.

How many items are needed to measure an individual's quality of life (QOL)? This article describes the strengths and weaknesses of single items and summated scores (from multiple items) as QOL measures. We also address the use of single global measures vs multiple subindices as measures of QOL. The primary themes that recur throughout this article are the relationships between well-defined research objectives, the research setting, and the choice single item vs summated scores to measure QOL. The conceptual framework of the study, the conceptual fit with the measure, and the purpose of the assessment should all be considered when choosing a measure of QOL. No "gold standard" QOL measure can be recommended because no "one size fits all." Single items have the advantage of simplicity at the cost of detail. Multiple-item indices have the advantage of providing a complete profile of QOL component constructs at the cost of increased burden and of asking potentially irrelevant questions. The 2 types of indices are not mutually exclusive and can be used together in a single research study or in the clinical setting.

Chronic Disease↗