Search PubMed⌕ Search

Biomedical subjects

Jakob B Bjorner

Publications and source records attributed to Jakob B Bjorner.

16 recordsLinked to original sources

Multidimensional computerized adaptive testing of the EORTC QLQ-C30: basic developments and evaluations.

OBJECTIVE: Self-report questionnaires are widely used to measure health-related quality of life (HRQOL). Ideally, such questionnaires should be adapted to the individual patient and at the same time scores should be directly comparable across patients. This may be achieved using computerized adaptive testing (CAT). Usually, CAT is carried out for a single domain at a time. However, many HRQOL domains are highly correlated. Multidimensional CAT may utilize these correlations to improve measurement efficiency. We investigated the possible advantages and difficulties of multidimensional CAT. STUDY DESIGN AND SETTING: We evaluated multidimensional CAT of three scales from the EORTC QLQ-C30: the physical functioning, emotional functioning, and fatigue scales. Analyses utilised a database with 2958 European cancer patients. RESULTS: It was possible to obtain scores for the three domains with five to seven items administered using multidimensional CAT that were very close to the scores obtained using all 12 items and with no or little loss of measurement precision. CONCLUSION: The findings suggest that multidimensional CAT may significantly improve measurement precision and efficiency and encourage further research into multidimensional CAT. Particularly, the estimation of the model underlying the multidimensional CAT and the conceptual aspects need further investigations.

Adult↗

An evaluation of a patient-reported outcomes found computerized adaptive testing was efficient in assessing osteoarthritis impact.

BACKGROUND AND OBJECTIVES: Evaluate a patient-reported outcomes questionnaire that uses computerized adaptive testing (CAT) to measure the impact of osteoarthritis (OA) on functioning and well-being. MATERIALS AND METHODS: OA patients completed 37 questions about the impact of OA on physical, social and role functioning, emotional well-being, and vitality. Questionnaire responses were calibrated and scored using item response theory, and two scores were estimated: a Total-OA score based on patients' responses to all 37 questions, and a simulated CAT-OA score where the computer selected and scored the five most informative questions for each patient. Agreement between Total-OA and CAT-OA scores was assessed using correlations. Discriminant validity of Total-OA and CAT-OA scores was assessed with analysis of variance. Criterion measures included OA pain and severity, patient global assessment, and missed work days. RESULTS: Simulated CAT-OA and Total-OA scores correlated highly (r = 0.96). Both Total-OA and simulated CAT-OA scores discriminated significantly between patients differing on the criterion measures. F-statistics across criterion measures ranged from 39.0 (P < .001) to 225.1 (P < .001) for the Total-OA score, and from 40.5 (P < .001) to 221.5 (P < .001) for the simulated CAT-OA score. CONCLUSIONS: CAT methods produce valid and precise estimates of the impact of OA on functioning and well-being with significant reduction in response burden.

Adaptation, Psychological↗

Burnout among employees in human service work: design and baseline findings of the PUMA study.

AIM: To present the theoretical framework, design, methods, and baseline findings of the first Danish study on determinants and consequences of burnout, and the impact of workplace interventions in human service work organizations. METHOD: A 5-year prospective intervention study comprising 2,391 employees from different organizations in the human service sector: social security offices, psychiatric prison, institutions for severely disabled, hospitals, and homecare services. Data were collected at baseline and at two follow-ups. The authors developed a new burnout tool (the Copenhagen Burnout Inventory) covering work-related, client-related, and personal burnout. The study includes potential determinants of burnout (e.g. the psychosocial work environment, social relations outside work, lifestyle factors, and personality aspects) and consequences of burnout (e.g. poor health, low job satisfaction, turnover, and absenteeism). Here, the focus is on the description of the study population at baseline, including associations of work burnout with psychosocial work environment scales and absence. RESULTS: Response rate at baseline was 80.1%. Midwives and homecare workers had high levels on both work- and client-related burnout. Prison officers had the highest level on client-related burnout. Supervisors and office assistants had low levels on both scales. Work burnout showed the highest correlations with job satisfaction (r = -0.51), quantitative demands (r = 0.48), role-conflicts (r = 0.44), and emotional demands (r = 0.42). Sickness absence was 13.9 vs 6.0 days among participants in the highest and lowest work burnout quartile, respectively. CONCLUSION: The findings indicate that study design and methods are adequate for the upcoming prospective analyses of aetiology and consequences of burnout and of the impact of workplace interventions.

Allied Health Personnel↗

The development of the EORTC QLQ-C15-PAL: a shortened questionnaire for cancer patients in palliative care.

This study aimed at developing a shortened version of the EORTC QLQ-C30, one of the most widely used health-related quality of life questionnaires in oncology, for palliative care research. The study included interviews with 41 patients and 66 health care professionals in palliative care to determine the appropriateness, relevance and importance of the various domains of the QLQ-C30. Item response theory methods were used to shorten scales. Patients and health care professionals rated pain, physical function, emotional function, fatigue, global health status/quality of life, nausea/vomiting, appetite, dyspnoea, constipation, and sleep as most important. Therefore, these scales/items were retained in the questionnaire. Four scales were shortened without reducing measurement precision. Important dimensions not covered by the questionnaire were identified. The resulting 15-item EORTC QLQ-C15-PAL is a 'core questionnaire' for palliative care. Depending on the research questions, it may be supplemented by additional items, modules or questionnaires.

Activities of Daily Living↗

Item response theory was used to shorten EORTC QLQ-C30 scales for use in palliative care.

BACKGROUND AND OBJECTIVE: The goal was to develop a shortened version of the EORTC QLQ-C30 for use in palliative care. We wanted to keep as few items as possible in each scale while still being able to compare results with studies using the original scales. We examined the possibilities of shortening the physical functioning, cognitive functioning, fatigue, and nausea and vomiting scales. STUDY DESIGN AND SETTING: The shortening was based on 2,366 (physical functioning) and 10,815 (three other scales) observations, respectively. We used item response theory to construct scoring algorithms for predicting scores on the original scales. RESULTS: Evaluations showed that a three-item physical scale, a two-item fatigue scale, and a one-item nausea or vomiting scale predicted the scores on the original scales with excellent agreement and had measurement abilities similar to the original scales with no loss or only a little loss in power to detect group differences. The results of the cognitive functioning scale indicated problems when predicting scores from a shortened version. CONCLUSION: Given the favorable results for the physical functioning, fatigue, and nausea or vomiting scales we expect that the shortened versions of these scales will be included in the abbreviated version of the EORTC QLQ-C30 for palliative care.

Cognition↗

Development of a computer-adaptive test for depression (D-CAT).

Depression is one of the most prevalent mental health problems and measuring depressive symptoms becomes increasingly important in science as well as medical practice. Computer Adaptive Tests (CAT) based on the Item Response Theory (IRT) promise to enhance measurement precision and reduce respondent's burden. Our aim was to develop a CAT application to measure depressive symptoms. Three thousand two hundred seventy psychosomatic patients answered an overall of 11 mental health questionnaires at the University Clinic in Berlin. Three independent reviewers rated 144 items out of these questionnaires as indicative of depressive symptoms. All items underwent six empirical steps to analyze unidimensionality, local independence and item discrimination. Finally 64 items could be used to calculate item parameters applying a Generalized Partial Credit Model (GPCM). CAT scores were estimated using an 'expected a posteriori' algorithm (EAP). Two simulation experiments showed that for theta values within the range of 2SD around the mean (98% of the cases), the latent trait can be estimated out of approximately six items with a predefined standard error of [Symbol: see text] 0.32 (reliability rho [Symbol: see text] 0.90). The CAT-scores correlated high with scores of all depression items (r = 0.95), with the Beck Depression Inventory (r = 0.79) and with a CES-D 8 item short form (r = 0.76). We conclude that the Depression-CAT measures depressive symptoms with high precision and low respondent burden.

Adolescent↗

Scoring based on item response theory did not alter the measurement ability of EORTC QLQ-C30 scales.

BACKGROUND AND OBJECTIVES: Most health-related quality-of-life questionnaires include multi-item scales. Scale scores are usually estimated as simple sums of the item scores. However, scoring procedures utilizing more information from the items might improve measurement abilities, and thereby reduce the needed sample sizes. We investigated whether item response theory (IRT)-based scoring improved the measurement abilities of the EORTC QLQ-C30 physical functioning, emotional functioning, and fatigue scales. METHODS: Using a database of 13,010 subjects we estimated the relative validities of IRT scoring compared to sum scoring of the scales. RESULTS: The mean relative validities were 1.04 (physical), 1.03 (emotional), and 0.97 (fatigue). None of these were significantly larger than 1. Thus, no gain in measurement abilities using IRT scoring was found for these scales. Possible explanations include that the items in the scales are not constructed for IRT scoring and that the scales are relatively short. CONCLUSION: IRT scoring of the three longest EORTC QLQ-C30 scales did not improve measurement abilities compared to the traditional sum scoring of the scales.

Adult↗

The Short-Form Headache Impact Test (HIT-6) was psychometrically equivalent in nine languages.

BACKGROUND AND OBJECTIVE: This study examined the psychometric properties and equivalence of the six-item Headache Impact Test (HIT-6) across 11 languages in 14 countries. METHODS: A multicenter, international cross-sectional study conducted in a primary care setting. Data obtained from 1,171 adults from 14 countries who consulted their primary care physician for headache completed the HIT-6 questionnaire and a headache survey were included in this analysis. Item-level statistics (e.g., range of response choices used by participants), item-scale statistics (e.g., item-total correlations), scale level statistics (e.g., internal consistency reliability), and tests of differential item functioning were conducted to examine the psychometric properties of all HIT-6 translations and their comparability across translations. RESULTS: Across languages, missing data were low, item-scale correlations were high, reliability was adequate, and item-level statistics were generally comparable. We found only minor differential item functioning, suggesting that the HIT-6 translations are equivalent to the U.S. English form. CONCLUSIONS: Psychometric analyses indicate that most HIT-6 translations (Canadian English, French, Greek, Hungarian, UK English, Hebrew, Portuguese, German, Spanish, and Dutch) are comparable to U.S. English. Improvements may be needed in the Finnish and Slovakian translations and the appropriateness of using the HIT-6 in South Africa should be explored further.

Cross-Sectional Studies↗

The potential synergy between cognitive models and modern psychometric models.

Analyses of cognitive aspects of survey methodology (CASM) and psychometric analysis are two methods that are able to complement each other. We use concrete examples to illustrate how psychometric analyses can test hypotheses from CASM. The psychometrics framework recognizes that survey responses are affected by other factors than the concept being assessed, for example by cognitive factors and processes. Such factors are subsumed under the concept of measurement error. Possible sources of measurement error can be tested, e.g. by randomized experiments. A standard way to reduce measurement error is to ask several questions about the same concept and combine the answers into a multi-item scale that is more precise than the individual items. Techniques like structural equation models use the item correlations to assess the magnitude of measurement error and to test the assumptions behind the multi-item scale, e.g. the effect of common response choices and item time frames. A central problem in modern psychometrics is how to model the mapping of the continuous latent variable onto the item response choice categories. This is achieved by threshold models (e.g. item response models and structural equation models for categorical data). These models can, for example, analyze the impact of mode of administration, test whether the items function in the same way for all people (measurement invariance/differential item functioning) and examine the consistency of responses from any single person. Such analyses provide new possibilities for combining psychometrics and cognitive methods.

Attitude to Health↗

Use of differential item functioning analysis to assess the equivalence of translations of a questionnaire.

In cross-national comparisons based on questionnaires, accurate translations are necessary to obtain valid results. Differential item functioning (DIF) analysis can be used to test whether translations of items in multi-item scales are equivalent to the original. In data from 10,815 respondents representing 10 European languages we tested for DIF in the nine translations of the EORTC QLQ-C30 emotional function scale when compared to the original English version. We tested for DIF using two different methods in parallel, a contingency table method and logistic regression. The DIF results obtained with the two methods were similar. We found indications of DIF in seven of the nine translations. At least two of the DIF findings seem to reflect linguistic problems in the translation. 'Imperfect' translations can affect conclusions drawn from cross-national comparisons. Given that translations can never be identical to the original we discuss how findings of DIF can be interpreted and discuss the difference between linguistic DIF and DIF caused by confounding, cross-cultural differences, or DIF in other items in the scale. We conclude that testing for DIF is a useful way to validate questionnaire translations.

Adult↗

Applications of computerized adaptive testing (CAT) to the assessment of headache impact.

OBJECTIVE: To evaluate the feasibility of computerized adaptive testing (CAT) and the reliability and validity of CAT-based estimates of headache impact scores in comparison with 'static' surveys. METHODS: Responses to the 54-item Headache Impact Test (HIT) were re-analyzed for recent headache sufferers (n = 1016) who completed telephone interviews during the National Survey of Headache Impact (NSHI). Item response theory (IRT) calibrations and the computerized dynamic health assessment (DYNHA) software were used to simulate CAT assessments by selecting the most informative items for each person and estimating impact scores according to pre-set precision standards (CAT-HIT). Results were compared with IRT estimates based on all items (total-HIT), computerized 6-item dynamic estimates (CAT-HIT-6), and a developmental version of a 'static' 6-item form (HIT-6-D). Analyses focused on: respondent burden (survey length and administration time), score distributions ('ceiling' and 'floor' effects), reliability and standard errors, and clinical validity (diagnosis, level of severity). A random sample (n = 245) was re-assessed to test responsiveness. A second study (n = 1103) compared actual CAT surveys and an improved 'static' HIT-6 among current headache sufferers sampled on the Internet. Respondents completed measures from the first study and the generic SF-8 Health Survey; some (n = 540) were re-tested on the Internet after 2 weeks. RESULTS: In the first study, simulated CAT-HIT and total-HIT scores were highly correlated (r = 0.92) without 'ceiling' or 'floor' effects and with a substantial reduction (90.8%) in respondent burden. Six of the 54 items accounted for the great majority of item administrations (3603/5028, 77.6%). CAT-HIT reliability estimates were very high (0.975-0.992) in the range where 95% of respondents scored, and relative validity (RV) coefficients were high for diagnosis (RV = 0.87) and severity (RV = 0.89); patient-level classifications were accurate 91.3% for a diagnosis of migraine. For all three criteria of change, CAT-HIT scores were more responsive than all other measures. In the second study, estimates of respondent burden, item usage, reliability and clinical validity were replicated. The test-retest reliability of CAT-HIT was 0.79 and alternate forms coefficients ranged from 0.85 to 0.91. All correlations with the generic SF-8 were negative. CONCLUSIONS: CAT-based administrations of headache impact items achieved very large reductions in respondent burden without compromising validity for purposes of patient screening or monitoring changes in headache impact over time. IRT models and CAT-based dynamic health assessments warrant testing among patients with other conditions.

Computer Systems↗

Using item response theory to calibrate the Headache Impact Test (HIT) to the metric of traditional headache scales.

BACKGROUND: Item response theory (IRT) scoring of health status questionnaires offers many advantages. However, to ensure 'backwards comparability' and to facilitate interpretations of results, we need the ability to express the IRT score in the metrics of the traditional scales. OBJECTIVES: To develop procedures to calibrate IRT-based scores on the Headache Impact Test (HIT) into the metrics of the traditional headache scales. To assess the degree to which the calibrated HIT scores agree with the observed traditional scores and lead to the same conclusions in group comparisons. METHODS: We used telephone interview data (n = 1016) and Internet data (n = 1103) from general population surveys of recent headache sufferers. Analyses were conducted in four steps: (1) develop IRT models for all items, (2) for each IRT score level, calculate the expected score on each of the traditional scales (calibration), (3) adjust this calibrated score for measurement error in the IRT score, (4) for each of the traditional scales, assess agreement between calibrated HIT scores and observed scores using intraclass correlation (ICC) and evaluate the agreement of mean scores and the relative validity (RV) in discriminating among groups differing in migraine diagnosis, headache severity, and change in impact over time. RESULTS: For the traditional categorical questionnaire items (the Migraine Specific Questionnaire (MSQ) and the Headache Disability Inventory (HDI)) the calibrated HIT agreed with the observed traditional scores: ICC's were between 0.80 and 0.94. In RV analyses the maximum mean difference between the observed and expected scores was 1.7 points on a 0-100 scale for comparisons at one point in time. Analyses of change over time and analyses calibrating scores from the fixed-form HIT-6 to the metric of other questionnaires were also satisfactory although less precise. Analysis of non-standard questionnaire items (e.g. On how many days in the past 3 months did you have a headache, from the HIMQ and the MIDAS) required special IRT models. Agreement was less good: ICC's were between 0.56 and 0.61 and the maximum mean differences were 2.9 (on a 0-270 scale) and 3.8 (on a 0-450 scale) in RV analyses at one point in time. The ability of the calibrated scale scores to discriminate between groups was at least as good as the ability of the observed sum scales and often remarkably better. CONCLUSION: The theoretical advantage of IRT models in scale calibration is supported by our results. This approach to achieving comparability of new and widely-used scales and accelerating the accumulation of interpretation guidelines based on previous work warrant testing for measures of other generic and disease-specific concepts.

Adolescent↗

Calibration of an item pool for assessing the burden of headaches: an application of item response theory to the headache impact test (HIT).

BACKGROUND: Measurement of headache impact is important in clinical trials, case detection, and the clinical monitoring of patients. Computerized adaptive testing (CAT) of headache impact has potential advantages over traditional fixed-length tests in terms of precision, relevance, real-time quality control and flexibility. OBJECTIVE: To develop an item pool that can be used for a computerized adaptive test of headache impact. METHODS: We analyzed responses to four well-known tests of headache impact from a population-based sample of recent headache sufferers (n = 1016). We used confirmatory factor analysis for categorical data and analyses based on item response theory (IRT). RESULTS: In factor analyses, we found very high correlations between the factors hypothesized by the original test constructers, both within and between the original questionnaires. These results suggest that a single score of headache impact is sufficient. We established a pool of 47 items which fitted the generalized partial credit IRT model. By simulating a computerized adaptive health test we showed that an adaptive test of only five items had a very high concordance with the score based on all items and that different worst-case item selection scenarios did not lead to bias. CONCLUSION: We have established a headache impact item pool that can be used in CAT of headache impact.

Adolescent↗

The feasibility of applying item response theory to measures of migraine impact: a re-analysis of three clinical studies.

BACKGROUND: Item response theory (IRT) is a powerful framework for analyzing multiitem scales and is central to the implementation of computerized adaptive testing. OBJECTIVES: To explain the use of IRT to examine measurement properties and to apply IRT to a questionnaire for measuring migraine impact--the Migraine Specific Questionnaire (MSQ). METHODS: Data from three clinical studies that employed the MSQ-version 1 were analyzed by confirmatory factor analysis for categorical data and by IRT modeling. RESULTS: Confirmatory factor analyses showed very high correlations between the factors hypothesized by the original test constructions. Further, high item loadings on one common factor suggest that migraine impact may be adequately assessed by only one score. IRT analyses of the MSQ were feasible and provided several suggestions as to how to improve the items and in particular the response choices. Out of 15 items, 13 showed adequate fit to the IRT model. In general, IRT scores were strongly associated with the scores proposed by the original test developers and with the total item sum score. Analysis of response consistency showed that more than 90% of the patients answered consistently according to a unidimensional IRT model. For the remaining patients, scores on the dimension of emotional function were less strongly related to the overall IRT scores that mainly reflected role limitations. Such response patterns can be detected easily using response consistency indices. Analysis of test precision across score levels revealed that the MSQ was most precise at one standard deviation worse than the mean impact level for migraine patients that are not in treatment. Thus, gains in test precision can be achieved by developing items aimed at less severe levels of migraine impact. CONCLUSIONS: IRT proved useful for analyzing the MSQ. The approach warrants further testing in a more comprehensive item pool for headache impact that would enable computerized adaptive testing.

Adolescent↗

Trends in the Danish work environment in 1990-2000 and their associations with labor-force changes.

OBJECTIVES: The aims of this study were (i) to describe the trends in the work environment in 1990-2000 among employees in Denmark and (ii) to establish whether these trends were attributable to labor-force changes. METHODS: The split-panel design of the Danish Work Environment Cohort Study includes interviews with three cross-sections of 6067, 5454, and 5404 employees aged 18-59 years, each representative of the total Danish labor force in 1990, 1995 and 2000. In the cross-sections, the participation rate decreased over the period (90% in 1990, 80% in 1995, 76% in 2000). The relative differences in participation due to gender, age, and region did not change noticeably. RESULTS: Jobs with decreasing prevalence were clerks, cleaners, textile workers, and military personnel. Jobs with increasing prevalence were academics, computer professionals, and managers. Intense computer use, long workhours, and noise exposure increased. Job insecurity, part-time work, kneeling work posture, low job control, and skin contact with cleaning agents decreased. Labor-force changes fully explained the decline in low job control and skin contact to cleaning agents and half of the increase in long workhours, but not the other work environment changes. CONCLUSIONS: The work environment of Danish employees improved from 1990 to 2000, except for increases in long workhours and noise exposure. From a specific work environment intervention point of view, the development has been less encouraging because declines in low job control, as well as skin contact to cleaning agents, were explained by labor-force changes.

Adolescent↗