Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Reliability and validity of a new method of measuring posterior shoulder tightness.

STUDY DESIGN: Repeated measures of shoulder flexibility on nonimpaired subjects and intercollegiate baseball pitchers. OBJECTIVES: To present a new objective method of measuring posterior shoulder tightness, define the intratester and intertester reliability of the measurement, and assess its construct validity. BACKGROUND: Posterior shoulder tightness has been linked to anterior humeral head translation and decreased internal rotation. The reliability of an objective assessment of posterior shoulder tightness has yet to be established in the literature. METHODS AND MEASURES: Five repeat measurements were made using a standardized protocol on 21 nonimpaired subjects to determine intratester reliability. To determine intertester reliability, 2 testers (blinded to their measurement) each performed 1 measurement on 49 shoulders. Twenty-two intercollegiate baseball pitchers were measured once by 1 tester to evaluate the construct validity of the measurement. RESULTS: Measurements of posterior shoulder tightness performed by the same physical therapist had high reliability (ICC dominant = 0.92, nondominant = 0.95). Intertester measures revealed good reliability (ICC = 0.80). Pitchers had reduced dominant arm internal rotation and increased external rotation ROM compared to their other arm whereas nonimpaired subjects had less reduction in external rotation compared to the nondominant arm (pitchers: dominant, 109.7 degrees +/-2.4 degrees, nondominant, 98.9 degrees +/-1.6 degrees; nonimpaired subjects: dominant, 95.9 degrees +/-1.5 degrees, nondominant, 95.2 degrees +/-1.6 degrees) and internal rotation (pitchers: dominant, 50.0+/-2.0 degrees, nondominant, 69.5+/-2.5 degrees; nonimpaired subjects: dominant, 46.4+/-1.3 degrees, nondominant, 50.2+/-1.4 degrees). Pitchers had significantly greater posterior shoulder tightness compared to nonimpaired subjects (pitchers; dominant, 44.9+/-0.8 cm, nondominant, 37.5+/-0.7 cm, nonimpaired subjects; dominant, 32.9+/-0.8 cm, nondominant, 31.4+/-0.8 cm) and manifested a significant correlation between posterior shoulder tightness and internal rotation (r = -0.61) that was not evident in nonimpaired subjects. CONCLUSIONS: Measurement of posterior shoulder tightness using this technique is objective and reliable when done by the same physical therapist. Validity of this measurement is supported from the observation of athletes thought to have tight posterior structures. Further study is needed to determine the relationship of this measurement to patients diagnosed with shoulder impingement syndrome.

Adolescent↗

Reliability of the lateral pull test and tilt test to assess patellar alignment in subjects with symptomatic knees: student raters.

STUDY DESIGN: Test-retest reliability with blinded testers. OBJECTIVES: To determine the inter- and intra-rater reliability of the lateral pull test and patellar tilt test. BACKGROUND: If patellar malalignment can be detected by clinical examination, then condition-specific treatment interventions may be implemented in patients with patellofemoral pain syndrome. However, several clinical tests used to assess patellar mobility have recently been shown to have poor to fair reliability. Because the lateral pull test and the patellar tilt test are widely used clinically as diagnostic tests for patellofemoral pain syndrome but have not been previously tested for reliability, we examined these tests. METHODS AND MEASURES: Fifty-two subjects (age range, 21-48 years) provided 95 knees (19 symptomatic and 76 asymptomatic) for assessment of the lateral pull test. Two testers, blinded to the presence or absence of symptoms, independently performed the lateral pull test in random order. Fifty-five subjects (age range, 22-42 years) provided 99 knees (73 asymptomatic and 26 symptomatic) for assessment of the patellar tilt test. Three blinded testers independently performed the patellar tilt test in random order. All subjects were tested and retested within 3-5 days. A kappa (kappa) statistic was used to assess the agreement of findings within each tester and between testers. RESULTS: The kappa coefficients for intrarater reliability varied from 0.39 to 0.47 for the lateral pull test and from 0.44 to 0.50 for the patellar tilt test, while the coefficients for interrater reliability were 0.31 for the lateral pull test and varied from 0.20 to 0.35 for the tilt test. CONCLUSIONS: Repeated lateral pull tests and patellar tilt tests had fair intrarater and poor interrater reliability. Our results suggest that care must be taken in placing too much emphasis on these tests when making clinical decisions.

Adult↗

Reliability of 2 functional goniometric methods for measuring forearm pronation and supination active range of motion.

STUDY DESIGN: Test-retest reliability study. OBJECTIVES: To determine intra- and intertester reliability of the hand-held pencil (HHP) and the plumbline goniometer (PLG) methods for measuring active forearm pronation and supination motions in individuals with and without injuries. BACKGROUND: The distal forearm method has been considered the gold standard for measuring forearm pronation and supination motion. The HHP and PLG, however, are 2 more functional methods for measuring forearm motions, though limited information on the psychometric properties of these tests is currently available. METHODS AND MEASURES: Intra- and intertester reliability of the HHP and PLG methods were determined in 40 subjects of convenience (20 injured and 20 noninjured). Two testers performed 3 repeated measurements for each motion and method on all subjects. Intraclass correlation coefficients (ICC3,1 for intratester reliability, ICC2,3 for intertester reliability) and standard error of measurements (SEMs) were determined. RESULTS: The ICCs for the measurements of pronation and supination using the HHP and PLG methods were high (range, 0.86-0.98) for individuals with and without injuries, with the reliability for the PLG method being equal or slightly greater than the HHP method for the majority of pronation and supination measurements. Intratester ICCs were higher (SEMs were conversely lower) than intertester ICCs for nearly all measurements. The ICC values were generally the same or higher for individuals with injuries compared to individuals without injuries. CONCLUSIONS: The HHP and PLG are highly reliable methods for measuring functional forearm pronation and supination. Because plumbline goniometers are not commercially available and the instrumentation for the HHP method is readily accessible, clinicians should consider the latter as their method of choice for measuring functional forearm pronation and supination.

Adult↗

Reliability of the probability effect on event-related potentials during repeated testing.

The reliability of event-related potentials (ERPs) was studied in 10 healthy adults who were tested 8 times over 7-10 day intervals using a standard auditory oddball paradigm. The difference waveforms, obtained by subtracting the averaged waveforms for frequent trials from those obtained in rare trials, were designed to analyze the components of the ERPs, such as the P300, and to focus on the reliability of the probability effect on the ERPs. The between-session reliability (8 sessions) and the within-session reliability (order of blocks or of different visual procedures) were computed for the obtained difference waveforms. The between-session reliabilities, expressed as the intraclass correlations (r') for the P300 amplitude, area and latency, were 0.70, 0.61 and 0.65, respectively. The within-session reliability, presented as the Pearson correlation coefficients (r) for the three P300 measures were 0.43, 0.35 and 0.25 for different eyes. The values were 0.45, 0.39, 0.42 between the first and the second blocks (eyes-open) and 0.58, 0.47 and 0.29 (eyes-closed). These findings indicate that the P300 amplitude calculated from the difference waveforms may be the most stable marker for the between-session reliabilities. There were no significant differences in the P300 measures over the 8 sessions, suggesting that habituation may not occur with the difference waveform reflecting the probability effect on ERPs. The difference waveform may be useful in research on repetitive group ERPs.

Adult↗

Clinical reliability of shoulder function assessment in patients with rheumatoid arthritis.

A model for functional assessment and a dynamic test of the shoulder joint were designed and tested for normal variation and clinical inter- and intra-rater reliability. The functional assessments, which covered four common shoulder functions, were compared with assessments of pain, recordings of active motion range and the results of a Health Assessment Questionnaire, in eight patients with rheumatoid arthritis according to the ARA criteria. Intra-rater reliability was satisfactory for all four functions and inter-rater reliability was satisfactory for the hand-raising and hand-to-opposite-shoulder functions but less so for hand-behind-back and hand-to-neck. A second test-retest study in 15 patients, with a slight modification of one of the functional tests, confirmed the results and improved the reliability of the modified test. The reliability of the dynamic test and of the active motion range measurement was less satisfactory or not satisfactory. No significant correlation was found between shoulder functional assessment and the Fries index, but there were positive significant correlations between active motion range and shoulder functions. It is concluded that the method presented for evaluating shoulder functions has satisfactory reliability and in the first test-retest study was more reliable than conventional motion range measurement of the shoulder joint.

Adult↗

Reliability of corneal thickness and endothelial cell density measures.

PURPOSE: We evaluated the reliability and agreement between the Orbscan, an ultrasonic pachymeter (Humphrey 855), and the Konan SP 9000-LC in terms of central corneal thickness. The Konan was also used to study the reliability and agreement between endothelial cell density measures. METHODS: Twenty-five normal subjects were examined on two occasions (mean separation = 9 +/- 5 days) by a single examiner using all three instruments for central corneal pachymetry. The Konan Center Method and a manual counting method were performed by two examiners to determine endothelial cell density. Reliability and agreement were assessed by calculating the 95% limits of agreement (LoA) and intraclass correlation coefficients (ICC). RESULTS: For corneal pachymetry test-retest reliability, the 95% limits of agreement were -20 to +17 microm for the ultrasound, -27 to +22 microm for the Konan, and -13 to +13 microm for the Orbscan. There was fair-to-good agreement between the pachymeters (intraclass correlation coefficients range = 0.85 to 0.92). For endothelial cell density test-retest reliability, the 95% limits of agreement for the Konan Center Method was -498 to +530, and -482 to +333 cells/mm2 for examiners 1 and 2, respectively. The test-retest 95% limits of agreement for the manual overlaid grid method was -355 to +355, and -535 to +670 cells/mm2 for examiners 1 and 2, respectively. CONCLUSIONS: The reliability and agreement of the Orbscan and Konan corneal pachymeters was good, although the reliability of the Konan for estimating endothelial cell density was fair, at best.

Adult↗

Applicability of telemedicine for assessing patients with schizophrenia: acceptance and reliability.

BACKGROUND: Telemedicine holds promise for providing expert psychiatric consultation to underserved populations, but has not been quantitatively studied in schizophrenia or any other major mental disorder. This study was conducted to assess the reliability and acceptance of video-conferencing equipment in the assessment of patients with schizophrenia. METHOD: We assessed reliability of the Brief Psychiatric Rating Scale (BPRS), Scale for the Assessment of Positive Symptoms (SAPS), and Scale for the Assessment of Negative Symptoms (SANS) under three conditions: (1) in person, (2) by videoconferencing at low (128 kilobits per second [kbs]) bandwidth, (3) by videoconferencing at high (384 kbs) bandwidth. All 45 patients met DSM-IV criteria for schizophrenia. All patients and the two interviewers rated various aspects of the study interviews against previous live psychiatric interviews. RESULTS: Total scores on both the BPRS and SAPS were assessed equally reliably by the three media. Total score on the SANS was less reliably assessed at the low bandwidth, as were several specific negative symptoms of schizophrenia that depend heavily on nonverbal cues. Video interviews were well accepted by patients in both groups, although patients in the high bandwidth group were more likely to prefer the video interview to a live interview. CONCLUSION: Global severity of schizophrenia and overall severity of positive symptoms were reliably assessed by videoconferencing technology. Higher bandwidth resulted in more reliable assessment of negative symptoms and was preferred over low bandwidth, although patients' and raters' acceptance of video was good in both conditions. Videoconsultation appears to be a reliable method of assessing schizophrenic patients in remote locations who have limited access to expert consultation.

Adult↗

Reliability and validity of the self-assessment of occupational functioning.

OBJECTIVE: Two studies examined the reliability and validity of the Self-Assessment of Occupational Functioning (SAOF), a 23-item self-assessment of perceptions of strengths, and weaknesses relative to occupational functioning, grounded in the Model of Human Occupation. METHOD: The first study examined the test-retest reliability of the SAOF, and involved 37 college students without disabilities who completed the SAOF twice. The second study, which involved 39 young persons hospitalized with psychiatric disorders, examined internal consistency reliability of the SAOF, and examined correlations between SAOF scores and composite scores on the Self-Perception Profile, a widely used measure of perceived competence. In addition, data from both studies were combined to examine the ability of the SAOF to discriminate between the college students without disabilities and the young persons with psychiatric disorders. RESULTS: Kappa and intraclass correlation coefficients (ICCs) were used to examine test-retest reliability and Cronbach's alpha was used to examine internal consistency. Acceptable levels of test-retest (ICCs) and internal consistency (Cronbach's alpha) reliability were found for the subscale and total scores of the SAOF. However, test-retest reliability (kappa) was lower than desirable for many of the individual SAOF items. The young persons with psychiatric disorders had lower item, subscale, and total scores on the SAOF than did the college students without disabilities. In addition, a discriminant analysis predicting group membership (college students without disability vs. young persons with psychiatric disorder) correctly classified 76.6% of the participants based on the four subscale scores of the SAOF. CONCLUSION: The SAOF has the potential to be a reliable and valid clinical assessment; however, additional research is needed.

Adolescent↗

Reliability of the foot posture index and traditional measures of foot position.

Repeatable measures are essential for clinicians and researchers alike. Both need baseline measures that are reliable, as intervention effects cannot be accurately identified without consistent measures. The intrarater and interrater reliability of the new Foot Posture Index and current podiatric measures of foot position were assessed using a same-subject, repeated-measures study design across three age groups. The Foot Posture Index total score showed moderate reliability overall, demonstrating better reliability than most other current measures, although navicular height (normalized for foot length) was the single most reliable measure in adults. None of the tested measures exhibited adequate reliability in young children, and, with less-than-desirable reliability being demonstrated, most measures need to be interpreted accordingly when repeated measures are involved.

Adolescent↗

The Norwegian version of the Quality of Life Scale (QOLS-N). A validation and reliability study in patients suffering from psoriasis.

The aim of this study was to adapt, validate, and test for reliability the Quality of Life Scale in Norwegian (QOLS-N) for patients suffering from psoriasis. Two hundred and eighty-two patients with psoriasis were included in the study. Self-reported health was measured using the SF-36. Disease severity was also measured in 95 patients using the Psoriasis Area and Severity Index (PASI). The reliability of the QOLS-N was computed using the internal consistency reliability (Cronbach's alpha) and the test-retest reliability test. Face and content validity and construct discriminant ability of the QOLS-N were assessed. The results indicated that the QOLS-N has highly satisfactory rates of test-retest reliability (r = 0.83) and internal consistency reliability (alpha 0.86). As expected, the QOLS-N had a lower correlation with physical health (r = 0.24, p < 0.000) and self-reported symptoms (r = -0.20, p < 0.001), and a higher correlation with mental health (r = 0.52, p < 0.000). The correlation with disease severity was not significant (-0.06). The results reported in the present paper are in accordance with those derived in other validation studies. The QOLS-N seems to be a reliable and valid measure of global quality of life in patients suffering from psoriasis.

Activities of Daily Living↗

Reliability and sensitivity of diagnostic tests for primary Sjögren's syndrome.

OBJECTIVE: To investigate whether diagnostic tests for primary Sjögren's syndrome (pSS) are reproducible when repeated after one year (reliability). To evaluate whether the sensitivity of the diagnostic tests increases with repeated testing. METHODS: A structured interview investigating the subjective sensation of dry eyes and dry mouth, and the diagnostic tests Schirmer I, unstimulated whole saliva collection (UWSC), serological tests for antinuclear antibodies (ANA), for anti-Ro/SSA and anti-La/SSB antibodies as well as Waaler's test for rheumatoid factor, were performed twice with a one year interval in 66 patients with pSS. Reliability was given as the percentage of positive tests remaining positive at the second examination, while sensitivity was given as the percentage of patients with positive tests. RESULTS: Highest reliability was obtained for the sensation of dry mouth (98.2%) and sensation of dry eyes (96.4%), and anti-SSA/SSB antibodies (93.3%). Lowest reliability was obtained for rheumatoid factor at cutoff titer 1:32 (70.6%) and positive Schirmer I in one eye (77.4%). The reliability for ANA was 80% at cutoff titer 1:32, and increased to 93.3% at cutoff titer 1:128. UWSC had a reliability of 84.2%. The pooled sensitivity for all the tests increased significantly (p < 0.05) compared to the examination, which had the lowest sensitivity. CONCLUSION: The diagnostic tests for pSS are generally highly reliable when performed twice with a one year interval. The gain in sensitivity by repeating the tests is limited, being most marked for Schirmer I.

Adult↗

Radiological scoring methods in ankylosing spondylitis: reliability and sensitivity to change over one year.

Our aim was to compare reliability and sensitivity to change of different radiological scoring methods in ankylosing spondylitis (AS). Two trained observers scored 30 AS radiographs twice with an interval of 4 weeks. The same two observers scored 187 AS radiographs in pairs, at baseline and after one year followup, to measure change and agreement on change. The sacroiliac (SI) joints were scored in 5 grades by the New York method and the SASSS (Stoke Ankylosing Spondylitis Spine Score). Hips were graded 0-5 (according to Larsen). Cervical and lumbar spine were graded (0-4, Bath Ankylosing Spondylitis Radiological Index, BASRI), and scored in detail (0-72, SASSS). SASSS of the cervical and lumbar spine scored on the anterior sites of the vertebrae proved most reliable, with both intra and interobserver intraclass correlation coefficients (ICC) between 0.87 and 0.97. BASRI was only moderately reliable, with Cohen's kappa ranging between 0.50 and 0.82 for intra, and 0.38-0.64 for interobserver reliability. Similarly, SI joint scores (New York, SASSS) showed intraobserver kappa between 0.56 and 0.84, and interobserver reliability with kappa between 0.37 and 0.47. Larsen hip scores proved unreliable: moderate intraobserver kappa of 0.47-0.58 and low interobserver kappa of 0.29. After retraining, interobserver kappa did not improve (0.45 and 0.17). In retrospect, a one year period was too short to measure sensitivity to change. Observers agreed that no change occurred in up to 89% of cases. A measurable change of deterioration or improvement occurred rarely. We conclude that in AS, only the SASSS method for the spine and the BASRI reached good reliability. Other methods for spine, SI joints, and hips were moderately reliable at best. There was moderate to good agreement on no change between the observers. No method showed change over a period of one year in a considerable number of patients.

Arthrography↗

Good test--retest reliability for standard and advanced false-belief tasks across a wide range of abilities.

Although tests of young children's understanding of mind have had a remarkable impact upon developmental and clinical psychological research over the past 20 years, very little is known about their reliability. Indeed, the only existing study of test-retest reliability suggests unacceptably poor results for first-order false-belief tasks (Mayes, Klin, Tercyak, Cicchetti, & Cohen, 1996), although this may in part reflect the nonstandard (video-based) procedures adopted by these authors. The present study had four major aims. The first was to re-examine the reliability of false-belief tasks, using more standard (puppet and storybook) procedures. The second was to assess whether the test-retest reliability of false-belief task performance is equivalent for children of contrasting ability levels. The third aim was to explore whether adopting an aggregate approach improves the reliability with which children's early mental-state awareness can be measured. The fourth aim was to examine for the first time the test-retest reliability of children's performances on more advanced theory-of-mind tasks. Our results suggest that most standard and advanced false-belief tasks do in fact show good test-retest reliability and internal consistency, with very strong test-retest correlations between aggregate scores for children of all levels of ability.

Aptitude↗

Interexaminer reliability of transrectal ultrasound for estimating prostate volume.

PURPOSE: We investigated the interexaminer reliability of transrectal ultrasound measurement of total prostate and transition zone volume among 3 examiners with various levels of experience. MATERIALS AND METHODS: A total of 121 patients 39 to 82 years old (average plus or minus standard deviation 60.7 +/- 10.3) from a single urology clinic volunteered to participate. Patients with prostate cancer, previous prostate surgery or recent invasive prostatic examination were excluded from study. Each individual was examined independently by each of 3 examiners with various levels of experience, including an attending urologist, a PGY-2 resident in the second year of general surgery before urology training and a PGY-4 resident in the second year of urology training. Transrectal ultrasound was performed in each case by each examiner in pre-specified random order. RESULTS: Mean total prostate and transition zone volume was 35.9 +/- 27.2 and 15.6 +/- 18.8 ml., respectively. Interexaminer agreement or reliability of the ultrasound measurements was high for total prostate and transition zone volume (intraclass correlation 0.96, 95% confidence interval [CI] 0.95 to 0.97 and 0.93, 95% CI 0.90 to 0.95, respectively). For individual prostatic dimensions reliability estimates were 0.78 to 0.86, while for transition zone dimensions reliability was 0.85 to 0.90. Total prostate volume reliability was higher for prostate volume greater than 40 ml. versus smaller prostates (intraclass correlation 0.95, 95% CI 0.90 to 0.97 versus 0.77, 95% CI 0.67 to 0.84). Mean differences in transrectal ultrasound measurements by different examiners were highest for the resident with least experience. CONCLUSIONS: The reliability of transrectal ultrasound measured total prostate and transition zone volume is high for examiners with different levels of experience at this institution. Reliability in patients without prostate cancer appears to be better for larger volume prostates and for examiners with more experience.

Adult↗

[Mathematical model for study the reliability of organism function in extreme conditions of high altitude].

The health of a person and his capacity for work under conditions of high mountains in many respects is determined by the reliability of the function of physiological systems. Mathematical methods of both the reliability theory and mathematical simulation of basic functional systems are proposed to be used for investigating the reliability. It is shown that the most suitable reliability model for living systems is a chain model-a successive connection of links representing separate functional systems of organism. Besides, the weakest links determining the reliability of functioning of the whole organism under the extreme conditions of high mountains even for a healthy person are-respiration, blood circulation, thermoregulation and psychophysiological systems. Quantitative characteristics of the reliability of these systems are determined through the main indicators. The influence of non-sufficient contents of oxygen in respiration mixture, low atmospheric pressure, low temperatures of the environment are simulated by computer models of organism. An analyses of modeling data shows that moderate physical loading improves indicators of organisms adaptively to external conditions of high mountains and promotes the increasing of persons capacity for work and the reliability of his functioning.

Adaptation, Physiological↗

[Determination of reliability of psychometric tests in psychiatry using canonical correlation].

Test results (raw scores) are composed of an unknown true score and an error term. The error term can be estimated by means of test reliability which is defined by the ratio of true variance and obtained variance. Different estimates of reliability either based on single measurements (e.g. Cronbach's coefficient, split half reliability, Kuder Richardson method) or two measurements (test/retest, inter- or intrarater reliability) are available. Parallel test reliability depends on the correlation of two different tests obtained in one session. Canonical correlation methods allow an extension of the parallel test situation and split half technique. Two or more tests are performed in a sample of subjects. Randomized subsets are correlated using canonical correlation technique. The objective of this study is to estimate the homogeneity of test batteries. 94 patients (64 f, 30 m; age: 54-89 ys.) supposed to have dementia were tested using the clocktest (CT, scores: 1-5), MMSE (mini mental state examination) and SKT (Syndrom Kurztest). Four (i, j: 1-4) subsets of 20 patients each were determined by random and the following characteristics were calculated: Empiric correlation coefficient for n = 94 (R), canonical correlation coefficient (Rcan), eigenvalues (EV) and redundancy (Rnd) of corresponding variable sets. The results of canonical analysis showed canonical correlation coefficients in order of 0.8 to 0.9 (p-values < 0.001). This high internal consistency can be interpreted as a measure of reliability of the test batteries. In conclusion, canonical correlation based on parallel tests splitted in subsets gives information on consistency, i.e. reliability, of test batteries in addition to conventional correlation methods.

Adult↗

Reliability of accelerometry-based activity monitors: a generalizability study.

INTRODUCTION: Numerous studies have examined the validity of accelerometry-based activity monitors but few studies have systematically studied the reliability of different accelerometer units for assessing a standardized bout of physical activity. Improving understanding of error in these devices is an important research objective because they are increasingly being used in large surveillance studies and intervention trials that require the use of multiple units over time. METHODS: Four samples of college-aged participants were recruited to collect reliability data on four different accelerometer types (CSA/MTI, Biotrainer Pro, Tritrac-R3D, and Actical). The participants completed three trials of treadmill walking (3 mph) while wearing multiple units of a specific monitor type. For each trial, the participant completed a series of 5-min bouts of walking (one for each monitoring unit) with 1-min of standing rest between each bout. Generalizability (G) theory was used to quantify variance components associated with individual monitor units, trials, and subjects as well as interactions between these terms. RESULTS: The overall G coefficients range from 0.43 to 0.64 for the four monitor types. Corresponding intraclass correlation coefficients (ICC) ranged from 0.62 to 0.80. The CSA/MTI was found to have the least variability across monitor units and trials and the highest overall reliability. The Actical was found to have the poorest reliability. CONCLUSION: The CSA/MTI appeared to have acceptable reliability for most research applications (G values above 0.60 and ICC values above 0.80), but values with the other devices indicate some possible concerns with reliability. Additional work is needed to better understand factors contributing to variability in accelerometry data and to determine appropriate calibration protocols to improve reliability of these measures for different research applications.

Adult↗

Measuring impairment caused by leprosy: inter-tester reliability of the WHO disability grading system.

This paper reports the results of a study on the inter-tester reliability of the WHO disability grading system. The WHO disability grading system is the most frequently used method of grading impairment in leprosy patients. With this method, a grade of 0-2 is assigned to each of six individual body sites (both eyes, hands and feet). The maximum grade of any of these sites is used as an overall indicator of the person's impairment status. To date, the WHO disability grading scale has not been subjected to reliability testing. The reliability of the grading system depends on the operational definitions of the grades, the way the tester interprets these definitions and the skill of the tester. It is therefore important that the definitions are unambiguous and leave as little room as possible for multiple interpretations. Three testers with varying degrees of experience did paired assessments on a total of 150 leprosy patients in the Leprosy Mission Hospital Purulia, India, using recently published operational definitions of the WHO disability grades. For every patient, they determined the maximum grade (minimum 0, maximum 2), and calculated the impairment sum-score (EHF score), adding up the six grades for eyes, hands and feet (minimum 0, maximum 12). The weighted Kappa statistic (Kw) was used as the coefficient of inter-tester reliability. A kappa of 0 represents agreement no better than chance, and 1.0 complete (chance-corrected) agreement. Kw values of > or = 0.80 are considered very good and adequate for monitoring and research. Weighted Kappa analysis yielded a reliability coefficient of 0.89 (95%CI 0.84-0.94) for the maximum grade and a Kw of 0.97 (95%CI 0.96-0.98) for the EHF score. We concluded that, when using standard operational definitions, the WHO disability grading system can be used reliably in the hands of both experienced and inexperienced testers, provided adequate training has been given. Reliability should be evaluated further in a field setting, when used by primary health care workers. It is recommended that the 'WHO disability grading' be renamed 'WHO impairment grading', using the terminology as defined by the International Classification of Functioning, Disability and Health (ICF).

Disability Evaluation↗