Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Indirect ultrasound measurement of humeral torsion in adolescent baseball players and non-athletic adults: reliability and significance.

Accurate clinical interpretation of shoulder rotation range requires knowledge of the contribution of humeral torsion. This study compares the reliability of indirectly measuring humeral torsion using ultrasound visualisation with direct palpation and compares the degree of humeral torsion in a non-throwing adult population with a population of elite adolescent baseballers. The reliability of a novel method of indirectly measuring humeral torsion using palpation and ultrasound was established prior to using the ultrasound method to determine the amount of humeral torsion in both humeri of a group of 16 non-throwing subjects and 36 elite adolescent baseball players. Excellent inter-tester reliability was found for the ultrasound method of measuring humeral torsion in each arm (ICC 2,1: 0.98 and 0.94) but poor reliability for the direct palpation method (ICC 2,1: 0.51 and 0.49). Using the ultrasound method, side-to-side differences in humeral torsion ranged from 0 to 13 degrees for the non-throwing group and 0 to 29 degrees for the baseball players. This side-to-side difference was significant in the baseball players (p<0.001) but not significant in the non-throwing group (p=0.43). Whilst side-to-side differences in humeral torsion were noted between all subjects irrespective of arm dominance, only throwers demonstrated consistently greater humeral retrotorsion in their throwing arm. A reliable clinical tool for estimating humeral torsion, such as that employed in this study, will allow for a more valid prescription of interventions for the treatment of shoulder dysfunction.

Adolescent↗

Reliability of clinician-based (GRBAS and CAPE-V) and patient-based (V-RQOL and IPVI) documentation of voice disorders.

This study examined the reliability of two methods for documenting voice quality by clinicians and compared the methods for documenting patients' perceptions of voice quality. It involved a prospective reliability study and a retrospective chart review. Reliability of two clinician-based voice assessment protocols-Grade, Roughness, Breathiness, Asthenia, Strain (GRBAS) and Consensus Auditory Perceptual Evaluation-Voice (CAPE-V)-was evaluated. These two protocols were then compared after use in voice assessments of 42 males and 61 females performed by a certified speech-language pathologist specializing in the assessment of voice disorders. In addition, two patient-based scales (Voice Related Quality of Life, or V-RQOL, and Iowa Patient's Voice Index, or IPVI) obtained from the same patients were compared with each other and with the clinician-based scales. Reliability of clinicians' ratings of overall severity of dysphonia using GRBAS and CAPE-V scales was very good (r>0.80). Agreement between V-RQOL Total scores and IPVI ratings of the patient's perceptions of impact of dysphonia was less strong (Spearman's r=-0.76). There was relatively weak agreement between patient-based and clinician-based scales. Clinician's perceptions of dysphonia appeared to be reliable and unaffected by rating tool, as indicated by the high level of agreement between the two rating systems when they were used together. The CAPE-V system appeared to be more sensitive to small differences within and among patients than the GRBAS system. The V-RQOL and IPVI approaches to documenting patient's perceptions of dysphonia agreed less well possibly due to differences in patient dependence on voice and on interpretation of the rating tool items. The differences between clinician-based and patient-based data support the conclusion that clinicians and patients experience and consider dysphonia very differently.

Adolescent↗

Test-retest reliability of cervicocephalic kinesthetic sensibility in three cardinal planes.

The test-retest reliability of both the head-to-neutral head position (NHP) and head-to-target repositioning tests in three cardinal planes has been examined in this study. Twenty young adults underwent both head repositioning tests and retests with 10 min rest intervals. Root mean square error (RMSE, total error), constant error (CE, directional bias), variable error (VE, variability), and standard error of measurement (SEM) were calculated from the position data recorded by an ultrasound-based motion analysis system. Intra-class correlation coefficients (ICC) were used to examine reliability. The results showed fair to excellent reliability of RMSE during head-to-NHP (ICC=0.45-0.80) and head-to-target tests (ICC=0.42-0.90), except during the head-to-NHP test (ICC=0.29) from a head extended position. Low reliability of VE associated with the neck motion toward left side bending indicated a direction-dependent effect. The SEM of RMSE (0.7-2.6 degrees), CE (0.3-4.0 degrees) and VE (0.4-1.5 degrees) indicated an acceptable range of error. The present study indicated acceptable and reliable RMSE measurements with a motion analysis system in healthy young adults. Furthermore, examining the CE and VE could contribute to the interpretation of whether the subject performed the reposition tests with directional bias and repositioning variability, respectively.

Adult↗

Measurement of participation in myotonic dystrophy: reliability of the LIFE-H.

Person-perceived social participation is an important aspect to measure in rehabilitation in neuromuscular disorders. The objective was to document the test-retest and inter-rater reliability of the Assessment of Life Habits (LIFE-H), in a sample of individuals with myotonic dystrophy (DM1). Twenty-eight participants with myotonic dystrophy aged between 29 and 75 years (mean: 52.7) were recruited. The LIFE-H questionnaire was administered at three different occasions. The LIFE-H demonstrates high test-retest and inter-rater reliability (ICCs: 0.80-0.92) for the total score and subscores (daily activities and social roles). Moderate to high test-retest reliability (ICCs: 0.76-0.92) was found for most categories (8/10) of the LIFE-H. Similar results were obtained for inter-rater reliability (ICCs: 0.68-0.93). Moderate to high agreement using the Bland and Altman method was obtained for most categories. The LIFE-H, a measure of person-perceived social participation, demonstrates adequate test-retest and inter-rater reliability for clinical and research purposes in myotonic dystrophy.

Activities of Daily Living↗

Eyelid position measurement in Graves' ophthalmopathy: reliability of a photographic technique and comparison with a clinical technique.

PURPOSE: To evaluate intraobserver reliability and interobserver reliability of a computer-based digital image measurement of eyelid position in Graves' ophthalmopathy and to compare digital image measurement with clinical measurement. DESIGN: Cross-sectional study. PARTICIPANTS: Eighty-four eyes of 42 patients with mild to moderate bilateral Graves' ophthalmopathy. METHODS: Digital images were created from 35-mm color slides of both eyes of participants and projected onto a 15-inch flat-screen computer monitor. Three observers (2 oculoplastic surgeons and 1 ophthalmology resident) independently recorded eyelid fissure height, margin-reflex distance, and inferior scleral show for each eye. MAIN OUTCOME MEASURES: Intraobserver reliability and interobserver reliability of eyelid parameter measurements, as described by the intraclass correlation coefficient (ICC) and Bland-Altman plots. Agreement between digital image measurements of the investigators and clinical measurements taken on the same day as the photographs also was assessed. RESULTS: Excellent intraobserver agreement was found for the measurement of all eyelid parameters for all 3 investigators (ICC range, 0.93-0.99). Interobserver agreement for all eyelid parameters was also excellent for all investigators (ICC, 0.86-0.97). Agreement between the photographic and clinical measurements for eyelid parameters was fair to moderate (ICC range, 0.38-0.62). CONCLUSION: Measurement of several eyelid parameters in Graves' ophthalmopathy patients from computer-based digital images is reliable. Associations between photographic and clinical measurements for all parameters are weaker. Relative to clinical measurements, the photographic technique offers the advantages of potential for masking and ease of transmission that might be useful in clinical trials.

Cross-Sectional Studies↗

The ophthalmic clinical evaluation exercise: reliability determination.

PURPOSE: Reliable and valid tools must be developed to assess the core residency competencies identified by the Accreditation Council for Graduate Medical Education. The Ophthalmic Clinical Evaluation Exercise (OCEX) is a tool designed to assess the ophthalmology resident's competence in patient care. The OCEX has been shown to have face and content validity. This study will determine the degree to which the OCEX is reliable and has construct validity. PARTICIPANTS: Ninety-four academic ophthalmology teaching faculty from ophthalmology residency programs across the country. METHODS: Participants reviewed a video compact disc of the same resident and new patient encounter and then completed the OCEX. A scoring rubric was provided. RESULTS: Results indicate that the OCEX is a reliable tool for faculty to use to assess residency competency. The coefficient alpha statistic (a measure of reliability/internal consistency) for the OCEX as a whole was 0.81. The alpha statistics for 3 of 4 subscales that comprise the OCEX (i.e., interviewing skills = 0.65, interpersonal skills/professionalism = 0.73, case presentation = 0.70) were lower than the OCEX as a whole, but were acceptable for new scales. However, the alpha for the examination subscale (i.e., 0.27) was extremely low. Interrater reliability assessment shows that of 33 individual OCEX items, 31 (94%) had at least 85% of the raters rating the student in 1 of 2 consecutive rating categories. CONCLUSIONS: The OCEX shows both reliability and validity and, therefore, meets the Accreditation Council for Graduate Medical Education criteria for an acceptable assessment tool.

Accreditation↗

An investigation to examine the inter-tester and intra-tester reliability of the Rolimeter knee tester, and its sensitivity in identifying knee joint laxity.

PURPOSE: The purpose of this study is to evaluate the Rolimeter knee tester (Aircast, Europe) as reliable and clinically sensitive tool for identifying and quantifying knee joint laxity utilising a sample of both known ACLD and normal knees. METHODS: Thirty matched subjects (15 known ACLD and 15 normal subjects) were tested for knee joint laxity using the Rolimeter. Each subject was measured at both 90 degrees and 30 degrees of knee flexion, by each of the six investigators. This was then repeated again by all six investigators so that inter-tester and intra-tester reliability could be examined. RESULTS: Results showed that there was good reliability between testers, and intra-tester reliability was good for both left and right knees in both 90 degrees and 30 degrees of flexion. Results also demonstrated a high level of sensitivity for determining knee joint laxity in ACLD compared to normal knees. CONCLUSION: The Rolimeter knee tester is a reliable device for quantifying knee joint laxity, and is sensitive enough to identify anterior cruciate ligament deficiency.

Adult↗

Reliability and factorial validity of the Observer Alexithymia Scale-Chinese translation.

The purpose of the present study was to develop a Chinese translation of the Observer Alexithymia Scale (OAS-C) and evaluate its reliability and factorial validity. The original English-version of the Observer Alexithymia Scale (OAS) was translated into Chinese and given to 468 Chinese undergraduate students. Students were asked to rate a person (other than themselves) whom they knew well (e.g., a parent, sibling, another relative, or friend). We evaluated internal consistency, test-retest and inter-rater reliability, and factorial validity. Average OAS-C scores were slightly higher than, but comparable to, OAS scores in the normative samples (English-speaking/nonclinical). The OAS-C showed adequate internal consistency (Cronbach's alpha coefficient was 0.84, and the mean inter-item correlation coefficient was 0.14), good stability (test-retest reliability with a 2-week interval was 0.90), and inter-rater reliability (intra-class correlation coefficient was 0.78). Moreover, the OAS five-factor model (Distant, Uninsightful, Somatizing, Humorless, and Rigid) was confirmed: incremental fit index=0.905, comparative fit index=0.904, and root mean square error of approximation=0.086; each represented an adequate model fit. The OAS-C appears to be a reliable and valid observer-rated alexithymia measure. We recommend that researchers collect both self- and observer-rated alexithymia data and, when possible, obtain observer reports from more than one person.

Adolescent↗

Experiences of discrimination: validity and reliability of a self-report measure for population health research on racism and health.

Population health research on racial discrimination is hampered by a paucity of psychometrically validated instruments that can be feasibly used in large-scale studies. We therefore sought to investigate the validity and reliability of a short self-report instrument, the "Experiences of Discrimination" (EOD) measure, based on a prior instrument used in the Coronary Artery Risk Development in Young Adults (CARDIA) study. Study participants were drawn from a cohort of working class adults, age 25-64, based in the Greater Boston area, Massachusetts (USA). The main study analytic sample included 159 black, 249 Latino, and 208 white participants; the validation study included 98 African American and 110 Latino participants who completed a re-test survey two to four weeks after the initial survey. The main and validation survey instruments included the EOD and several single-item discrimination questions; the validation survey also included the Williams Major and Everyday discrimination measures. Key findings indicated the EOD can be validly and reliably employed. Scale reliability was high, as demonstrated by confirmatory factor analysis, Cronbach's alpha (0.74 or greater), and test-re-test reliability coefficients (0.70). Structural equation modeling demonstrated the EOD had the highest correlation (r=0.79) with an underlying discrimination construct compared to other self-report discrimination measures employed. It was significantly associated with psychological distress and tended to be associated with cigarette smoking among blacks and Latinos, and it was not associated with social desirability in either group. By contrast, single-item measures were notably less reliable and had low correlations with the multi-item measures. These results underscore the need for using validated, multi-item measures of experiences of racial discrimination and suggest the EOD may be one such measure that can be validly employed with working class African Americans and Latino Americans.

Adult↗

The intra- and inter-assessor reliability of measurement of functional outcome by lameness scoring in horses.

The objective of this study was to assess the reliability of lameness scoring in horses. One veterinary surgeon examined nineteen lame horses on four occasions. Gait was recorded by camcorder, and scored from 0 to 10 ranging from sound to non-weight bearing lameness. A global score of overall change in lameness during the study was also determined for each horse. To measure intra-assessor reliability of the scoring systems, one veterinary surgeon scored videotapes of the horses' gaits on two occasions. To measure inter-assessor reliability, three veterinary surgeons viewed the videotapes, assigning individual lameness scores plus global scores to each horse. Reliability of individual lameness scoring was good intra-assessor, but only just within our acceptable limit inter-assessor. However, global scoring of change in lameness throughout the study was found to be reliable overall. Since clinician scoring is commonly used to assess lameness in horses, this is an important finding, fundamental to future clinical studies.

Animals↗

Comparison of extended field of view and dual image ultrasound techniques: accuracy and reliability of distance measurements in phantom study.

This study was undertaken to investigate and compare the accuracy and reliability of dual image and extended field-of-view (EFOV) ultrasound (US) techniques in distance measurements using acoustic phantoms. Ten tissue phantoms were constructed and were scanned twice with an interval of 3 days by two operators. Measurements of various known distance (ranging from 4.6 to 7.2 cm) in the phantoms were made with dual image and EFOV US. Results showed that both dual image and EFOV US have a high accuracy and reliability in distance measurements, with the EFOV US (r = 0.997 to 0.998, reproducibility = 99.8%, repeatability = 98.2 to 99.8%) being slightly more accurate and reliable than dual image US (r = 0.948 to 0.981, reproducibility = 94.6%, repeatability = 89.6 to 97.9%). EFOV US has a higher accuracy and reliability than dual image US in distance measurements. However, the dual image US is a useful alternative with a high accuracy and reliability when EFOV US is not available.

Data Display↗

Test-retest reliability of self-reported mammography in women veterans.

BACKGROUND: Mammography self-report is used to monitor screening and evaluate intervention trends; however, few studies have examined reliability. METHODS: Reliability of self-reported lifetime number of mammograms, most recent mammogram date, and predictors of reliability were assessed using data from Project H.O.M.E. The study population was 2,494 women 52 years and over, listed in the U.S. National Registry of Women Veterans, with no history of breast cancer, who completed both baseline (2000-2002) and year 1 (2002-2003) surveys. RESULTS: Reliability of lifetime number of mammograms was 60.9% for exact consistency and 79.9% for consistency within one mammogram. Thirty-five percent was exactly consistent in reporting mammogram date; 55.6% was consistent within 3 months. Completing both surveys by mail and reporting fewer lifetime mammograms at baseline were positively associated with consistency of reporting lifetime number. White race/ethnicity, having a Bachelor's degree, reporting a health care provider's recommendation for a mammogram, having a screening mammogram, completing both surveys by mail, and being in the maintenance or action stages of change were associated with consistency in reporting date. CONCLUSIONS: Reliability varies with the measure of self-reported mammography. Likewise, predictors show different patterns of association with different definitions. Our findings call attention to the need for explicit definitions and measures of mammography use.

Adult↗

Reliability and validity of the Child and Adolescent Trial for Cardiovascular Health (CATCH) Food Checklist: a self-report instrument to measure fat and sodium intake by middle school students.

OBJECTIVE: To develop a scoring algorithm and evaluate the reliability and validity of scores from the Child and Adolescent Trial for Cardiovascular Health (CATCH) Food Checklist (CFC) as measures of total fat, saturated fat, and sodium intake in middle school students. DESIGN: Randomized, controlled trial in which participants were assigned to 1 of 3 study protocols that varied the order of CFC and 24-hour dietary recall administration. Criterion outcomes were percent energy from total fat, percent energy from saturated fat, and sodium intake in milligrams. SUBJECTS/SETTING: A multiethnic sample (33% ethnic and racial minorities) of 365 seventh-grade students from 8 schools in 4 states. STATISTICAL ANALYSES: Multivariable regression models were used to calibrate the effects of individual food checklist items; bootstrap estimates were used for cross-validation; and kappa statistics, Pearson correlations, t tests, and effect sizes were employed to assess reliability and validity. RESULTS: The median same-day test-retest reliability kappa for the 40 individual CFC food items was 0.85. With respect to item validity, the median kappa statistic comparing student choices to those identified by staff dietitians was 0.54. Test-retest reliability coefficients ranged from 0.84 to 0.89 for CFC total nutrient scores. Correlations between CFC scores and 24-hour recall values were 0.36 for total fat, 0.36 for saturated fat, and 0.34 for sodium; CFC scores were consistent with hypothesized gender differences in nutrient intake. APPLICATIONS/CONCLUSIONS: The CFC is a reliable and valid tool for measuring fat, saturated fat, and sodium intake in middle school students. Its brevity and ease of administration make the CFC a cost-effective way to measure middle school students' previous day's intake of selected nutrients in school surveys and intervention studies.

Adolescent↗

Interobserver reliability of digital and endovaginal ultrasonographic cervical length measurements.

OBJECTIVE: Our purpose was to prospectively evaluate the interobserver reliability of digital and endovaginal ultrasonographic cervical length measurements. STUDY DESIGN: Forty-three women were recruited from our antepartum clinic to participate in this study. Two independent and blinded digital cervical examinations were performed by the first author and a second examiner. Instructions were given to estimate the cervical length in millimeters. After micturition endovaginal ultrasonographic cervical length measurements were performed by two independent, blinded registered diagnostic medical sonographers. Cervical lengths were compared with the Student t test and Pearson's correlation coefficient. A kappa statistic was calculated for interobserver reliability at three levels of agreement +/- 1 mm, +/- 4 mm, and +/- 10 mm. Data are expressed as means +/- SD. RESULTS: Digital cervical lengths were not different between the two examiners (18.7 +/- 4.8 mm, 20.5 +/- 6.2 mm) nor between the two ultrasonographic measurements (38.6 +/- 6.1 mm, 39.2 +/- 5.4 mm). The digital cervical lengths agreed (+/- 1 mm) 35% of the time (R2 0.10, p = 0.02). The endovaginal ultrasonographic measurements agreed (+/- 1 mm) 74% of the time with a stronger correlation (R2 0.53, p = 0.0001). The kappa statistic for interobserver variability was marginal for both digital and endovaginal cervical length measurements when agreement was defined as +/- 1 mm. Endovaginal ultrasonography was significantly more reliable than digital examination when agreement between examiners was defined as either +/- 4 mm or +/- 10 mm. CONCLUSION: Although both digital and endovaginal ultrasonographic cervical length measurements show correlation between examiners, endovaginal ultrasonography is significantly more reliable when agreement is defined as > or = +/- 4 mm. Serial cervical length measurements to predict preterm labor will be enhanced by the interobserver reliability of endovaginal ultrasonography.

Cervix Uteri↗

Reliability of retinal photography in the assessment of retinal microvascular characteristics: the Atherosclerosis Risk in Communities Study.

PURPOSE: Retinal microvascular characteristics, as graded from retinal photography, have been shown to predict stroke. We evaluated the reliability of retinal photographic grading in the Atherosclerosis Risk in Communities Study. DESIGN: Cohort study. METHODS: Retinal photographs were taken of all subjects who attended the third Atherosclerosis Risk in Communities Study examination (1993 to 1995). These were graded using standardized protocols. Focal retinal characteristics were graded using a "light box" system. Generalized retinal arteriolar narrowing was quantified from computer-assisted measurements of digitized photographs. Two sub-studies were conducted to investigate the reliability of these grading methods. In the Individual Variability Study, selected subjects (n = 206) had two retinal photographs taken on one day, and a further one or two photographs taken 3 weeks later. In the Grader Variability Study, a stratified random sample of photographs had repeat retinal grading (n = 495 photographs for light box grading; n = 276 photographs for computer-assisted grading). RESULTS: Reliability of the computer-assisted quantification of generalized retinal arteriolar narrowing was high in both studies (reliability coefficients 0.64 to 0.69 for Individual Variability Study, and 0.79 to 0.83 for the Grader Variability Study). There was more variability for focal abnormalities graded using the light box system. Variability for Individual Variability Study (same individuals, repeat photographs) tended to be greater than for the Grader Variability Study (same photographs, repeat gradings). CONCLUSION: Retinal microvascular characteristics, especially computer-assisted quantification of generalized retinal arteriolar narrowing, can be ascertained reliably by standardized photographic grading methods, supporting the validity of their associations with cardiovascular disease. However, these characteristics appear to vary somewhat between eyes and over time in a single individual.

Aged↗

Reliability of the electronic early treatment diabetic retinopathy study testing protocol in children 7 to <13 years old.

PURPOSE: To assess the test-retest reliability of the electronic Early Treatment Diabetic Retinopathy Study (E-ETDRS) visual acuity algorithm using the computerized Electronic Visual Acuity (EVA) tester in children 7 to <13 years old. DESIGN: Test-retest reliability study. METHODS: This multicenter study involved 245 subjects at four clinical sites. As the main outcome measure, visual acuity was measured twice using the E-ETDRS testing protocol on the EVA system, which uses a programmed handheld device to communicate with a personal computer and a 17-inch monitor at a 3-m test distance. RESULTS: Test-retest reliability was high (r =.94 for right eyes and 0.96 for left eyes) and for both right and left eyes, 89% of retest scores were within 0.1 logarithm of the minimal angle of resolution (logMAR) (five letters) of the initial test score and 99% of retests were within 0.2 logMAR (10 letters). Reliability was high across the age range of 7 to <13 years. Based on 95% confidence level estimates, a change in visual acuity of 0.2 logMAR (10 letters) from a previous acuity measure is unlikely to result from measurement variability. CONCLUSIONS: The E-ETDRS protocol using the EVA has high test-retest reliability in children 7 to <13 years of age. Potential advantages include better standardization across multiple sites, the ability to directly capture data electronically with an automatic acuity score calculation, the reduction of potential bias by limiting the tester's role, and the requirement of only a single testing distance for measurements from 20/800 to 20/12. This computerized testing method should be considered when visual acuity is used as an outcome measure in eye research involving children 7 to <13 years old.

Adolescent↗

Reliability and validity of the objective structured clinical examination in assessing surgical residents.

The purpose of this research was to assess reliability and construct validity of the objective structured clinical examination (OSCE) for evaluating the clinical skills of surgical residents. Reliability refers to precision of the examination and construct validity to the degree to which the examination can discriminate between different levels of training. Twenty-seven second postgraduate year surgical residents took a 38-station OSCE representing seven surgical specialties and that tested history-taking, physical examination, problem-solving, technical skills, and attitudes. A couplet methodology was used wherein a patient encounter was followed by written questions aimed at testing problem-solving and patient management capabilities. Thirty-six standardized patients were trained and 36 surgeons served as examiners marking from structured checklists. Overall reliability, Cronbach's alpha, was 0.89. Construct validity was examined by comparing the scores of the residents with those of a group of graduates of foreign medical schools applying for a "pre-internship" program. For 17 of 19 stations that both groups took, the residents performed significantly better (p less than 0.01). Individual station validity was significant for 32 of 38 stations (r = 0.36 to 0.82, p less than 0.05). The examinations took 3.83 hours at a cost of $5,293 (Canadian dollars). The OSCE has been shown to be a reliable method of assessing clinical skills of surgical residents, construct validity has been established, and inter-item validity confirmed. Reliabilities achieved exceed those traditionally required for both acceptance and promotion decisions.

Analysis of Variance↗

Evaluation of interrater reliability for posture observations in a field study.

This paper examines the interrater reliability of a quantitative observational method of assessing non-neutral postures required by work tasks. Two observers independently evaluated 70 jobs in an automotive manufacturing facility, using a procedure that included observations of 18 postures of the upper extremities and back. Interrater reliability was evaluated using percent agreement, kappa, intraclass correlation coefficients and generalized linear mixed modeling. Interrater agreement ranged from 26% for right shoulder elevation to 99 for left wrist flexion, but agreement was at best moderate when using kappa. Percent agreement is an inadequate measure, because it does not account for chance, and can lead to inflated measures of reliability. The use of more appropriate statistical methods may lead to greater insight into sources of variability in reliability and validity studies and may help to develop more effective ergonomic exposure assessment methods. Interrater reliability was acceptable for some of the postural observations in this study.

Analysis of Variance↗