Search PubMedSearch

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Effects of cognitive impairment on the reliability of geriatric assessments in nursing homes.

OBJECTIVE: To explore the relationship between an elderly subject's cognitive status and the reliability of multidimensional assessment data. DESIGN: Survey, with cognitive status as the independent variable and interrater reliability as dependent variable. SETTING: Medicare/Medicaid-certified nursing homes. PARTICIPANTS: 147 residents age 65 or older. MEASUREMENTS: Dual assessments of elderly nursing home residents were performed by nurse assessors using the Health Care Financing Administration's new Minimum Data Set for Nursing Home Resident Assessment and Care Screening (MDS). Assessments were classified on the basis of residents' cognitive status, and levels of disagreement between assessors were analyzed. MAIN RESULTS: Overall assessment reliability, agreement concerning a resident's activities of daily living status, and the reliability of estimates of his or her communication skills and sensory abilities were significantly affected by a resident's cognitive status. The presence of cognitive impairment made these measurements less reliable--especially those related to communication skills, vision, and hearing. CONCLUSIONS: Assessments of residents suffering from cognitive impairment were significantly less reliable than assessments of cognitively intact residents. However, these differences in reliability were not uniform across all assessment domains. When treating the cognitively impaired elderly, clinicians must exercise caution in their reliance on standardized measurements that may be less reliable for this population.

Activities of Daily Living

Reliability of a standardized and expanded Brief Psychiatric Rating Scale: a replication study.

This study aimed to determine the replicability of the interrater reliability coefficients obtained with a standardized and expanded Brief Psychiatric Rating Scale (BPRS-E) in a 1991 psychometric evaluation. Furthermore, intrarater reliability was assessed. At item level, interrater concordance turned out to be satisfactory for most of the BPRS-E items. However, only a few of the items reached acceptable chance-corrected coefficients. In contrast to the previous study, the anxiety-depression subscale met the standard of acceptable interrater reliability in the present study. As in the 1991 study, the 10-item psychotic disintegration scale as well as BPRS-18 global scores met (or closely approximated) this standard. The 6 additional items of BPRS-E did not contribute to the scale's reliability. Joining the samples of the 1991 and replication studies (to cover the range of symptoms' severity and heterogeneity more fully) did not improve interrater reliability. Intrarater reliability coefficients were globally comparable to interrater reliability coefficients. In all, the results of this replication study suggest that only the anxiety-depression subscale, the 10-item psychotic disintegration scale and the BPRS-18 global scale can be used reliably in unselected groups of psychiatric inpatients in acute distress.

Adolescent

Factors affecting reliability coefficients of health attitude scales.

This study determined the minimum number of health attitude items and minimum sample size required to achieve maximum scale reliability coefficients, using different methods of estimating reliability. A 54-item alcohol attitude scale was administered to 700 participants. The scale produced .96 and .91 reliability coefficients, using the Cronbach Alpha (CA) and the Split-half (S-B) methods, respectively. A computer program randomly selected groups of participants and items from the pool of participants and items using different increments. A matrix of coefficients of reliability for both methods was calculated for different groups of items and sample size. To replicate the study, a 30-item cancer attitude scale was administered to more than 1,000 representative participants and produced reliability coefficients of .94 (using CA) and .82 (using S-B). The same computer and statistical procedures were repeated for the second data set. Results from both analyses consistently demonstrated that sample size has an insignificant effect on the coefficient values of reliability. Reliability increased as the number of items reached 18. Adding more items only negligibly increased the coefficients. Overall, the CA method consistently produced higher coefficient values of reliability compared to the S-B method.

Attitude to Health

Reliability of the National Institutes of Health Stroke Scale. Extension to non-neurologists in the context of a clinical trial.

BACKGROUND AND PURPOSE: The reliability of the National Institutes of Health Stroke Scale (NIHSS) has been established through testing its use in live and videotaped patients. This reliability testing has primarily focused on the use of the scale by neurologists. We sought to determine the reliability of the NIHSS as used by non-neurologists in the context of a clinical trial. METHODS: In anticipation of the initiation of a randomized trial of a new therapy for patients with acute ischemic stroke, 30 physician investigators (30% of whom were not neurologists) and 29 non-physician study coordinators were trained in the use of the NIHSS at an informational and training conference using standardized videotaped patient examinations. A series of 4 patients were rated initially. After 3 months, the same 4 patients were rerated, providing a measure of intraobserver reliability. An additional series of 4 new patients were also rated after 3 months and, with the initial 4 ratings, provided data for assessment of interobserver reliability. RESULTS: Overall, 28% of the raters had previous experience with the NIHSS, and 22% had previously used the videotapes as used in the present trial. The coefficients of determination (r2) were each greater than .95 when the means of the two ratings of the same 4 cases were compared between (1) neurologists and other types of physicians, (2) physicians and study coordinators, (3) raters who had prior experience with the NIHSS and those without prior experience, and (4) raters who had used the videotapes in the past and those who had never viewed the tapes. The calculated r2s were greater than .98 for the initial rating of the first 4 cases and for the later rating of the 4 new cases. The slopes of the regression lines were all near 1, indicating that the raters were similarly calibrated. The intraclass correlation coefficients were .93 and .95, reflecting high levels of intraobserver and interobserver reliability. CONCLUSIONS: These data extend the previously demonstrated reliability of the NIHSS to non-neurologists and show that both a variety of physician investigators and nurse study coordinators can be rapidly trained to reliably apply the scale in the context of an actual clinical trial.

Cerebrovascular Disorders

The reliability of Form 90: an instrument for assessing alcohol treatment outcome.

OBJECTIVE: Project MATCH is a randomized clinical trial consisting of five outpatient and five aftercare units at nine sites. Of importance in this multisite trial examining the efficacy of client-treatment matching was the cross- and within-site reliability of the structured interview used to assess alcohol treatment outcomes, the Form 90. Evaluation of the reliability of Form 90 is the subject of this article. METHOD: The reliability of Form 90 was evaluated in two test-retest studies. The cross-site reliability study consisted of 70 paired test-retest interviews conducted by different interviewers. Clients for this study were recruited from inpatient, outpatient and college settings. The within-site reliability study had a total of 108 paired test-retest interviews, with 54 of the retests conducted by different interviewers and 54 by the same interviewer. Clients for this study were most often presenting for alcohol treatment at the nine sites and were selected to be representative of the larger Project MATCH sample. RESULTS: Good-to-excellent reliability was found for all key summary measures of alcohol consumption and psychosocial functioning, and most frequently used illicit drugs had moderate reliability. No decay in consistency of self-reported drinking was found at more distal points from dates of test-retest interviews. Application of 68% confidence intervals for primary alcohol consumption measures suggests that trained researchers and clinicians can obtain consistent information regarding client drinking. CONCLUSIONS: Form 90 appears to be a reliable instrument for alcohol treatment assessment research when interviewers have received careful training and supervision in its use.

Adult

Testing reliability of plaque and gingival indices. Two methods.

This investigation was undertaken to compare two methods of interexaminer and intraexaminer reliability in the evaluation of Plaque and Gingival Indices prior to a study of toothbrushing. Inter-/intraexaminer reliabilities were compared using a projected slide series consisting of 40 slides of clinical examples of gingival inflammation and plaque accumulation. Time between assessments was three weeks. Using the slide technique, intraexaminer reliability was established for: (1) Gingival Indices and (2) Plaque Indices. Interexaminer reliability was also established for Gingival Indices. Interexaminer reliability could not be established for Plaque Indices on the first assessment but was established on the post-assessment. Intraexaminer reliability was also determined through clinical examinations of patients. A third clinician was used to manipulate the tissue while investigators evaluated bleeding on provocation and plaque accumulation. Significant results were established for the Gingival Indices and Plaque Indices. Results of this investigation suggest that significant inter-/intraexaminer reliabilities may be obtained for gingival indices using the slide technique. In addition, the clinic technique appeared useful for assessing interexaminer reliability for Gingival Indices. Plaque Indices using the slide technique required more practice than those using the clinic technique.

Dental Health Surveys

Intramachine and intermachine reliability for selected dynamic muscle performance tests.

The Cybex 6000 isokinetic dynamometer is a new isokinetic device for which no published reports of reliability have been presented in the literature. In addition, the manufacturer not only claims that the new Cybex 6000 is reliable but that torque data obtained from the Cybex 6000 are consistent with data obtained from past Cybex systems, such as the Cybex II. The purpose of this study was to investigate the intramachine reliability of the Cybex 6000 to itself and the intermachine reliability of the Cybex 6000 and the Cybex II. Data on peak torque, work, and power were collected using the Cybex 6000, and data on peak torque were obtained using the Cybex II for knee flexion and extension in 20 volunteers (10 males, 10 females). Subjects were tested three times, twice on the Cybex 6000 and once on the Cybex II, approximately 1 week apart across a 3-week period of time at angular velocities of 60, 180, and 300 degrees/sec. Data were analyzed using intraclass correlations. Results indicated that the majority of test-retest correlation coefficients for all parameters for intramachine reliability of the Cybex 6000 were above .90. Comparing peak torque obtained with the Cybex 6000 to that obtained with the Cybex II (intermachine reliability), correlation coefficients ranged from .72 to .89. In conclusion, information obtained on the Cybex 6000 appears to be quite reliable in a test-retest situation using the same equipment and moderately reliable when compared to the Cybex II. Clinical implications for these results are discussed.

Adult

Relationship of the pelvic angle to the sacral angle: measurement of clinical reliability and validity.

There is a need to better document the reliability and validity of assessment measures used in physical therapy. Studies documenting the reliability of measurement of the pelvic angle and its relationship to sacral motion are presently inconclusive. The purpose of this study was twofold. First, we wanted to determine the reliability and validity of a goniometric measurement of the pelvic angle. We also wanted to test the hypothesis that there is a relationship between the pelvic angle and the sacral angle. Intertester and intratester reliability of goniometric pelvic angle measurements of 23 healthy young adults were examined using three different raters. Radiographic measurements of the pelvic and sacral angle using two raters and goniometric measurement of the pelvic angle using a single rater were taken from 15 patients with low back pain who had been referred for X-rays. Intraclass correlation coefficients (ICCs) of intratester reliability for goniometric measurements of the pelvic angle were .93, .96, and .96. The intertester reliability was .95. The ICCs for intratester reliability for radiological measurements were .92 and .95 for the sacral angle and .98 for both measurements of the pelvic angle. Intertester reliability coefficients were .86 and .88, respectively. The Pearson correlation coefficients for the goniometric and radiological measurements of the pelvic angle were .85 and .68. A comparison of the radiological and goniometric measurements of the pelvic angle with the sacral angle demonstrated low average correlations of .43 and .58, respectively. The results indicate a high level of correlation between and within testers for goniometric measurements of the pelvic angle but only a fair correlation between goniometric and radiological measurements of the pelvic angle.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult

Consultation competence in general practice: testing the reliability of the Leicester assessment package.

BACKGROUND: An acceptable assessment must be both valid and reliable; the face validity of the Leicester assessment package has already been established. AIM: This study set out to test the reliability of the Leicester assessment package, and the factors influencing it, when used by multiple assessors to assess performance in general practice consultations. METHOD: Six randomly selected course organizer assessors simultaneously used the package to conduct independent assessments of the performance of five doctors of widely varying abilities in consultation with six simulated patients. The scores allocated were subjected to generalizability analysis. RESULTS: The mean scores allocated for consultation performance of individual doctors ranged from 51% to 70%, with the lower scores being allocated to the less experienced doctors. Scores of each assessor across the cases were examined for internal consistency and five of the six assessors consistently scored the doctors with an alpha coefficient of the minimum accepted level of 0.80 or greater. The other assessor had a consistency of only 0.22. Measurements of consistency within cases between markers indicated that the first case produced unreliable results (alpha coefficient 0.25) but all other cases were scored consistently. Two independent assessors scoring eight consultations are the requisite numbers to achieve acceptable levels of reliability in a formal assessment process; seven consultations produce the minimum acceptable generalizability coefficient of 0.80 plus the first 'non-counting' consultation. CONCLUSION: Required levels of reliability can be achieved when the package is used by multiple markers assessing the same consultations over a wide range of consultation performance. To achieve reliability only two hours of assessment time are required using the Leicester package compared with the previously regarded minimum of 32 hours. Although assessors can produce reliable scores with minimal training, intra-assessor reliability cannot be taken for granted and all assessors should be trained and calibrated before being sanctioned to conduct assessments, particularly for regulatory purposes. The Leicester assessment package has now been shown to be valid, reliable, feasible and easy to use in practice. It can, therefore, be recommended for use in both formative and summative assessment of consultation competence in general practice.

Communication

Chiropractic biophysics lateral cervical film analysis reliability.

OBJECTIVE: To determine the degree to which the geometric line drawings used in Chiropractic Biophysics Technique (CBP) on lateral cervical radiographs are reliable. DESIGN: A blind, delayed repeated measures design was used. Three examiners were presented radiographs in random order. All identifying marks were removed prior to each examiner's individual marking and measurement. Each examiner was blinded as to how the previous examiners marked and measured the radiographs. SETTING: Primary care private chiropractic clinic. PATIENTS PARTICIPANTS: Sixty-five subject films were provided from the patient records of a primary care private chiropractic clinic. The 65 radiographs qualified for inclusion in the study based on two criteria: C1 through C7 had to be clearly visible, and there had to be no identifying artifacts. MAIN OUTCOME MEASURES: Anterior head translation in millimeters, atlas plane to horizontal, Ruth Jackson's cervical stress lines, and five relative rotation angles for C2-C3, C3-C4, C4-C5, C5-C6, C6-C7. Inter- and intrareliability of the three examiners were statistically analyzed. RESULTS: Intraexaminer for a) C1 to horizontal reliability was .98-.99 with confidence intervals of .96-.99, b) absolute rotation angle from C2 to C7 reliability was .82-.95 with confidence intervals of .80-.99, c) anterior head translation [+Sz] reliability was .86-.99, with confidence intervals of .74-.99, d) relative rotation angle reliability ranges were (C2-C3) .99, and (C3-C4) .98-.99, (C4-C5) .88-.99, (C5-C6) .80-.99, and (C6-C7) .94-.98. Interexaminer reliabilities across examiners ranged from a) Winer:.89-.99 and b) Bartko: .72-.96. CONCLUSIONS: The reliabilities for intra- and interexaminer were all greater than .70, indicating that these measurements in CBP technique would be considered accurate enough to provide measurements for future clinical studies. The data indicated that the C6-C7 relative rotation angle was the least reliable measurement. This might be due to the very small angles found at this level.

Analysis of Variance

Intra- and interexaminer reliability of the chiropractic biophysics lateral lumbar radiographic mensuration procedure.

OBJECTIVE: To determine the intra- and interexaminer reliability of a specific method of mensuration commonly used to evaluate the positional configuration of the lumbopelvic spine viewed on lateral lumbar radiographs. DESIGN: A blind, repeated-measures design was used. Lateral lumbopelvic radiographs were presented to each of three examiners in random order. Each film was marked and measurements were recorded. The films were cleaned of all markings and randomized again for a second run by each examiner. Each examiner's measurements were unavailable to the other examiners. SETTING: Private, primary-care chiropractic clinic. MAIN OUTCOME MEASURES: Anterior/posterior thoracic translation in millimeters, Ferguson's sacral-plane angle to horizontal, arcuate line angle to horizontal, L1 to L5 absolute rotation angle and four relative rotation angles for L1-L2, L2-L3, L3-L4 and L4-L5. Intra- and interrelibility of the three radiographic examiners were analyzed. RESULTS: Intraexaminer reliability for (a) L1-L5 absolute rotation angle was .98, with confidence intervals included in the range of 0.95-0.99, (b) anterior/posterior thorax translation [+/- Sz] was .97-.99, with confidence intervals included in the range of 0.94-1.00, (c) arcuate angle (AA) .40-.81, with confidence intervals included in the range of 0.07-0.90, (d) Ferguson's angle (FA) was .91-.97, with confidence intervals included in the range of 0.82-0.98, (e) relative rotation angle reliability ranges were L1-L2, .84-.94; L2-L3, .80-.85; L3-L4, .78-.89; L4-L5, .87-.92. Interexaminer reliabilities for the three examiners ranged from .66-.98. CONCLUSION: With the exception of the arcuate angle measurement, the reliabilities for all other measurements were at least .78. Those measurements with reliabilities approaching .80 or better would be considered accurate enough for use in future clinical studies. The arcuate angle measurement may have been least reliable because of the subjective nature of the method of affixing a best-fit line to a radiographic landmark that often takes on the appearance of a mild curvature. Establishing reliability is an important first step toward evaluating these and other similar radiographic measurements that have yet to be examined for their validity.

Analysis of Variance

Spatial information content and reliability of hippocampal CA1 neurons: effects of visual input.

The effects of darkness on quantitative spatial firing characteristics of 235 hippocampal CA1 "complex spike" (CS) cells were studied in young and old Fischer-344 rats during food-motivated performance of a randomized, forced-choice task on an eight-arm radial maze. The room lights were turned on or off on alternate blocks of all eight arms. In the dark, a lower proportion of CS cells had "place fields," and the fields were less specific and less reliable than in the light. A small number of cells had place fields unique to the dark condition. Like CS cells, Theta cells showed a reduction in spatially related firing in the dark. The specificity and reliability of the place fields under both light and dark conditions were similar for both age groups. Increasing the salience of the environment, by increasing the light level and the number of visual cues in the light condition, did not affect the specificity or reliability of the place fields. Even though all rats had substantial prior experience with the environment, and were placed on the maze center under normal illumination before the first dark trial, the correlation between the firing pattern in the light and dark increased after the rat first traversed the maze in the light. Thus, even after considerable experience with the environment over days, experiencing the illuminated environment from different locations on a given day was a significant factor affecting subsequent location and reliability of place fields in darkness. While the task was simple and errors rare, rats that made fewer errors (i.e., re-entries into the previously visited arm) also had more reliable place cells, but no such correlation was found with place cell specificity. Thus, the reliability of spatial firing in the hippocampus may be more important for spatial navigation than the size of the place fields per se. Alternatively, both spatial memory and place field reliability may be modulated by a common variable, such as attention.

Action Potentials

Reliability of self-reported breast screening information in a survey of lower income women.

BACKGROUND: Self-reported behavior is widely used to estimate the prevalence of breast cancer screening and to evaluate programs for promoting screening, but detailed studies of reliability have not previously been performed. METHODS: Reliability was assessed by comparing responses to questions about screening behavior from repeat personal interviews of 382 women age 40 and older living in low-income census tracts of two Florida communities. Reliability was assessed using Pearson's correlation (r) and kappa (kappa) coefficients. RESULTS: Estimated reliabilities were kappa = 0.38 for "ever had clinical breast examination," kappa = 0.82 for "ever had mammogram," kappa = 0.65 for "mammogram in past year," r = 0.54 for "date of last mammogram," and r = 0.72 for "number of mammograms." The dates of last mammogram reported at the two interviews agreed within 1 month for 64% of the women, while the dates of last clinical breast examination agreed within 1 month for 50% of the women. Reliability of "ever had mammogram" was significantly related to demographic variables. CONCLUSIONS: Women reliably report ever having mammography, but information about timing and frequency has lower reliability. The results have implications for breast screening research because measurement error affects the precision of estimates and the sample sizes needed to detect program effects.

Adult

Reliability of psychophysiological responding as a function of trait anxiety.

This study examined the temporal stability of three psychophysiological responses (frontal electromyographic activity, hand surface temperature, and heart rate) recorded over four sessions (days 1, 2, 8, and 28) on 34 subjects, 17 with high Spielberger Trait Anxiety Inventory scores and 17 with low scores. Each session consisted of a 20-minute adaptation period, a baseline condition, and two stressors (one cognitive, the other physical). Two forms of reliability coefficients were employed, intraclass correlations and Pearson Product Moment; the two types of reliability coefficients arrived at the same conclusions. Results indicated that reliability coefficients for the two anxiety groups did not differ on frontal EMG or heart rate responses; however, hand surface temperature responding was considerably less reliable for high anxious individuals than low anxious individuals. Reliability coefficients on absolute scores were, for the most part, reliable. Treating the responses as relative measures (percent change from baseline or simple change scores from baseline) produced smaller and less reliable coefficients. Magnitudes of the three physiological responses did not significantly differ as a function of high or low trait anxiety. Findings are discussed in terms of their clinical, as well as basic psychophysiological, importance.

Adult

Reliability and validity of clinical outcome measurements of osteoarthritis of the hip and knee--a review of the literature.

High reliability and validity of clinical rating schemes is crucial for their use as outcome measurements of treatment of hip and knee osteoarthritis. In this paper, we review the empirical evidence on the reliability and validity of commonly used clinical scores. Clinical scores and related reliability and validity studies were identified by systematic literature search. Scores were classified according to the type and joint. Reliability and validity studies were characterized according to design, population, number and qualification of observers, number of measurements, time interval between repeat measurements and results. Reliability and validity studies were reported for only 6 and 15 of the 45 identified clinical scores, respectively. Although comparisons are difficult due to differences in study design, relatively high reliability was reported for most measurements of pain, stiffness, and physical function, while results are less conclusive for clinical signs. Most validity studies focused on the correlation between various scores. Correlation was generally found to be high for overall numerical ratings, but scores often differed with respect to the interpretation of these ratings. Validity has been more comprehensively studied for Lequesne's scores, WOMAC, and ILAS, and these scores have shown satisfactory responsiveness to different treatment effects. Overall, knowledge on reliability and validity of clinical scores of hip and knee osteoarthritis is limited, underlining the need for further properly designed and conducted studies.

Hip

Reliability of the Health Utilities Index--Mark III used in the 1991 cycle 6 Canadian General Social Survey Health Questionnaire.

This study presents information on the test-retest reliability of the Health Utility Index--Mark III (HUI) system used in cycle 6 of the Canadian General Social Survey (GSS). The HUI system used in this reliability study consists of an eight-attribute health status classification system (HSCS) and a function for generating a summary score of health-related quality of life. To estimate test-retest reliability, a stratified random sample of individuals (n = 506) completing GSS telephone interviews during August and September, 1991 were interviewed again 1 month later. Weighting adjustments based on the probability of selection were invoked during the analyses to provide unbiased estimates of test-retest reliability for all GSS respondents in the August-September period. The results indicate that the individual questions, attributes and provisional index scores generally provided reliable information on health status in the GSS. The exceptions to this were limitations in speech and dexterity which were reported very infrequently. Kappa estimates of test-retest reliability for individual questions varied from 0.184 to 0.766. For the eight attributes, kappa estimates varied from 0.137 to 0.728. Using the provisional index scores to quantify health overall, a test-retest reliability of 0.767 was obtained (intra-class correlation coefficient).

Activities of Daily Living

Factors affecting the reliability of ratings of students' clinical skills in a medicine clerkship.

OBJECTIVE: To determine the overall reliability and factors that might affect the reliability of ratings of students' clinical skills in a medicine clerkship. DESIGN: A nine-item instrument was used to evaluate students' clinical skills. Raters were also asked to provide a grade of each student's overall clinical performance. Generalizability studies were performed to estimate the reliability of the ratings. The effects of rater experience and clerkship setting were investigated by regression analysis. SETTING: Teaching hospitals and community-based sites in three Northwestern states. PARTICIPANTS: All students (328) who had completed the 12-week clerkship in internal medicine at one medical school during the academic years 1987-1989. Raters included attending physicians, chief residents, and other residents. RESULTS: Seven observations were needed to provide a reliable rating of the overall clinical grade. More observations were needed to obtain reliable ratings for individual items, ranging from seven observations needed for the rating of data gathering skills to 27 observations needed for the rating of interpersonal relationships with patients. Rater experience and clerkship setting (i.e., teaching hospitals vs. community-based clinics) were found, in general, not to affect significantly the ratings received by students. CONCLUSIONS: Reliable ratings of students' overall clinical skills, including overall clinical grades, can be achieved by collecting a minimum of seven observations. More observations are needed to measure reliably the interpersonal aspects of clinical performance. These findings support the use of performance ratings to evaluate clinical skills and knowledge of students in clerkship settings.

Clinical Clerkship

Reliability of goniometric measurements and visual estimates of ankle joint active range of motion obtained in a clinical setting.

We examined intratester and intertester reliability for goniometric measurements of ankle dorsiflexion (ADF) and ankle plantar flexion (APF) active range of motion (AROM). Parallel-forms intratester reliability for ankle AROM measurements obtained by the universal goniometer (UG) and by visual estimation (VE) and intertester reliability for VE of ADF and APF were examined. Repeated measurements were obtained on 38 patients with orthopedic problems by 10 physical therapists in a clinical setting. For intratester reliability of measurements obtained with UG, intraclass correlation coefficients (ICC) for all physical therapists were 0.64 to 0.92 (median, 0.825) for ADF and 0.47 to 0.96 (median, 0.865) for APF. Intertester reliability was quantified with use of ICC. ICCs for measurements obtained by UG were 0.28 for ADF and 0.25 for APF; ICC of VE for ADF was 0.34 and was 0.48 for APF. ICC for parallel-forms intratester reliability obtained with UG and VE ranged from 0 to 0.94 (median, 0.58) for ADF and 0 to 0.86 (median, 0.625) for APF. Thus, a physical therapist should use a goniometer when making repeated measurements of ankle joint AROM. Considerable inconsistency exists when two or more physical therapists make repeated goniometric and visual measurements of ankle motion on the same subject. Physical therapists may erroneously conclude that a patient's AROM has changed because of treatment when the change could be attributed to a lack of intertester reliability.

Adolescent