Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Interexaminer reliability of observations in physical examinations of the neck.

The purpose of this study was to collect data on interexaminer reliability of a set of tests representative of the clinical examination of a patient with neck and radicular pain. A conventional neurological examination, palpations, and tests for the provocation or relief of radicular symptoms were performed on 52 patients by two independent raters. Good reliability was obtained in the atrophy inspection of the small muscles of the hand, in the sensitivity tests for touch and pain, and in the neck compression and axial manual traction tests. Fair reliability was obtained in muscle strength testing and in the estimation of the range of motion, and poor reliability was obtained for many palpations. Poor standardization of examination procedures and changes in the patients' attention were considered the main factors affecting reliability. Better operational definitions and procedures, such as the standardization of palpation pressure and traction force, are suggested for future studies.

Adolescent↗

Reliability of a diabetic foot evaluation.

The purpose of this study was to establish the interrater and intrarater reliability of various ankle and foot measures common to a diabetic evaluation. Bilateral biomechanical, sensory, and wound-size measurements were obtained in 31 subjects with diabetes mellitus. Twenty-five subjects were retested by the initial examiner to determine intratester reliability, and all subjects were retested by another examiner to determine intertester reliability. Both examiners participated in an extensive training period prior to the initiation of this study to minimize variability between and within measurers. Intraclass correlation coefficients for interrater and intrarater measurements ranged from .58 to .89 and from .74 to .99, respectively. The results of this study indicate that ankle and foot measurements common to a diabetic evaluation can be taken reliably between testers. We believe extensive examiner training in these clinically relevant measures can improve reliability between testers.

Adult↗

Intrasession and intersession reliability of hand-held dynamometer measurements taken on brain-damaged patients.

Recent reports have characterized force measurements obtained with hand-held dynamometers from brain-damaged patients as being highly reliable. The purposes of this two-part study were to replicate essential parts of those studies and to further examine the reliability of these measurements in a clinical context. Repeated force measurements were taken from the nonparetic and paretic limbs of brain-damaged patients during the same testing session (Part 1) and during two testing sessions separated by two days (Part 2). The intratester intraclass correlation coefficients (ICCs) for all measurements taken during a single session ranged from .88 to .98. The ICCs for repeated measurements taken two days apart from the paretic limbs ranged from .90 to .98. The ICCs for repeated measurements taken two days apart from the nonparetic limbs ranged from .31 to .93. The ICCs for repeated measurements taken two days apart from the combined data for all limbs ranged from .79 to .97. Hand-held dynamometer measurements taken on brain-damaged patients appear to be highly reliable when taken during the same testing session. When repeated measurements are separated by a longer time interval, the measurements taken from the paretic limbs continue to be highly reliable, whereas most measurements taken on the nonparetic limbs exhibit poor reliability.

Adult↗

Interrater and test-retest reliability of two pediatric balance tests.

The purpose of this study was to examine the interrater and test-retest reliability of a one-leg balance test and a tiltboard balance test. Twenty-four normally developing children aged 4 through 9 years participated in the study. Time and quality of balance on one leg and degrees of tilt on a tiltboard prior to postural adjustment were measured. Both tests were completed with eyes open and with eyes closed. Interrater reliability was examined using two raters. Test-retest reliability, with a one-week interval between test and retest, was examined for a subgroup consisting of 12 children. Spearman rank-order correlation coefficients were used as indexes of both interrater and test-retest reliability for time and degrees of tilt. To supplement the correlation coefficients, the magnitudes of difference between raters' scores and between test and retest scores were calculated. Spearman coefficients were moderate to high for one-leg balance when scores for both feet were combined for both eyes-open and eyes-closed conditions. The magnitude of difference between scores was low, indicating good agreement between raters and across time. Interrater and test-retest reliabilities of quality of one-leg standing balance were examined by calculating percentages of agreement and Cohen's Kappa statistics. Results of these analyses revealed the need for further study. The Spearman coefficients for the interrater tiltboard test were high; however, the test-retest coefficients were low. The magnitudes of difference between scores were small for the two raters, but large for test and retest. These results are important to consider when using these tests for initial evaluation or for monitoring patient progress.

Biomechanical Phenomena↗

Reliability of clinical measurements of forward bending using the modified fingertip-to-floor method.

The purpose of this study was to examine the intratherapist and intertherapist reliability of measurements obtained with a modified version of the fingertip-to-floor method of assessing forward bending. With the modified fingertip-to-floor (MFTF) method, patients stand on a stool and forward bend so that measurements can be taken on patients who are able to touch the floor or reach beyond the level of the floor. Randomly paired physical therapists took repeated MFTF measurements on 73 patients with low back pain. Intraclass correlation coefficients (ICCs) were calculated for intratherapist and intertherapist reliability. The ICC value for intratherapist reliability was .98, and the ICC value for intertherapist reliability was .95. The results of this study suggest that measurements of forward bending obtained on patients with low back pain using the MFTF method are highly reliable.

Adolescent↗

Reliability and concurrent validity of the Metrecom for length measurements on inanimate objects.

BACKGROUND AND PURPOSE: The purpose of this study was to test the reliability and concurrent validity of length measurements (distance between two points) produced by a computer-interfaced, three-dimensional digitizer called the Metrecom. METHODS: A total of 344 points were marked on the surfaces of five different inanimate objects and were digitized in pairs with the Metrecom. On the first three objects, each of two testers digitized each point twice for each of two testing modes (Line Length and 3-D Digitizer); on the last two objects, only one mode (3-D Digitizer) was used. Intraclass correlation coefficients and Pearson Product-Moment Correlation Coefficients were used to assess reliability and concurrent validity, respectively. RESULTS: All of the correlation coefficients were > or = .99. In further analysis of the results, a repeated-measures, one-way analysis of variance (used with reliability data) and repeated-measures t tests (used with validity data) were used to test for differences between repeated measures. After adjustment of the alpha level for the total number of comparisons, two of the t tests for validity comparisons were significant. CONCLUSION AND DISCUSSION: The results indicate that the Metrecom provides reliable length measurements (distance between two points) on inanimate objects and that two different test modes produce consistent measurements. Further study of the validity and reliability of length measurements obtained with the Metrecom on humans under applied conditions is needed before the results of this study can be generalized to applied settings.

Humans↗

Interrater reliability of craniosacral rate measurements and their relationship with subjects' and examiners' heart and respiratory rate measurements.

BACKGROUND AND PURPOSE: The evaluation of craniosacral motion is an approach used by physical therapists and other health professionals to assess the causes of pain and dysfunction, but evidence for the existence of this motion is lacking and the reproducibility of the results of this palpatory technique has not been studied. This study examined the interexaminer reliability of craniosacral rate and the relationships among craniosacral rate and subjects' and examiners' heart and respiratory rates. SUBJECTS: Participants were 12 children and adults with histories of physical trauma, surgery, or learning disabilities. Three physical therapists with expertise in craniosacral therapy were the examiners. METHODS: One of three nurses recorded heart and respiratory rates of both subject and examiner. The examiner then palpated the subject to determine craniosacral rate and reported the findings to the nurse. Each subject was examined by each of the three examiners. RESULTS: Reliability was estimated using a repeated-measures analysis of variance and the intraclass correlation coefficient (2,1). Significant differences among examiners and the scatter plot of rates showed lack of agreement among examiners. The ICC was -.02. The correlations between subject craniosacral rate and subject and examiner heart and respiratory rates were analyzed with Pearson correlation coefficients and were low and not statistically significant. DISCUSSION AND CONCLUSIONS: Measurements of craniosacral motion did not appear to be related to measurements of heart and respiratory rates, and therapists were not able to measure it reliably. Measurement error may be sufficiently large to render many clinical decisions potentially erroneous. Further studies are needed to verify whether craniosacral motion exists, examine the interpretations of craniosacral assessment, determine the reliability of all aspects of the assessment, and examine whether craniosacral therapy is an effective treatment. [Wirth-Pattullo V. Hayes KW. Interrater reliability of craniosacral rate measurements and their relationship with subjects' and examiners' heart and respiratory rate measurements.

Adolescent↗

Reliability of pain and stiffness assessments in clinical manual lumbar spine examination.

BACKGROUND AND PURPOSE: The purpose of this study was to determine the intertherapist reliability of judgments of stiffness and pain at L-1 to L-5 made using posteroanterior (PA) central pressure testing. SUBJECTS: Three pairs of manipulative physical therapists with a minimum of 5 years of experience were asked to rate pain and stiffness in a total of 90 patients with low back pain. METHODS: Each pair of therapists assessed 30 patients within their own clinic, using their preferred technique to perform an examination using the PA central pressure test at the five lumbar levels. Each pair of therapists recorded their ratings of pain and stiffness. Reliability of judgments was evaluated by intraclass correlation coefficients (ICC) and percentage of exact agreement scores. RESULTS: The ICC values for pain judgments for the group as a whole ranged from .67 to .72, with agreement scores ranging from 31% to 43%. The ICC values for stiffness judgments ranged from .03 to .37, with agreement scores ranging from 21% to 29%. CONCLUSION AND DISCUSSION: Judgments of stiffness made by experienced manipulative physical therapists examining patients in their own clinics were found to have poor reliability, whereas pain judgments had good reliability. Further investigation of this test is required in order to develop a more reliable method of assessing PA stiffness.

Adult↗

Influence of examiner experience and gender on interrater reliability of KT-1000 arthrometer measurements.

BACKGROUND AND PURPOSE: Measurements of the integrity of knee ligaments are used to diagnose injuries as well as to document the state of recovery. Many factors, such as gender and experience of the examiner, are capable of influencing the reliability of such measurements. The purpose of this study was to determine the effects on interrater reliability of measurements obtained using the KT-1000 arthrometer of experience, gender, and leg tested. SUBJECTS: Two experienced examiners (1 male, 1 female) and two inexperienced examiners (1 male, 1 female) tested 22 subjects with unilateral anterior cruciate ligament (ACL) pathology. METHODS: The leg with an ACL injury and the uninjured leg of each subject were evaluated by all four examiners within one test session using 67-N, 89-N, maximum manual, and active anterior drawer tests. RESULTS: Greater anterior displacement values were found in the legs with ACL injury than in the uninjured legs. Reliability estimates, as assessed by intraclass correlation coefficients (2,k) and measurement error (SEM), suggest that therapist experience may be a more important factor influencing reliability than gender. CONCLUSION AND DISCUSSION: Given the magnitude of the errors obtained for tests routinely conducted in the clinic using the KT-1000 arthrometer, we recommend that repeated measurements should be taken by the same examiners whenever possible. [Ballantyne BT, French AK, Heimsoth SL, et al. Influence of examiner experience and gender on interrater reliability of KT-1000 arthrometer measurements.

Adolescent↗

Reliability of measurements obtained with four tests for patellofemoral alignment.

BACKGROUND AND PURPOSE: A series of patellofemoral (PF) alignment tests have been described that are used to determine when and how PF taping techniques should be applied. The reliability of measurements obtained with these tests has not been reported. The purpose of this study was to determine the intertester reliability of measurements obtained with four PF alignment tests: medial/lateral displacement, medial/lateral tilt, medial/lateral rotation, and anterior tilt. SUBJECTS: Twelve physical therapists from four clinics served as testers. A total of 66 patients were evaluated. METHODS: Paired testers performed all four PF alignment tests on the same patient. The intertester reliability of judgments for each of the PF alignment tests was determined by a kappa correlation coefficient. RESULTS: Kappa correlation coefficients ranged from .10 to .36 for the four PF alignment tests. CONCLUSION AND DISCUSSION: These findings suggest that the reliability of measurements obtained with the PF alignment tests described in this report ranged from poor to fair. Potential factors affecting the reliability of these measurements are discussed. Alternative methods for deciding when and how to apply PF taping techniques are also discussed.

Adolescent↗

Reliability of the Gross Motor Performance Measure.

BACKGROUND AND PURPOSE: The reporting of reliability coefficients and the method of their determination is expected of test developers. The purpose of this study was to estimate the interrater, intrarater, and test-retest reliability of the Gross Motor Performance Measure, a measure of quality of movement designed to accompany the Gross Motor Function Measure. SUBJECTS: Subjects were 28 children (25 with cerebral palsy, 2 nondisabled, 1 with head injury) between the ages of 1 and 10 years. METHODS: Reliability data were obtained from assessments of 19 therapists. RESULTS: Intraclass correlation coefficients for reliability varied from .92 to .96 for the total scores and from .84 to .94 for the five attribute scores. CONCLUSION AND DISCUSSION: When the Gross Motor Performance Measure was administered by therapists who are familiar with the Gross Motor Function Measure and had a 1-day training workshop, reliability of the total scores was above recommended minimums. Scores of single attributes were less reproducible.

Cerebral Palsy↗

Intertester reliability of a modified version of McKenzie's lateral shift assessments obtained on patients with low back pain.

BACKGROUND AND PURPOSE: McKenzie described a two-step process for assessing patients with low back pain for a lateral shift. The purpose of this study was to determine whether reliable judgments about lateral shifts could be obtained. SUBJECTS: Forty-nine patients with low back pain were each examined separately by two randomly paired physical therapists. METHODS: Assessments of the presence and direction of lateral shifts (step 1) were obtained by use of a simple instrument. The relevance of the lateral shifts to the patients' pain complaints (step 1) also was assessed by use of the side-glide test sequence. RESULTS: Generalized kappa coefficients were calculated to determine reliability. The kappa value for the two-step process of lateral shift assessment was .16. The percentage of agreement was 47%. CONCLUSION AND DISCUSSION: Each step in this two-step process was examined separately for possible sources of error. The kappa value for determinations of the presence and direction of lateral shifts was .00, indicating very poor reliability. The kappa value for the determination of the presence of a positive side-glide test sequence was .74, indicating high reliability. The role of lateral shift assessment in the McKenzie system should be reconsidered, given the strong research evidence for poor reliability of determinations of the presence and direction of lateral shifts.

Adolescent↗

The interrater reliability of force measurements using a modified sphygmomanometer in elderly subjects.

BACKGROUND AND PURPOSE: Physical therapists working with elderly people require an instrument that provides reliable force measurements and can be used in a clinical setting. The modified sphygmomanometer has been identified as potentially fulfilling these requirements, yet there is an absence of research on the reliability of measurements taken with this instrument on elderly patients. This study was undertaken to investigate the interrater reliability of force measurements, in a group of elderly subjects, using a modified sphygmomanometer. SUBJECTS: Thirty-six hospitalized subjects (mean age=75.28 years, SD=9.43, range=62-95) participated in the study. METHODS: With the modified sphygmomanometer, 3 examiners evaluated the isometric force of the elbow extensors and hip extensors using a break test and a make test, respectively. RESULTS: Intraclass correlation coefficients (2,1) reflecting reliability were .87 for the elbow extensors and .65 for the hip extensors. The estimation of the components of variance for hip extensors revealed that these results were due in part to the raters but that random error contributed to a much larger extent. CONCLUSION AND DISCUSSION: The modified sphygmomanometer appears to be practical to use, and the high correlations found in this study for the elbow extensors suggest that reliable measurements can be obtained with this instrument. Further research is needed, however, to specify the manner in which the modified sphygmomanometer can be used when assessing different muscle groups.

Aged↗

An investigation of the reliability and validity of posteroanterior spinal stiffness judgments made using a reference-based protocol.

BACKGROUND AND PURPOSE: The reliability and criterion-related validity of ratings of posteroanterior (PA) spinal stiffness made using reference values for comparison have not been investigated. In this study, mechanical reference stimuli for points on an 11-point rating scale were used to determine whether using a reference scale may be feasible. Subjects. Five different raters took part in 2 studies in which they rated 40 subjects who were asymptomatic for low back pain. METHODS: The interrater reliability of ratings was evaluated with intraclass correlation coefficients (ICCs) and standard errors of the measurement (SEMs). Criterion-related validity was evaluated by correlating judgments of PA spinal stiffness assessed manually with measurements of PA spinal stiffness provided by a mechanical device, the "Stiffness Assessment Machine" (SAM). RESULTS: Although the reliability indices were generally high, with ICCs reaching .77 and with SEMs as low as 0.72 points, the evidence for criterion-related validity (i.e., the ability of the examiner to judge spinal stiffness levels) was not strong, with correlations reaching only .56. CONCLUSIONS AND DISCUSSION: The reference-based protocol allows for more reliable measures of PA stiffness judgments than previous protocols have; however, the human ratings are not highly correlated with the SAM measures. The protocol will have clinical value if judgments made using it are shown to be reliable in clinically relevant subjects and to have validity for clinical management of patients.

Biomechanical Phenomena↗

Reliability of measurements obtained with the Timed "Up & Go" test in people with Parkinson disease.

BACKGROUND AND PURPOSE: The Timed "Up & Go" Test (TUG) is used to measure the ability of patients to perform sequential locomotor tasks that incorporate walking and turning. This study investigated the retest reliability, interrater reliability, and sensitivity of scores obtained with the TUG in detecting changes in mobility in subjects with idiopathic Parkinson disease (PD). SUBJECTS: The performance of 12 people with PD was compared with that of 12 age-matched comparison subjects without PD. METHODS: The subjects with PD completed 5 trials of the TUG after withdrawal of levodopa for 12 hours ("off" phase of the medication cycle) as well as an additional 5 trials 1 hour after levodopa was administered ("on" phase of the medication cycle). They were scored on the Modified Webster Scale at both sessions. The comparison subjects also performed 5 TUG trials. All trials were videotaped and timed by 2 experienced raters. The videotape was later rated by 3 experienced clinicians and 3 inexperienced clinicians. RESULTS: For the subjects with PD, within-session performance was highly consistent, with correlations (r) ranging from.80 to.98 for the "off" phase and from.73 to.99 for the "on" phase. The performance of the comparison subjects across the 5 trials was also highly consistent (r=.90-.97). Comparisons showed differences between trials 1 and 2 on the TUG for both groups. Removal of data for trial 1 (the practice trial) further enhanced retest reliability. There was close agreement in TUG scores among raters despite different levels of experience (intraclass correlation coefficient [3,1]=.87-.99). Mean TUG scores were different between the "on" and "off" phases of the levodopa cycle and between subjects with PD and comparison subjects during the "on" phase. CONCLUSION AND DISCUSSION: Retest reliability and interrater reliability of the TUG measurements were high, and the measurements reflected changes in performance according to levodopa use. The TUG can also be used to detect differences in performance between people with PD and elderly people without PD.

Activities of Daily Living↗

Validity and reliability of an Italian version of the revised Leeds disability questionnaire for patients with ankylosing spondylitis.

OBJECTIVE: The purpose of the present study was to produce an Italian version of the Revised Leeds Disability Questionnaire (LDQ) in a group of patients with ankylosing spondylitis, and to examine the psychometric properties of this version, evaluating its internal consistency, external validity and reliability. METHODS: The LDQ was administered to 60 Caucasian patients affected by ankylosing spondylitis (50 males, 10 females, mean age 46.1 +/- 14.2 yr, range 22-74, median disease duration 4.5 yr, range 1-24) together with the Italian version of the Stanford Health Assessment Questionnaire (HAQ), and anthropometric measurements. Thirty patients completed the questionnaire after a 10-day interval. Internal consistency was evaluated with Cronbach's alpha coefficient of reliability. Construct validity of the LDQ was evaluated using the correlation between the HAQ and anthropometric measurements. Test-retest reliability was assessed with the intraclass correlation coefficient. RESULTS: All patients completed the validation study. The questionnaire was internally consistent (alpha=0.90). A significant correlation was recorded between the LDQ and the HAQ score (rho=0.841, P<0.01) and the anthropometric measurements. Test-retest reliability showed a good correlation coefficient (intraclass correlation=0.97). CONCLUSION: The Italian LDQ is a valid and reliable instrument for detecting and measuring functional disability in patients with ankylosing spondylitis. Our results confirm the utility of this questionnaire as a valid and feasible functional measure for patients with ankylosing spondylitis.

Adult↗

Evaluations of aesthetic results in breast reconstruction: an analysis of reliability.

This study evaluated the reliability of three commonly used measures of aesthetic outcomes of breast surgery: a four-point ordinal scale of overall aesthetics, five four-point subscales, and a visual analogue scale. Fifty patients were randomly selected from women who underwent breast reconstruction surgery at University of Michigan hospitals between July 1989 and May 1993. Postoperative photographs of these patients were provided to three plastic surgeons, who were asked to rate the photographs using the three methods. The same process was repeated 4 weeks later. Intrarater and interrater reliability ranged from poor to good for the three methods, with the subscales showing the highest reliability. The lowest reliability occurred for those scales with the least-explicit rating criteria. Without explicit criteria, raters must develop and use their own criteria, which are likely to differ for each rater. Separating the various components of the aesthetic results of breast surgery into different subscales helps make the rating criteria more explicit. Scales with demonstrated reliability are critical for ensuring comparability of results across studies.

Breast↗

Performance assessment of community-based physicians: evaluating the reliability and validity of a tool for determining CME needs.

PURPOSE: To evaluate the reliability, validity, and feasibility of the Physician Assessment in Medical Practice (PAMP) as a means of determining the CME needs of practicing, community-based physicians. METHOD: A group of 45 randomly selected community-based physicians (19 certified family physicians and 26 general practitioners) affiliated with the Department of Family Medicine at Bruce Rappaport Faculty of Medicine, Technion Institute of Technology, Haifa, Israel, volunteered to participate in the study, conducted in 1997. All participants took a ten-station, performance-based examination designed to closely represent the physicians' work settings. At each station, a different medical problem was presented by a standardized patient. Physician-candidates' performances were assessed by physician-examiners using global ratings. The following performance domains were assessed: information gathering, diagnosis and management plan, and communication skills. A CME needs assessment score was determined for each of the participants and a CME level to meet the needs of the physician was recommended. RESULTS: Overall reliability of the examination was high (.87), with domain reliabilities ranging from.76 to.87. Reliability of the examiners' judgments of the physicians' competence was.66. All the stations' validity scores were significant, and differences in performances between family physicians and general practitioners demonstrated construct validity of the test results. Overall, the cost of running the examination was U.S. $250 per physician-candidate. CONCLUSIONS: Using the PAMP to determine CME needs of community-based physicians was found to reliable, valid, and feasible, and the cost per physician-candidate was not excessive. Performance results provided indepth information for use by both the individual physician and providers of CME programs.

Adolescent↗