Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Reliability of single-item ratings of quality in higher education: a replication.

Single-item ratings of the quality of instructors or subjects are widely used by higher education institutions, yet such ratings are commonly assumed to have inadequate psychometric properties. Recent research has demonstrated that reliability of such ratings can indeed be estimated, using either the correction for attenuation formula or factor analytic methods. This study replicates prior research on the reliability of single-item ratings of quality of instruction, using a different, more student-focussed approach to teaching and learning evaluation than used by previous researchers. Class average data from 1,097 classes, representing responses from 59,815 students, were analysed. At the "class" level of analysis, both methods of estimation suggested the single item of quality had high reliability: .96 using the correction for attenuation formula, and .94 using the factor analytic method. An alternative method of calculating reliability, which takes into account the hierarchical nature of the data, likewise suggested high estimated reliability (.92) of the single-item rating. These results indicate the suitability of the overall class rating for quality improvement in higher education, with a large sample.

Education↗

Reliability of measurements obtained by use of an instrument designed to indirectly measure iliotibial band length.

Ober's test is widely used to indirectly assess iliotibial band length on patients with painful conditions of the lower extremity. The validity of Ober's test is questioned on several accounts: there is no method for standardizing the position of the pelvis during the measurement; the scale used to describe the results is nominal and appears to have been arbitrarily determined; and, finally, the reliability of judgments made using Ober's test has not been reported. The purposes of this study were to develop a method for quantifying an indirect measurement of iliotibial band length and to examine the intratester and intertester reliability of measurements obtained. Iliotibial band lengths of 10 patients (N = 10) with anterior knee pain were measured twice by each of two examiners. Data were analyzed using the intraclass correlation coefficient (ICC) and standard error of the measurement (SEM). The ICC values were 0.94 and 0.73 for intratester and intertester reliability, respectively. The SEM values were 1 degree and 2 degrees for intratester and intertester reliability, respectively. Measurements obtained with the modified method are reliable when taken on young patients with anterior knee pain.

Adolescent↗

Reliability of the figure-of-eight method of ankle measurement.

Physical therapists need a reliable method by which to measure ankle girth following injury so that there can be clinical quantification of the volume of edema. The purpose of this study was to examine the intratester and intertester reliability of the ankle figure-of-eight method for the measurement of ankle size. Fifty healthy subjects were positioned on a plinth in a long sitting position. Four measurements were made by each of the three testers for a total of 12 measurements per subject. The intraclass correlation coefficient was 0.99 for intertester reliability and 0.99 for intratester reliability. These results support the use of the ankle figure-of-eight method as a reliable tool for measuring ankle girth.

Ankle Injuries↗

Investigation of the validity and reliability of four objective techniques for measuring forward shoulder posture.

Clinicians often rely on visual inspection and descriptive terms to documents a patient's forward shoulder posture. The purpose of this study was to assess the validity and intrarater reliability of four objective techniques to measure forward shoulder posture. Subjects were 25 males and 24 females. Subjects had a lateral cervical spine radiograph taken, from which the horizontal distance from the C7 spinous process to the anterior tip of the left anterior acromion process was measured. Subjects then proceeded twice through a random order of four measurements: the Baylor square, the double square, the Sahrmann technique, and scapular position. These results were then used to determine the intrarater reliability of each technique. Multiple regression analyses were performed on each measure's mean scores to determine both the correlation with and the predictive value for the radiographic measurement. The intraclass correlation coefficients for intrarater reliability ranged from .89 to .91. The correlation coefficients ranged from -.33 to .77, and the coefficients of determination ranged from .10 to .59 (N = 49). The researchers demonstrated clinical reliability for each technique; however, validity compared with the radiographic measurement could not be established. These techniques may have clinical value in objectively measuring change in a patient's shoulder posture as a result of a treatment program. Before any of these measures could be universally recommended in clinical practice, future research is necessary to establish interrater reliability and assess each technique's ability to detect postural changes over time.

Adult↗

Reliability of lower extremity functional performance tests.

Clinicians routinely have used functional performance tests as an evaluation tool in deciding when an athlete can safely return to unrestricted sporting activities. These practitioners assumed that these tests provide a reliable measure of lower extremity performance; however, little research has been reported on the reliability of these measures. The purpose of this investigation was to determine the reliability of lower extremity functional performance tests. Five male and 15 female volunteers were evaluated using the single hop for distance, triple hop for distance, 6-m timed hop, and cross-over hop for distance as described by Noyes (10). One clinician measured each subject's performance using a standardized protocol and retested subjects in the same manner approximately 48 hours later. The order of testing was randomly determined. Subjects' average and individual scores on each functional performance test were used for statistical analysis. Intraclass correlation coefficients (ICCs) and standard error of measurement (SEM) values based on average day 1 and day 2 scores were used to estimate the reliability of each functional performance test. Intraclass correlation coefficients were .96, .95, and .96, and SEMs were 4.56 cm, 15.44 cm, and 15.95 cm, respectively, for the single hop, triple hop, and cross-over hop for distance tests. An ICC of .66 and SEM of .13 seconds for the 6-m timed hop resulted from limited variability between measurements; however, its small SEM value inferred that the inconsistency of measurement would occur in an acceptably small range. A repeated measures analysis of variance revealed no significant difference ( p > .05) between individual trial scores except for the single hop for distance. We concluded that this difference represented a learning effect not found with the other tests. The results of this investigation demonstrate that clinicians can use functional performance testing to obtain reliable measures of lower extremity performance when using a standardized protocol.

Adult↗

Intrarater reliability of selected clinical outcome measures following anterior cruciate ligament reconstruction.

STUDY DESIGN: Single group repeated measures following anterior cruciate ligament (ACL) reconstruction. OBJECTIVES: The purpose of this study was to evaluate the intrarater reliability of selected clinical outcome measures in patients having ACL reconstruction. BACKGROUND: Several investigations have reported the reliability of isokinetic testing and knee ligament arthrometry. Fewer studies have examined the reliability of lower extremity functional tests, with most of these studies evaluating normal subjects. METHODS AND MEASURES: Fifteen physically active males with unilateral ACL-reconstructed knees were evaluated with the KT-1000, Biodex isokinetic dynamometer, and 3 functional hop tests on 5 occasions. RESULTS: Intraclass correlation coefficients (ICCs) revealed good to high intrarater reliability (ICC > 0.80) of the functional hop tests and isokinetic peak torque values ICCs were higher for the involved limb than the uninvolved limb using the scores from the KT-1000 Manual Maximum Test. CONCLUSIONS: The outcome measures examined in this investigation have been shown to be reliable in patients with ACL reconstructions, and support previous investigations in nonimpaired populations. Further research is needed to examine the validity of these postoperative outcome measures in patients with ACL reconstructions.

Adolescent↗

Reliability by surgical status of self-reported outcomes in patients who have shoulder pathologies.

STUDY DESIGN: A test-retest design was used to evaluate the reliability of the self-report sections of 4 shoulder pain and disability scales. OBJECTIVE: The objective of the study was to compare interitem consistency and test-retest reliability by surgical status (postoperative versus nonoperative) and to evaluate the effect of surgical status in the prediction of retest scores. BACKGROUND: Patients and healthcare providers evaluate shoulder status based on self-evaluations of pain and disability. Shoulder outcome measures have been developed that include self-reports, but the properties of these measures have not been assessed by surgical status. METHODS AND MEASURES: A questionnaire containing self-report sections of 4 shoulder scales was administered to study participants twice with 1 week between administrations. The outcome measures examined were the: (1) University of California at Los Angeles (UCLA) Shoulder Score; (2) Constant-Murley Scale (CMS); (3) American Shoulder and Elbow Society (ASES) Shoulder Index; and (4) Shoulder Pain and Disability Index (SPADI). Intraclass correlation coefficients (ICC) were calculated to estimate the test-retest reliability of each of the scales and subscales. The interitem consistencies of the multi-item subscales were assessed using Cronbach's alpha. The effect of surgical status on shoulder outcome scale reliability was evaluated using a general linear models approach. RESULTS: The interitem consistency estimates for the multi-item scales were high with both operative and nonoperative participants (0.88 to 0.96). With the exception of the satisfaction subscale of the UCLA Shoulder Score for the nonsurgical group, the estimated intraclass coefficients ranged from 0.51 to 0.91. The prediction of UCLA-satisfaction and ASES-disability, pain, and total retest scores was improved with the addition of surgical status into a regression model. CONCLUSIONS: The examined scales exhibited good internal consistency across surgical status. The postsurgical sample's reproducibility estimates tended to be higher than those of the nonsurgical sample. Reliability of shoulder outcome scales can be affected by patient surgical status.

Adult↗

Reliability and concurrent validity of the figure-of-eight method of measuring hand size in patients with hand pathology.

STUDY DESIGN: Methodological study using correlational methods. OBJECTIVE: To determine the intratester and intertester reliability and concurrent validity of the figure-of-eight method of measuring hand size in patients with hand pathology. BACKGROUND: Measuring edema is an important component of the physical examination of patients with conditions affecting the hand. The figure-of-eight method of measuring hand size has been suggested as an alternative to volumetry. The reliability and concurrent validity of the figure-of-eight method has been established in individuals without hand pathology, but not in patients with conditions involving the hand. METHODS AND MEASURES: Participants were 24 patients with conditions affecting the hand, 9 with bilateral involvement. Two testers performed 3 figure-of-eight measurements of hand size each. A third tester performed 2 volumetric measurements. Intraclass correlation coefficient (ICC3,1) was used to determine intratester reliability of both measurement procedures. ICC2,3 was used to examine intertester reliability of the figure-of-eight method. Pearson product moment correlation coefficients examining the association between the 2 methods were used to establish concurrent validity of the figure-of-eight technique. RESULTS: Intratester ICCs for figure-of-eight and volumetric methods were 0.98 to 0.99. The intertester ICC for the figure-of-eight method was 0.99. Pearson correlation coefficients examining the relationship between the 2 methods were 0.92 to 0.94. CONCLUSION: The figure-of-eight method is a reliable and valid measure of hand size in individuals with conditions affecting the hand.

Adult↗

Reliability and responsiveness of the lower extremity functional scale and the anterior knee pain scale in patients with anterior knee pain.

STUDY DESIGN: Prospective methodological study of repeated measures using a sample of consecutive patients. OBJECTIVE: To determine the test-retest reliability and responsiveness of the Anterior Knee Pain Scale (AKPS) and the Lower Extremity Functional Scale (LEFS) in patients with anterior knee pain. BACKGROUND: Anterior knee pain is one of the most common orthopedic complaints affecting the knee. Yet there is currently no self-report outcome measure that has well-established reliability and responsiveness, specifically for this population. As a result, clinicians and researchers may be making inappropriate conclusions regarding patient outcomes by using questionnaires that are misleading. METHODS AND MEASURES: This multisite study involved 30 patients from 4 outpatient physical therapy clinics in Dallas, TX (24 women, 6 men; age range, 16-50 years; mean+/-SD age, 35.2+/-9.1 years). Patients receiving physical therapy for a chief complaint of anterior knee pain completed the AKPS and LEFS at their initial appointment and again 2 to 3 days later. Upon completion of physical therapy, the patients completed the AKPS, LEFS, and a global rating of change form. The treating therapist also completed a global rating of change form at the patient's final visit. The mean of the patient's and therapist's global rating of change was used as the criterion measure of change. RESULTS: Test-retest reliability was high for both questionnaires (ICC2,1 = 0.95 for the AKPS and 0.98 for the LEFS). A significant correlation was found between the criterion measure of change and both questionnaires. Receiver-operating characteristic curve analysis revealed that both questionnaires were moderately responsive with the area under the curve slightly higher for the LEFS (0.77) than the AKPS (0.69). CONCLUSION: The LEFS and the AKPS both demonstrated high test-retest reliability and appear to be moderately responsive to clinical change in patients with anterior knee pain. Reliability and responsiveness were slightly higher in the LEFS than the AKPS. Further research is needed to determine if these measures could be modified, or new measures created, to produce an even more sensitive tool for this population.

Adolescent↗

Reliability of function-related tests in patients with shoulder pathologies.

STUDY DESIGN: Nonexperimental. OBJECTIVE: To investigate the intertester and intratester reliability of a battery of function-related tests in patients with shoulder pathologies and associated reduced range of motion. BACKGROUND: A battery of function-related tests has the potential to complement assessment of functional limitation in patients who have shoulder pathologies. METHODS AND MEASURES: Three function-related tests (hand to neck, hand to scapula, and hand to opposite scapula) were conducted on 46 patients with shoulder pathologies, and 46 age- and gender-matched control subjects. The tests were performed by 2 independent physiotherapists to test intertester reliability. Intratester reliability was examined by investigating the reproducibility of the tests performed twice, with 3 to 5 days between tests, by the same physiotherapist. Comparison of the scores on the function-related tests between patients and controls was evaluated. A correlation matrix was calculated to test the level of association among the tests. RESULTS: Intratester and intertester reliability on the 3 tests (weighted K) varied from 0.83 to 0.90. The patient's test performances were decreased in comparison to the control group. The correlation matrix demonstrated a level of associations among the 3 tests varying from r = 0.64 to r = 0.66. CONCLUSION: The results of this study indicate that function-related tests are reliable and could be used in clinical practice to document reduced function of the shoulder. The level of association among the tests indicates that each test measured different aspects of shoulder function.

Adult↗

Concurrent criterion-related validity and reliability of a clinical device used to assess lateral patellar displacement.

STUDY DESIGN: Repeated-measures, within-subject design. OBJECTIVE: To assess the concurrent criterion-related validity and reliability of a clinical device to quantify lateral patellar displacement. BACKGROUND: Excessive lateral displacement of the patella is an impairment that is widely associated with patellofemoral pain and/or pathology. Currently, no valid or reliable clinical method to assess lateral patellar displacement has been described in the literature. METHODS AND MEASURES: A total of 26 individuals (14 asymptomatic and 12 symptomatic; mean +/- SD age, 27 +/- 4 years) participated in the validity portion of this study, while an additional 10 asymptomatic volunteers (mean +/- SD age, 28 +/- 5 years) participated in the reliability portion. Lateral displacement of the patella was assessed using a custom-designed patellofemoral arthrometer (PFA) and was compared to actual position of the patella as determined by magnetic resonance imaging (MRI). Both PFA and MRI measurements of lateral patellar displacement were made with the knee extended and the quadriceps contracted. The intraclass correlation coefficient (ICC) was used to assess the level of agreement between the PFA and MRI measurements, as well as the intrarater and interrater reliability of the PFA measurements. RESULTS: The ICC assessing the level of agreement between the MRI and PFA measures of lateral patellar displacement was good (0.86). Excellent intratester (ICC, 0.96 and 0.97) and intertester reliability (ICC, 0.92) were demonstrated. CONCLUSION: Our results suggest that reasonable estimations of lateral patellar displacement can be obtained using the PFA.

Adult↗

The Penn shoulder score: reliability and validity.

STUDY DESIGN: Psychometric evaluation of a cross-sectional survey. OBJECTIVES: The purpose of this study was to examine the psychometric properties of reliability and validity of the Penn Shoulder Score (PSS). BACKGROUND: Shoulder outcome measures are used to assess patient self-report levels of pain, satisfaction, and function. The PSS is a 100-point shoulder-specific self-report questionnaire consisting of 3 subscales of pain, satisfaction, and function. This scale has been utilized in the literature. However, the measurement properties of reliability and validity, including responsiveness, of the PSS subscales and overall scale need to be established. METHODS AND MEASURES: Patients (n = 40) with shoulder disorders undergoing a course of outpatient physical therapy completed the PSS at initial visit and again within 72 hours to assess test-retest reliability. The Constant Shoulder Score (CSS) and the American Shoulder and Elbow Surgeons Shoulder Score (ASES) were also completed at the initial visit and compared to the PSS to assess convergent construct validity. A separate cohort of patients (n = 109) completed the PSS at initial visit and 4 weeks later. These scores were used to assess internal consistency and responsiveness. RESULTS: Reliability analysis revealed a test-retest ICC2,1 of 0.94 (95% CI, 0.89-0.97). Internal consistency analysis revealed a Cronbach alpha of 0.93. The standard error of measurement (SEM) was +/- 8.5 scale points (based on a 90% CI) and the minimal detectable change (MDC) was +/- 12.1 scale points (based on a 90% CI). The minimal clinically important difference (MCID) for improvement was 11.4 points. Pearson product moment correlation coefficients between the PSS and the CSS and ASES were 0.85 and 0.87, respectively. Responsiveness analysis revealed an effect size of 1.01 and a standardized response mean of 1.27. CONCLUSIONS: This study has demonstrated that the PSS is a reliable and valid measure for reporting outcome of patients with various shoulder disorders.

Adult↗

Cobalamin absorption determined by the stool spot test. Reliability in patients with uremia or disorders of the ileum.

A cobalamin absorption test, the stool spot test (SST), which makes use of radioactive cobalamin and a nonabsorbable isotope, 51Cr-trichloride, has been shown to produce reliable results in patients with pernicious anemia and in healthy controls. The reliability of the SST in patients with bowel disorders and in patients with decreased renal function was investigated by comparing with both whole-body counting and the Schilling test. Fourteen patients with bowel disorders and eight patients with uremia joined the trial. The SST correlated highly significantly with the whole-body counting method. However, the precision of the SST was poor in patients with decreased bowel transit time and inferior to that in the uremic patients. In one of two patients with decreased bowel transit time the two isotopes were shown to have different transit times, thus invalidating the test. In patients with uremia the SST was significantly more reliable than the Schilling test. It is concluded that the SST is reliable also in patients with uremia but may not be reliable in patients with intestinal disorders and decreased bowel transit time. In these patients collection of larger stool samples is recommended.

Feces↗

Diagnosing comorbidity in substance abusers: a comparison of the test-retest reliability of two interviews.

This study examines the test-retest reliability of two interview schedules (computer- and clinician-administered) in diagnosing lifetime comorbidity in treated substance abusers. The Computerized Diagnostic Interview Schedule (C-DIS) and the Structured Clinical Interview for DSM-III-R (SCID) were both administered to 173 substance abusers after random assignment to one of two groups. Within 1 to 2 weeks, subjects in the first group repeated the C-DIS and subjects in the second group were reinterviewed by a different clinician, blind to the results of the initial SCID. Both instruments showed good to excellent reliability for DSM-III-R psychoactive substance use disorders with kappas ranging from .50 to .89 for individual disorders. However, the reliability of comorbid other mental disorders was substantially poorer on both instruments, particularly the SCID. C-DIS kappas ranged from -.05 for generalized anxiety to .70 for simple phobia. SCID kappas ranged from .31 for panic disorder to .83 for antisocial personality disorder. Anxiety disorders as a category, some phobic disorders, and antisocial personality disorder showed acceptable levels of test-retest reliability on both instruments. There was a trend for borderline or threshold cases to account for some of the disagreement on the C-DIS. Differences of opinion between clinicians on organicity accounted for some of the disagreements on panic disorder and major depression. The C-DIS, unlike the SCID, tended to diagnose more disorders at initial interview, perhaps a result of its tedious probe structure. Neither instrument should be administered only once to provide a reliable lifetime diagnostic profile of comorbidity in substance abusers.

Adolescent↗

The validity and reliability of an asthma knowledge questionnaire used in the evaluation of a group asthma education self-management program for adults with asthma.

This paper reports aspects of the validity and reliability of an asthma general knowledge questionnaire for adults (AGKQA), developed as one of the outcome measures for a randomized controlled effectiveness trial of an asthma education program for adults with asthma. It also illustrates how study data can provide a valuable, generally neglected opportunity to assess the validity and reliability of a measure where resources are not available for the extensive investigation of these properties prior to administration of the measure in a study. The AGKQA was demonstrated to have good content and face validity. Construct validity assessed using the principal components method for factor analysis suggested the scale was unidimensional. Criterion-related validity, assessed using the contrasted groups method, demonstrated a significant difference (p < 0.0001) in total score and for 68% of item responses for the adults with and without direct experience of asthma. The Kuder-Richardson 20 reliability coefficient for internal consistency calculated using responses at baseline, immediately, and 12 months post-intervention were, respectively, 0.56, 0.80, and 0.75, indicating excellent reliability. The AGKQA is an acceptably valid and reliable measure for the assessment of program content mastery that it was designed to test.

Adolescent↗

Methods for measuring maximal isometric grip strength during short and sustained contractions, including intra-rater reliability.

The purposes of this study were to develop methods for measuring maximal isometric grip strength during short and sustained contractions in a laboratory setting, and to evaluate the test-retest reliability of these methods in short- and long-term perspectives. Eleven healthy men and women were assessed on four occasions. Maximal voluntary isometric grip strength (MVC) was measured in standardized and optional positions, and sustained maximal isometric strength (SMVC) in the standardized position. The results indicated that three trials in a session might be insufficient to obtain a true measure of MVC. The within-session and test-retest reliability of the described multi-trial procedure was considered satisfactory. The mean score of the last three trials tended to show the highest short-term and long-term variability. There were no clear differences between scores obtained in standardized and optional positions. The standardized position seemed more consistently to yield higher test-retest reliability and lower variability over time. The described method for measuring SMVC, expressed as area and peak score, had high test-retest reliability and an acceptable degree of short-term and long-term variability. The time taken to reach the peak score was not a reliable measure.

Adult↗

Inter-rater and intra-rater reliability of disability ratings based on the modified D Code of the ICIDH.

Since application of the International Classification of Impairments, Disabilities, and Handicaps (ICIDH) in its full form proved to be impractical, the use of selected parts, or of instruments based on the ICIDH, has been suggested. We have assessed the inter- and intra-rater reliability of disability ratings using a screening instrument, based on the D Code of the ICIDH. The level of functional abilities of all new patients (n = 39) seen by a resident at our outpatient clinic in an 11-week period was rated on a four-point scale for each of 28 items in five ability categories. Independent ratings were done by a resident and his supervisor (inter-rater reliability). Repeated rating by the resident was used to assess intra-rater reliability. Calculation of the inter- and intra-rater reliability was based both on the method of the relative agreement and on the measure kappa. The results showed relative agreements of greater than or equal to 82%, and kappa values of greater than or equal to 0.71 were found. This study shows that in a teaching hospital very satisfactory inter-rater and intra-rater reliabilities using a short instrument based on the D Code of the ICIDH can be achieved. It can therefore be recommended as a method for global disability rating.

Disability Evaluation↗

The reliability of the items of the Functional Assessment Measure (FAM): differences in abstractness between FAM items.

The reliability of the Functional Assessment Measure (FIM+FAM) is an important issue with its increased use in the measurement of neurological disability and rehabilitation outcome. Although the Motor items have good reliability ratings, the Cognitive items are more difficult to complete and their reliability is not as good. This study tests the suggestion that this might be due to the Cognitive items being more abstract. A keyword from each of four Motor items was compared with a keyword from four Cognitive items. Abstractness was measured by measuring the 'imageability' of each keyword. The Motor items were found to have a significantly higher mean imageability rating than the Cognitive items. Thus, there is support for the suggestion that abstractness contributes to the poorer reliability of the Cognitive items. These results led to the proposal that the reliability of the Cognitive items might be improved by various methods of increasing the tangibility of these measures (e.g. subdivision of broad categories of disabilities, enhancing item descriptions, training raters to increase their recognition of relevant observations, and using specific assessment tasks to elicit relevant behaviours).

Activities of Daily Living↗