Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Joint angle measurement: a comparative study of the reliability of goniometry and wire tracing for the hand.

OBJECTIVES: To compare the inter- and intra-rater reliability of goniometry and wire tracing in the assessment of finger joint angles: metacarpo-phalangeal (MCPJ), proximal (PIPJ) and distal interphalangeal joints (DIPJ). DESIGN: Twenty occupational therapists and 20 physiotherapists with a range of clinical experience were recruited from nine different centres. Using a masked goniometer and wire tracing they carried out repeated assessments of the MCPJ, PIPJ and DIPJ of a normal subject fixed in two different positions. RESULTS: The two assessment methods did not produce comparable angle measurements. Goniometry showed greater inter- and intra-rater reliability than wire tracing. Regardless of the assessment tool, the repeatability coefficient indicated that DIPJ measurement was less reliable than the other joints. Clinical and specialist experience did not affect reliability. CONCLUSION: Although both goniometry and wire tracing show limitations as reliable assessment tools, it is recommended that where possible goniometry should be used.

Equipment Design↗

The reliability and validity of knee-specific and general health instruments in assessing acute patellar dislocation outcomes.

BACKGROUND: The most reliable and valid instruments for assessing patient outcome after patellar dislocation have not been identified. HYPOTHESIS: Knee-specific and general health instruments will differ in validity and reliability for patients with patellar dislocation. STUDY DESIGN: Prospective cohort study. METHODS: Subjects consisted of 153 patients with acute patellar dislocation (110 with first-time dislocations and 43 with a history of patellofemoral subluxation or dislocation). We administered the modified International Knee Documentation Committee form, Kujala, Fulkerson, Lysholm, Tegner, Short Form 36, and Musculoskeletal Function Assessment instruments on two separate occasions (test-retest reliability). Validity was assessed by comparing scores of the two groups and by comparing scores of patients with and without recurrent subluxations/dislocations during follow-up. RESULTS: The knee-specific instruments yielded the highest test-retest reliability. The knee-specific and general health instruments identified higher disability levels in the patients with a history of patellofemoral problems than in those with first-time dislocations. The general health instruments identified higher disability levels in patients with patellar dislocation than published norms. The Fulkerson and Lysholm scales were the only instruments to differentiate between patients with and without recurrent subluxations/dislocations. CONCLUSIONS: Knee-specific scales yielded higher reliability coefficients and stronger validity than did general health instruments. Knee-specific, general health, and activity level instruments are complementary and in combination provide a more complete assessment for patients with patellar dislocation.

Adolescent↗

Reliability of stress radiography for evaluation of posterior knee laxity.

BACKGROUND: Although stress radiography has been recommended for quantifying posterior tibial displacement in knees with posterior cruciate ligament insufficiency, the intratester reliability and intertester reliability of this measurement method have not been evaluated. HYPOTHESIS: Stress radiography is a reproducible measurement method in the assessment of posterior knee laxity in patients with posterior cruciate ligament lesions. STUDY DESIGN: Cohort study (diagnosis); Level of evidence, 2. METHODS: Stress radiographs of 787 patients with suspected posterior cruciate ligament lesions taken using the Telos device were evaluated independently by 3 testers: 2 of the testers were clinically experienced in the evaluation of stress radiographs, and 1 tester was a novice tester. Change in mean, standard error of measurement with calculated confidence intervals, and intra-class correlation coefficients were determined to assess intratester and intertester reliability. RESULTS: There was no significant intratester change in mean. Intratester standard error of measurement was 1.03 mm; 95% confidence intervals were+/-2.02 mm for a single measurement and+/-2.86 mm for a change in measurement. The intratester intra-class correlation coefficient was 0.95. Intertester reliability revealed a significant change in mean between the experienced testers and the novice tester (P<.001). There was no substantial difference for the standard error of measurement of each tester. The mean intertester standard error of measurement was 1.41 mm; 95% confidence intervals were+/-2.77 mm for a single measurement and+/-3.91 mm for a change in measurement. The intertester intraclass correlation coefficient was 0.91. CONCLUSION: Stress radiography was found to be a measurement method with a useful reliability for evaluation of posterior laxity in patients with posterior cruciate ligament lesions. The reproducibility of stress radiography may be influenced by multiple variables, and standardized methods are needed to minimize measurement error.

Adolescent↗

Relative and absolute reliability of the KT-2000 arthrometer for uninjured knees. Testing at 67, 89, 134, and 178 N and manual maximum forces.

We assessed the reliability of the KT-2000 knee arthrometer at 67, 89, 134, and 178 N and at manual maximum forces on 30 college students who were free from present or previous knee injuries. Two examiners tested all subjects on two occasions. Anterior laxity (P < 0.0001) and side-to-side difference (P < 0.05) significantly increased as force increased. There was a significant difference (P < 0.0001) between testers for anterior laxity but not for side-to-side difference. We used intraclass correlation coefficients to estimate relative reliability. Anterior laxity intraclass correlation coefficients (2,1) between testers ranged from 0.81 to 0.86 and within tester correlations ranged from 0.92 to 0.95. Intraclass correlation coefficients for between testers for side-to-side differences ranged from 0.38 to 0.58 and within tester correlations ranged from 0.53 to 0.64. Subject-to-subject variability needs to be taken into account when interpreting intraclass correlation coefficient values. Our absolute reliability estimates (95% confidence intervals) were small, indicating little variability. Our data demonstrate the KT-2000 arthrometer to be reliable. Researchers should present both relative and absolute reliability estimates, although we believe absolute estimates are of greater clinical value. Side-to-side differences are better discriminators than individual absolute values. We recommend that a < 3 mm side-to-side difference be used to indicate stable knees.

Adult↗

Reliability of a utilization review instrument in a large field study.

One important question for a utilization management program is whether the utilization review instrument is consistent or stable when used on many occasions by the same abstractor (intrarater reliability) or by several abstractors (inter-rater reliability). As part of a nationwide study of inappropriate utilization of inpatient services by the Department of Veterans Affairs, we conducted a thorough investigation of the inter-rater reliability of a widely used utilization review instrument by 27 nurse abstractors. All abstractors were extensively trained, both by the developers of the instrument and by use of practice medical records. A standard protocol for resolving questions was implemented, with immediate communication of decisions to abstractors. The results of three reliability assessments, conducted immediately after formal training, after several weeks of reviewing practice records, and midway through review of the study records, demonstrated good to excellent reliability, both when comparing the nurse abstractors with a physician gold standard and among themselves. Therefore, with appropriate training and monitoring, utilization management programs in large hospitals, multihospital systems, and other health care organizations needing to examine inpatient utilization should feel confident that they can achieve reviews that would be in close agreement with physician and other nurse abstractors. Such confidence should increase the acceptability of utilization management programs.

Abstracting and Indexing↗

Reliability of drug utilization evaluation as an assessment of medication appropriateness.

OBJECTIVE: To test the reliability of drug utilization evaluation (DUE) applied to medications commonly used by the ambulatory elderly. METHODS: A DUE model was developed for four domains: (1) justification for use, (2) critical process indicators, (3) complications, and (4) clinical outcomes. DUE criteria specific to use in the elderly were developed for angiotensin-converting enzyme (ACE) inhibitors and histamine2 (H2)-antagonists, and consensus was reached by an external expert panel. After pilot testing, two clinical pharmacists independently evaluated these medications, applying the DUE criteria and rating each item as appropriate or inappropriate. Interrater and intrarater reliability was assessed by using kappa statistics. RESULTS: In a sample of 208 ambulatory elderly veterans, 42 (20.2%) were taking an ACE inhibitor and 56 (26.9%) an H2-antagonist. The interrater agreement for individual domains, represented by kappa statistics, were 0.10-0.58 and 0-0.83 for ACE inhibitors and H2-antagonists, respectively. The kappa statistic for overall agreement, which considered ratings from all criteria across all domains, was 0.24 for ACE inhibitors and 0.18 for H2-antagonists. Intrarater reliability was assessed 3 months later, and kappa statistics were 0.61-0.65 (0.49 overall) and 0-0.96 (0.81 overall) for ACE inhibitors and H2-antagonists, respectively. CONCLUSIONS: Intrarater reliability for DUE was good to excellent. However, interrater reliability exhibited only marginal reproducibility, particularly where evaluators were required to use subjective judgement (i.e., complications, clinical outcomes). DUE may not be a suitable standard for assessing medication appropriateness in ambulatory elderly patients.

Aged↗

The concept of reliability in emergency medicine.

Despite the fact that the United States boasts one of the most advanced health care systems in the world, this system is highly "unreliable" and fraught with error. This article is an introduction to the concept of "reliability" in emergency medicine. It suggests ways in which the health care system could promote increased reliability of operations and processes in the emergency department by using reliability principles and tools that have proven successful in other high-risk settings. Through comparisons to aviation and nuclear power, this article illustrates the differences in culture between emergency medicine and other high-risk organizations and points to the qualities that promote reliability. Finally, a specific model for reliability in the emergency department, operations, and clinical processes is proposed.

Efficiency, Organizational↗

Roger A. Mann Award . The reliability of angular measurements in hallux valgus deformities. .

The purpose of this study was to determine the intraobserver and inter-observer reliability of physicians on a repetitive basis in making angular measurements of hallux valgus deformities. The hallux valgus angle, the 1-2 intermetatarsal angle, and the distal metatarsal articular angle and the assessment of congruency/subluxation of the first MTP joint were evaluated on a repetitive basis. Physicians were provided with a series of black and white photographs of radiographs with a hallux valgus deformity. Three different sets of photographs randomly ordered were sent at a minimum interval of six weeks to the participants. Participating physicians were extremely reliable in the measurement of the 1-2 metatarsal angle. 96.7% of the photographs were repeatedly measured within a range of 5 degrees or less. The angular measurements to determine the hallux valgus angle were slightly less reliable, but 86.2% of photos were repeatedly measured within a range of 5 degrees or less. In the measurement of the distal metatarsal articular angle, 58.9% of photographs were repeatedly measured within a range of 5 degrees or less. There was a wide range within physician evaluators who recognized very few congruent joints (2 of 21) and those who recognized several congruent joints (11 of 21). Most physicians appeared to be internally consistent in the assessment of MTP congruency; however, some photographs were much more difficult to assess than others. This study validates the reliability of the measurement of the hallux valgus and the 1-2 metatarsal angle. The interobserver reliability in the measurement of the distal metatarsal articular angle is questioned.

Adult↗

A meta-analysis of outcome rating scales in foot and ankle surgery: is there a valid, reliable, and responsive system?

BACKGROUND: Rating scales that are valid, reliable, and responsive communicate the severity of a functional problem, facilitate the accurate study of treatment modalities, and provide a common language for those involved in research. The purpose of this study was to determine which outcome rating scales are currently used in the foot and ankle literature and to identify rating scales with proven reliability, validity, and responsiveness. METHOD: A meta-analysis of the foot and ankle literature from 1990 to 2001 was done. All referenced rating scales were reviewed to determine if any had proven to be reliable, valid, or responsive. RESULTS: Forty-nine different rating scales were identified. The most frequently referenced scales were the subscales of the American Orthopaedic Foot and Ankle Society (AOFAS). No rating scale was identified that demonstrated reliability, validity, and responsiveness in patients with a variety of foot and ankle conditions. CONCLUSIONS: The development of a reliable, valid, and responsive rating scale would have value not only in assessing patient outcomes but also in reporting the results of clinical studies in foot and ankle surgery.

Ankle↗

Clinicians' assessment of the hindfoot: a study of reliability.

BACKGROUND: Static measurements of position of the hindfoot and clinical assessment of motion of the hindfoot often are used in the assessment of foot function and manufacturing of orthoses. However, the reliability and validity of static measurements and dynamic observation and assessment of the hindfoot are controversial. The purpose of this investigation was to examine reliability of static and dynamic assessments of the hindfoot in a setting that reproduced clinical conditions. METHODS: Twenty-four healthy participants were evaluated by four experienced clinicians for four commonly used static measurements and dynamic assessment of hindfoot function. The protocol was repeated 2 weeks later. RESULTS: Results indicated that reliability of results, both intertester and from test to test were poor to fair for static measurements of the hindfoot (r = 0.075 to r = 0.755, p < 0.05). The error estimates associated with these measures were high; subtalar neutral position and resting calcaneal stance position both demonstrated measurement errors of more than 4 degrees (95% confidence intervals-4.1 degrees and 5.1 degrees, respectively). Retest reliability of dynamic assessments were considered reasonable for only one clinician (kappa = 0.55). Intertester agreement was poor among all clinicians. CONCLUSION: Clinicians taking static measurements demonstrated large errors that do not reflect the precision that has been assumed in clinical theory using these measurements. The availability of static assessments did not improve dynamic assessment. This poor reliability calls into question the importance placed on static and dynamic measurements of the hindfoot in clinical decision-making.

Ankle↗

Reliability of an in-shoe pressure measurement system during treadmill walking.

We examined the reliability of in-shoe foot pressure measurement using the Pedar in-shoe pressure measurement system for 25 participants walking at treadmill speeds of 0.89, 1.12, and 1.34 meters/sec. The measurement system uses EMED insoles, which consist of 99 capacitive sensors, sampled at 50 Hz. Data were collected for 20 seconds at two separate times while participants walked at each gait speed. Differences in some of the loading variables across speed relative to the total foot and across the different anatomical regions were detected. Different anatomical regions of the foot were loaded differently with variations in walking speed. The results indicated the need to control speed when evaluating loading parameters using in-shoe pressure measurement techniques. Coefficients of reliability were calculated. Variables such as peak force for the total foot required two steps to achieve a coefficient of reliability of 0.98. To achieve excellent reliability (> 0.90) in the peak force, force time integral, peak pressure, and pressure time integral across the total foot and the seven regions, a maximum of eight steps was needed. In general, timing variables, such as the instant of peak force and the instant of peak pressure, tended to be the least reliable measures.

Adult↗

Reliability of the WAIS-III subtests, indexes, and IQs in individuals with substance abuse disorders.

Reliability of the WAIS-III for 100 male patients with substance abuse disorders was determined. Means for age and education were 46.06 were (SD = 8.81 years) and 12.70 years (SD = 1.51 years). There were 63 Caucasians and 37 African Americans. Split-half coefficients for the 11 subtests (Digit Symbol-Coding, Symbol Search, and Object Assembly were omitted) ranged from .92 for Vocabulary and Digit Span to .77 for Picture Arrangement. The median subtest reliability coefficient was .86. Composite reliabilities were excellent for the Indexes (.94 to .95) and IQs (.94 to .97), with all coefficients > or = .94. Using the Fisher z test to compare correlation coefficients from independent samples, none of the reliability estimates differed significantly from those reported for the WAIS-III standardization sample. Similar findings emerged when reliabilities were determined separately for Caucasian and African American participants.

Black or African American↗

Measuring deception: test-retest reliability of physicians' self-reported manipulation of reimbursement rules for patients.

This study examined the test-retest reliability of physicians' self-reported manipulation of reimbursement rules for patients. The test-retest reliability of self-report of three specific tactics were examined: (1) exaggerating the severity of patients' conditions, (2) changing a patient's official (billing) diagnosis, and (3) reporting signs or symptoms that patients did not have. The reliability of a scaled summary measure of physicians' manipulation of reimbursement rules was also assessed. Overall, the authors found high levels of test-retest agreement across all three items and the summary measure. These findings suggest that self-report can be used to produce reliable data on this controversial issue. Specifically, the three items reported here can be used to produce a reliable summary measure of physicians' manipulation of reimbursement rules to help patients obtain care that physicians perceive as necessary.

Deception↗

MS-CANE: a computer-aided instrument for neurological evaluation of patients with multiple sclerosis: enhanced reliability of expanded disability status scale (EDSS) assessment.

Kurtzke's EDSS remains the most widely-used measure for clinical evaluation of MS patients. However, several studies have demonstrated the limited reliability of this tool. We introduce a computerized instrument, MS-CANE (Multiple Sclerosis Computer-Aided Neurological Examination), for clinical evaluation and follow up of patients with multiple sclerosis (MS) and to compare its reliability to that of conventional Expanded Disability Status Scale (EDSS) assessment. We developed a computerized interactive instrument, based on the following principles: structured gathering of neurological findings, reduction of compound notions to their basic components, use of precise definitions, priority setting and automated calculations of EDSS and functional systems scores. An expert panel examined the consistency of MS-CANE with Kurtzke's specifications. To determine the effect of MS-CANE on the reliability of EDSS assessment, 56 MS patients underwent paired conventional EDSS and MS-CANE-based evaluations. The inter-observer agreement in both methods was determined and compared using the kappa statistic. The expert panel judged the tool to be compatible with the basic concepts of Kurtzke's EDSS. The use of MS-CANE increased the reliability of EDSS assessment: Kappa statistic was found to be 0.42 (i.e. moderate agreement) for conventional EDSS assessment versus 0.69 (i.e. substantial agreement) for MS-CANE (P=0.002). We conclude that the use of this tool may contribute towards a standardized and reliable assessment of EDSS. Within clinical trials, this could increase the power to detect effects, thus reducing trial duration and the cohort size required. Multiple Sclerosis (2000) 6 355 - 361

Algorithms↗

Reliability and validity assessments of measures of social networks, social support and control--results from the Malmö Shoulder and Neck Study.

The reliability and validity of methods to assess social networks, social support and control were investigated in a population of 12,009 females and males born between 1926 and 1945 (the "Malmö Shoulder and Neck Study"). This study demonstrated an overall reliability with kappa coefficients between 0.70 and 0.47, but the reliability was more varying among females and lower in the youngest age group. The analysis of the construct validity indicated that the different indices measure different aspects of the psychosocial environment, but both theoretical and methodological problems were identified, when the validity of multidimensional concepts are to be determined. The validity of such indices can best be judged by combining quantitative and qualitative methods. Potential validity problems must be kept in mind when these indices are used in epidemiological research. The results from the reliability analysis call for repeated assessments and the sample size must be adjusted vis-a-vis the reliability.

Aged↗

The reliability of the ankle-brachial index in the Atherosclerosis Risk in Communities (ARIC) study and the NHLBI Family Heart Study (FHS).

BACKGROUND: A low ankle-brachial index (ABI) is associated with increased risk of coronary heart disease, stroke, and death. Regression model parameter estimates may be biased due to measurement error when the ABI is included as a predictor in regression models, but may be corrected if the reliability coefficient, R, is known. The R for the ABI computed from DINAMAP readings of the ankle and brachial SBP is not known. METHODS: A total of 119 participants in both the Atherosclerosis Risk in Communities (ARIC) study and the NHLBI Family Heart Study (FHS) had repeat ABIs taken within 1 year, using a common protocol, automated oscillometric blood pressure measurement devices, and technician pool. RESULTS: The estimated reliability coefficient for the ankle systolic blood pressure (SBP) was 0.68 (95% CI: 0.57, 0.77) and for the brachial SBP was 0.74 (95% CI: 0.62, 0.83). The reliability for the ABI based on single ankle and arm SBPs was 0.61 (95% CI: 0.50, 0.70) and the reliability of the ABI computed as the ratio of the average of two ankle SBPs to two arm SBPs was estimated from simulated data as 0.70. CONCLUSION: These reliability estimates may be used to obtain unbiased parameter estimates if the ABI is included in regression models. Our results suggest the need for repeated measures of the ABI in clinical practice, preferably within visits and also over time, before diagnosing peripheral artery disease and before making therapeutic decisions.

Ankle↗

Balance in single-limb stance in healthy subjects--reliability of testing procedure and the effect of short-duration sub-maximal cycling.

BACKGROUND: To assess balance in single-limb stance, center of pressure movements can be registered by stabilometry with force platforms. This can be used for evaluation of injuries to the lower extremities. It is important to ensure that the assessment tools we use in the clinical setting and in research have minimal measurement error. Previous studies have shown that the ability to maintain standing balance is decreased by fatiguing exercise. There is, however, a need for further studies regarding possible effects of general exercise on balance in single-limb stance. The aims of this study were: 1) to assess the test-retest reliability of balance variables measured in single-limb stance on a force platform, and 2) to study the effect of exercise on balance in single-limb stance, in healthy subjects. METHODS: Forty-two individuals were examined for test-retest reliability, and 24 individuals were tested before (pre-exercise) and after (post-exercise) short-duration, sub-maximal cycling. Amplitude and average speed of center of pressure movements were registered in the frontal and sagittal planes. Mean difference between test and retest with 95% confidence interval, the intraclass correlation coefficient, and the Bland and Altman graphs with limits of agreement, were used as statistical methods for assessing test-retest reliability. The paired t-test was used for comparisons between pre- and post-exercise measurements. RESULTS: No difference was found between test and retest. The intraclass correlation coefficients ranged from 0.79 to 0.95 in all stabilometric variables except one. The limits of agreement revealed that small changes in an individual's performance cannot be detected. Higher values were found after cycling in three of the eight stabilometric variables. CONCLUSIONS: The absence of systematic variation and the high ICC values, indicate that the test is reliable for distinguishing among groups of subjects. However, relatively large differences in an individual's balance performance would be required to confidently state that a change is real. The higher values found after cycling, indicate compensatory mechanisms intended to maintain balance, or a decreased ability to maintain balance. It is recommended that average speed and DEV 10; the variables showing the best reliability and effects of exercise, be used in future studies.

Adult↗

The use of the SF-36 questionnaire in adult survivors of childhood cancer: evaluation of data quality, score reliability, and scaling assumptions.

BACKGROUND: The SF-36 has been used in a number of previous studies that have investigated the health status of childhood cancer survivors, but it never has been evaluated regarding data quality, scaling assumptions, and reliability in this population. As health status among childhood cancer survivors is being increasingly investigated, it is important that the measurement instruments are reliable, validated and appropriate for use in this population. The aim of this paper was to determine whether the SF-36 questionnaire is a valid and reliable instrument in assessing self-perceived health status of adult survivors of childhood cancer. METHODS: We examined the SF-36 to see how it performed with respect to (1) data completeness, (2) distribution of the scale scores, (3) item-internal consistency, (4) item-discriminant validity, (5) internal consistency, and (6) scaling assumptions. For this investigation we used SF-36 data from a population-based study of 10,189 adult survivors of childhood cancer. RESULTS: Overall, missing values ranged per item from 0.5 to 2.9 percent. Ceiling effects were found to be highest in the role limitation-physical (76.7%) and role limitation-emotional (76.5%) scales. All correlations between items and their hypothesised scales exceeded the suggested standard of 0.40 for satisfactory item-consistency. Across all scales, the Cronbach's alpha coefficient of reliability was found to be higher than the suggested value of 0.70. Consistent across all cancer groups, the physical health related scale scores correlated strongly with the Physical Component Summary (PCS) scale scores and weakly with the Mental Component Summary (MCS) scale scores. Also, the mental health and role limitation-emotional scales correlated strongly with the MCS scale score and weakly with the PCS scale score. Moderate to strong correlations with both summary scores were found for the general health perception, energy/vitality, and social functioning scales. CONCLUSION: The findings presented in this paper provide support for the validity and reliability of the SF-36 when used in long-term survivors of childhood cancer. These findings should encourage other researchers and health care practitioners to use the SF-36 when assessing health status in this population, although it should be recognised that ceiling effects can occur.

Adolescent↗