Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

The validity and reliability of the graphic rating scale and verbal rating scale for measuring pain across cultures: a study in Egyptian and Dutch women with rheumatoid arthritis.

OBJECTIVE: To compare the validity and reliability of a graphic rating scale (GRS) and a verbal rating scale (VRS) for measuring pain intensity in young female Egyptian and Dutch patients with rheumatoid arthritis (RA). METHODS: Data were obtained in a cross-cultural study of 42 Egyptian and 30 Dutch female outpatients with stable RA. Construct validity was assessed by correlating the scales with other core measures of disease activity in RA. Test-retest reliability was assessed over a 1-week interval. RESULTS: The GRS and the VRS were strongly intercorrelated in the total study cohort and in the Egyptian and Dutch subgroups. In the individual subgroups, only the GRS demonstrated the expected pattern of correlations with other disease activity measures. Test-retest reliability of the GRS was adequate in both Egyptian and Dutch patients (intraclass correlation coefficient 0.78 vs. 0.83, respectively), whereas reliability of the VRS was unsatisfactory in the Egyptian subgroup (weighted kappa 0.60 vs. 0.82 in the Netherlands). DISCUSSION: The study confirmed that the GRS and VRS were reliable and valid in the total study cohort. Within the individual countries, the GRS seemed to perform better than the VRS.

Adult↗

Reliability and agreement of urodynamics interpretations in a female pelvic medicine center.

OBJECTIVE: To estimate the reliability and interobserver consistency of urodynamic interpretations of female bladder and urethral function. METHODS: Three urogynecologists and three female urologists at a tertiary care medical center reviewed masked, abstracted clinical and urodynamic information from 100 charts, selected for adequate completeness from a consecutive series of 135 women referred for urodynamic testing. For each of the 100 cases, the reviewers assigned International Continence Society filling and voiding phase diagnoses, and overall clinical diagnoses. Raw agreement proportions and weighted kappa chance-corrected agreement statistics (kappa) were used jointly to describe both reliability and interobserver agreement. Reliability was estimated from duplicate reviews, masked and separated by at least 4 months, of each case by each physician. Interobserver agreement was estimated from comparisons of all pairs of responses from different physicians. RESULTS: For clinical diagnosis of stress incontinence (present, absent, indeterminate), the within- and across-physician weighted kappa's were, respectively, 0.78 and 0.68. Corresponding results were 0.40 and 0.13 for detrusor overactivity without incontinence, 0.58 and 0.38 for detrusor overactivity with incontinence, and 0.51 and 0.26 for voiding dysfunction. Standard errors of each kappa were between 0.023 and 0.043. CONCLUSION: In our group, lower urinary tract diagnoses of stress urinary incontinence from both clinical and urodynamic data demonstrated substantial reliability and interobserver agreement. However, by conventional interpretation of kappa-statistics, reliability of diagnoses of detrusor overactivity or voiding dysfunction was only moderate, and interobserver agreement on these diagnoses was no better than fair. Urodynamic interpretations may not be satisfactorily reproducible for these diagnoses.

Diagnostic Techniques, Urological↗

Reliability analysis for radiographic measurement of limb length discrepancy: full-length standing anteroposterior radiograph versus scanogram.

Patients with limb length discrepancy (LLD) often have associated angular deformities requiring a standing full-length radiograph of the lower limb in addition to a scanogram. The purpose of our study was to determine the intraobserver and interobserver reliability of measuring LLD with both techniques, using computed radiography. The LLD was measured on 70 supine scanograms and standing anteroposterior radiographs of the lower extremity by 5 blinded observers on 2 separate occasions. Intraclass correlation coefficient (ICC) and mean absolute difference (in millimeters) was calculated to assess intraobserver and interobserver reliability and found to be excellent for both radiographic techniques. Intraobserver ICC and mean absolute difference was 0.975 to 0.995 and 1.5 to 2.6 mm for scanogram and 0.939 to 0.996 and 1.5 to 4.6 mm for the standing radiograph, respectively. Repeated measurements for both radiographic studies were within 5 mm of the first measurement greater than 90% and within 10 mm greater than 95% of times. Interobserver ICC and mean absolute difference was 0.979 and 2.6 mm for scanogram and 0.968 and 3.0 mm for the standing radiograph. The reliability was excellent irrespective of age, sex, and underlying diagnosis other than Blount disease, which had good reliability. A standing anteroposterior radiograph of the lower extremity should be the imaging modality of choice when evaluating patients with limb length inequality who may have angular deformities because it allows a comprehensive evaluation of the extremity and is as reliable as a scanogram for measuring LLD. This approach may decrease the radiation exposure and financial burden involved in assessing patients with unequal limb lengths.

Adolescent↗

Physical impairment index: reliability, validity, and responsiveness in patients with acute low back pain.

STUDY DESIGN: Cohort study of patients with acute low back pain undergoing physical therapy. OBJECTIVES: Examine the reliability and validity of the Physical Impairment Index in a group of patients with acute low back pain and determine the responsiveness and minimum detectable change of the index and its component tests. SUMMARY OF THE BACKGROUND DATA: The Physical Impairment Index was originally described as a reliable and valid means of assessing physical impairment in patients with low back pain. The psychometric properties of the index have not been reported in patients with acute low back pain, nor has its responsiveness been examined. METHODS: Seventy-eight patients with acute (<3 weeks duration) low back pain participating in a clinical trial were assessed at baseline and after 4 weeks. Interrater reliability of the index was examined in a subgroup of 20 patients. Validity was examined through correlations with concurrent measures of pain, disability, and psychosocial variables. Changes in the index over 4 weeks were used to assess responsiveness and minimum detectable change. RESULTS: Interrater reliability of the index was high (intraclass correlation coefficient = 0.89), and its validity was generally supported by the pattern of correlations. The index was more responsive to change than any of its component tests but was less responsive than the Oswestry disability score. The minimum detectable change on the index was approximately 1 point. CONCLUSIONS: The Physical Impairment Index appears to be a reliable and valid measure of physical impairment for patients with acute low back pain and may be useful as an adjunct outcome measure for studies involving these patients. Further research on patients with chronic pain is needed before it can be advocated for outcomes research with this population.

Acute Disease↗

Reliability, validity, and responsiveness of the short form 12-item survey (SF-12) in patients with back pain.

STUDY DESIGN: Secondary analysis of data collected from spine patients' normal clinic visits from 1998 to 2001. OBJECTIVE: To evaluate the reliability, validity, and responsiveness of the short form 12-item survey in patients with back pain. SUMMARY OF BACKGROUND DATA: The reliability, validity, and responsiveness of the short form 12-item survey in patients with back pain has not been previously evaluated. METHODS: Patients were asked to complete a comprehensive computerized survey questionnaire during their regular clinic visits. A total of 2520 patients who indicated in their first surveys that they had back pain were included in the study of the reliability and validity of the short form 12-item survey. Of these, 506 patients completed another survey within 3-6 months of follow-up and were included in the responsiveness evaluation. RESULTS: The two summary scales of the short form 12-item survey, physical component summary and mental component summary, demonstrated internal consistency reliability, with Cronbach alpha for both scales exceeding the recommended level of 0.70. Correlation of physical component summary and mental component summary with six other measures theoretically related or unrelated to these scales performed as expected without exception, demonstrating the construct validity of the short form 12-item survey. The responsiveness of the short form 12-item survey was supported by several pieces of evidence. First, the changes in physical component summary and mental component summary scores were correlated with the changes in back pain intensity. Second, for patients whose back pain improved, there was a significant increase in the follow-up physical component summary and mental component summary scores as compared to the baseline. Third, small to moderate effect size was observed for patients whose back pain became improved or became worse. CONCLUSIONS: The short form 12-item survey demonstrated good internal consistency reliability, construct validity, and responsiveness in patients with back pain.

Academic Medical Centers↗

Inter- and intraobserver reliability of computed tomography in assessment of thoracic pedicle screw placement.

STUDY DESIGN: Reliability study of computed tomography imaging in 12 cadaver specimens instrumented with titanium or stainless steel thoracic pedicle screws. OBJECTIVE: To evaluate inter- and intraobserver reliability of computed tomography scan in determining the accuracy of thoracic pedicle screw placement and to identify the differences in observers' agreement when viewing stainless steel versus titanium screws. SUMMARY OF BACKGROUND DATA: Computed tomography is often used to assess the accuracy of pedicle screw placement. Accuracy of screw placement is important in the thoracic spine where pedicle morphometry increases the difficulty of screw placement (vital structures are at increased risk). The current literature lacks a critical evaluation of computed tomography reliability among observers. METHODS: Twelve adult cadavers were instrumented with thoracic pedicle screws. Nine cadavers were instrumented with titanium screws and three with stainless steel screws. The spines were imaged with computed tomography. Three observers used a grading scale to score the extent of pedicle violation and independently scored the placement of each pedicle screw on three separate occasions. Interobserver and intraobserver agreement were determined by using the kappa statistic. RESULTS: The mean kappa score for interobserver agreement for all 12 specimens (including titanium and stainless steel screws) was 0.51, which correlates with a moderate degree of agreement. Although the interobserver kappa statistics for titanium (kappa = 0.53) and stainless screws (kappa = 0.44) showed a moderate degree of agreement, the intraobserver reliability was substantial (kappa = 0.63). The mean intraobserver kappa for titanium screws was 0.63 and for stainless steel screws was 0.62. CONCLUSIONS: Our data show that interobserver agreement is moderate and intraobserver agreement is substantial when computed tomography is used to assess placement of thoracic pedicle screws. We conclude that computed tomography is reliable for evaluating thoracic pedicle screw placement throughout the thoracic spine.

Adult↗

Interrater reliability of scoring of pain drawings in a self-report health survey.

STUDY DESIGN: Study of interrater reliability. OBJECTIVE: To assess the interrater reliability of data from pain drawings scored by multiple raters and the consistency of the subsequent classification of cases of widespread pain. SUMMARY OF BACKGROUND DATA: In large health surveys, pain drawings used to capture self-reported pain, and to classify cases of widespread pain, are often scored by several raters. The reliability of multiple rater scoring of pain drawings has not been investigated. METHODS: As part of a postal survey sent to adults 50 years and older, subjects were asked to shade their pain on a blank body manikin. The first 50 pain drawings in which respondents had shaded pain were selected for this study. Eight nonclinical staff were trained to score pain drawings using transparent templates divided into 50 body areas. Interrater reliability was assessed by comparing the scoring of "pain" or "no pain" for all 50 areas of each pain drawing. RESULTS: Complete scoring agreement among all raters was observed for at least 78% of pain drawings across all body areas (kappa > 0.60). The raters had complete agreement in 42 of 50 areas in 90% or more of pain drawings. From the raters' scoring of pain areas, there was complete agreement on the presence or absence of widespread pain for 49 of 50 pain drawings (98% agreement, Kappa = 0.98). CONCLUSIONS: This study shows that multiple raters, with training and guidelines, can reliably score pain drawings, and high consistency in the subsequent classification of cases of widespread pain can be obtained from such data.

Art↗

Computer-assisted algorithms improve reliability of King classification and Cobb angle measurement of scoliosis.

STUDY DESIGN: Interobserver and intraobserver reliability study of improved method to evaluate radiographs of patients with scoliosis. OBJECTIVE: To determine the reliability of a computer-assisted measurement protocol for evaluating Cobb angle and King et al classification. SUMMARY OF BACKGROUND DATA: Evaluation of scoliosis radiographs is inherently unreliable because of technical and human judgmental errors. Objective, computer-assisted evaluation tools may improve reliability. METHODS: Posteroanterior preoperative radiographic images of 27 patients with adolescent idiopathic scoliosis were each displayed on a computer screen. They were marked 3 times in random sequence by each of 5 evaluators (observers) who marked 70 standardized points on the vertebrae and sacrum in each radiograph. A computer program (Spine 2002;27:2801-5) that identified curves, calculated Cobb angles, and generated the King et al classification automatically analyzed coordinates of these points. The interobserver and intraobserver variability of the Cobb angle and King et al classification evaluations were quantified and compared with values obtained by unassisted observers. RESULTS: Average Cobb angle intraobserver standard deviation was 2.0 degrees for both the thoracic and lumbar curves (range 0.1 to 8.3 degrees for different curves). Interobserver reliability was 2.5 degrees for thoracic curves and 2.6 degrees for lumbar curves. Among the 5 observers, there was an inverse relationship between repeatability and time spent marking images, and no correlation with image quality or curve magnitude. Kappa values for the variability of the King et al classification averaged 0.85 (intraobserver). CONCLUSIONS: Variability of Cobb measurements compares favorably with previously published series. The classification was more reliable than achieved by unassisted observers evaluating the same radiographs. The same principles may be applicable to other radiographic measurement and evaluation procedures.

Algorithms↗

Reliability of a novel classification system for thoracolumbar injuries: the Thoracolumbar Injury Severity Score.

STUDY DESIGN: Prospective study of 5 spine surgeons rating 71 clinical cases of thoracolumbar spinal injuries using the Thoracolumbar Injury Severity Score (TLISS) and then re-rating the cases in a different order 1 month later. OBJECTIVE: To determine the reliability of the TLISS system. SUMMARY OF BACKGROUND DATA: The TLISS is a recently introduced classification system for thoracolumbar spinal column injures designed to simplify injury classification and facilitate treatment decision making. Before being widely adopted, the reliability of the TLISS must be studied. METHODS: A total of 71 cases of thoracolumbar spinal trauma were distributed on CD-ROM to 5 attending spine surgeons, including clinical/radiographic data, details of the TLISS, and a scoring sheet in which cases would be scored using the system. The surgeons were later assigned the task with the cases reordered. Intraobserver and interobserver reliability was calculated for TLISS components, total score, and surgeon's treatment decision using the Cohen unweighted kappa coefficients and Spearman rank-order correlation. RESULTS: Interrater reliability assessed by generalized kappa coefficients was 0.33 +/- 0.03 for injury mechanism, 0.91 +/- 0.02 for neurologic status, 0.35 +/- 0.03 for posterior ligamentous complex status, 0.29 +/- 0.02 for TLISS total, and 0.52 +/- 0.03 for treatment recommendation. Respective results using the Spearman correlation were 0.35 +/- 0.04, 0.94 +/- 0.01, 0.48 +/- 0.04, 0.65 +/- 0.03, and 0.51 +/- 0.04. Surgeons agreed with the TLISS recommendation 96.4% of the time. Intrarater kappa coefficients were 0.57 +/- 0.04 for injury mechanism, 0.93 +/- 0.02 for neurologic status, 0.48 +/- 0.04 for posterior ligamentous complex status, 0.46 +/- 0.03 for TLISS total, and 0.62 +/- 0.04 for treatment recommendation. Respective results using the Spearman correlation were 0.70 +/- 0.04, 0.95 +/- 0.02, 0.59 +/- 0.05, 0.77 +/- 0.04, and 0.59 +/- 0.05. CONCLUSIONS: The TLISS has good reliability and compares favorably to other contemporary thoracolumbar fracture classification systems.

Lumbar Vertebrae↗

Reliabilities of and correlations among five standard methods of assessing the sagittal alignment of the cervical spine.

STUDY DESIGN: The reliabilities of and correlations among 5 standard methods of assessing cervical sagittal alignment were evaluated. OBJECTIVE: To investigate the reliabilities of and correlations among 5 standard methods of assessing cervical sagittal alignment. SUMMARY OF BACKGROUND DATA: Although various cervical sagittal alignment assessment methods are widely used, their relative reliability and intercorrelation have not been reported. METHODS: From 442 lateral cervical radiographs, 40 with lordotic, 40 with straight or sigmoid, and 40 with kyphotic alignment were selected. Two orthopedic surgeons independently evaluated the sagittal alignment in each group twice using CCL, C1-C7 Cobb, C2-C7 Cobb, sagittal tangent, and the Ishihara methods. Intraobserver and interobserver reliabilities were confirmed and the correlations among the 5 methods were measured. RESULTS: Intraobserver and interobserver reliabilities for all 5 methods were good. In the lordotic group, the correlations among all 5 methods were consistently strong (r = 0.731 to 0.922). In the straight or sigmoid group, the correlations were weak to moderate among the CCL, C2-C7 Cobb, sagittal tangent, and Ishihara methods but tended to be weak between these 4 methods and the C1-C7 Cobb method (r = -0.245 to 0.777). In the kyphotic group, the correlations were also weak to moderate among the same 4 methods, and were statistically insignificant between them and the C1-C7 Cobb. CONCLUSIONS: The correlations among the CCL, C1-C7 Cobb, C2-C7 Cobb, sagittal tangent, and Ishihara methods are strong when lordosis is retained; otherwise, they are moderate to poor. In the kyphotic group, C1-C7 Cobb has no significant correlation with the other 4 methods.

Adolescent↗

Reliability and validity of 2 single-item measures of psychosocial stress.

BACKGROUND: Practical limitations in epidemiologic research may necessitate use of only a few questions for assessing the complex phenomenon called "stress." The objective of this study was to evaluate the measurement characteristics of 2 single-item measures on the amount of stress and the ability to handle stress. METHODS: We selected 218 adults age 50 to 76 years living in western Washington state from a large prospective cohort study of lifestyle factors and cancer risk to evaluate the 3-month test-retest reliability and intermethod reliability of the stress questions. To assess the latter, we compared 2 single-item measures on stress with 3 more fully validated multi-item instruments on perceived stress, daily hassles, and life events, which assessed the same underlying constructs as the single-item measures. RESULTS: The test-retest reliabilities for the single-item stress measures were good (kappa and intraclass correlations between 0.66 and 0.74). The intermethod reliabilities comparing the 2 single-item stress measures with 3 multi-item instruments were moderate (r = 0.31-0.46) and comparable to correlations observed among the 3 multi-item instruments (r = 0.25-0.47). CONCLUSIONS: The 2 single-item stress measures are reliable at measuring stress with validity similar to longer questionnaires. Single-item measures offer a practical instrument for assessing stress in large prospective epidemiologic studies that lack space for longer instruments.

Aged↗

How many patients are needed to provide reliable evaluations of individual clinicians?

PURPOSE: The purpose of this study was to determine how many patients are needed to provide reliable patient ratings of care at the individual clinician level. SETTING AND SOURCES OF DATA: The study was conducted in an academic medical center and was based on analysis of 34,985 patients who completed a 50-item survey rating the care received during a recent outpatient visit to a physician or midlevel provider. STUDY DESIGN: Analyses of patient satisfaction surveys was done to: 1) confirm the dimensions of satisfaction with outpatient care in an existing measure, and 2) determine the number of patients required to provide reliable estimates of clinician care for single items and an 11-item composite scale. PRINCIPAL FINDINGS: Factor analysis showed that the survey measured 2 dimensions of satisfaction: 1) clinician care, and 2) features of visiting the office. The 11-item clinician care scale had high reliability (Cronbach's alpha=0.97). The number of patients needed to achieve reliability of 0.80 at the clinician level was 66 for the 11-item scale and ranged from 52 to 91 for individual items. For primary care physicians only, the comparable number of patients per clinician was 77 for the 11-item scale and ranged from 50 to 147 across items. CONCLUSIONS: For the survey items that we analyzed, the answer to the question "How many patients are needed to obtain useful and reliable feedback?" is at least 50, but varies by item type (global vs. specific) and by number of items (composite scale or single-item rating) and by the conditions of use (for self-assessment and learning or reward and punishment).

Academic Medical Centers↗

Reliability of the Berg Balance Scale and balance master limits of stability tests for individuals with brain injury.

PURPOSE: The purpose of this study was to determine the test-retest reliability of 2 measurement tools used to examine balance in persons with brain injury. METHODS: Five participants (ages 20-32) were recruited from a transitional living facility in Galveston, Texas. Each participant performed the Berg Balance Scale (BBS) and Balance Master Limits of Stability Test (BMLOST) sequence in random order on the same day and 1 week later. RESULTS: The test-retest reliability for the BBS was excellent (ICC2,1 = 0.986). Reliability of the BMLOST ranged from poor for weight shifts (ICC2,1 = 0.228) to good, for movement time (ICC2,1 = 0.825) and path sway (ICC2,1 = 0.846). DISCUSSION: The BBS, a function based test, was found to have greater test-retest reliability than the BMLOST, a strategy level test. The scores of the BBS were high and may suggest the potential of ceiling effects. CONCLUSION: Preliminary test-retest reliability for both the BBS and aspects of the BMLOST for persons with brain injury was demonstrated. The value of administering the tests in isolation or combination for patients with brain injury is yet to be determined.

Adult↗

Reliability of clinical measures used to assess patients with peripheral vestibular disorders.

PURPOSE: The purposes of this research were to (1) determine test-retest reliability of clinical measures of self reported disability and subjective complaints, gait, and fall risk; and (2) establish normal variability for each of these measures based on test-retest variability in people with peripheral vestibular disorders. METHODS: Sixteen patients with confirmed peripheral vestibular disorders performed 2 trials of each of the measures within a single physical therapy session. The measures included rating of disability, percent of day affected by dizziness, head movement induced dizziness, preferred gait speed, gait deviations, and Dynamic Gait Index. In order to assess test-retest reliability of the measures intraclass correlation coefficients (ICC) were calculated. RESULTS: All measurement tools demonstrated excellent reliability (ICC 3,1 = 0.86 - 1.00) except for head movement induced dizziness (ICC 3,1 = 0.48). For each measure we report normal variability as tested within a single session. DISCUSSION: Clinical measures commonly used in the assessment of vestibular patients were found to have excellent test-retest reliability, except for the subjective measure of head movement-induced dizziness. CONCLUSION: Incorporation of valid and reliable assessments in clinical practice is critical in order to demonstrate the effectiveness of therapeutic intervention.

Adult↗

Reliability and criterion-related validity of self-report of syphilis.

OBJECTIVE: To determine the test-retest reliability, sensitivity and specificity, and criterion-related validity of the Risk Behavior Assessment (RBA) syphilis questions. The RBA is a standardized instrument that has been used in several studies of STDs in drug users. METHODS: For the test-retest reliability study, 219 injection drug users completed the RBA twice within a 48-hour period. To determine criterion-related validity, 207 individuals, who also completed the RBA, were tested with the rapid plasma reagin test (RPR), and 206 individuals were also tested with the Serodia Treponema pallidum particle agglutination test (TP-PA). RESULTS: The test-retest reliability for the question "How many times have you been told by a doctor or a nurse that you had syphilis?" was 0.78. The test-retest reliability for the question "In what year were you last treated for syphilis?" was 0.89. For the comparison of self-report with the RPR test, the sensitivity of self-report was 46.2% and the specificity was 95.7%. For the comparison of self-report with the TP-PA test, the sensitivity of self-report was 37% and the specificity was 97.7%. CONCLUSIONS: Self-reports of syphilis infection history were found to have good reliability, excellent specificity, and moderate sensitivity. These characteristics need to be taken into account in any study using these self-report items.

Adult↗

Reliability and validity of the gross motor function classification system for cerebral palsy.

PURPOSE: The purposes of this study were to evaluate interrater reliability using videotapes and criterion-related and construct validity of the Gross Motor Function Classification System (GMFCS), aspects of reliability and validity not previously published. METHODS: Two experienced pediatric physical therapists rated 30 videotapes of children with cerebral palsy (CP) or Down syndrome (DS) to test interrater reliability. Criterion-related validity was evaluated by comparing GMFCS levels with tests of motor and nonmotor development. Construct validity was assessed by comparing GMFCS trends over time in children with CP and DS. RESULTS: Interrater reliability was 0.84. Correlation was higher between GMFCS level and tests of motor development than GMFCS level and tests of nonmotor development. The GMFCS level remained relatively stable in children with CP but tended to improve in children with DS. CONCLUSIONS: This study extends reliability and validity of the GMFCS, supporting its use in clinical practice and research.

Journal Article↗

Interrater reliability of early intervention providers scoring the alberta infant motor scale.

PURPOSE: This study was designed to examine the interrater reliability of early intervention providers scoring of the Alberta Infant Motor Scale (AIMS) and to examine whether training on the AIMS would improve their interrater reliability. METHODS: Eight early intervention providers were randomly assigned to two groups. Participants in Group 1 scored the AIMS on seven videotapes of infants prior to receiving training and after training on another set of seven videotapes of infants. Participants in Group 2 scored the AIMS on all 14 videotapes of the infants after receiving training. RESULTS: Overall interrater reliability before and after training was high with intraclass correlation coefficients ranging from 0.98 to 0.99. Detailed examination of the results showed that training improved the reliability of the supine subscale in a subgroup of infants between the ages of five and seven months. Training also had an effect on the classification of infants as normal or abnormal in their motor development based on their percentile rankings. CONCLUSION: The AIMS manual provides sufficient information to attain high interrater reliability without training, but revisions regarding scoring are strongly recommended.

Journal Article↗

Is patient-reported function reliable for monitoring postacute outcomes?

OBJECTIVE: A major challenge in the development of a comprehensive measurement system to evaluate effectiveness across a broad range of postacute care settings is the stability and consistency of outcomes measures across respondents and settings. The objective of this study was to investigate the test-retest and subject-proxy reliability of activity scores for use in a new postacute care outcome instrument using an interview format across different care settings. DESIGN: Twenty-five subjects were randomly selected from a larger study of 485 individuals and were interviewed on two occasions within 1 to 4 days to assess self-reported test-retest reliability of summary scores of the Activity Measure-Post-Acute Care item pool. Proxy reliability was tested by interviewing the primary physical or occupational therapist or family member using an identical questionnaire in addition to the subject in 45 patients. RESULTS: Test-retest and subject-proxy reliability was acceptable for the three domains of the activity construct: physical and movement, personal and instrumental, and applied cognition with intraclass correlation coefficients of the summary scores of each of the three domains ranging between 0.91 and 0.97 for test-retest and 0.68 and 0.90 for subject-proxy. CONCLUSIONS: Reliability is adequate to justify use of these activity scales across respondents and settings.

Adult↗