Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Rehabilitative placement of poststroke patients: reliability of the Clinical Practice Guideline of the Agency for Health Care Policy and Research.

OBJECTIVE: To determine the interrater reliability of the United States Agency for Health Care Policy and Research (AHCPR) Clinical Practice Guideline Number 16 for rehabilitative placement of poststroke patients. DESIGN: Pairs of rehabilitation professionals, highly trained in the Guideline, rated the appropriateness of rehabilitative placements. SETTING: Acute care hospitals in three regions of the country. PATIENTS: Sixty patients with moderate-to-severe stroke. MEASURES: Numerous factors affecting appropriate placement according to the Guideline were abstracted from medical records or obtained by direct evaluation of patients. RESULTS: Good reliability was attained for home and nursing facility placement with rehabilitation services but with no multidisciplinary rehabilitation program (intraclass correlation coefficient = .73 and .60, respectively). Serious reliability problems were found for placements in low-intensity outpatient rehabilitation and high-intensity inpatient rehabilitation programs. Chief sources of unreliability were ambiguous or missing data in hospital medical records, complexities in the Guideline, and raters' tendencies to follow their own clinical judgments. More than one type of placement was appropriate for 65% of patients. CONCLUSIONS: Reliable placement guidelines are possible, but aspects of the Guideline require additional development. Evidence of demonstrated reliability and validity will be required to resolve disputes between rehabilitation professionals and payers regarding appropriate levels and types of rehabilitation and to guide patients and their families.

Adult↗

Interexaminer reliability of the palpation of trigger points in the trunk and lower limb muscles.

OBJECTIVES: To determine the interexaminer reliability of palpation of three characteristics of trigger points (taut band, local twitch response, and referred pain) in patients with subacute low back pain, to determine whether training in palpation would improve reliability, and whether there was a difference between the physiatric and chiropractic physicians. DESIGN: Reliability study. SETTING: Whittier Health Campus, Los Angeles College of Chiropractic. PARTICIPANTS: Twenty-six nonsymptomatic individuals and 26 individuals with subacute low back pain. INTERVENTION: Twenty muscles per individual were first palpated by an expert and then randomly by four physician examiners. MAIN OUTCOME MEASURES: Palpation findings. RESULTS: Kappa scores for palpation of taut bands, local twitch responses, and referred pain were .215, .123, and .342, respectively, between the expert and the trained examiners, and .050, .118, and .326, respectively, between the expert and the untrained examiners. Kappa scores for agreement for palpation of taut bands, twitch responses, and referred pain were .108, -.001, and .435, respectively, among the nonexpert, trained examiners, and -.019, .022, and .320, respectively, among the nonexpert, untrained examiners. CONCLUSIONS: Among nonexpert physicians, physiatric or chiropractic, trigger point palpation is not reliable for detecting taut band and local twitch response, and only marginally reliable for referred pain after training.

Adult↗

Interrater reliability of judgments of the centralization phenomenon and status change during movement testing in patients with low back pain.

OBJECTIVE: To determine the interrater reliability of judgments of status change, including the centralization phenomenon during examination of the lumbar spine, and to determine the effects of training and experience on reliability. DESIGN: A videotape study of judgments by physical therapists and physical therapy students of the results of movement testing during the examination of patients with low back pain. SETTING: Outpatient physical therapy clinic. PATIENTS: Patients receiving physical therapy treatment for low back pain. INTERVENTION: Patients with low back pain were videotaped while performing a variety of movement tests including single, repeated, and sustained movements. Forty licensed physical therapists and 40 physical therapy students were provided with operational definitions of the three potential judgments of status change with movement testing; centralization, peripheralization, status quo. All therapists and students viewed the videotape and made a judgment regarding the patient's status change in response to the test. MAIN OUTCOME MEASURE: Percentage agreement and kappa coefficient values for agreement. RESULTS: Interrater reliability was excellent for the total sample of examiners (kappa = .793), for the licensed physical therapists (kappa = .823), and for the students (kappa = .763). CONCLUSIONS: Judgments of status change are reliable when operational definitions are provided. Clinical experience does not appear to substantially improve reliability.

Adult↗

Reliable serial measurement of cognitive processes in rehabilitation: the Cognitive Log.

OBJECTIVE: To evaluate the reliability and utility of a brief quantitative measure of cognitive recovery, the Cognitive Log (Cog-Log), developed for daily use with rehabilitation inpatients to provide information about the recovery of higher neurocognitive processes including verbal recall, attention, working memory, motor sequencing, and response inhibition. DESIGN: Descriptive study of the Cog-Log's normative scores, reliability (interrater, internal consistency), and validity as shown by its relationship to standard neuropsychologic measures. SETTING: Inpatient rehabilitation hospital affiliated with a large university medical center. PARTICIPANTS: One hundred fifty neurorehabilitation inpatients with acquired brain injury; 83 young adults without acquired brain injury were included to provide normative data. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: The Cog-Log; standardized neuropsychologic measures of memory (Wechsler Memory Scale-Revised, Rey Auditory Verbal Learning Test), language, attention (Wechsler Adult Intelligence Scale-Revised), and reasoning (Trail Making Test). RESULTS: Reliability analysis showed strong interrater reliability across items (Spearman r, .749-1.00) and high internal consistency (Cronbach alpha=.778). Factor analysis of the Cog-Log using principal components extraction revealed a unitary factor (eigenvalue=3.48). Cog-Log items designed to measure working memory and immediate and delayed verbal memory were most strongly predictive of performance on similar standardized neuropsychologic measures administered on the same day. CONCLUSION: The Cog-Log appears to be a reliable and efficient tool for measuring ongoing neurocognitive recovery during inpatient rehabilitation.

Adolescent↗

Reliability, validity, and responsiveness of the modified Kapandji index for assessment of functional mobility of the rheumatoid hand.

OBJECTIVE: To determine the reliability, validity, and responsiveness of the modified Kapandji index (MKI). DESIGN: Prospective study. A cohort of patients planned for surgery of the wrist and/or fingers was evaluated within 48 hours before surgery and at least 6 months after surgery. SETTING: Patients were in hospitalized or private care in France. PARTICIPANTS: Patients with rheumatoid arthritis according to criteria of the American College of Rheumatology. Forty-two patients (36 women; mean age, 57.5y; range, 22-80y) were included in the reliability study. Fifty patients (42 women; mean age, 54.18y; range, 19-77y) were included in the validity study. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Clinical outcome measures included the MKI, the overall mobility score of the wrist and fingers, the finger mobility score, a visual analog scale (VAS) of pain in the hands and wrists, morning stiffness duration, total score of tenderness, total score of swelling, grip and pinch strength, the Hand Functional Index (HFI), and the Cochin rheumatoid hand disability scale. Reliability was studied with the intraclass correlation coefficient (ICC) and the Bland and Altman method. Convergent and divergent validity were assessed with the Spearman correlation coefficient. Responsiveness was assessed by the paired t test, the effect size, and the standardized response mean (SRM). RESULTS: Interobserver reliability was good with an ICC of.90, and the Bland and Altman analysis showed homogeneous distribution of the differences, with no systematic trend. The MKI correlated well with the other mobility measures (HFI, the finger mobility score measured with the finger goniometer), indicating a good convergent validity, and the expected divergent validity with the other outcome measures (grip and pinch strength, total score of swelling, total score of Ritchie Articular Index, Cochin scale, VAS of pain) was observed. The 50 patients in the validity study were evaluated twice, before and after surgery, at a mean interval +/- standard deviation of 7.16+/-2.10 months (range, 6-15mo). Thirty-six patients (72%) were very satisfied or satisfied with the results of surgery, 7 (14%) were not satisfied or dissatisfied, and 7 (14%) were dissatisfied or very dissatisfied. The SRM and effect size values of the MKI were -.19 and -.10, respectively. Individual changes in the score had the best correlation (r(s)=.51) with overall patient satisfaction. CONCLUSIONS: The MKI has excellent validity and reliability. Individual changes in the score are clinically relevant. This index can be used in clinical practice and in therapeutic trials; it needs further study concerning its use for hand surgery.

Activities of Daily Living↗

Strapped versus unstrapped technique of the prone press-up for measurement of lumbar extension using a tape measure: differences in magnitude and reliability of measurements.

OBJECTIVES: To determine (1) the reliability of the prone press-up to measure lumbar extension using a strap and not using a strap to control pelvic movement in experienced clinicians and students and (2) if a difference exists between the magnitude of lumbar extension range of motion between the strapped and unstrapped condition. DESIGN: Prospective study. SETTING: Academic laboratory. PARTICIPANTS: Convenience sample of 63 unimpaired volunteers (mean age +/- standard deviation, 25.95+/-5.75 y). INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Lumbar extension was measured in the prone position by using a tape measure to measure the perpendicular distance of the sternal notch to the support surface while using a strap and not using a strap to control pelvic movement. All measurements were performed independently by 2 groups of examiners (1 experienced group, 1 student group) and repeated to determine intrarater and interrater reliabilities. RESULTS: Intrarater and interrater reliability were good or excellent for all methods and all measurement group comparisons (intraclass correlation coefficient range, .82-.91). Additionally, the amount of lumbar extension, as measured by the prone press-up, during the strapped condition was significantly greater than with the unstrapped condition. CONCLUSION: Use of a tape measure while the subject performs a prone press-up appears to be a reliable method for the measurement of lumbar extension. This technique is reliable whether the examiner is experienced or inexperienced and whether or not the subject has the pelvis secured with a strap.

Adult↗

Health-related fitness test battery for adults: aspects of reliability.

OBJECTIVE: In two studies, the reliability of 3 balance, 2 flexibility, and 4 muscular strength tests proposed as test items were investigated in a health-related fitness (HRF) test battery for adults. DESIGN: Methodological study. SETTING: A health promotion research institute. SUBJECTS: In study A, volunteers (n=42) from two worksites participated. In study B, a population sample (n=510) of 37-to 57-year-old men and women was selected. MAIN OUTCOME MEASURES: Intraclass correlation coefficient of repeated measures was used to assess inter-rater reliability. The degree of measurement error was expressed as the standard error of measurement. The mean difference with 95% confidence intervals between the testing days or test trials was used to assess test-retest or trial-to-trial reproducibility. The coefficient of variation(CV=[SD/mean] x 100%) from day to day was also calculated. RESULTS: The following tests appeared to provide acceptable reliability as methods for field assessment of HRF: standing on one leg with eyes open for balance, side-bending of the trunk for spinal flexibility, modified push-ups for upper body muscular function, and jump and reach and one leg squat for leg muscular function. CONCLUSIONS: This reliability assessment provided useful information on the characteristics of potential test items in a HRF test battery for adults and on the limitations of its practical use. Testers must be properly trained to ensure reliable assessment of HRF of adults.

Adult↗

The Timed "up and go" test: reliability and validity in persons with unilateral lower limb amputation.

OBJECTIVE: To determine the interrater and intrarater reliability and the validity of the Timed "up and go" test as a measure for physical mobility in elderly patients with an amputation of the lower extremity. DESIGN: To test interrater reliability, the test was performed for two observers at different times of the same day in an alternating order. To test intrarater reliability, the patients performed the test for one observer on two consecutive visits with an interval of 2 weeks. To test validity, the results of the Timed "up and go" test were compared with the results on the Sickness Impact Profile, 68-item version (SIP68), and the Groningen Activity Restriction Scale (GARS). PATIENTS: Thirty-two patients, age 60 yrs or older, with unilateral transtibial or transfemoral amputation because of peripheral vascular disease. RESULTS: The Timed "up and go" test showed good intrarater and interrater reliability (r = .93 and .96, respectively). A moderate relationship exists between the Timed "up and go" test and the GARS, a good relationship exists with the "physical subscales" of the SIP68, and there is no relationship with the "mental subscales" of the SIP68. CONCLUSIONS: The Timed "up and go" test is a reliable instrument with adequate concurrent validity to measure the physical mobility of patients with an amputation of the lower extremity.

Activities of Daily Living↗

Twelve-month test-retest reliability of the structured clinical interview for DSM-III-R personality disorders in cocaine-dependent patients.

This study examined 12-month test-retest reliability of the Structured Clinical Interview for DSM-III-R Personality Disorders (SCID-II) in cocaine-dependent patients. Thirty-one patients completed the SCID-II during the second week of hospitalization for cocaine dependence, and again 12 months later. In both interviews, patients were asked to answer questions about their personality during the several years preceding admission to the hospital. Test-retest reliability, as measured by kappa, was relatively poor at .46. However, reliability of negative diagnoses (the absence of a disorder at both time points) was higher than reliability of positive diagnoses (the presence of a disorder at both time points). Reasons for the difficulty in attaining long-term test-retest reliability of axis II diagnoses in cocaine-dependent patients are discussed.

Adult↗

Intermittent explosive disorder-revised: development, reliability, and validity of research criteria.

The study of human aggression has been hindered by the lack of reliable and valid diagnostic categories that specifically identify individuals with clinically significant displays of impulsive aggressive behavior. DSM intermittent explosive disorder (IED) ostensibly identifies one such group of individuals. In its current form, IED suffers from significant theoretical and psychometric shortcomings that limit its use in clinical or research settings. This study was designed to develop a revised criteria set for IED and present initial evidence supporting its reliability and validity in a well characterized group of personality disordered subjects. Accordingly, research criteria for IED-Revised (IED-R) were developed. Clinical, phenomenologic, and diagnostic data from 188 personality disordered individuals were reviewed. IED-R diagnoses were assigned using a best-estimate process. The reliability and construct validity of IED-R were examined. IED-R diagnoses had high interrater reliability (kappa = .92). Subjects meeting IED-R criteria had higher scores on dimensional measures of aggression and impulsivity, and had lower global functioning scores than non-IED-R subjects, even when related variables were controlled. IED-R criteria were more sensitive than DSM-IV IED criteria in identifying subjects with significant impulsive-aggressive behavior by a factor of four. We conclude that in personality disordered subjects, IED-R criteria can be reliably applied and appear to have sufficient validity to warrant further evaluation in field trials and in phenomenologic, epidemiologic, biologic, and treatment-outcome research.

Adult↗

Diagnostic accuracy and test-retest reliability of nonword repetition and digit span tasks administered to preschool children with specific language impairment.

UNLABELLED: To assess diagnostic accuracy and test-retest reliability, two forms of a nonword repetition task were administered to 22 preschool children with specific language impairment (SLI) and to 22 age- and gender-matched children with normal language (NL). Results were compared with performance on a digit span task and norm-referenced test scores. Nonword repetition scores provided excellent sensitivity and specificity for discriminating between groups. Scores on both nonword repetition and digit span tasks improved significantly from first to second administrations for both groups, but remained relatively stable at the third administration. The SLI group appeared to benefit more from repetition than the NL group. Acceptable levels of test-retest reliability were achieved for the digit span task, but not for the NL group on the nonword repetition task. These preliminary findings suggest that with further refinement to improve test-retest reliability, nonword repetition holds promise as a diagnostic measure for SLI in preschool children. EDUCATIONAL OBJECTIVES: As a result of this activity, the participant will be able to (1) describe the content and administration of nonword repetition tasks; (2) explain why evidence of test-retest reliability is necessary before a measure may be considered reliable for diagnostic purposes; and (3) accurately compare the sensitivity and specificity of the nonword repetition task utilized in this study to standardized language test scores.

Child, Preschool↗

The validity and reliability of a Computerized Dementia Screening Test developed in Korea.

OBJECTIVE: This study was done to verify the validity and the reliability of the newly developed Computerized Dementia Screening Test (CDST) to be easily used in the primary care setting of Korea. DESIGN: Comparison of the results of CDST between 103 healthy control subjects and 41 patients who were diagnosed as having mild cognitive impairment or early dementia, having a clinical dementia rate of 0.5-1 from one health examination center and two neurology clinics in university hospitals. MEASUREMENTS: In order to estimate the criterion-related validity, logistic regression analysis for dementia was done using the four individual test results of CDST, age and educational level. The correlation between Korean Mini-Mental State Examination (K-MMSE) and the predicted probability of mild cognitive impairment and early dementia from the logistic regression was measured to verify its validity. The reliability of CDST was measured by test-retest reliability. RESULTS: The sensitivity and specificity of CDST were 75.6% and 94.2%, respectively, if the cut-off point was set to be 0.5 in the logistic regression model. The Pearson's Correlation Coefficient between K-MMSE and the predicted probability of mild cognitive impairment and early dementia from the logistic regression analysis was 0.59 (P<0.001). The overall test-retest reliability using the predicted probability of dementia from the logistic regression analysis of CDST was 0.89 (P=0.01). CONCLUSION: The validity and reliability of CDST is adequate for use as a screening tool to identify mild cognitive impairment and early dementia in Korean primary care.

Aged↗

Interexaminer reliability and validity of a three-dimensional model to assess prostate volume by digital rectal examination.

OBJECTIVES: To evaluate the interexaminer reliability and accuracy compared with transrectal ultrasound (TRUS) of a three-dimensional (3D) model and other scales to improve the estimation of prostate volume by digital rectal examination (DRE). METHODS: Volunteers from a urology clinic (n = 121) were examined independently by three examiners with different levels of experience in randomized order. During DRE, the examiners estimated the prostate size in increments of 5 g, using various rating scales and a 3D sizing model, without access to the findings of the other investigators. TRUS was then performed by each examiner. RESULTS: The 121 volunteers were 39 to 82 years old, with a mean +/- SD total TRUS prostate size of 35.9 +/- 27.2 g. The DRE size estimates ranged from 15 to 100 g across all examiners and patients. The interexaminer reliability across examiners for the best DRE prostate size estimates (in grams) was 0.78 (95% confidence interval 0.70 to 0.84), and the correlation coefficients (r(s)) with the TRUS volume ranged from 0.61 to 0.72 for the three examiners. A 3D model showed good reliability (intraclass correlation coefficient 0.86, 95% confidence interval 0.75 to 0.93), and correlated well with the TRUS volume (r(s) = 0.67 to 0.75). Other scales showed fair reliability (0.58 to 0.68) and correlated with the TRUS measurements (0.57 to 0.67). The area under the receiver operating characteristic curve to identify prostate volumes greater than 40 g ranged from 0.78 to 0.90 for DRE estimates (in grams) and 0.69 to 0.89 for the 3D model. CONCLUSIONS: DRE size estimates and TRUS volume were moderately to highly correlated in men without prostate cancer. A 3D sizing model showed comparable reliability and correlation with TRUS. Although the DRE estimates generally tend to underestimate the TRUS-measured prostate volume, these tools may be useful in identifying men with enlarged prostate glands.

Adult↗

Long-term mechanical reliability of multicomponent inflatable penile prosthesis: comparison of device survival.

OBJECTIVES: To determine the mechanical reliability of multicomponent inflatable penile prosthesis, comparing five different types of devices, as well as the two-piece versus three-piece as a group. METHODS: We followed 83 patients with two-piece and 283 patients with three-piece inflatable penile prostheses for a mean time of 66 months. At a cutoff of 63 months, mechanical complication rates were reviewed and statistically analyzed. RESULTS: Thirty-one device-related complications occurred, and all were secondary to fluid leakage. The Mentor Alpha-1 prosthesis was significantly better than the Mentor Mark-II in terms of mechanical reliability (P = 0.01). A trend was noted toward the AMS 700 Ultrex inflatable penile prosthesis having fewer mechanical complications than the Mentor Mark-II (P = 0.06). In addition, a trend toward all three-piece prostheses being more mechanically reliable than the two-piece was noted (P = 0.08). The Mentor Alpha-1 device had a higher cumulative proportional survival (0.957) than all other devices (0.842 for AMS 700 Ultrex, 0.839 for AMS 700 CX, 0.783 for Mentor GFS, and 0.750 for Mentor Mark-II). CONCLUSIONS: As a group, a trend was noted toward the three-piece prosthesis having better mechanical reliability than the two-piece prosthesis. Comparisons between the individual types of prostheses showed thatthe Mentor Alpha-1 device was significantly more mechanically reliable than the Mentor Mark-II device, and a trend was noted toward the AMS 700 Ultrex device having fewer mechanical complications than the Mentor Mark-II. The Mentor Alpha-1 prosthesis had the highest cumulative proportional survival.

Adult↗

An hypothesis about redundancy and reliability in the brains of higher species: analogies with genes, internal organs, and engineering systems.

The phenomena of behavioral resistance to massive brain damage and behavioral recovery from brain damage suggest there is redundancy in neural tissue. This paper uses basic concepts from probability theory and reliability engineering, as a first step toward more rigorously establishing the plausibility of the redundancy hypothesis. Exponential effects in the relevant formulas lead to results that are intuitively surprising. Thus, within a broad range of parametric assumptions related to lifespan and number of neurons or neural subsystems, it appears that the human brain may be at least twice as large as it would have to be for short-term survival. Simple reliability models suggest that redundancies are in parallel connections of smallest subsystems, such as individual neurons. Other implications of the basic formulas concern the relation between backed-up subsystem reliability and lifetime usage frequency for each subsystem, and the evolution of approximately equal allocation of lifetime reliability among components of a system. In addition, the paper briefly reviews more complex reliability engineering approaches. Redundancy as a reason for neural mass action is compared to other theoretical reasons for mass action in sensorimotor function and learning. Relationships of the present hypothesis to other theories of recovery from brain damage and to theories of regressive trophic phenomena in ontogeny are briefly discussed; it is suggested that as stages of ontogeny progress, both redundancy and flexibility in simpler behavioral functions are traded away for a larger, more differentiated repertoire of complex functions and memories.

Animals↗

Reliability of detection of lumbar lateral shift.

BACKGROUND AND PURPOSE: The poor reliability of lateral shift detection has been attributed to lack of rater training, biologic variation, and test reactivity. This study aimed to remove the potential confounding arising from biological variation and test reactivity and control the level of rater experience/training in making judgments of lateral shift. SUBJECTS: One hundred forty-eight raters with 3 levels of clinical physical therapy experience and training in the McKenzie method participated. METHOD: The raters viewed photographic slides of 45 patients with low back pain. Slides were judged on a numerical scale for presence and direction of a shift. Intrarater reliability was evaluated using the intraclass correlation coefficient (ICC) and interrater reliability was evaluated using both the ICC and kappa statistic. RESULTS: Reliability of shift judgments was only moderate for all groups (eg, ICC [2,1] values ranged from 0.48 to 0.64). CONCLUSION: Lateral shift judgements have only moderate reliability, even when trained raters judge stable stimuli. We propose that the photo model employed can be used to explore the source of error in this process.

Adult↗

Interexaminer reliability in physical examination of the cervical spine.

BACKGROUND: Most of the studies of physical examinations of the cervical spine have shown poor reliability. PURPOSE: To assess the interexaminer reliability in physical examinations of the cervical spine. SUBJECTS: Forty-eight subjects, age range 18 to 63 years. METHODS: Two physiotherapists independently evaluated a number of clinical tests of passive general and intersegmental movement. RESULTS: Acceptable kappa/kappa (w) values were obtained in several of the clinical tests of passive general motion range but in few of the clinical tests of passive intersegmental movement. More clinical tests had acceptable reliability and less bias in symptomatic subjects than asymptomatic subjects. CONCLUSION: Many of the clinical tests of passive general motion range were shown to be reliable. The increased number of acceptable kappa (w) values obtained in the symptomatic subjects indicates that further studies of the reliability of the clinical tests of passive intersegmental movement should be performed on patients.

Adolescent↗

Reliability of visual field results over repeated testing.

Fifty-one normal subjects, 337 with ocular hypertension, and 55 patients with glaucoma underwent C-30-2 testing on the Humphrey Field Analyzer on at least three occasions over a 6-year period. The time between tests was approximately 1 year. Using the manufacturer's standard for a reliable field (false-positive and false-negative rates, less than 33%; fixation losses, less than 20%), no trends in the proportion of reliable fields or the component indices were observed over time. Four percent of normal subjects, 9% of those with ocular hypertension, and 8% of patients with glaucoma were unable to meet the reliability standard every time they were tested. This repeated lack of reliability was due almost exclusively to fixation losses. However, patients with glaucoma were more likely to have repeatedly high false-negative responses than those with ocular hypertension or normal subjects, providing further evidence that false-negative responses are more indicative of glaucoma than of patient reliability.

Adult↗