Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Retest reliability of scores on objective and projective measures of dependency: relationship to life events and interest interval.

The retest reliabilities of widely used objective and projective measures of dependency were assessed in a mixed-sex sample of undergraduates (54 women and 34 men). Subjects completed Hirschfeld and colleagues' (1977) Interpersonal Dependency Inventory (IDI) and Masling, Rabie, and Blondheim's (1967) Rorschach Oral Dependency (ROD) scale on two occasions separated by 16, 28, or 60 weeks. The IDI and ROD scale showed good retest reliability over 16 weeks in both men and women. However, the ROD scale did not show adequate retest reliability over longer periods in subjects of either sex. IDI scores showed excellent long-term retest reliability in women, but poor long-term retest reliability in men. Subjects' self-reports and impact ratings of life events experienced during the intertest period were unrelated to changes in subjects' IDI and ROD scale scores from Time 1 to Time 2, regardless of the intertest interval used.

Adolescent↗

Factors affecting the reliability of clinical judgments about the function of children's school-refusal behavior.

Conducted two studies to examine the interrater reliability, test-retest stability, and the effect of various clinician variables, such as years of clinical experience, theoretical orientation, and prior experience with children, on clinical judgments about the reinforcement functions of children's school-refusal behavior. Results indicated that the judgments by individual clinicians were of questionable reliability. Judgments aggregated across 3 clinicians yielded acceptable interrater and test-retest reliability in Study 1, but a greater number of clinicians were necessary to achieve acceptable reliability in Study 2. Years of clinical experience and training were the only clinician variables related to the reliability of judgments about reinforcement functions. Several recommendations for the clinical assessment of the function of children's school-refusal behavior are discussed.

Adolescent↗

Intrarater and interrater reliability of the MS functional composite outcome measure.

OBJECTIVE: To assess practice effects, and intrarater and interrater reliability of the MS functional composite (MSFC) outcome measure. BACKGROUND: To address the poor reliability and insensitivity to change of available MS clinical rating scales, the National MS Society's Clinical Outcomes Assessment Task Force developed the MSFC, a multidimensional quantitative clinical outcome measure that includes tests of leg function/ambulation (Timed 25-Foot Walk), arm function (Nine-Hole Peg Test), and cognitive function (Paced Auditory Serial Addition Test). METHODS: Ten patients with secondary progressive MS underwent six testing sessions over a 2-week period. The MSFC was administered by the same examining technician in the first five sessions and by the other technician in the sixth. Patients were reassessed by both technicians after 6 months (sessions 7 and 8). The MSFC score was calculated as the mean of the Z scores of the three components. A pooled dataset derived from secondary progressive MS patients in the placebo arms of previous clinical trials and natural history studies served as the reference population to standardize scores. RESULTS: Practice effects were evident initially but stabilized by the fourth administration. The intraclass correlation coefficient (ICC) was 0.97 for the MSFC for session 4 versus session 5 (intrarater reliability). The ICC was 0.95 for session 5 versus session 6 (interrater reliability), and was 0.96 for session 7 versus session 8 when patients were reassessed 6 months later. CONCLUSIONS: The MS functional composite (MSFC) outcome measure had excellent intrarater and interrater reliability when standardized procedures were used to train examining technicians and to assess patients. Prebaseline testing sessions should be included in clinical trials employing the MSFC to compensate for practice effects.

Humans↗

A videotaped CIBIC for dementia patients: validity and reliability in a simulated clinical trial.

BACKGROUND: The global impression of a clinician is an Food and Drug Administration--mandated primary outcome measure for clinical trials in dementia. Reliability and validity of these measures are not well established. METHODS: A videotaped version of the Clinician's Interview Based Impression of Change (CIBIC) was evaluated. Raters were informed that the videotaped interviews were taken at baseline and 6 to 12 months later, when in fact half of the interviews were shown in reverse order. Ratings on "true order" interviews were compared with ratings on "reverse order" interviews. In addition, ratings by neurologists experienced in dementia were compared with those of less experienced raters. RESULTS: Inter-rater reliability of the neurologists was poor when measured by absolute agreement on a 7-point scale (kappa = 0.18). With a less stringent 3-point scale (better, worse, or unchanged), inter-rater reliability was significantly better for the true order videos (kappa = 0.51) than for the reversed order videos (kappa = 0.12). Validity also was reduced in the reverse order group: neurologists rated 90% of subjects correctly in the "true order" group and 63% correctly in the "reversed order" group. The inter-rater reliability of the neurologists was greater than the less experienced raters, but the validity of the neurologists' ratings was only marginally better. CONCLUSIONS: The reliability and validity of the videotape CIBIC are reasonable when patients follow the expected course of gradual decline, but are poor when patients appear to improve. These findings suggest that global assessments should be modified as outcome measures in clinical trials with patients with dementia.

Aged↗

Reliability of a measure of prediagnosis physical activity for cancer survivors.

PURPOSE: This study was conducted to examine the test-retest reliability of a measure of prediagnosis physical activity participation administered to colorectal cancer survivors recruited from a population-based state cancer registry. METHODS: A total of 112 participants completed two telephone interviews, 1 month apart, reporting usual weekly physical activity in the year before their cancer diagnosis. Intraclass correlation coefficients (ICC) and standard error of measurement (SEM) were used to describe the test-retest reliability of the measure across the sample; the Bland-Altman approach was used to describe reliability at the individual level. The test-retest reliability for categorized total physical activity (active, insufficiently active, sedentary) was assessed using the kappa statistic. RESULTS: When the complete sample was considered, the ICC ranged from 0.40 (95% CI: 0.24, 0.55) for vigorous gardening to 0.77 (95% CI: 0.68, 0.84) for moderate physical activity. The SEM, however, were large, indicating high measurement error. The Bland-Altman plots indicated that the reproducibility of data decreases as the amount of physical activity reported each week increases. The kappa coefficient for the categorized data was 0.62 (95% CI: 0.48, 0.76). CONCLUSION: Overall, the results indicated low levels of repeatability for this measure of historical physical activity. Categorizing participants as active, insufficiently active, or sedentary provides a higher level of test-retest reliability.

Adult↗

Comparison of the validity and reliability of two image classification systems for the assessment of mammogram quality.

OBJECTIVE: To compare the reliability and validity of two classification systems used to evaluate the quality of mammograms: PGMI ('perfect', 'good', 'moderate' and 'inadequate') and EAR ('excellent', 'acceptable' and 'repeat'). SETTING: New South Wales (Australia) population-based mammography screening programme (BreastScreen NSW). METHODS: Thirty sets of mammograms were rated by 21 radiographers and an expert panel. PGMI and EAR criteria were used to assign ratings to the medio-lateral oblique (MLO) and cranio-caudal (CC) views for each set of films. Inter-observer reliability and criterion validity (compared with expert panel ratings) were assessed using mean weighted observed agreement and kappa statistics. RESULTS: Reliability: Kappa values for both classification systems were low (0.01-0.17). PGMI produced significantly higher values than EAR. Agreement between raters was higher using PGMI than EAR for the MLO view (77% versus 74%, P < 0.05), but was similar for the CC view. Dichotomized ratings ('acceptable' or 'needs repeating') did not improve reliability estimates. VALIDITY: Kappa values between raters and the reference standard were low for both classification systems (0.05-0.15). Agreement between raters and the reference standard was higher using PGMI than EAR for the MLO view (74% versus 63%), but was similar for the CC view. Dichotomized ratings of the MLO view showed slightly higher observer agreement. CONCLUSIONS: Both PGMI and EAR have poor reliability and validity in evaluating mammogram quality. EAR is not a suitable alternative to PGMI, which must be improved if it is to be useful.

Breast Neoplasms↗

Reliability of radiological measurements in the assessment of hip dysplasia in adults.

Radiographic measurements are commonly used to quantify the treatment results of hip dysplasia and assess further need of operative treatment. We investigated the interobserver and intraobserver reliability of the commonest radiographic techniques in the assessment of hip dysplasia in skeletally mature adults. Three observers independently analysed 100 hip radiographs of patients with hip dysplasia aged between 16 and 32 years. We measured centre-edge angle of Wiberg, acetabular angle of Sharp, acetabular index of the weightbearing zone, acetabular index of depth to width, ACM-angle, MZ-distance, acetabular head index, lateral subluxation and neck-shaft angle. In addition, the radiographs were reviewed a second time 3 months apart by two of the observers to assess intraobserver reliability. We found a high correlation (intraclass correlation coefficient) for interobserver reliability (0.76-0.87) and intraobserver reliability (0.70-0.92) for all radiographic measurements except acetabular index of depth to width, ACM-angle and MZ-distance. Depending on the clinical question we therefore recommend the use of one of the reliable measurements to assess the radiograph of a dysplastic hip.

Adolescent↗

Reliability of a measure of muscle extensibility in fullterm and preterm newborns.

PURPOSE: The purpose of this study was to examine the test-retest and inter-rater reliability of a measure of muscle extensibility developed by Tardieu, de la Tour, Bret, and Tardieu (1982) in fullterm and preterm newborns. METHOD: Twenty-one fullterm infants and twenty preterm infants were examined by two physical therapists. Each physical therapist measured AO (shortened position of the muscle belly and lengthened tendon) and AMax (maximum muscle belly and tendon length) of the gastrocnemius/ soleus muscle twice in succession. Reliability was assessed using intraclass correlation coefficients (2,2) and (3,2). RESULTS: Inter-rater reliability between the two examiners ranged from .86 to .97 and test-retest reliability on the two measures ranged from .91 to .98. CONCLUSIONS: The results suggest that this measure of muscle extensibility is reliable in the gastrocnemius/soleus muscle with fullterm and preterm newborns. Further research is needed to investigate if differences in muscle extensibility are present between fullterm and preterm infants and the relationship between muscle extensibility and active movement.

Ankle↗

Using the eating disorder examination in the assessment of bulimia and anorexia: issues of reliability and validity.

OBJECTIVE: The Eating Disorder Examination will be assessed according to its reliability and validity in the assessment of anorexia nervosa and bulimia nervosa. METHOD: A thorough review of the literature was conducted to judge the reliability and validity of the Eating Disorder Examination and its subscales. RESULTS: The review shows that the EDE and its subscales have good interrater reliability and internal consistency reliability. Similarly, high levels of discriminant validity, construct validity, and treatment validity in the assessment of eating disorders were also found. A summary of each study concerning the various types of reliability and validity will be provided. CONCLUSIONS: The EDE is considered to be the "gold standard" by which to identify eating disorders, so this tool used in conjunction with other behavioral measures will be imperative for clinical social work practice.

Anorexia↗

Are the Naranjo criteria reliable and valid for determination of adverse drug reactions in the intensive care unit?

BACKGROUND: The Naranjo criteria are frequently used for determination of causality for suspected adverse drug reactions (ADRs); however, the psychometric properties have not been studied in the critically ill. OBJECTIVE: To evaluate the reliability and validity of the Naranjo criteria for ADR determination in the intensive care unit (ICU). METHODS: All patients admitted to a surgical ICU during a 3-month period were enrolled. Four raters independently reviewed 142 suspected ADRs using the Naranjo criteria (review 1). Raters evaluated the 142 suspected ADRs 3-4 weeks later, again using the Naranjo criteria (review 2). Inter-rater reliability was tested using the kappa statistic. The weighted kappa statistic was calculated between reviews 1 and 2 for the intra-rater reliability of each rater. Cronbach alpha was computed to assess the inter-item consistency correlation. The Naranjo criteria were compared with expert opinion for criterion validity for each rater and reported as a Spearman rank (r(s)) coefficient. RESULTS: The kappa statistic ranged from 0.14 to 0.33, reflecting poor inter-rater agreement. The weighted kappa within raters was 0.5402-0.9371. The Cronbach alpha ranged from 0.443 to 0.660, which is considered moderate to good. The r(s) coefficient range was 0.385-0.545; all r(s) coefficients were statistically significant (p < 0.05). CONCLUSIONS: Inter-rater reliability is marginal; however, within-rater evaluation appears to be consistent. The inter-item correlation is expected to be higher since all questions pertain to ADRs. Overall, the Naranjo criteria need modification for use in the ICU to improve reliability, validity, and clinical usefulness.

Adverse Drug Reaction Reporting Systems↗

Examining group differences in reliability of multiple-component instruments.

A covariance structure modelling method for examining group differences in scale reliability of multi-component measuring instruments is proposed. The procedure is based on a set of appropriate parameter constraints imposed in a multiple-population structural model. Unlike tests of group differences in coefficient alpha that in general do not evaluate the discrepancies in scale reliability coefficients, the described approach permits one to examine differences in the latter quantities of actual interest when concerned with questions pertaining to reliability of composite scores. The method can be used to test whether a given measuring instrument has identical reliabilities in studied populations, and/or whether distinct instruments, or modes of administration of the same instrument, have equal reliabilities across groups. The proposed procedure is illustrated with a pair of examples.

Humans↗

Reliability and validity of generalizable skills instruments for students who are deaf, blind, or visually impaired.

The study examined the validity and reliability of four assessments, with three instruments per domain. Domains included generalizable mathematics, communication, interpersonal relations, and reasoning skills. Participants were deaf, legally blind, or visually impaired students enrolled in vocational classes at residential secondary schools. The researchers estimated the internal consistency reliability, test-retest reliability, and construct validity correlations of three subinstruments: student self-ratings, teacher ratings, and performance assessments. The data suggest that these instruments are highly internally consistent measures of generalizable vocational skills. Four performance assessments have high-to-moderate test-retest reliability estimates, and were generally considered to possess acceptable validity and reliability.

Blindness↗

Assessment of interrater and intrarater reliability in the evaluation of metered dose inhaler technique.

STUDY OBJECTIVE: To determine if a training session using videotaped metered dose inhaler (MDI) performances can result in high interrater and intrarater reliability of five evaluators assessing MDI technique. DESIGN: Five evaluators (three pharmacists, two pulmonary fellows) were trained to evaluate MDI technique during a 2-h training session. The training session consisted of verbal instruction and practical experience in evaluating MDI technique using video-taped MDI performances of six nonstudy subjects. After the training session, the evaluators independently observed the same videotaped MDI demonstrations of 14 subjects on two occasions separated by a 7- to 10-day interval. Interrater and intrarater reliability was determined for individual steps by calculating percent agreement and intraclass correlation (ICC) coefficient. RESULTS: Interrater. The interrater reliability for individual steps ranged from 29 to 86 percent (ICC coefficient = 0.13 to 0.81). Steps in which evaluators were in agreement for less than 9 of the 14 subjects were shaking the inhaler before inhalation, exhaling, continuing to inhale slowly, and adequate breath hold. Intrarater: The overall percent agreement by step ranged from 74 to 97 percent. Exhaling to functional residual volume (76 percent) and continuing to inhale slowly and deeply (74 percent) had the lowest overall agreement between the first and second observation day. The consistency of evaluating a step between the two observation days varied considerably depending on the step and evaluator. CONCLUSIONS: High interrater and intrarater reliability in MDI evaluation is difficult to obtain. Clinicians and researchers involved in MDI evaluation and education should be trained to achieve consistency. A single training session using videotaped MDI demonstrations was not adequate in achieving consistency among evaluators. To improve accuracy of research results, researchers should include at least two evaluators to assess MDI technique or take other measures to show and report reliability.

Administration, Inhalation↗

Reliability, repeatability, and sensitivity of the modified shuttle test in adult cystic fibrosis.

STUDY OBJECTIVES: The purpose of this study was to investigate the test-retest reliability, repeatability, and sensitivity of the modified shuttle test (MST) in adult patients with cystic fibrosis (CF). DESIGN: : Prospective study. SETTING: Adult CF Unit, Belfast City Hospital. PATIENTS: : Adult patients with CF. INTERVENTIONS: Test-retest reliability-none; sensitivity-inpatient IV antibiotic therapy for an acute exacerbation of respiratory disease. MEASUREMENTS: The test-retest reliability and repeatability of the MST was assessed by comparing performance on two consecutive MSTs performed in 12 patients with CF and stable disease. The sensitivity of the MST was assessed by measuring the change in MST performance after 2 weeks of IV antibiotic therapy in 24 patients admitted to hospital with acute exacerbations of their respiratory disease. RESULTS: In the assessment of test-retest reliability and repeatability (n = 12), there was a significant and strong correlation between trials for distance completed (Pearson's r = 0. 99; p < 0.01), peak heart rate (Pearson's r = 0.99; p < 0.01), peak arterial oxygen saturation (SaO(2); Pearson's r = 0.99; p < 0.01), and peak Borg rating of perceived breathlessness (Pearson's r = 0. 99; p < 0.01). The coefficients of repeatability for these variables were small (coefficient of repeatability: distance completed, 4 shuttles; peak heart rate, 6 beats/min; peak SaO(2), 4%; and peak Borg rating of perceived breathlessness, 0.9). In the assessment of sensitivity (n = 24), the standardized response mean (SRM) for distance completed on MST (SRM = 1.18) was the SRMs for spirometric measures of lung function (FEV(1), SRM = 0.96; FEV(1) percent predicted, SRM = 0.88). CONCLUSIONS: This study demonstrates that the MST is a reliable, repeatable, and sensitive measure of exercise capacity in adult CF. The MST may be of value in determining prognosis, evaluation for lung transplantation, exercise prescription, and establishing the impact of new treatments on the disability associated with CF.

Adolescent↗

Reliability of selected physical performance tests in young adult women.

The purposes of this investigation were to establish the reliability of selected physical performance tests in women athletes and nonathletes and to determine performance differences between groups. Fifty women (25 athletes, 25 nonathletes) performed 5 tests in 2 sessions. The performance tests included the figure-eight hop test, up-and-down hop test, side-to-side hop test, hexagon hop test, and zigzag run test. Intraclass correlation coefficients (ICC [2, 1]) were calculated for trial-to-trial, intertester, and day-to-day reliability. Independent t-tests with Bonferroni adjustment (alpha = 0.01) were used for each individual test to compare differences between groups. All tests showed good reliability values (ICC > or = 0.76) in the nonathlete group for all conditions and varied reliability values (0.48-0.99) among conditions in the athlete group. The independent t-tests showed a statistically significant group effect (t > or = 3.041; p < or = 0.004) for all tests. The results showed that these physical performance tests are reliable measurement tools in the female population.

Adult↗

Reliability of power output during short-duration maximal-intensity intermittent cycling.

The aims of the present study were: (a) to determine the number of familiarization trials required to establish a high degree of reliability in measures of power output during maximal intermittent cycling; and (b) to examine the reliability of those same measures after familiarization had been established. On separate days over a 3-week period, 2 groups of 7 recreationally active men completed 8 trials of 1 of 2 maximal (20 x 5-second) intermittent cycling tests with contrasting recovery periods (10-seconds or 30-seconds). Significant (p < 0.05) between-trial differences were detected in post-hoc tests involving trials 1 and 2 only. Within-subject test-retest reliability was therefore assessed across trials 3-8. Apart from values of maximum power output in Protocol 1 (10-second recovery periods), all remaining measures of power output showed high degrees of within-subject test-retest reliability (coefficient of variation: 2.4-3.7%). The results of the present study indicate that in subjects unfamiliar with maximal intermittent cycling, high degrees of reliability in many performance measures can be achieved following the completion of 2 familiarization trials.

Adult↗

Reliability of unfamiliar, multijoint, uni- and bilateral strength tests: effects of load and laterality.

To better understand the reliability of unfamiliar multijoint strength tests, 16 resistance-trained men performed maximum velocity uni- (1 leg [1L]) and bilateral (2 legs [2L]) lifts on an unfamiliar semiprone leg squat machine with loads equivalent to 40 and 70% of maximum isometric force on 2 separate occasions. Peak force was highly reproducible between testing occasions at the heavy load under both uni- and bilateral conditions (intraclass correlation coefficient [ICC](1L70%) = 0.91, ICC(2L70%) = 0.92), was slightly reduced in the light load bilateral condition (ICC(2L40%) = 0.85), and was significantly (p < 0.05) reduced in the light load unilateral condition (ICC(1L40%) = 0.57). Test reliability was not related to total load lifted (2L 70% > 1L 70% > 2L 40% > 1L 40%) or to the peak force developed during the tests (2L 70% > 1L 70% = 2L 40% > 1L 40%), but it was somewhat related to the time taken to attain peak force (2L 70% = 1L 70% > 2L 40% > 1L 40%). To obtain reliable strength data from athletes, more familiarization seems to be needed when they perform modified versions of common multijoint strength tests, or unfamiliar strength tests, under light load, unilateral conditions. The marked differences in reliability resulting from variation in loading conditions suggests that the reliability of a test needs to be reestablished when it is modified, before it is used to assess athlete/subject strength performance.

Adult↗

Reliability of rate of velocity development and phase measures on an isokinetic device.

Isokinetic dynamometers have been measured for torque and force reliability in the past, but little research has been performed on rate of velocity development (RVD) measures. The purpose of this study was to determine the reliability of RVD measures on an isokinetic device at slow and fast speeds. Twenty volunteers performed 5 repetitions of concentric knee extension at 1.04 and 4.18 rad.s(-1) on a Kin-Com isokinetic dynamometer. Each subject was identically posttested 7 days later. Data were separated into 3 velocity range-of-motion (ROM) phases of RVD, load range (LR), and deceleration (DCCROM). Analyses of variance (ANOVAs) were performed to analyze the mean data between day 1 and day 2, while intraclass correlation coefficients (ICCs) were performed for reliability. Results at 1.04 rad.s(-1) demonstrated a low but significant (p < 0.05) ICC value (0.58) only for LR, while at 4.18 rad.s(-1) RVD (0.87), LR (0.83), and DCC (0.55) all exhibited significant ICC values. Percent error for high-speed testing ranged from 1.4-3.19%. No variable exhibited a significant mean difference between testing days. These results collectively point to moderate to high phase reliability for RVD measures at fast speeds of testing, while the slow speed showed very low reliability. Therefore, care should be exercised at slow speeds when comparing RVD measures from test to test.

Acceleration↗