Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Head-at-risk signs in Legg-Calvé-Perthes disease: poor inter- and intra-observer reliability.

BACKGROUND: The head-at-risk signs are used as prognostic indicators in Legg-Calvé-Perthes disease. These signs have been assessed only once regarding inter-observer reliability, however. Intra-observer reliability seems not to have been studied to date. METHOD: 76 anteroposterior pelvic radiographs of unilateral Legg-Calvé-Perthes disease were assessed by 5 observers on 2 occasions, in order to assess the inter- and intra-observer reliability in identifying head-at-risk signs. The observers included 1 consultant pediatric orthopaedic surgeon, 1 consultant radiologist, 2 specialist registrars and 1 senior house officer. Inter- and intra-observer reliabilities were assessed using the kappa coefficient. RESULTS: The intra-observer reliability was good for lateral subluxation and metaphyseal cystic changes, moderate for lateral calcification, and fair for Gage's sign and horizontal growth plate. The inter-observer reliability was moderate for lateral subluxation, fair for lateral calcification and metaphyseal cystic changes, and slight for Gage's sign and horizontal growth plate. INTERPRETATION: There was considerable variation in the diagnosis of the head-at-risk signs between observers. This makes the classification difficult to use in clinical practice.

Calcinosis↗

Low reliability of perceptual priming: consequences for the interpretation of functional dissociations between explicit and implicit memory.

In this study, three experiments are presented that investigate the reliability of memory measures. In Experiment 1, the well-known dissociation between explicit (recall, recognition) and implicit memory (picture clarification) as a function of age in a sample of 335 persons aged between 65 and 95 was replicated. Test-retest reliability was significantly lower in implicit than in explicit measures. In Experiment 2, parallel-test reliabilities in a student sample confirmed the finding of Experiment 1. In Experiment 3, the reliability of cued recall and word stem completion was investigated. There were significant priming effects and a dissociation between explicit and implicit memory as a function of levels of processing. However, the reliability of implicit memory measures was again substantially lower than in explicit tests in all test conditions. As a consequence, differential reliabilities of direct and indirect memory tests should be considered as a possible determinant of dissociations between explicit and implicit memory as a function of experimental or quasi-experimental manipulations.

Adult↗

How does a change in the administration method affect the reliability of the COOP/WONCA Charts? World Organization of National Colleges, Academies and Academic Associations of General Practitioners/Family Physicians.

BACKGROUND: An interviewer is often needed to administer the COOP/WONCA Charts to Chinese patients, and this may affect the reliability of results. OBJECTIVES: We aimed to find out the reliability of the COOP/WONCA Charts administered by an interviewer, and whether a change in the interviewer or administration method would affect the results. METHODS: We carried out a cross-sectional test-retest study on 487 Chinese adult patients attending a family medicine clinic in Hong Kong. The COOP/WONCA Charts were administered by the same interviewer, two different interviewers or self-completion and interviewer administration, on test and retest. The random, inter-observer and inter-method variances were compared with the inter-subject variance. The reliability coefficient of each COOP/WONCA Chart was calculated for each method of administration. RESULTS: Random errors could change the scores by 0.57-1.04, inter-observer variations could change the scores of four charts by 0.72-0.80, and a change in the method could change the physical fitness score by 1.79 and the daily activities score by 1.31, on a five-point scale. The reliability coefficients of the six COOP/WONCA Charts were 0.68-0.92 for one interviewer, 0.59-0.82 for two interviewers and 0.46-0.81 for two methods. CONCLUSION: The Chinese COOP/WONCA Charts were reliable in detecting real differences when administered by an interviewer. A change in the method of administration significantly decreased the reliability of the results. The use of more than one method of data collection in the same survey should be discouraged.

Adult↗

Reliability of a single-session isokinetic and isometric strength measurement protocol in older men.

BACKGROUND: The purposes of the current study were (a) to determine the test-retest reliability of a single-session isokinetic and isometric strength testing protocol in older healthy men, and (b) to compare the outcomes of the reliability measures derived from averaged torque scores with those derived from a single peak torque score. METHODS: In 19 men (mean age, 72 +/- 5 years), both lower limbs were assessed independently on 2 separate test days using the Biodex System 3 dynamometer. After completing a 5-minute warm-up, each man performed three submaximal knee extensions followed by five maximal contractions at 90 degrees /s (CON), 0 degrees /s (ISO), and -90 degrees /s (ECC). Average (best 3 of 5) and peak CON, ISO, and ECC torque, and CON work and CON power were determined. Peak CON work and CON power were recorded from the highest peak torque concentric contraction (HPTCC). RESULTS: Intraclass correlation coefficients ranging from 0.84 to 0.94 were found to have good reliability. The typical error as a coefficient of variation ranged from 8% to 10% for averaged measures and from 8% to 17% for peak torque and HPTCC. The ratio limits of agreement for average and peak CON, ISO, and ECC torque ranged from 23% to 33% and from 40% to 54% for average CON and HPTCC work and power. CONCLUSIONS: The test-retest reliability of a single-session isokinetic and isometric strength testing protocol in this group of older healthy men displayed good relative reliability (intraclass correlation coefficient > 0.84); however, because the typical error as a coefficient of variation and ratio limits of agreement (absolute reliability) were large, single-session testing is not recommended.

Aged↗

Assessing the reliability of a stage of change scale.

The purpose of this study was to assess the test-retest reliability of a scale measuring Prochaska's stages of change. Although structured questionnaire items are being increasingly used to segment target audiences according to Prochaska and DiClemente's stages of change, we could find only one report in the literature assessing the reliability of such scales. The unreliability of single-item or algorithm questionnaire scales might be why a number of studies show only minimal differences on some variables between individuals in different stages of change. A survey of the Perth metropolitan general population aged 16-69 years (N = 2629) was completed in August-September 1992 as part of a 3 year evaluation of the Western Australian Health Promotion Foundation. The consistency of respondents' responses was assessed across two questions measuring stages of change for the behaviours quitting smoking (n = 404), reducing alcohol consumption (n = 57) and doing more exercise (n = 704). Given the immediacy of the test-retest situation, the reliability results are moderately encouraging: kappa = 0.72, 0.73 and 0.52 for quitting smoking, reducing alcohol and doing more exercise, respectively. Health researchers should be aware of the probable moderate level of reliability if using the type of scale assessed in this study, when interpreting differences between individuals in different stages. In practice, several questionnaire items for classification purposes should be used so that internal reliability measures can be calculated. It is recommended that research be undertaken to devise more reliable scales for stages of change for the various health behaviours. It is noted that the attitude literature with respect to context and time specific intentions could be helpful in devising such scales.

Adolescent↗

The reliability of passive smoking histories reported in a case-control study of lung cancer.

A test-retest design has been used to examine the reliability of passive smoking histories reported in personal interviews. A total of 117 control subjects initially interviewed in a lung cancer case-control study conducted in metropolitan Toronto, Canada, between 1983 and 1984 were reinterviewed on average six months later. Responses to initial screening questions used to detect a person's exposure to passive smoke were more reliable for residential than for occupational exposure. Respondents also more reliably reported residential exposure to spouse's passive smoke than to the passive smoke of others at home. Quantitative measures of exposure to passive smoke, i.e., number and duration of exposure, were even less reliably reported. Nonsmoking respondents gave the most reliable information. The low reliability of self-reported duration of exposure to passive smoke is consistent with the inability of several studies to detect a significant dose-response relation with lung cancer risk when measures of dose that depend solely on duration are used.

Data Collection↗

The reliability of the structured interview for schizotypy-revised.

We investigated the reliability of the Structured Interview for Schizotypy-Revised (SIS-R). The original interview (SIS) was developed by Kendler. We revised the SIS, primarily by standardizing the rating procedures. Operational definitions and explicit criteria for rating were given. We introduced a four-point scale and provided clear criteria for rating severity of symptoms and signs (frequency, duration, and level of conviction) to operationalize schizotypal features. We divided schizotypal signs of global affect and global organization of speech into three separate signs of affect and five separate signs of thinking and speech. The main goal of this study was the assessment of test-retest reliability of the SIS-R. A robust test-retest design using different interviewers at both times, with a mean interval of 19 days, was used. The sample consisted of 42 psychiatric patients, almost all with personality disorders. The strong linear-weighted kappa statistic was used to evaluate reliability. The first conclusion is that most schizotypal symptoms can be reliably assessed with the SIS-R. The second conclusion is that most schizotypal signs do not reach sufficient levels of reliability. After unreliable items are excluded, the shortened SIS-R is a reliable research instrument for measuring schizotypal features (as far as it concerns our mixed samples). It covers all three dimensions of schizotypy.

Adult↗

Reliability of measuring trunk motions in centimeters.

A method of measuring trunk motion and two related motions using a tape measure and a stepstool was developed by physical therapists at our hospital. The purpose of this study was to assess the reliability of this method. Three repetitions of six motions performed by 24 subjects were each measured by three physical therapist raters on two separate days. The motions were forward bending, backward bending, right side bending, right rotation, right straight leg raising, and right prone knee bending. Reliability, standard deviation, and standard error were calculated for each motion. Only forward bending exhibited good single measurement reliability. Reliability coefficients for all motions were higher for the average of three successive measurements or for the measurement of a motion on successive days by the same rater. Measurements of rotation and straight leg raising, despite the improvement, continued to have low reliability. Analysis of variance was used to determine the significance of the differences between means for each motion across three raters, three repetitions, and two days. By looking at the analysis of variance and reliability estimates together, the authors identified two types of constant error affecting the data.

Adult↗

Reliability of observational measures of the Movement Assessment of Infants.

This study was conducted to examine the reliability of the Movement Assessment of Infants (MAI), a recently published neuromotor assessment tool. Interobserver and test-retest reliability data were collected on 27 full-term and 26 preterm 4-month-old infants. Reliability coefficients (Pearson r) were calculated for both the total-risk scores and the section-risk scores on the MAI. The total-risk score was calculated by summing the questionable or abnormal ratings on each of the 65 test items. For each of the four sections of the test, tone, primitive reflexes, automatic reactions, and volitional movement, an individual section-risk score was computed in a similar manner. Fair reliability was demonstrated for the total-risk scores (interobserver: r = .72; test-retest: r = .76). Section-risk score coefficients yielded a wide range of values for both interobserver and test-retest reliabilities (poor to good reliability). These measures provide needed technical data for therapists using this test and will assist the authors of the MAI in their attempts to improve the clinical validity of this assessment tool.

Evaluation Studies as Topic↗

Reliability of the attraction method for measuring lumbar spine backward bending.

The distraction method is one method used to measure forward bending of the spine. Although this technique, which requires the use of a tape measure held over the spine and the location of anatomical landmarks, appears to be highly practical, previous studies have not examined its use for measuring backward bending. The purpose of our study was to determine the reliability of a similar technique, the attraction method, for measuring backward bending of the lumbar spine and to examine whether subjects with low back pain (LBP) could perform similar motion as subjects without LBP. Two groups composed of 100 subjects each, one with "significant" limiting low back pain (SLBP) and the other without "significant" limiting low back pain (NSLBP), were evaluated twice by a physical therapist to assess intrarater reliability. To assess interrater reliability, 11 subjects from the NSLBP Group were evaluated by a second therapist. For the total sample of 200 subjects, the intraclass correlation coefficient (ICC) for intrarater reliability was .95; for the SLBP Group, the ICC was .93; and for the NSLBP Group, the ICC was .90. For the sample of 11 NSLBP Group subjects examined for interrater reliability, the ICC was .94. Using a Kolmogorov-Smirnov test, we found the distribution for backward bending of the two groups to be significantly different. The attraction method, thus, appears to be a reliable method for measuring backward bending of the lumbar spine.

Adolescent↗

Intrarater reliability of manual muscle testing and hand-held dynametric muscle testing.

Physical therapists require an accurate, reliable method for measuring muscle strength. They often use manual muscle testing or hand-held dynametric muscle testing (DMT), but few studies document the reliability of MMT or compare the reliability of the two types of testing. We designed this study to determine the intrarater reliability of MMT and DMT. A physical therapist performed manual and dynametric strength tests of the same five muscle groups on 11 patients and then repeated the tests two days later. The correlation coefficients were high and significantly different from zero for four muscle groups tested dynametrically and for two muscle groups tested manually. The test-retest reliability coefficients for two muscle groups tested manually could not be calculated because the values between subjects were identical. We concluded that both MMT and DMT are reliable testing methods, given the conditions described in this study. Both testing methods have specific applications and limitations, which we discuss.

Adult↗

Reliability of the Modified Motor Assessment Scale and the Barthel Index.

Many physical therapists use descriptive and functional assessments of motor recovery for patients with stroke. The purpose of this study was to establish the reliability of two such assessments. The Modified Motor Assessment Scale (MMAS) assesses motor recovery; the Barthel Index assesses functional independence. Interrater and intrarater reliability were determined for the total scores and individual item ratings using videotaped MMAS and Barthel Index assessments of seven patients with stroke. Therapists viewed and rated the videotaped assessments on two occasions separated by one month. The intrarater reliability results were higher than the interrater reliability results for total scores, and both results were acceptable statistically. Interrater and intrarater reliability of the individual item ratings were also determined. The MMAS and Barthel Index are reliable assessments of motor recovery and function for patients with stroke. Physical therapists are encouraged to use the two scales to document changes in the motor recovery and functional independence of patients with stroke.

Activities of Daily Living↗

Reliability of clinical measurements of lumbar lordosis taken with a flexible rule.

The purpose of this study was to examine the intratester and intertester reliability of lumbar lordosis measurements taken with a flexible rule. Two physical therapists (Tester 1 and Tester 2) took measurements on 40 subjects without low back pain (LBP) and on 40 subjects with LBP. Intraclass correlation coefficients (ICCs) were used to determine the degree of agreement between repeated measurements taken by the same therapist and between measurements taken by the two therapists. The ICC values for intratester reliability of Tester 1 were .84 for subjects without LBP and .94 for subjects with LBP. The ICC values of Tester 2 were .73 for subjects without LBP and .83 for subjects with LBP. Intertester reliability generally was poor, with ICC values of .41 for subjects without LBP and .50 for subjects with LBP. The results suggest that measurements of lumbar lordosis with a flexible rule may be reliable if taken by the same physical therapist. The degree of reliability, however, may vary from therapist to therapist. The intertester reliability of these measurements appears to be poor, but these conclusions must be interpreted carefully because of the limited number of therapists participating in this study.

Adult↗

Intrarater reliability of manual muscle test (Medical Research Council scale) grades in Duchenne's muscular dystrophy.

The purpose of this study was to document the intrarater reliability of manual muscle test (MMT) grades in assessing muscle strength in patients with Duchenne's muscular dystrophy (DMD). Subjects were 102 boys, aged 5 to 15 years, who were participating in a double-blind, multicenter trial to document the effects of prednisone on muscle strength in patients with DMD. Four physical therapists participated in the study. Two identical (duplicate) evaluations were performed within 5 days of each other by the same examiner initially and after 6 and 12 months of treatment. A total of 18 muscle groups were tested on each patient, 16 of them bilaterally, using a modification of the Medical Research Council scale. Reliability of muscle strength grades obtained for individual muscle groups and of individual muscle strength grades was analyzed using Cohen's weighted Kappa. The reliability of grades for individual muscle groups ranged from .65 to .93, with the proximal muscles having the higher reliability values. The reliability of individual muscle strength grades ranged from .80 to .99, with those in the gravity-eliminated range scoring the highest. We conclude the MMT grades are reliable for assessing muscle strength in boys with DMD when consecutive evaluations are performed by the same physical therapist.

Adolescent↗

Reliability of lumbar isometric torque in patients with chronic low back pain.

In this study, the test-retest reliability of lumbar isometric strength testing in patients with chronic low back pain (CLBP) was assessed. Isometric torque measurements were obtained from 89 patients with CLBP at seven different angles of lumbar flexion. Because previous studies have demonstrated significant strength differences between male and female subjects, separate data analyses were performed for each gender. Results indicated moderate to high reliability for patients with CLBP when tested at individually determined angles of flexion within their idiosyncratic range of motion (ROM) (female subjects: r = .59-.96, P less than .05, SEE = 12.0-24.2 N.m; male subjects: r = .71-.93, P less than .05, SEE = 25.1-62.1 N.m). For comparison with previously published data on asymptomatic controls, an additional set of analyses was conducted for subjects with full lumbar ROM. Similar reliability was demonstrated for this subsample (female subjects: r = .57-.93, P less than .05, SEE = 12.4-27.9 N.m; male subjects: r = .63-.93, P less than .05, SEE = 34.2-44.2 N.m). The authors concluded that isometric lumbar extension torque could be reliably measured in patients with CLBP at multiple positions within the full ROM, although reliability decreased at the most extended positions. The demonstrated reliability will allow researchers to assess treatment effects and group differences without undue concern for artifact attributable to measurement error.

Back Pain↗

Reliability of the scores for the finger-to-nose test in adults with traumatic brain injury.

BACKGROUND AND PURPOSE: The purpose of this study was to determine the intrarater and interrater reliability of measurements of three clinical features of coordination based on the performance of the "finger-to-nose" test. SUBJECTS: Thirty-seven persons with traumatic brain injury (26 male, 11 female), aged 17 to 64 years (mean = 29.1, SD = 9.9), participated in the study. METHODS: Each subject's performance was videotaped and evaluated for the right and left upper extremities (UEs) (two trials each) with respect to the following variables: time of execution, degree of dysmetria, and degree of tremor (four-point ordinal ratings). One year later, five experienced physical therapists (including the original investigator) independently rated each patient's videotaped performance in the same manner as described above. RESULTS: Intraclass correlation coefficients (ICC[3,1]) for intrarater reliability were .971 and .986 and ICCs for interrater reliability were .920 and .913 for right and left UEs, respectively, for the time of execution. A generalized Kappa statistic of .54 was calculated for the scoring of dysmetria (both UEs), and Kappa statistics calculated for the scoring of tremor were .18 and .31 for right and left UEs, respectively. Interrater reliability was lower for the scoring of these variables and varied from .36 to .40 for dysmetria and from .27 to .26 for tremor (right and left UEs, respectively). CONCLUSION AND DISCUSSION: These results indicate that physical therapists demonstrate low reliability in assessment of the presence of dysmetria and tremor using videotaped performances of the finger-to-nose test. The results suggest, however, that therapists reliably measure the time of execution of this test. If the limitations associated with therapists' capacity for objective measurement of subjective phenomena cannot be overcome (eg, by establishment of more definitive scoring criteria for the measures of dysmetria and tremor), then therapists should seek alternative methods of evaluation of UE coordination.

Adolescent↗

Reliability of clinical pressure-pain algometric measurements obtained on consecutive days.

BACKGROUND AND PURPOSE: Algometers have been used to measure muscle and other soft tissue tenderness. The purpose of this study was to investigate (1) "normal" pressure-pain threshold (PPT) in the biceps brachii muscle, (2) the reliability of repeated measurements of PPT in subjects without pain over 3 consecutive days, (3) the reliability of measurements of PPT between examiners, and (4) the number of measurements required to obtain a best estimate of PPT. SUBJECTS: Thirty-five subjects participated in the study. METHODS: Pain-pressure threshold of the biceps brachii muscle was measured using a Fischer algometer. Three test trials were done on each subject on each of 3 days by each of two examiners. Intraclass correlation coefficients (ICCs) and graphical methods were used to analyze the results. RESULTS: The ICCs revealed almost perfect reliability for measurements of PPT within and across 3 days and substantial reliability between examiners. The best estimate of PPT was obtained using the mean of the second and third trials each day. Graphical methods demonstrated that agreement between examiners was greatest at low mean pain thresholds. There was no effect for order of examiner. CONCLUSION AND DISCUSSION: The PPT is a reliable measure, and repeated algometry does not change pain threshold in healthy muscle over 3 consecutive days. The PPT can be used to evaluate the development and decline of experimentally induced muscle tenderness. Reliability is enhanced when all measurements are taken by one examiner.

Adult↗

Reliability of measurements obtained with the modified Ashworth scale in the lower extremities of people with stroke.

BACKGROUND AND PURPOSE: Abnormal muscle tone is a common motor disorder following stroke, which may require rehabilitation. The Modified Ashworth Scale is a 6-point rating scale that is used to measure muscle tone. The interrater and intrarater reliability of measurements obtained with the scale remain equivocal. The purpose of this study was to investigate the reliability of measurements obtained with the scale in the lower limb of patients with stroke. SUBJECTS: Twenty patients were tested 2 weeks after their stroke, and 12 patients were tested 12 weeks after their stroke. METHODS: Gastrocnemius, soleus, and quadriceps femoris muscles on the hemiplegic side were tested. RESULTS: Interrater reliability for 2 raters was poor, with a Kendall tau-b correlation for the combined muscle group of.062 (P=.461). For intrarater reliability, the Kendall tau-b correlation was.567 (P<.001). The agreement within one rater occurred mostly on the grade of 0. DISCUSSION AND CONCLUSION: The Modified Ashworth Scale yielded reliable measurements in the lower limb for a single examiner, and agreement was best on the grade of 0. The reliability between examiners was not good, which may bring into question the validity of measurements obtained with the scale.

Aged↗