Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Reliability of seven measures of social intelligence in a sample of adolescents with mental retardation.

This study evaluates the reliability of seven measures, selected to assess the social-cognitive variables hypothesized by Greenspan to define social intelligence. Responses from 75, 30 and 20 adolescents with mental retardation were used to assess each test's internal, interrater, and test-retest reliabilities, respectively. Interrater reliability coefficients were high to very high (.76 to .98), internal reliabilities were moderate to very high (.66 to .90), and test-retest reliabilities were moderate to high (.54 to .74). Internal and test-retest reliability coefficients compared favourably with those reported for the subtests of the Revised Wechsler Intelligence Scale for Children.

Activities of Daily Living↗

The sickness impact profile: SIP68, a short generic version. First evaluation of the reliability and reproducibility.

In previous research a short version of the Sickness Impact Profile (SIP136) was developed, containing 68 items. This SIP68 is intended as a short generic alternative to the original SIP. High reliability of the SIP68 was reported when it was extracted from the SIP136. This paper is a report on the first reliability testing of the SIP68 administered as an independent instrument without the context of the SIP136. To establish the test-retest reliability and the internal consistency of the new instrument, 51 patients of an outpatient department of rheumatology completed the SIP68 twice, with an interval of 48 hours. To compare the performance of the independent SIP68 with the SIP68 extracted from the SIP136, the SIP136 also was completed two times by the same 51 respondents. Test-retest reliability for both administration types was assessed by means of the intraclass correlation coefficient and the Jaccard's similarity ratio. Internal consistency was assessed by means of Cronbach's alpha. The reliability appears to be high in both the independent SIP68 as well as the extracted SIP68. Moreover, the reliability of the independent SIP68 appears to be as high as for the SIP136. These findings were very encouraging, indicating that the SIP68 may very well serve as a generic alternative to the SIP136.

Activities of Daily Living↗

Clinical evaluation of anterior vaginal wall support defects: interexaminer and intraexaminer reliability.

OBJECTIVE: The purpose of this study was to determine the interobserver and intraobserver reliability of the clinical examination of anterior vaginal wall support defects. STUDY DESIGN: Sixty-three patients with at least stage II anterior vaginal wall prolapse were prospectively evaluated with a standardized examination to detect anterior vaginal wall support defects. Interobserver reliability was assessed with a duplicate examination performed by a blinded second examiner. Intraobserver reliability was assessed with a second examination performed at least 3 weeks later by 1 of the original 2 examiners. Examination reliability for the 4 types of defects (central, right lateral, left lateral, and superior) was evaluated with the kappa statistic. RESULTS: The inter- and intraexaminer reliability of the clinical examination for central, superior, and right and left paravaginal defects was poor; all kappas were less than 0.50. Overall interexaminer agreement was 42% with a kappa of 0.16 (95% CI, 0-0.32). Overall intraexaminer agreement was 46% with a kappa of 0.16 (95% CI, 0-0.45). Reliability was noted to improve with increasing stage of prolapse. CONCLUSION: The clinical examination of anterior vaginal wall support defects displays poor interexaminer and intraexaminer agreement.

Aged↗

The reliability of multiple objective measures of surgery and the role of human performance.

BACKGROUND: There is a need for reliable and valid objective methods of technical skills in surgery. Six-bench surgical top stations have been combined to assess basic surgical trainees (BSTs) objectively. The current study examines its reliability and validity across repeat sittings. METHODS: Eleven surgical trainees (6 senior BSTs and 5 higher surgical trainees [HSTs]) undertook 5 sittings of the 6-station assessment designed to be completed within 90 minutes. The 6 stations consisted of knot tying, suturing, closure of enterotomy, excision of sebaceous cyst, laparoscopic task, and instrument examination. Methods of analysis employed were motion analysis, observation with criteria, and inbuilt simulation metrics. RESULTS: On analysis 3 knot tying and suturing stations exhibited significant differences in either time or movement; any difference was over by the second run. The intertest reliabilities were .66, .74, .55, .51, and .65 for the 5 runs. The intratest reliability across repeated sittings varied from .56 to .96. The inter-rater reliability for video assessment varied from .77 to .94. CONCLUSION: The assessment is reliable and valid across repeated sittings. Its use in assessment of basic technical skills needs to be encouraged.

Clinical Competence↗

Durability, value, and reliability of selected electric powered wheelchairs.

OBJECTIVE: To compare the durability, value, and reliability of selected electric powered wheelchairs (EPWs), purchased in 1998. DESIGN: Engineering standards tests of quality and performance. SETTING: A rehabilitation engineering center. SPECIMENS: Fifteen EPWs: 3 each of the Jazzy, Quickie, Lancer, Arrow, and Chairman models. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Wheelchairs were evaluated for durability (lifespan), value (durability, cost), and reliability (rate of repairs) using 2-drum and curb-drop machines in accordance with the standards of the American National Standards Institute and Rehabilitation Engineering and Assistive Technology Society of North America. RESULTS: The 5 brands differed significantly (P<or=.05) in durability, value, and reliability, except in terms of reliability of supplier repairs. The Arrow had the highest durability, value, and reliability in terms of the number of consumer failures, supplier failures, repairs, failures, consumer repairs and failures, and supplier repairs and failures. The Lancer had the poorest durability and reliability, and the Chairman had the lowest value. K0014 wheelchairs (Arrow, Permobil) were significantly more durable than K0011 wheelchairs (Jazzy, Quickie, Lancer). No significant differences in durability with respect to rear-wheel-drive (Arrow, Lancer, Quickie), mid-wheel-drive (Jazzy), or front-wheel-drive (Chairman) wheelchairs were found. CONCLUSIONS: The Arrow consistently outperformed the other wheelchairs in nearly every area studied, and K0014 wheelchairs were more durable than K0011 wheelchairs. These results can be used as an objective comparison guide for clinicians and consumers, as long as they are used in conjunction with other important selection criteria. Manufacturers can use these results as a guide for continued efforts to produce higher quality wheelchairs. Care should be taken when making comparisons, however, because the 5 brands had different features. Purchased in 1998, these models may be used for several more years. In addition, problem areas in these models may still be present in newer models.

Electricity↗

Intratester and intertester reliability of neck isometric dynamometry.

OBJECTIVE: To evaluate the reproducibility of measurement for maximum voluntary isometric contractions of the cervical musculature in different movements. DESIGN: Repeated test-retest measurements. SETTING: A department of physiotherapy. PARTICIPANTS: Thirty-three healthy subjects (17 men, 16 women; age range, 19-63 y) for the intraexaminer study and 10 healthy subjects (4 men, 6 women; age range, 20-37 y) for the interexaminer study. INTERVENTIONS: Maximum isometric strength in sitting and standing for flexion, extension, lateral flexion, and rotation using a custom isomyometer device. Three tests, performed 5 to 8 days apart, to assess intraexaminer reliability. Two examiners, each performing 1 trial, measuring on the same day to assess interexaminer reliability. MAIN OUTCOME MEASURES: Intraexaminer and interexaminer reliability of neck muscle strength. RESULTS: The standing position showed better reproducibility than the sitting position. The intraclass correlation coefficient (ICC1,3) was above .84 for all tests in any movement and position and above .93 when the first test was excluded. The standard error (SE) of measurement (<16.5 N; <.13 N-m for rotation) and smallest detectable difference (SDD) (<20.1%) were also small. For interexaminer reliability, the ICC(2,1) ranged from.88 to.94 and the SE from 10.7 to 20.8 N (<1.15 N-m for rotation); the SDD was less than 29.8% (except right rotation, which was 38.8%). CONCLUSIONS: A reliable protocol for measuring neck strength has been developed. Standing position and a full practice session produces more reliable measurements.

Adult↗

Reliability of the Dynamic Gait Index in individuals with multiple sclerosis.

OBJECTIVES: To determine if the Dynamic Gait Index (DGI) is a reliable tool for assessing balance in people with multiple sclerosis (MS) and to determine the validity of the DGI by using the 6.1-m timed walk. DESIGN: Instrument reliability test: physical therapists viewed a videotape of 10 subjects with MS performing the DGI and scored their gait by using DGI criteria. Two weeks after the first session, therapists' viewed the videotape again and scored subjects' gait to establish interrater reliability. SETTING: Hospital-based outpatient rehabilitation clinic. PARTICIPANTS: Eleven physical therapists and 10 people with MS. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Total DGI scores and each of the 8 DGI items were compared between and within raters (physical therapists). Time to walk 6.1m was compared with the total DGI score to examine concurrent validity. RESULTS: Interrater reliability for total DGI scores was .983, with each of the 8 items ranging from .910 to .976 (intraclass correlation coefficient, P<.05). Intrarater reliability for total DGI scores ranged between .760 and .986 (Pearson bivariate analysis, P<.05). An inverse relationship of -.801 (Pearson bivariate analysis, P<.01) existed between the total DGI scores and the 6.1-m walk. CONCLUSIONS: The DGI is a reliable functional assessment tool that correlated inversely with timed walk, showing its concurrent validity.

Adult↗

The reliability, distribution, and responsiveness of the Postural Control and Balance for Stroke Test.

OBJECTIVES: To determine the inter- and intrarater reliability of the Postural Control and Balance for Stroke (PCBS) test and to assess its distribution and responsiveness to changes during 1-year follow-up. DESIGN: Intrarater reliability of the PCBS test was assessed by comparing the repeat ratings of videotaped test performances by each of the 5 raters. Interrater reliability was assessed by comparing the ratings of the videotaped test performances between the raters. SETTING: Hospital neurologic ward and outpatient department of physiotherapy as well as health centers in Finland. PARTICIPANTS: Fifty stroke patients (age range, 42-89 y) were measured 7, 120, and 360 days poststroke for the study of distribution and responsiveness and 19 patients (age range, 55-85 y) were measured during a period between 7 and 60 days poststroke for the reliability study. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: The distributions of the scores of the PCBS test, intraclass correlation coefficient (ICC) and weighted kappa values, and Wilcoxon matched-pairs tests. RESULTS: The PCBS test had limited floor and ceiling effects at 7, 120, and 360 days poststroke. The differences between scores 7 and 120 days poststroke were significant (P<.001). The differences between scores 120 and 360 days poststroke were not significant (P>.05). The Cronbach alpha for all the items combined was .96. The ICC values for the interrater and intrarater reliability of the PCBS test were .94 and .96, respectively. CONCLUSIONS: The PCBS test showed an acceptable level of reliability and the responsiveness results indicated a good level before 120 days but not between 120 and 360 days after stroke.

Aged↗

Assessing walking ability in subjects with spinal cord injury: validity and reliability of 3 walking tests.

OBJECTIVE: To assess the validity and reliability of 3 timed walking tests (Timed Up & Go [TUG], 10-meter walk test [10MWT], 6-minute walk test) in subjects with spinal cord injury (SCI). DESIGN: Cross-sectional study and repeated assessments. SETTING: The SCI center of a university hospital in Switzerland. PARTICIPANTS: Validity was assessed by using the data of 75 patients with SCI, and reliability was determined with 22 patients with SCI. INTERVENTION: Patients performed the timed tests and the Walking Index for Spinal Cord Injury II (WISCI II) on the same day. Three measurements within 7 days were taken to assess reliability. MAIN OUTCOME MEASURES: The measures were scatterplots, correlation coefficients ( r ), and the Bland-Altman plot. Validity was determined in patients with different walking abilities. RESULTS: Overall, correlation of the 3 timed walking tests was excellent with each other (| r |>.88) and moderate with the WISCI II (| r |>.60). The correlation between the timed tests for patients with poor walking ability remained high (| r |>.70) but decreased in WISCI II (| r |<.35). High correlation coefficients ( r >.97) were found for intra- and interrater reliability. However, TUG and 10MWT reliability were negatively influenced by a poor walking function. CONCLUSIONS: The 3 timed tests are valid and reliable measures for assessing walking function in patients with SCI.

Adolescent↗

The reliability and validity of the physiological cost index in healthy subjects while walking on 2 different tracks.

OBJECTIVE: To investigate the reliability and validity of the Physiological Cost Index (PCI) scores, as a measure of energy expenditure, when healthy subjects walk on 2 different tracks (20-m and 12-m figure eight tracks). DESIGN: Intra- and interrater reliability and construct validity. SETTING: Physiotherapy division of a university in London, UK. PARTICIPANTS: Forty healthy subjects (15 men, 25 women; mean age +/- standard deviation, 34.5+/-12.6 y). INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Heart rate (in beats/min) and speed (in m/min) were used to calculate the PCI (in beats/m). Rate of oxygen consumption (VO2, in mL x kg(-1) x min(-1)) and oxygen cost (EO2, in mL x kg(-1) x m(-1)) were used as criterion estimates of energy cost EO2. Pearson correlation coefficients between the PCI, components of the PCI, EO2, and VO2 were used to quantify validity. Intrarater reliability was assessed in all participants and interrater reliability was assessed on a subset of 13 subjects using intraclass correlation coefficients and Bland-Altman plots. RESULTS: Intrarater (r=.73, r=.79) and interrater (r=.62, r=.66) reliability were acceptable between PCI scores from 20-m and 12-m tracks, respectively. Correlations between VO2 and EO2 with PCI were weak. PCI scores from the 20-m track were significantly lower than those on the 12-m track (P=.002). Subjects walked significantly faster on the 20-m track (P<.001). Results suggest a large difference in PCI scores would be necessary to indicate a "true" alteration in performance (52% for 20-m track, 43.4% for the 12-m track). CONCLUSIONS: The PCI is reliable but not valid as a measure of the energy cost of walking in healthy subjects, on either track. The 20-m track is recommended for clinical use because it enables subjects to walk at a faster pace.

Adult↗

Inter- and intrarater reliability of a back range of motion instrument.

OBJECTIVE: To examine the interrater and intrarater reliability of a back range of motion (BROM) instrument when measuring lumbar spine active planar motions and pelvic inclination. DESIGN: Single-group repeated measures for inter- and intrarater reliability. SETTING: Academic institution. PARTICIPANTS: Ninety-one participants (61 women, 30 men; mean age, 28 y) without a current complaint of low back pain volunteered. INTERVENTION: Two examiners measured pelvic inclination and all lumbar motions by using the BROM device. Subjects alternated between examiners for 4 complete trials; examiners remained blinded to the measurements. MAIN OUTCOME MEASURES: Intraclass correlation coefficients (ICCs) were used to determine intrarater and interrater reliability. Regression analysis was performed to determine the role palpation played in sagittal plane measurement error. RESULTS: Intrarater reliability for side bending was good (ICC range, .85-.83), lumbar forward flexion and pelvic inclination was good to fair (ICC range, .84-.79), and extension and rotation was fair to poor (ICC range, .76-.58). Interrater reliability was fair to poor for all lumbar motions and for pelvic inclination (ICC range, .79-.55). Less than 2% of the variation in sagittal plane measurements was explained by consistency of palpation for device placement. CONCLUSIONS: The BROM provides a reliable means of measuring lumbar forward flexion, side bending, and pelvic inclination when performed by the same examiner in asymptomatic subjects.

Adult↗

The Cumberland ankle instability tool: a report of validity and reliability testing.

OBJECTIVE: To test the Cumberland Ankle Instability Tool (CAIT), a 9-item 30-point scale, for measuring severity of functional ankle instability. DESIGN: Cross-sectional study. SETTING: General community. PARTICIPANTS: Volunteer sample of 236 subjects. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Concurrent validity by comparison with the Lower Extremity Functional Scale (LEFS) and a visual analog scale (VAS) of global perception of ankle instability by using the Spearman rho. Construct validity and internal reliability with Rasch analysis using goodness-of-fit statistics for items and subjects, separation of subjects, correlation of items to the total scale, and a Cronbach alpha equivalent. Discrimination score for functional ankle instability by maximizing the Youden index and tested for sensitivity and specificity. Test-retest reliability by intraclass correlation coefficient, model 2,1 (ICC(2,1)). RESULTS: There were significant correlations between the CAIT and LEFS (rho=.50, P<.01) and VAS (rho=.76, P<.01). Construct validity and internal reliability were acceptable (alpha=.83; point measure correlation for all items, >0.5; item reliability index, .99). The threshold CAIT score was 27.5 (Youden index, 68.1); sensitivity was 82.9% and specificity was 74.7%. Test-retest reliability was excellent (ICC(2,1)=.96). CONCLUSIONS: CAIT is a simple, valid, and reliable tool to measure severity of functional ankle instability.

Adult↗

Discovering reliable protein interactions from high-throughput experimental data using network topology.

OBJECTIVE: Current protein-protein interaction (PPI) detection via high-throughput experimental methods, such as yeast-two-hybrid has been reported to be highly erroneous, leading to potentially costly spurious discoveries. This work introduces a novel measure called IRAP, i.e. "interaction reliability by alternative path", for assessing the reliability of protein interactions based on the underlying topology of the PPI network. METHODS AND MATERIALS: A candidate PPI is considered to be reliable if it is involved in a closed loop in which the alternative path of interactions between the two interacting proteins is strong. We devise an algorithm called AlternativePathFinder to compute the IRAP value for each interaction in a complex PPI network. Validation of the IRAP as a measure for assessing the reliability of PPIs is performed with extensive experiments on yeast PPI data. All the data used in our experiments can be downloaded from our supplementary data web site at . RESULTS: Results show consistently that IRAP measure is an effective way for discovering reliable PPIs in large datasets of error-prone experimentally-derived PPIs. Results also indicate that IRAP is better than IG2, and markedly better than the more simplistic IG1 measure. CONCLUSION: Experimental results demonstrate that a global, system-wide approach-such as IRAP that considers the entire interaction network instead of merely local neighbors-is a much more promising approach for assessing the reliability of PPIs.

Algorithms↗

Reliability of laterality effects in a dichotic listening task with words and syllables.

Large and reliable laterality effects have been found using a dichotic target detection task in a recent experiment using word stimuli pronounced with an emotional component. The present study tested the hypothesis that the magnitude and reliability of the laterality effects would increase with the removal of the emotional component and variations in word frequency. Thirty-two participants completed both a dichotic syllable detection task and a dichotic word detection task. In both tasks, stimuli were pronounced in a neutral tone of voice. Each task was completed twice to allow the estimation of test-retest reliability. Results failed to confirm the hypothesis since they were generally similar to those obtained in the previous study. A significant right ear advantage (REA) was found only with the word task. Although no ear advantage was found for the syllable task, somewhat better reliability was demonstrated compared to that obtained with words. The present findings suggest that including an emotional component does not reduce the reliability or magnitude of auditory laterality effects. In fact, the emotional component may have forced participants to focus on the language aspects of the stimuli. This might partly account for the reduced reliability in words and the absence of REA in syllables. Motivational factors inherent to the within-subject design used here are also discussed.

Adult↗

Reliability of Greek version Gross Motor Function Classification System.

The Gross Motor Function Classification System for Cerebral Palsy (GMFCS), a reliable and valid system, has been widely utilized for objective classification of the patterns of motor disability in children with cerebral palsy. The objective of this study was to produce a Greek version of the instrument, with the same construct as the original one and to investigate the reliability of application of the Greek version GMFCS. Translation and back translation was made by two of the authors, one of whom did not know the original English text. The final translation was fixed by consensus. Two physicians were trained and given practice in the use of the GMFCS and its application to clinical documentation. The raters classified children with cerebral palsy according to GMFCS - Greek version. The reliability was assessed with the weighted kappa statistic. The sample consisted of 47 boys and 47 girls, mean age 5.4 years. The overall weighted Kappa was 0.80 (95% CI=0.67-0.94). Weighted Kappa for level I was 0.91 (95% CI=0.74-1.09), for level II, 0.78 (95% CI=0.62-0.95), for level III, 0.85 (95% CI=0.68-1.02), for level IV, 0.85 (95% CI=0.67-1.03) and for level V, 0.84 (95% CI=0.66-1.03). The inter-rater reliability was lowest at level II. Percent agreement was 75%. Results of this study suggest that GMFCS - Greek version can be used reliably to classify patients with CP from clinical documentation. These results further support use of the GMFCS in clinical settings and for research. Investigation is needed to further assess the reliability and to determine the validity of the Greek version of the GMFCS.

Age Factors↗

Reliability of the Eating Disorder Examination-Questionnaire in patients with binge eating disorder.

This study examined the test-retest reliability of the Eating Disorder Examination-Questionnaire (EDE-Q) in patients with binge eating disorder (BED). Short-term (mean days = 4.8; SD = 3.6) test-retest reliability of the EDE was examined in a sample of 86 patients with BED. Test-retest reliability was excellent for objective bulimic episodes (correlation = .84), but poor to unacceptable for subjective bulimic episodes and objective overeating episodes (correlations = .51 and .39, respectively). Test-retest reliabilities were good for the EDE-Q scales (correlations = .66 to .77), albeit somewhat variable for the individual EDE-Q items (.54 to .78). These findings support the reliability of the EDE-Q for patients with BED. The EDE-Q has utility for assessing the number of binge eating episodes (objective bulimic episodes) and associated features of eating disorders in patients with BED. The results for subjective bulimic episodes are consistent with previous studies in suggesting that these eating behaviors may not be reliable indicators of eating disorders for patients with BED.

Adult↗

Skin elasticity meter or subjective evaluation in scars: a reliability assessment.

UNLABELLED: Various methods are available for evaluating the elasticity of scars. However, the reliability and validity of these methods have been sparsely examined. The aim of this study was to examine the reliability of the subjective evaluation of scar pliability, while at the same time testing the reliability of the measurements of a non-invasive suction device (Cutometer Skin Elasticity Meter 575) on scars. Four observers assessed 49 scar areas of 20 patients with a subjective assessment of pliability. Subsequently, each observer measured the scar areas with the Cutometer. The intraclass correlation coefficients (ICC) of the elasticity (Ue) and extension (Uf) parameters of the Cutometer were acceptable (r = 0.76 and 0.74, respectively) when a single observer carried out the measurements. The subjective assessment of pliability needs to be completed by two or more observers to make the evaluation reliable (r = 0.79). The concurrent validities between the subjective pliability-assessment and each of the Cutometer parameters were statistically significant and ranged from r = 0.29-0.53. The correlations between each of the Cutometer parameters were high and statistically significant (r > or = 0.71). CONCLUSION: A single observer can reliably use the Cutometer for the elasticity measurements of scars. Furthermore, either Ue or Uf, instead of all five elasticity values provided by the Cutometer, can be adequately used for the elasticity measurements of scars. The subjective assessment of pliability of scars can only be assessed reliably when completed by two or more observers. The concurrent validity showed that all Cutometer parameters, except for visco-elasticity (Uv), and the subjective assessment of pliability measured the same characteristic of a scar.

Adolescent↗

Reliability, effect size, and responsiveness of health status measures in the design of randomized and cluster-randomized trials.

BACKGROUND: New health status survey instruments are often described by their psychometric (measurement) properties, such as Validity, Reliability, Effect Size, and Responsiveness. For cluster-randomized trials, another important statistic is the Intraclass Correlation (ICC) for the instrument within clusters. Studies using better instruments can be performed with smaller sample sizes, but better instruments may be more expensive in terms of dollars, opportunity cost, or poorer data quality due to the response burden of longer instruments. METHODS: We defined the psychometric statistics in terms of a mathematical model, and examined the power of a two-sample test as a function of the test-retest Reliability, Effect Size, Responsiveness, and Intraclass Correlation of the instrument. We examined the "cost-effectiveness" of using a one-item versus a five-item measure of mental health status. FINDINGS: Under the standard model for measurement error, the psychometric statistics are all functions of the same error term. They are also functions of the setting in which they were estimated. In randomized trials, power is a function of Reliability and sample size, and a less reliable instrument can achieve the desired power if N is increased. In cluster-randomized trials, adequate power may be obtained by increasing the number of clusters per treatment group (and often the number of persons per cluster), as well as by choosing a more reliable instrument. The one-item measure of mental health status may be more cost-effective than the five-item measure in some situations. CONCLUSION: If the goal is to diagnose or refer individual patients, an instrument with high Validity and Reliability is needed. In settings where the sample sizes are large or can be increased easily, any valid instrument may be cost-effective. It is likely that many published values of psychometric statistics are accurate only in settings similar to that in which they were estimated.

Cost-Benefit Analysis↗