Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

A comparison of reliability estimation using trials-to-criterion and sequential probability ratio testing.

The purpose of this study was to compare estimates of test reliability obtained from two sequential testing plans--trials-to-criterion (TTC) and sequential probability ratio (SPR) testing--when reliability is defined as the consistency of classification. Data from a golf chip test given to 110 beginning golf students (n = 80 males; n = 30 females) at the University of Wisconsin were used for analysis. Test specifications for the SPR test were alpha = beta = .05, theta 0 = .70, and theta 1 = .50. Two mastery levels for the TTC test were examined, .70 and .60, with success criteria ranging from R = 6 to R = 12. For each sequential testing plan, both P and kappa were calculated to estimate reliability. Results for the total group and for gender indicated that reliability was higher with the SPR test when the mastery level was .70, while reliability was similar under both plans at a mastery level of .60. Median test lengths for the group were 21 for the SPR test and an average of 12 across all R values for the TTC test. Misclassification error rates for the TTC test, however, were substantially higher than under the SPR test, particularly for false nonmaster errors. These data suggest that SPR testing would be the preferred approach when misclassification errors are of primary importance, such as to determine minimal competency for certification. However, TTC testing is a viable alternative for classroom tests because of ease of administration and shorter test length.

Educational Measurement↗

Test-Retest Reliability of the Computerized CIDI (CIDI-Auto): Substance Abuse Modules.

This paper discusses the reliability of the Alcohol and Substance Abuse modules of the CIDI-Auto in two countries, Australia and Puerto Rico, and two languages, English and Spanish. CIDI-Auto is a computer-assisted version of the CIDI. Reliability estimates for DSM/ICD are presented at the diagnostic and symptom level. In total, 286 subjects, ages 17-60 years, who had at least 12 drinks of alcohol in their lifetime participated in the study. Adequate to good test-retest reliability estimates were obtained, with no major differences by nosology, site, substance, or time. Harmful use/abuse showed lower kappas than dependence. Reliability estimates for dependence ranged from 0.70 to 0.95. For harmful use, kappa's ranged from 0.45 to 0.66. The findings are encouraging; CIDI-Auto produces reliable classification across two settings and in two languages with an instrument that has good coverage of different manifestations of the illness.

Journal Article↗

The Refractive Status and Vision Profile (RSVP): Translation into Persian, reliability and validity.

PURPOSE: To translate and test the reliability and validity of a Persian translation of the Refractive Status and Vision Profile (RSVP), a vision-related quality of life questionnaire, in Iran. METHODS: Forward & backward translation, committee review and pilot testing were performed to develop a final Iranian version of the RSVP. Seventy-three consecutive patients with refractive error before or after refractive surgery at the LASIK ward of Farabi Eye Hospital completed the questionnaire. A convenience sample of 14 patients completed the questionnaire twice within one week. Reliability was measured by internal consistency (Cronbach's alpha) and the intraclass correlation coefficient for test-retest reliability. Validity was evaluated by correlation between the different RSVP subscales, known groups comparison analysis, and correlation between the subscales versus global items and traditional clinical measures. RESULTS: Internal consistency was high (Cronbach's alpha : 0.71-0.92; except for the subscale expectations, alpha : 0.6). Test-retest reliability of subscales and the overall RSVP scale, as estimated by the intraclass correlation coefficient, was high except for optical problems and glare. Comparisons between pre- and post-operative groups of patients showed significantly higher (worse) scores for concern, physical/social functioning, and the overall score in the pre-operative group. Almost all subscales showed desirable inter-scale correlations. CONCLUSION: The Iranian version of the RSVP is a reliable and valid measure of vision-related quality of life in patients with refractive error.

Adult↗

Reliability and reproducibility of a Chinese-language visual function assessment.

OBJECTIVE: The purpose of the study was to validate a Chinese-language visual function assessment within the context of a routine cataract surgery practice and to assess the contribution of the method of questionnaire administration. DESIGN: The visual function assessment (VFA) was translated into Chinese. Two groups of study subjects were recruited: Chinese who did not speak English and Chinese conversant in English. Consecutive preoperative cataract patients of Chinese ancestry presenting to an urban ophthalmology practice were recruited. The questionnaire was administered in person or by telephone interview. Pre-operative visual acuity was recorded. Visual function scores were analyzed to assess reliability and correlation with visual acuity. RESULTS: Among the 186 potential study subjects, 155 patients completed the study The Chinese-language visual function assessment had good internal consistency (Cronbach alpha = 0.97, inter-item correlations = 0.43 to 0.96) . Reliability (with regard to the English version) and test-retest reproducibility of the Chinese questionnaire were strong with intraclass correlation coefficients greater than 0.60. The method of administration contributed to the measures of reliability and reproducibility. CONCLUSION: These results show that a Chinese-language version of the VFA questionnaire is reliable and valid. In industrialized countries with large Chinese-speaking populations and newly developed countries of East and Southeast Asia, the visual function assessment may be helpful in assisting routine clinical patient evaluation and cross-cultural outcome assessment programmes. Our findings also suggest that self-administered visual function assessments may be more reliable and valid than interview-generated assessments.

Aged↗

A test-retest reliability study of the Barthel Index, the Rivermead Mobility Index, the Nottingham Extended Activities of Daily Living Scale and the Frenchay Activities Index in stroke patients.

PURPOSE: To assess the test-retest reliability of a range of outcome measures in stroke patients. METHOD: Twenty-two patients > 1 year post-stroke were tested twice at an interval of 1 week using the Barthel Index (BI); the Rivermead Mobility Index (RMI); the Nottingham Extended Activities of Daily Living Scale (NEADL); and the Frenchay Activities Index (FAI). The mean difference (bias) and reliability coefficient (random error) were calculated for the total scores. Percentage agreement and the kappa coefficient were used to analyse individual items. RESULTS: The mean differences and reliability coefficients were BI 0.4 +/- 2.0, RMI 0.3 +/- 2.2, the NEADL 0.6 +/- 5.6, FAI -0.6 +/- 7.1. There was little bias between assessments. The performance of the BI and RMI were better with lower random error. The NEADL and FAI did not perform as well having larger random error components. Percentage agreements were generally high especially for the BI (>75%) and RMI (>85%), but there was considerable variation in the kappa coefficients. CONCLUSION: Measurement of basic activities of daily living and mobility as measured by the BI and RMI is reliable post-stroke. Measurements used to assess extended activities of daily living were less reliable in this study.

Activities of Daily Living↗

A new stroke activity scale-results of a reliability study.

PURPOSE: To investigate the internal consistency, inter-rater and intra-rater reliability of a disability stroke activity scale (SAS) for stroke patients. Its intended use is as a measure of motor function at the level of disability in stroke patients. METHOD: Twelve stroke in-patients were video-recorded performing the five activities from the SAS. Seven senior physiotherapists, experienced in stroke care, independently rated the recordings on two occasions, three weeks apart, using the SAS. Twelve hospital inpatients participated in the study. The subjects were aged between 48 and 86 and were between 6 and 87 days post stroke. RESULTS: Reliability for total scores was found to be excellent (generalizability correlation co-efficient (GCC) values> or =0.95) and reliability for individual item scores was good (kappa> or =0.7). Internal consistency reliability using Cronbach's alpha was also good (0.68 at time 1 and 0.68 at time 2). CONCLUSION: The stroke activity scale is a reliable instrument for hospital stroke patients. It can be administered in less than 10 minutes and requires minimal equipment and training. Further work on the validity and responsiveness of the SAS is in progress.

Activities of Daily Living↗

Development and reliability of the General Motor Function Assessment Scale (GMF)--a performance-based measure of function-related dependence, pain and insecurity.

PURPOSE: To develop a scale for assessment of three components-dependence, pain and insecurity - related to motor functions of importance for activities of daily living among older rehabilitation patients and to establish its clinical practicality and reliability. METHOD: A General Motor Function Assessment Scale (GMF) with the above aims was constructed. Clinical practicality was explored by questionnaires to 14 physiotherapists. Inter-rater and test-retest reliability was tested on patients in three different forms of geriatric rehabilitation (n=20-25) and analysed by percentage agreement (PA) and a non-parametric statistical method, which provide measures of the random disagreement separately from the systematic part of the disagreement. RESULTS: In the clinical test the GMF was found to be time efficient and clinically adequate. Analysis of reliability showed overall high values of PA (PA> or =70) and of the rank-order agreement coefficient (r(a)>0.82), and low degrees of systematic disagreement. CONCLUSIONS: GMF was found to be a clinically useful assessment scale in geriatric rehabilitation. The statistical analyses indicted a high degree of reliability. Comparison of these results with reliability of comparable rating scales is difficult on account of the statistical methods used in other studies, which commonly do not take into account the non-metric properties of the data.

Activities of Daily Living↗

Reliability and stability of the Roland Morris Disability Questionnaire: intra class correlation and limits of agreement.

PURPOSE: To analyse test-retest reliability and stability of the Dutch language version of the Roland Morris Disability Questionnaire (RMDQ) in a sample of patients (n = 30) suffering from Chronic Low Back Pain (CLBP). METHOD: Patients filled out the Dutch language version of the RMDQ questionnaire twice, before starting the rehabilitation programme, with a 2-week interval. Intra Class Correlations (ICC), (one way random) was used as a measure for reliability and the limits of agreement were calculated for quantifying the stability of the RMDQ. An ICC of 0.75 or more was considered as an acceptable reliability. No criteria for limits of agreement were available. However, smaller limits of agreement indicate more stability because it indicates that the natural variation is small. RESULTS: The Dutch RMDQ showed good reliability, with an ICC of 0.91. Calculating limits of agreement to quantify the stability, a large amount of natural variation ( +/- 5.4) was found relative to the total scoring range of 0 to 24. CONCLUSION: The Dutch RMDQ proves to be a reliable instrument to measure functional status in CLBP patients. However, the natural variation should be taken into account when using it clinically.

Adult↗

Measuring social participation: reliability of the LIFE-H in older adults with disabilities.

PURPOSE: Much more attention should be paid to instruments documenting social participation as this area is increasingly considered a pivotal outcome of a successful rehabilitation. The purpose of this study was to document the reliability of a participation measure, the Assessment of Life Habits (LIFE-H), in older adults with functional limitations. METHODS: Eighty-four individuals with physical disabilities living in three different environments were assessed twice with the LIFE-H, an instrument that documents the quality of social participation by assessing a person's performance in daily activities and social roles (life habits). RESULTS: The intraclass correlation coefficients (ICC) computed for intrarater reliability exceeded 0.75 for seven out of the 10 life habits categories. For interrater reliability, the total score and daily activities subscore are highly reliable (ICC </=0.89), and the social roles subscore is moderately reliable (ICC = 0.64). 'Personal care' is the category with the highest ICC, and for five other categories ICCs are moderate to high (< 0.60). CONCLUSION: LIFE-H is a valuable addition to instruments that mostly emphasize the concepts of function or functional independence. It is particularly meaningful to evaluate the participation of older adults in significant social role domains such as recreation and community life. It may be considered among the instruments having the best fit with the ICF definition of participation (the person's involvement in a life situation) and a majority of its related domains.

Activities of Daily Living↗

Reliability of the seated postural control measure for adult wheelchair users.

PURPOSE: To evaluate the test-retest and interrater reliability of the Seated Postural Control Measure for Adults 1.0 (SPCMA 1.0). METHOD: The participants were evaluated first by two raters and then, 3 weeks later, by one rater. Section 1 (one item, seven-point scale) evaluates the adult's overall ability to control its posture in a sitting position. Sections 2 and 3 (22 items each, scored on a seven-point scale), evaluate the adult's postural alignment in a static position and the changes in postural alignment induced by a dynamic activity. RESULTS: For the test-retest reliability, the intraclass correlation coefficient (ICC) of section 1 was excellent (0.95) and moderate to good for sections 2 and 3 (0.60 - 0.62) and their subsections (0.47 - 0.78). For interrater reliability, the three sections had good to excellent ICCs (0.68 - 0.93) and their subsections had moderate to good ICCs (0.41 - 0.69). A large range was observed in Kappa coefficients (test-retest and interrater reliability) for the item analysis of the sections 2 and 3, due to a lack of variability in some items. CONCLUSIONS: The results confirm that the SPCMA is reliable as a whole. Suitable information has been obtained for the development of the SPCMA 2.0 and, although further psychometric testing is needed, the latter should improve clinical evaluation of seated postural control in adult wheelchair users.

Activities of Daily Living↗

Reliability and validity of the Japanese-language version of the physical performance test (PPT) battery in chronic pain patients.

PURPOSE: To prepare a Japanese-language version of the Physical Performance Test (PPT) Battery and assess its reliability and validity. METHOD: Activity limitations by pain were evaluated by means of the Japanese-language version of the PPT Battery in 82 patients with chronic pain in the limbs and trunk. Two self-report questionnaires, one related to sensory evaluation of pain, and the other related to affective evaluation of pain, and the Functional Independence Measure (FIM), which evaluates activities of daily living, were simultaneously administered to the subjects. RESULTS: The results for reliability showed that the ICC values for inter-rater reliability and intra-rater reliability were 0.91 or more for every item. The results for validity showed significant associations between the scores for all of the items on the Japanese-language version of the PPT Battery and the total scores on the FIM (p < 0.01). Significant associations were found between 5 of the 8 items on the Japanese-language version of the PPT Battery and affective state due to the pain. CONCLUSIONS: The Japanese-language version of the PPT Battery was shown to possess adequate reliability and validity as a scale for evaluating the activity limitations of patients with chronic limb or trunk pain. The results also suggested that it might be possible to improve the activity limitations of patients with chronic pain by improving their affective state in response to the pain.

Activities of Daily Living↗

Reliability of bench tests of interface pressure.

Determination of an appropriate wheelchair cushion to optimize loading on buttock tissue is crucial to pressure ulcer prevention. Standardized test methods aim to simplify selection by helping clinicians and users identify a class or category of cushions that will meet the important medical need of adequate pressure distribution. The objective of this project was to determine the test-retest reliability of interface pressure measurements taken using bench tests as opposed to human subject tests. Ten wheelchair cushions were tested following the methods for interface pressure measurement as defined in a draft International Organization for Standardization document. Dispersion index, contact area, percent force in the ischial regions, peak pressure index, and seating pressure index-standard deviation are reliable measures. Average pressure is reliable but not very volatile between cushions. The data also indicate that peak pressure, seating pressure index-skew (SPI-sk), and the other five percent force regions are not reliable. Certain bench interface pressure variables were found to have adequate intralaboratory repeatability. Interlaboratory reliability must also be tested. If a bench interface pressure test is used to indicate cushion performance, its validity should also be studied. Research is underway to relate interface pressure variables to clinical measurements of wheelchair users. Once validity is shown, standardized test results can then be used by clinicians to simplify and improve the wheelchair cushion selection process.

Equipment Design↗

Test-retest reliability of a self-administered musculoskeletal symptoms and job factors questionnaire used in ergonomics research.

The purpose of this study was to investigate the test-retest reliability of questionnaire items related to musculoskeletal symptoms and the reliability of specific job factors. The type of questionnaire items described in the present study have been used by several investigators to assess symptoms of musculoskeletal disorders and problematic job factors among workers from a variety of occupations. Employees at a plastics molding facility were asked to complete an initial symptom and jobs factors questionnaire and then complete an identical questionnaire either two or four weeks later. Of the 216 employees participating in the initial round, 99 (45.8%) agreed to participate in the retest portion of the study. The kappa coefficient was used to determine repeatability for categorical outcomes. The majority of the kappa coefficients for the 58 questionnaire items were above 0.50 but ranged between 0.13 and 1.00. The section of the questionnaire having the highest kappa coefficients was the section related to hand symptoms. Interval lengths of two and four weeks between the initial test and retest were found to be equally sufficient in terms of reliability. The results indicated that the symptom and job factors questionnaire is reliable for use in epidemiologic studies. Like all measurement instruments, the reliability of musculoskeletal questionnaires must be established before drawing conclusions from studies that employ the instrument.

Adult↗

The test-retest reliability of a revised version of the Readiness for Interprofessional Learning Scale (RIPLS).

The original version of the Readiness for Interprofessional Learning Scale (RIPLS) was published by Parsell and Bligh in 1999. The only aspect of reliability considered by the authors was the internal consistency. A revised version for use with undergraduate students was published in 2005 (McFadyen et al., 2005). That paper also reported internal consistency of the revised version. Subsequently a sample from one professional group (n = 65) was used to assess test-retest reliability, over a one week period, of each of the 19 items and of the sub-scale totals, using Weighted Kappa and the intra-class correlation (ICC) respectively, and these results are reported in the present paper. The test-retest reliability of the individual items using Weighted Kappa was satisfactory, with the exception of two items (Items 11 and 12). The ICC results for the sub-scale totals were all in excess of 0.60 with the exception of sub-scale two. This revised version of RIPLS would appear to have good reliability in three of its sub-scales but further research, with larger samples, is required before the fourth sub-scale can be reliably assessed.

Clinical Competence↗

Interrater and test-retest reliability of a fixed condition design fluency test.

Despite its potential as a unique neuropsychological test, the emergence of a psychometrically sound research foundation for Jones-Gotman and Milner's (1977) Design Fluency Test (DFT) has been constrained by the lack of consistent administration and scoring practices and limited information about its reliability. Here we describe an approach to administering and scoring the fixed condition DFT that is modeled on Jones-Gotman and Milner's original method and that clarifies procedural ambiguities. Results include interrater and long-term test-retest reliability analyses using this approach. First, based on five raters who scored 50 DFT protocols, good to excellent intra-class correlation coefficients were obtained for all DFT scores. Second, in a broadly representative sample of 87 healthy adults who were tested twice over an average of 5 1/2 years, the test-retest reliabilities for total and novel design scores ranged from good to excellent. This study demonstrates that the fixed condition DFT can be scored reliably using these procedures and that the reliability coefficients for DFT total and novel designs scores are comparable to those of other commonly used neuropsychological tests.

Adult↗

Reliability of the prospective data collection protocol of the Swedish Spine Register: test-retest analysis of 119 patients.

BACKGROUND: The Swedish Lumbar Spine Register has been collecting patient-based data since 2000, and more than 80% of all spinal units in Sweden are now including their patients. In a few years, it will produce useful clinical information just as arthroplasty registers have, but to permit proper interpretation of data in the future, the reliability of the protocol must be tested. METHODS: Between January 2000 and March 2003, a sample of 122 patients was asked to fill in the questionnaire twice: 63 preoperatively and 59 postoperatively. Test-retest reliability was calculated with intra-class correlation coefficient (ICC) or weighted kappa when appropriate. RESULTS: Test-retest interval varied (range 0-235 days); in the "worst case scenario", the lowest ICC for SF-36 was 0.62 for the postoperative RE. Other values were above 0.70; for non-SF variables, ICC was in the range 0.79-0.89. Kappa values for the ordinal outcomes were high (0.74-0.91). INTERPRETATION: When separate reliability analysis was performed according to the time interval, a 0-2 days interval produced a significant memory effect; after 3 weeks, the reliability seemed to drop in the preoperative group, whereas results were reproducible up to 9 weeks postoperatively. The protocol studied can reliably detect postoperative improvements between large groups of patients such as in a register.

Adult↗

The reliability and validity of the Health of the Nation Outcome Scales: validation in relation to patient derived measures.

OBJECTIVE: The Health of Nation Outcome Scales (HoNOS) was developed to assess mental health outcomes. The aim of the studies is to examine the psychometric properties, reliability and validity of the HoNOS. METHOD: Three studies were conducted within St John of God Hospitals in New South Wales, Australia. They examined the reliability and the validity of the HoNOS. The first study examined the interrater reliability of the HoNOS, before and after staff training in the use of the HoNOS. The second study examined the validity of the HoNOS with the Symptom Checklist 90 Revised (SCL90-R) and the third study examined the validity of the HoNOS with the Short-Form 36 (SF-36). RESULTS: The first study showed an improvement in the interrater reliability (IRR) of the HoNOS due to training. However, a generally unsatisfactory IRR (range 0.50-0.65) was achieved. The second study found no correlation between the SCL90-R and the HoNOS on admission (r = 0.04) and discharge (r = 0.06). The third study found no significant correlation between the Mental Component Score of the SF-36 and the HoNOS on admission (r = -0.033) nor on discharge (r = -0.104). CONCLUSIONS: The HoNOS has at best moderate interrater reliabilities. Further, the validity of the HoNOS is under question, that is, it does not correlate with a major measure of mental health symptoms, nor with a major measure of health status. As such, it is concluded that the psychometric properties of the HoNOS do not warrant its use as a routine measure.

Catchment Area, Health↗

Reliability of the circadian rhythm of water and macronutrient-rich diets intake in dietary choice.

Rats with ad libitum water and the ability to self-select among three macronutrient-rich diets--carbohydrate (CHO), protein (PRO), and lipid (LIP)--show a circadian rhythmicity in their ingestion. The aim of the present study was to determine whether this circadian rhythmicity is reliable from day to day. Eight rats were offered ad libitum water and a choice of three isoenergetic diet rations providing carbohydrate, protein, and lipid. Water and food intake was recorded every 3 h for 7 days. The reliability of the circadian rhythm of water and food intake was assessed by the Intraclass Correlation Coefficient (ICC) and the test-retest reliability using the Pearson's Correlation Coefficient (r). The results showed that the circadian rhythm of water, CHO, and PRO intake are strongly reliable. However, the circadian rhythm of LIP intake is less reproducible. Among the three reliable parameters-water, CHO, and PRO, the circadian rhythm of water intake was the most reproducible over 7 days. This suggests that water intake may be used as a marker of circadian rhythmicity in ingestive behavior.

Animals↗