Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

A standardized protocol for measurement of range of movement of the shoulder using the Plurimeter-V inclinometer and assessment of its intrarater and interrater reliability.

OBJECTIVE: To develop a standardized protocol for measurement of shoulder movements using a gravity inclinometer designed for use in clinical trials, and to assess its intra- and interrater reliability in a group of manipulative physiotherapists. METHODS: After instruction, 6 manipulative physiotherapists independently assessed 8 movements of the shoulder, including total and glenohumeral flexion (TF, GHF), total and glenohumeral abduction (TA, GHA), external rotation in neutral (ERN) and abduction (ERA), internal rotation in abduction (IRA), and hand behind back (HBB), in random order in 6 patients with shoulder pain and stiffness according to a 6 x 6 Latin square design using the standardized protocol. The assessments were then repeated. Analysis of variance was used to partition total variability into components of variance in order to calculate intraclass correlation coefficients (ICCs). RESULTS: The intra- and interrater reliability of the different movements varied widely. Reliability was higher for TF and TA than for the corresponding glenohumeral movements (e.g., intrarater ICCs: TF = 0.80, GHF = 0.65, TA = 0.75, GHA = 0.62). Interrater reliability was higher in the second round suggesting a practice effect (e.g., round 1, 2 interrater ICCs TF = 0.62, 0.82; TA = 0.62, 0.88; ERN = 0.85, 0.95). CONCLUSION: The measurement of the active range of TF, TA, ERN, and HBB, measured by manipulative physiotherapists following the standardized protocol, has intra- and interrater reliability acceptable for use as an outcome measure in clinical trials assessing interventions for shoulder pain.

Humans↗

Revised British Isles Lupus Assessment Group 2004 index: a reliable tool for assessment of systemic lupus erythematosus activity.

OBJECTIVE: To test the interrater reliability of the revised British Isles Lupus Assessment Group 2004 (BILAG-2004) index for the assessment of systemic lupus erythematosus (SLE) activity. METHODS: Patients with SLE were recruited from 11 centers. Two physician raters separately assessed the patients' disease activity using the BILAG-2004 index in routine clinical practice. Scores ranged from A (for very active disease) to E (for inactivity). Two reliability exercises were performed. Changes were made to the index after the first exercise (E1), and additional training was provided to the raters before the second exercise (E2). E1 and E2 involved 12 and 14 raters, respectively. Interrater reliability was assessed using kappa statistics and intraclass correlation coefficients. Levels of agreement and the extent of major disagreement were also examined. Major disagreement was defined as a score difference between raters of A versus C, D, or E or B versus D or E. RESULTS: For each exercise, 97 patients were recruited. In E1, the mean age of the patients was 42.3 years (range 18.5-82.2 years), 89.7% were women, and 74.2% were white, 8.2% were Afro-Caribbean, and 13.4% were South Asian, and in E2, the mean age was 43.7 years (range 17.7-75 years), 90.7% were women, and 68% were white, 15.5% were Afro-Caribbean, and 11.3% were South Asian. The mean disease duration was 9.4 years (range 0-32.1 years) for patients in E1 and 10 years (range 0-34.8 years) in E2. There was improvement in the interrater reliability and the level of agreement from E1 to E2. Further improvement was achieved after removal of poorly performing items. CONCLUSION: The BILAG-2004 index is a reliable tool to assess SLE activity. The use of a well-defined glossary and training of raters are essential to ensure the optimal performance of the index.

Activities of Daily Living↗

Reliability of the self-report version of the panic disorder severity scale.

The interview-administered Panic Disorder Severity Scale (PDSS) has been demonstrated to be a reliable and valid measurement of panic disorder severity. The purpose of this paper is to report on the reliability and usefulness of the self-report form (PDSS-SR). The PDSS and PDSS-SR were administered to 108 psychiatric outpatients with and without panic disorder. The internal consistency of the instruments was analyzed with Cronbach's alpha. Test-retest reliability was assessed on 25 of these subjects by using intraclass correlation coefficient (ICC). In addition, decrease in panic symptoms was measured after cognitive-behavioral treatment (CBT) treatment on 27 subjects. Cronbach's alpha was.917 for PDSS-SR and.923 for the PDSS. The PDSS-SR had good test-retest reliability (ICC=0.81) and was sensitive to change with treatment, with a mean decrease of 7.3 (S.D.=5.1). The self-report version of the PDSS is a reliable format, which could be useful in clinical and research settings.

Adult↗

Intraurethral sonography and the test-retest reliability of urethral sphincter measurements in women.

PURPOSE: The purpose of this prospective study was to determine the test-retest reliability of urethral sphincter morphologic measurements obtained with intraurethral sonography. METHODS: The cross-sectional urethral sphincter anatomy of 29 asymptomatic nulliparous women was studied in a blinded fashion. Each patient returned for a repeat examination on a different day. At the point of maximal rhabdosphincter thickness, the urethral diameter and circumference and the longitudinal smooth muscle and rhabdosphincter thickness, diameter, circumference, and area were measured using the ultrasound scanner's integrated software. For each measured variable, the reliability between patients was assessed with a paired t test. Intraclass correlation coefficients were calculated to assess the reliability of each intraurethral sonographic measurement obtained from the same patient. RESULTS: On test-retest analysis, the differences for each measured variable between patients were not statistically significant (p > 0.05). Of the measurements obtained from the same patient, however, longitudinal smooth muscle thickness (rho = 0.44; p = 0.006), diameter (rho = 0.49, p = 0.003), circumference (rho = 0.49, p = 0.003), and area (rho = 0.43; p = 0.009) were significantly correlated. CONCLUSIONS: The urethral longitudinal smooth muscle layer is the only structure that can be measured reliably using sonography for diagnostic use. Sonographic measurements of the rhabdosphincter may not be reliable because the outer portion of that structure lies outside the depth of penetration of a 12.5-MHz transducer.

Cross-Sectional Studies↗

Reliability of reported age at onset for Parkinson's disease.

An individual's age at onset of Parkinson disease (PD) can be collected through a variety of sources, including medical records, family report, and clinical observation. The most common source of PD age at onset information in the research setting is family-report, which is then typically used to classify a subject as juvenile, young, or late age at onset. The reliability of the family-reported age at onset of PD has not been rigorously examined. The present study used data from individuals diagnosed with PD to evaluate the reliability of age at onset information by comparing data obtained from three sources: 1) the subject's medical records, 2) a Family History Questionnaire, and 3) a Subject History Questionnaire. Among the 149 subjects with data for all three age at onset sources, the estimated reliability was R = 0.94. Similar reliability was observed when the sample was stratified based on gender, age at examination, disease duration, first symptom of PD, and years of education. The three measures of age at onset of PD show excellent agreement, strengthening confidence in the reliability of the reported age of clinical onset for PD.

Adult↗

Reliability and validity of a new global dyskinesia rating scale in the MPTP-lesioned non-human primate.

Behavioral rating scales for dyskinesia in the non-human primate are frequently used to assess the efficacy of new treatments and to provide a clinical correlative with neurochemical and neuropathological changes. Although a large variety of different scales have been used in non-human primate studies, there is no single standardized scale, and none have been evaluated for reliability and validity. We are reporting a new global non-human primate dyskinesia rating scale (GPDRS) for the squirrel monkey, developed in the context of an independent study of dyskinesia. In this report we demonstrate the reliability and validity of this scale. The GPDRS is a single-item scale with well-defined points and brevity allowing for rapid and easy application for assessing the overall degree of dyskinesia. In this study, seven MPTP-lesioned and four non-lesioned (control) non-human primates were videotaped following treatment with either levodopa or water. To test inter- and intra-rater reliability, three examiners rated the videotape independently at two different time points and these assessments were compared. The validity of the scale was tested in two phases. First, examiners rated the videotape using the GPDRS and the Abnormal Involuntary Movement Scale (AIMS), a scale commonly used to rate dyskinesia in the non-human primate, and the ratings from each scale were compared. Second, validity was tested in the context of an independent dyskinesia study, in which the scale was used to distinguish between two treatment groups. The GPDRS was shown to have high inter- and intra-rater reliability and to be valid for the assessment of dyskinesia in the squirrel monkey. In this report we also demonstrate the inter- and intra-rater reliability of the AIMS.

1-Methyl-4-phenyl-1,2,3,6-tetrahydropyridine↗

Reliability of spatiotemporal gait outcome measures in Huntington's disease.

Gait impairments are very important in Huntington's disease (HD), because loss of independence in gait is an important predictor of nursing home placement. Given this importance, it is imperative to test reliable and sensitive outcome measures that can be tested easily in various clinical environments. Here, we examined the test-retest reliability of gait outcome measures using the GAITRite instrumented carpet. We tested 12 subjects with HD and 12 age-matched controls in two separate sessions. At each session, subjects walked across the GAITRite carpet at a comfortable speed. We used the intraclass correlation coefficient (ICC) and coefficient of variation (CoV) to measure test-retest reliability. Reliability was very high for all outcome measures (velocity, cycle time, stride length, cadence, and base of support), as seen by high ICC scores (0.86 to 0.95) and low CoV scores (0.042-0.102). In addition, the performance across the two subject groups was very different, indicating that the GAITRite is sensitive enough to distinguish between populations. Given that the GAITRite is a relatively inexpensive and portable piece of equipment, it can be used in a wide variety of clinical settings and clinical trials. Our data on high test-retest reliability and sensitivity extends the utility of the GAITRite to the HD population.

Adult↗

Reliability of patient completion of the historical section of the Unified Parkinson's Disease Rating Scale.

Evaluating the reliability of self-assessed disability in patients with Parkinson's disease (PD) is important for therapeutic trials epidemiologic surveys. This is the first study to examine the interrater reliability between physician and patient in the historical section of the Unified Parkinson's Disease Rating Scale (UPDRS). Thirty consecutive subjects with idiopathic PD self-administered the historical section of a modified UPDRS. This instrument was then readministered to these 30 subjects by a neurologist. Interobserver reliability was assessed with the weighted kappa statistic (kw). The kw for each of the 17 items in the historical section of the UPDRS ranged from 0.63 to 1.0 (moderate to excellent agreement). Total kw = 0.83 (95% confidence interval = 0.79-0.87). There were no correlations between kw and age, disease duration, dose of levodopa, Hoehn and Yahr score, Schwab and England score, or modified Mini-Mental State score. Nondemented subjects with PD may reliably assess their level of disability by self-administering the historical section of the UPDRS. This has important implications for the reliable use of self-administered disability instruments for clinical research.

Activities of Daily Living↗

Inter- and intraobserver reliability of walking-track analysis used to assess sciatic nerve function in rats.

The inter- and intraobserver reliability of the walking-track analysis of sciatic nerve function in the rat was assessed. Twenty-five walking tracks were assessed on three different occasions by four observers. Whereas the interobserver reliability was found to be excellent (r = 0.92), the intraobserver reliability was only satisfactory (r = 0.53 to r = 0.76). Walking-track analysis provides a noninvasive technique to assess function recovery in the rat, with excellent intraobserver reliability demonstrated. The lower intraobserver reliability suggests limitations to this new measurement technique.

Animals↗

Test-retest reliability estimation of functional MRI data.

Functional magnetic resonance imaging (fMRI) data are commonly used to construct activation maps for the human brain. It is important to quantify the reliability of such maps. We have developed statistical models to provide precise estimates for reliability from several runs of the same paradigm over time. Specifically, our method extends the premise of maximum likelihood (ML) developed by Genovese et al. (Magn Reson Med 1997;38:497-507) by incorporating spatial context into the estimation process. Experiments indicate that our methodology provides more conservative estimates of true positives compared to those obtained by Genovese et al. The reliability estimates can be used to obtain voxel-specific reliability measures for activated as well as inactivated regions in future experiments. We derive statistical methodology to determine optimal thresholds for region- and context-specific activations. Empirical guidelines are also provided on the number of repeat scans to acquire in order to arrive at accurate reliability estimates. We report the results from experiments involving a motor paradigm performed on a single subject several times over a period of 2 months.

Brain↗

Estimating test-retest reliability in functional MR imaging. II: Application to motor and cognitive activation studies.

Functional magnetic resonance imaging (fMRI) using blood oxygenation contrast has rapidly spread into many application areas. In this paper, a new statistical model is used to evaluate the reliability of fMRI activation in a finger opposition motor paradigm for both within-session and between-session data and in a working memory paradigm for between-session data. A slice prescription procedure for between-session reproducibility is introduced. Estimates are made for the probabilities of correctly and falsely classifying voxels as active or inactive and receiver operator characteristic curves are generated. In the motor paradigm, estimated between-session reliability was found to be somewhat reduced relative to within-session reliability; however, this includes additional sources of variation and may not reflect intrinsically lower reliability. After matching false-positive classification probabilities, between-session reliability was found to be nearly identical for both motor and cognitive activation paradigms.

Brain↗

Reliable surrogate outcome measures in multicenter clinical trials of Duchenne muscular dystrophy.

We studied the reliability of a series of endpoints in an evaluation of subjects with Duchenne muscular dystrophy (DMD). The endpoints included quantitative muscle tests (QMTs), timed function tests, forced vital capacity (FVC), and manual muscle tests (MMT). Thirty-one ambulatory subjects with DMD (mean age 8.9 years; range 5-16 years) were evaluated at eight sites by 15 newly trained evaluators as a test of interrater reliability of outcome measures. Both total QMT score [intraclass correlation coefficient (ICC) 0.96] and individual QMT assessments (ICC 0.85-0.96) were highly reliable. Forced vital capacity and all timed function tests were also highly reliable (ICC 0.97-0.99). MMT was the least reliable assessment method (ICC 0.61). These data suggest that primary surrogate outcome measures in large multicenter clinical trials in DMD should use QMT, FVC, or time function tests to obtain maximum power and greatest sensitivity.

Adolescent↗

Reliability, validity, and gender differences in the quality of life index of the SEAPI-QMM incontinence classification system.

PURPOSE: To evaluate reliability and validity of the SEAPI-QMM 15-item quality of life index and assess differences between male and female patients with urinary incontinence. MATERIALS AND METHODS: Twice pre- and once post-treatment, 315 patients (102 men, 213 women) with incontinence and 35 without incontinence completed the self-directed SEAPI-QMM quality of life index. A voiding diary reported frequency of incontinence episodes with number of pads or type of protection used daily for incontinence. In 30%, the Nottingham Health Profile (NHP) was administered to further validate the measure. RESULTS: Cronbach's alpha coefficient for the index was 0.91. Domain-specific alpha coefficients ranged from 0.88 to 0.73. Test-retest reliability scores at 5 days gave a reliability coefficient of 0.93. Split half reliability was 0.89. Correlation of the index with the NHP was 0.78 for women, 0.72 for men. Mean scores before and after treatment with medical or surgical management were significantly different in both genders and were sensitive to the presence or absence of use of protection and the type of protection chosen in men. Men with incontinence (61%) reported a high level of impact in the sexuality domain compared to 7% of women. CONCLUSIONS: The SEAPI quality of life index has a high degree of reliability relating to stability and internal consistency across a wide age range in both genders. There are differences between men and women in life domains most frequently affected by urinary incontinence.

Adult↗

Use of health records in research: reliability and validity issues.

Data extracted from health records are commonly used in studies to address a variety of questions raised by health researchers. However, concern about the reliability and validity of such data generally is limited to an assessment of interrater reliability. Less attention has been paid to the reliability of the health record itself, and to the validity of both the health record and the data extracted from it. This article reviews the distinctions and overlaps among these types of reliability and validity and the factors that influence the validity and reliability of research data obtained from health records. Recommendations to investigators who use health record data in their research projects are offered.

Data Collection↗

Reliability and validity of the fine motor scale of the Peabody Developmental Motor Scales-2.

This study examined the test-retest reliability, inter-rater reliability, convergent validity and discriminant validity of the Fine Motor Scale of the Peabody Developmental Motor Scales-second edition (PDMS-FM-2). Participants included two groups of 18 children between the ages of 4 and 5 years with and without mild fine motor problems. The PDMS-FM-2 was administered twice to 12 children and rated by two occupational therapists. The PDMS-FM-2 results were compared with scores on the Movement Assessment Battery for Children (M-ABC). In addition, the scores of the children with and without fine motor problems were compared. For the test-retest reliability and the inter-rater reliability, correlation coefficients varied from r = 0.84 to r = 0.99. These results suggest that PDMS-FM-2 has excellent test-retest and inter-rater reliability. Convergent validity with the fine motor section of the M-ABC and discriminant validity have been confirmed. Only 39% of the children in the group with problems in fine motor activities had fine motor problems according to the PDMS-FM-2. This finding seems to indicate that the PDMS-FM-2 may not be sensitive enough for this population.

Adolescent↗

Inter- and intra-tester reliability of the Balance Performance Monitor in a non-patient population.

BACKGROUND AND PURPOSE: There is a growing interest in the measurement and evaluation of balance deficits and a number of instruments for measurement are now available. However, few data exist that accurately describe the reliability when using these measurement tools. This study was designed to evaluate the inter- and intra-tester reliability of using the Balance Performance Monitor (BPM) (SMS Healthcare) in a non-patient population. METHODS: A total of 58 subjects (mean age 29.83 years (+/- 9.44 years)) and three testers participated in two separate experiments. Intra Class Correlation Coefficients (ICCs) and coefficients of variation were used to describe the reliability of two different protocols for positioning subjects on the footplates of the BPM. RESULTS: Measurements of weight distribution showed high and significant inter- and intra-tester reliability for both protocols (ICCs ranging from 0.720 to 0.868). Sway measurements showed more limited reliability (ICCs ranging from 0.183 to 0.775). Coefficients of variation were low for weight distribution measurements and high for sway measurements. CONCLUSIONS: Taking the mean of three measurements is recommended for both the weight distribution and the sway measurements as it has shown to produce acceptable measurement results.

Adult↗

Intra- and inter-rater reliability of an 11-test package for assessing dysfunction due to back or neck pain.

BACKGROUND AND PURPOSE: The intra- and inter-rater reliability of 11 tests assembled by physiotherapists for clinical purposes was investigated. Forty-five patients and 23 healthy volunteers participated in the study. METHOD: Twenty-one patients were tested simultaneously and independently by two physiotherapists to determine inter-rater reliability for two raters. Twenty-four patients and 11 healthy volunteers were tested by one physiotherapist three times in a week to determine intra-rater reliability over time. Twelve healthy volunteers were tested by three different physiotherapists in a week to determine inter-rater reliability for three raters. RESULTS: Inter-rater agreement for two simultaneous raters was clinically acceptable. Repeatability on three test occasions was clinically acceptable in six of the 11 tests. There were no systematic differences between occasions. CONCLUSIONS: Intra- and inter-rater reliability was acceptable for six of the 11 tests in the form described here: three gait tests, two functional lifting tests and a functional muscular endurance test in the right leg. If these tests are to be used as outcome measures, account must be taken of the size of the typical fluctuation in measurements shown. The repeatability figures given may be used as guidelines for interpreting the clinical value of possible changes in test values.

Adult↗

Balance assessment in patients with peripheral arthritis: applicability and reliability of some clinical assessments.

BACKGROUND AND PURPOSE: Many individuals with peripheral arthritis blame decreased balance as a reason for limiting their physical activity. It is therefore important to assess and improve their balance. The purpose of the present study was to evaluate the applicability and the reliability of some clinical balance assessment methods for people with arthritis and various degrees of disability. METHOD: To examine the applicability and reliability of balance tests, 65, 19 and 22 patients, respectively, with peripheral arthritis participated in sub-studies investigating the applicability, inter-rater reliability and test-retest stability of the following methods: walking on a soft surface, walking backwards, walking in a figure-of-eight, the balance sub-scale of the Index of Muscle Function (IMF), the Timed Up and Go (TUG) test and the Berg balance scale. RESULTS: For patients with moderate disability walking in a figure-of-eight was found to be the most discriminative test, whereas ceiling effects were found for the Berg balance scale. Patients with severe disability were generally able to perform the TUG test and the Berg Balance Scale without ceiling effects. Inter-rater reliability was moderate to high and test-retest stability was satisfactory for all methods assessed. CONCLUSIONS: Applicable and reliable assessment methods of clinical balance were identified for individuals with moderate and severe disability, whereas more discriminative tests need to be developed for those with limited disability.

Adolescent↗