Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

The interrater reliability of physical signs in patients with eating disorders.

OBJECTIVE: To evaluate the interrater reliability of five common signs of eating disorders. METHODS: Eating disorder patients with anorexia nervosa, bulimia nervosa, and eating disorders not otherwise specified (ED-NOS), at various stages of recovery, were evaluated for the presence or absence of lanugo hair, acrocyanosis, parotid hypertrophy, hypercarotinemia, and Russell's sign. Patients were examined by two physicians with similar experience and training. Results are analyzed for reliability using the kappa statistic. RESULTS: Kappa scores were as follows, indicating marginal reproducibility of results: lanugo hair (kappa = 0.606), acrocyanosis (kappa = 0.014), parotid hypertrophy (kappa = 0.266), hypercarotinemia (kappa = 0.101) , and Russell's sign (kappa = 0.140). CONCLUSION: The interrater reliability for individual items ranged from poor to moderate. Overall, there is marginal interrater reliability for the five common signs of eating disorders assessed.

Adult↗

Reliability and validity of a chronic care facility adaptation of the Clinical Dementia Rating scale.

OBJECTIVE: This study investigated the reliability and validity of a chronic care facility adaptation of the Clinical Dementia Rating scale (CDR-CC). METHOD: Sixty-two residents in a chronic care facility participated in an inter-rater and 1 month test-retest reliability study. The instrument was validated against the Mini-Mental State Examination (MMSE). RESULTS: Inter-rater and 1 month test-retest reliability for the global CDR-CC score were excellent (intraclass correlation coefficients 0.99 and 0.92, respectively). The CDR-CC domain and global scores were negatively correlated with the MMSE. CONCLUSIONS: The CDR-CC is a global assessment tool that reliably and validly measures cognitive and functional impairment in a chronic care setting.

Activities of Daily Living↗

Ultrasound anthropometric reliability.

Method errors and reliabilities were estimated for seven sonographic measurements in pregnancies of 106 women examined between January and July 1989. Teams of two experienced sonographers replicated the following measurements: biparietal diameter (BPD), occipital-frontal diameter (OFD), anterior-posterior diameter (APD), transabdominal distance (TAD), and femur diaphysis length (FDL). Multilevel modeling procedures were used to estimate the variance components. Significant (p < 0.01) covariates in the fixed part of the model included an increase in error with greater parity, estimated menstrual age (EMA), and maternal abdominal wall thickness (taken at the umbilicus). Intraobserver reliability ranged from 85.2% (AC) to 99.3% (FDL); interobserver reliability ranged from 80.8% (TAD) to 92.4% (FDL). Method errors, describing the expected error for 68% of the measurements taken, ranged from 0.8 mm to 7.7 mm (intraobserver) and from 1.2 mm to 7.8 mm (interobserver). These results suggest that large error components should be considered in the interpretation of the reliability of ultrasonographically obtained measurements.

Adolescent↗

Reliability and validity of plain radiographs to assess angulation of small finger metacarpal neck fractures: human cadaveric study.

To quantify reliability and validity of plain radiographs for assessing the degree of small finger metacarpal neck fracture angulation, we created typical two-fragments fractures in 30 adult cadaveric specimens. Reliability and validity of different radiographic measurement methods were determined by the intraclass correlation coefficient (ICC) and the Bland and Altman graphical approach. Intraobserver and interobserver reliability was high with any radiographic measurement method. Mean ICCs values (95% confidence intervals) varied from 0.76 (0.56-0.88) to 1.00 (0.99-1.00). The graphical approach confirmed good agreement. Validity was substantial when the fracture angle was measured between the line along the longitudinal axis of the metacarpal shaft and the line from the center of the metacarpal head to the fracture site on lateral radiographs. Mean ICCs values varied from 0.70 (0.36-0.86) to 0.79 (0.5-0.90). The graphical analysis also indicated good agreement. In contrast, considerable lack of validity was observed when the angle was measured on oblique radiographs. Although the mean ICCs values varied from 0.68 (0.12-0.88) to 0.74 (0.05-0.90), suggesting substantial correlation, the graphical analysis provided evidence for poor validity. There was systematic bias with oblique radiographs consistently producing higher readings (up to 35 degrees ). In summary, reliability and validity are good only when the degree of small finger metacarpal neck fracture angulation is measured after drawing lines on lateral radiographs. Oblique radiograph measurements consistently produce higher readings.

Adult↗

Test-retest reliability of UPDRS-III, dyskinesia scales, and timed motor tests in patients with advanced Parkinson's disease: an argument against multiple baseline assessments.

The primary objective of this study was to assess the intra-rater reliability of the motor section of the Unified Parkinson's Disease Rating Scale (UPDRS-III) in patients with advanced Parkinson's disease (PD). The secondary objective was to assess the intra-rater reliability of standard timed motor tests and dyskinesia scales to determine the necessity of multiple baseline core evaluations before surgery for PD. We carried out two standardized preoperative core evaluations of patients with advanced PD scheduled to undergo deep brain stimulation. Patients were examined in the defined off and on conditions by the same rater. UPDRS-III, timed tests, and dyskinesia scores from the two evaluations were compared using Wilcoxon Signed Ranks tests and intraclass correlation coefficients (ICC). Differences in UPDRS-III scores for the two visits were clinically and statistically nonsignificant, and the ICC was 0.9. Similarly, there were no significant differences in timed motor tests or dyskinesia scores, with a median ICC of 0.8. The results indicate that previous findings of high test-retest reliability of UPDRS-III in early untreated PD patients can now be extended to those with advanced disease complicated by motor fluctuations. In addition, test-retest reliability of dyskinesia scales and timed motor tests was high. Taken together, these findings challenge the need for multiple baseline assessments as currently stipulated in core assessment protocols for surgical intervention in PD.

Dyskinesias↗

Reliability and validity of the Beck depression inventory in patients with Parkinson's disease.

We evaluated the validity, reliability, and potential responsiveness of the Beck Depression Inventory (BDI) in patients with Parkinson's disease (PD). In part 1 of the study, 92 patients with PD underwent a structured clinical interview for DSM major depression and based on this patients were considered depressed (PD-D) or nondepressed (PD-ND). Subsequently, patients filled in the BDI. In part 2, a postal survey consisting the BDI was performed in 185 PD patients and 112 controls. Test-retest reliability was assessed in 60 PD patients. The factor analysis revealed a cognitive-affective and a somatic factor. Cronbachs alpha for the BDI was 0.88. Mean BDI indicated significant differences (P<0.001) between the PD and control group, between the PD-ND and PD-D group, and between PD-ND and control group. In part 1, the receiver operating characteristic curves showed that the area under the curve for the total BDI was 0.88. A cutoff was calculated for the BDI (14/15) that had the highest sum of sensitivity (0.71) and specificity (0.90). In part 2, the test-retest reliability for the BDI total score was 0.89 (intraclass correlation coefficient). The smallest real difference was 3.3 for the total BDI. The BDI is a valid, reliable, and potential responsive instrument to assess the severity of depression in PD. However, an adjusted cutoff is recommended.

Aged↗

Inter-rater reliability of the Brief Psychiatric Rating Scale and the Groningen Social Disabilities Schedule in a European multi-site randomized controlled trial on the effectiveness of acute psychiatric day hospitals.

The objectives of this study were to report the inter-rater reliability of the Brief Psychiatric Rating Scale (BPRS 4.0) and the Groningen Social Disabilities Schedule (GSDS-II) as assessed in a randomized controlled trial on the effectiveness of psychiatric day hospitals spanning five sites in countries of Central and Western Europe. Following brief training sessions, videotaped BPRS-interviews and written GSDS-vignettes were rated by clinically experienced researchers from all participating sites. Inter-rater reliability often proved to be poor for items assessing the severity of both psychopathology and social dysfunction, but findings suggest that both instruments allow for the assessment of the presence or absence of specific psychopathological symptoms or social disabilities. Inter-rater reliability at subscale level proved to be good for both instruments. Results indicate that, with a brief training session and proper use of the instruments, psychopathology and social disabilities can be reliably assessed within cross-national research studies. The results are of particular interest given that the need to conduct cross-national multi-site studies including countries with different cultural backgrounds increases.

Adolescent↗

Interrater reliability of the Psychiatric Research Interview for Substance and Mental Disorders in an HIV-infected cohort: experience of the National NeuroAIDS Tissue Consortium.

The interrater reliability of the Psychiatric Research Interview for Substance and Mental Disorders (PRISM) was assessed in a multicentre study. Four sites of the National NeuroAIDS Tissue Consortium performed blinded reratings of audiotaped PRISM interviews of 63 HIV-infected patients. Diagnostic modules for substance-use disorders and major depression were evaluated. Seventy-six per cent of the patient sample displayed one or more substance-use disorder diagnoses and 54% had major depression. Kappa coefficients for lifetime histories of substance abuse or dependence (cocaine, opiates, alcohol, cannabis, sedative, stimulant, hallucinogen) and major depression ranged from 0.66 to 1.00. Overall the PRISM was reliable in assessing both past and current disorders except for current cannabis disorders when patients had concomitant cannabinoid prescriptions for medical therapy. The reliability of substance-induced depression was poor to fair although there was a low prevalence of this diagnosis in our group. We conclude that the PRISM yields reliable diagnoses in a multicentre study of substance-experienced, HIV-infected individuals.

Adult↗

Estimating test-retest reliability in functional MR imaging. I: Statistical methodology.

A common problem in the analysis of functional magnetic resonance imaging (fMRI) data is quantifying the statistical reliability of an estimated activation map. While visual comparison of the classified active regions across replications of an experiment can sometimes by informative, it is typically difficult to draw firm conclusions by inspection; noise and complex patterns in the estimated map make it easy to be misled. Here, several statistical models, of increasing complexity, are developed, under which "test-retest" reliability can be meaningfully defined and quantified. The method yields global measures of reliability that apply uniformly to a specified set of brain voxels. The estimates of these reliability measures and their associated uncertainties under these models can be used to compare statistical methods, to set thresholds for detecting activation, and to optimize the number of images that need to be acquired during an experiment.

Brain↗

Reliability of hand-held dynamometry in spinal muscular atrophy.

We have assessed the reliability of hand-held myometry in 33 patients with spinal muscular atrophy (SMA), testing elbow flexion, handgrip, three-point pinch, knee flexion, knee extension, and foot dorsiflexion, and determining intraclass correlation coefficients (ICC). Interrater reliability was high for upper limbs, with an ICC of 0.92 for three-point pinch and 0.98 for elbow flexion and grip. For lower limbs interrater reliability was good with ICC >0.85 for all measures except foot dorsiflexion. Test-retest results were excellent with ICC >0.91 in all instances. Hand-held myometry is easily performed in SMA patients of various ages and muscle strengths, is a reliable measure of limb muscle strength, and can be used in longitudinal studies and clinical trials.

Adolescent↗

Clinical evaluator reliability for quantitative and manual muscle testing measures of strength in children.

Measurements of muscle strength in clinical trials of Duchenne muscular dystrophy have relied heavily on manual muscle testing (MMT). The high level of intra- and interrater variability of MMT compromises clinical study results. We compared the reliability of 12 clinical evaluators in performing MMT and quantitative muscle testing (QMT) on 12 children with muscular dystrophy. QMT was reliable, with an interclass correlation coefficient (ICC) of >0.9 for biceps and grip strength, and >0.8 for quadriceps strength. Training of both subjects and evaluators was easily accomplished. MMT was not as reliable, and required repeated training of evaluators to bring all groups to an ICC >0.75 for shoulder abduction, elbow and hip flexion, knee extension, and ankle dorsiflexion. We conclude that QMT shows greater reliability and is easier to implement than MMT. Consequently, QMT will be a superior measure of strength for use in pediatric, neuromuscular, multicenter clinical trials.

Child↗

Test-retest reliability of four questionnaires for patients with overactive bladder: the overactive bladder questionnaire (OAB-q), patient perception of bladder condition (PPBC), urgency questionnaire (UQ), and the primary OAB symptom questionnaire (POSQ).

AIMS: This study examined test-retest reliability of four patient-reported outcome measures for patients with overactive bladder (OAB): Overactive Bladder Questionnaire (OAB-q), Patient Perception of Bladder Condition (PPBC), Urgency Questionnaire (UQ), and Primary OAB Symptom Questionnaire (POSQ). METHODS: Patients recruited from urology clinics were scheduled for two visits 2 weeks apart and completed all questionnaires at both visits. A demographic form was completed at Visit 1; and a treatment effect scale was completed at Visit 2. Test-retest reliability was examined among stable patients using intraclass correlations (ICC), Spearman's correlations, paired t-tests, Feldt's statistic, and kappas. RESULTS: A total of 47 patients enrolled (mean age = 66.0 years, 74.5% female), with 46 completing both visits; 35 were classified stable. Statistically significant correlations were present between Visits 1 and 2 (P < 0.05) for all subscales of the OAB-q, UQ, and POSQ. Subscale ICCs were moderate to high (OAB-q > or = 0.83, UQ > or = 0.46, POSQ continuous items > or = 0.68). No significant differences between Visit 1 and 2 were noted, except for the OAB-q symptom bother scale (change of 5.8 points on a 100-point scale). The multi-item subscales of the OAB-q and the UQ demonstrated good internal consistency (Cronbach's alpha > or = 0.83 for all subscales) across both visits. Test-retest reliability of the PPBC was somewhat weaker than the other three measures, but still acceptable for use as a global, single-item outcome measure. CONCLUSIONS: The OAB-q, POSQ, and UQ demonstrated good test-retest reliability, with ICCs roughly equivalent or superior to those previously reported for 7-day micturition diaries. Findings suggest that the four measures examined in this study demonstrate the necessary reproducibility for use as outcome measures for OAB treatments.

Aged↗

Issues and approaches to estimating interrater reliability in nursing research.

Following a general discussion of the meaning of and need for reliability estimates, four different approaches to the estimation of interrater reliability are discussed: correlational techniques, comparison of means, percentage of agreement, and generalizability theory techniques. Data from a nursing research project are used to illustrate the method and to interpret reliability estimates obtained by each approach. The estimates vary widely, from a low of .03 (using a percentage of agreement technique), to a high of .68 (using a correlational approach). The usefulness and advantages of generalizability theory techniques, which afford the most comprehensive and complex means of assessing reliability, are emphasized.

Analysis of Variance↗

The process of rater training for observational instruments: implications for interrater reliability.

Although the process of rater training is important for establishing interrater reliability of observational instruments, there is little information available in current literature to guide the researcher. In this article, principles and procedures that can be used when rater performance is a critical element of reliability assessment are described. Three phases of the process of rater training are presented: (a) training raters to use the instrument; (b) evaluating rater performance at the end of training; and (c) determining the extent to which rater training is maintained during a reliability study. An example is presented to illustrate how these phases were incorporated in a study to examine the reliability of a measure of patient intensity called the Patient Intensity for Nursing Index (PINI).

Employee Performance Appraisal↗

Reliability and validity of the Leisure Satisfaction Scale (LSS--short form) and the Adolescent Leisure Interest Profile (ALIP).

This study aimed to evaluate the reliability and validity of the Leisure Satisfaction Scale (LSS short form) and the Adolescent Leisure Interest Profile (ALIP). The LSS and the ALIP are instruments that occupational therapists can use to evaluate the leisure activities that clients enjoy. Evaluation of leisure interest and participation will assist in creating goals for therapy to maximize a client's ability to participate in leisure activities. This study examined the test retest reliability and concurrent validity of the LSS and the ALIP using a sample of 37 adolescents between the ages of 13 and 17 with no known impairments. The assessments were administered individually or in small groups 7 to 17 days apart. Cronbach's alpha was used to determine the internal consistency. Pearson product moment correlations were calculated to examine the test retest reliability of the 60 subscales and the six question totals of the ALIP, as well as for the 6 subscales and total score of the LSS. Concurrent validity was evaluated between the 'How often?' question of the ALIP and the LSS (short form). Based on the study results, the ALIP and the LSS seem to have good test retest reliability levels when used with adolescents with no known physical or mental impairments. The concurrent validity between the two instruments was not supported, with many of the scores indicating only weak or no association to each of the subscales, suggesting that the assessments differ in some fundamental way. However, the evidence of some relationships between subscales may indicate some areas where the ALIP and the LSS are similar.

Adolescent↗

Assessing functional mobility in survivors of lower-extremity sarcoma: reliability and validity of a new assessment tool.

BACKGROUND: Reliability and validity of a new tool, Functional Mobility Assessment (FMA), were examined in patients with lower-extremity sarcoma. FMA requires the patients to physically perform the functional mobility measures, unlike patient self-report or clinician administered measures. PROCEDURE: A sample of 114 subjects participated, 20 healthy volunteers and 94 patients with lower-extremity sarcoma after amputation, limb-sparing, or rotationplasty surgery. Reliability of the FMA was examined by three raters testing 20 healthy volunteers and 23 subjects with lower-extremity sarcoma. Concurrent validity was examined using data from 94 subjects with lower-extremity sarcoma who completed the FMA, Musculoskeletal Tumor Society (MSTS), Short-Form 36 (SF-36v2), and Toronto Extremity Salvage Scale (TESS) scores. Construct validity was measured by the ability of the FMA to discriminate between subjects with and without functional mobility deficits. RESULTS: FMA demonstrated excellent reliability (ICC [2,1] >or=0.97). Moderate correlations were found between FMA and SF-36v2 (r = 0.60, P < 0.01), FMA and MSTS (r = 0.68, P < 0.01), and FMA and TESS (r = 0.62, P < 0.01). The patients with lower-extremity sarcoma scored lower on the FMA as compared to healthy controls (P < 0.01). CONCLUSION: The FMA is a reliable and valid functional outcome measure for patients with lower-extremity sarcoma. This study supports the ability of the FMA to discriminate between patients with varying functional abilities and supports the need to include measures of objective functional mobility in examination of patients with lower-extremity sarcoma.

Amputation, Surgical↗

Measuring clinical status in cystic fibrosis: internal validity and reliability of a modified NIH score.

We examined measurement properties of the NIH Clinical Score for Cystic Fibrosis (CF) as an index of disease status. This score is being employed as a research tool for defining study populations and as an outcome measure, yet there are no published data on its reliability or how its items contribute to the overall measure of disease status. Criteria for scoring some items in the original index lack specificity. In this study, we used a modified score to have more clearly specified criteria, while retaining the original weightings and structure. For 200 patients with CF in two centers, we analyzed the total NIH Score and its subscores for internal consistency, interrater reliability, and factor analysis. Internal consistency indicates how inter-related the items are. The pulmonary subscore and overall score had fairly high internal consistency. However, the general subscore had low internal consistency, suggesting that the items are not measuring a single element of disease status and should not be added. Factor analysis provides additional information on the underlying structure and relationships among items. Five factors (groups of items) were identified accounting for 85% of the consistent variance of 14 items. These factors were designated by items accounting for most of their variance: general pulmonary, weight, disability, psychosocial, and acute infiltrate. While inter-rater reliability for the overall index was high, individual items showed less agreement. The results indicate that most of the variability in the NIH Score is attributable to pulmonary items in the first factor. The analyses suggest a new scoring structure for the NIH Score; the general subscore items do not contribute to the reliability or account for significant variance. Therefore, they will likely require further refinement or be eliminated.

Child↗

Examination of movement in patients with long-lasting musculoskeletal pain: reliability and validity.

BACKGROUND AND PURPOSE: An examination method based on psychosomatic physiotherapy, with 24 standardized tests related to general aspects of mobility, flexibility and the ability to relax, is used in some pain and rehabilitation clinics in Scandinavia in order to document where--and to what degree--patients have aberrations within the domain 'movement'. The measurement properties of the movement tests have, however, not been investigated in patients with long-lasting musculoskeletal pain. The aims of the present study were, therefore, to investigate inter-tester reliability and validity (construct, discriminative and concurrent validity) related to movement. METHOD: The study design was cross-sectional. Reliability was examined by three physiotherapists examining 19 people. Construct validity was studied by means of structural equation modelling (SEM). Discriminative validity was examined by comparing movement data from 247 patients with long-lasting musculoskeletal pain, and 104 healthy subjects. The patient sample was categorized according to localized or widespread pain, and movement scores compared between the groups. Most patients filled in a psychological screening questionnaire (MMPI-2), plus information about pain intensity and function, and concurrent validity of the movement measures were examined by correlation. RESULTS: SEM results indicated a modified movement scale, consisting of 16 items in four subscales. Both the original and modified versions showed good reliability, and scores differed significantly between healthy subjects and patients, and between patients with localized versus widespread pain. A relationship was found between the movement tests and psychological characteristics, but mainly in patients with widespread pain. Significant relationships were found between the ability to relax and pain, and between all aspects of movement and function. CONCLUSIONS: Movement may be reliably and validly assessed with composite scores from 4 x 4 items. The method may be useful as a global screening instrument in order to examine where, and to what extent, patients with long-lasting pain problems have movement aberrations; findings to be addressed in treatment.

Adult↗