Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

[The Verona Medical Interview Classification System/Patient (VR-MICS/P): the tool and its reliability].

OBJECTIVE: To assess the reliability and to describe the categories and the procedure to apply the VR-MICS/P (Verona-Medical Interview Classification System/Patient). SETTING: The interviews used for the reliability study were audiotaped. Five general practitioners (GPs) working in two general practices in South-Verona recorded their consultations. SAMPLE: 50 interviews selected randomly from 120, 10 for each GP. The selection criterion for the participating patients was a GHQ-12 score of 3 and the consultation for a new illness episode. MAIN OUTCOME MEASURES: The VR-MICS/P classifies patients' verbal behaviours into 21 categories, 15 of them are defined by form (cue or statement) and content. PROCEDURE: Two trained raters classified 50 interviews. Before applying the classification system each interview is divided into units which are numbered to define doctor's and patient's sequence of speech. RESULTS: The reliability of VR-MICS/P was satisfactory (Kappa 0.85). Similarity Index (Dice, 1945) for categories varied between 0.71 and 0.94. Reliability for form and content classification was satisfactory too (Similarity Index between 0.81 and 0.89 and between 0.84 and 0.94, respectively). CONCLUSIONS: The VR-MICS/P is a reliable measure for describing patients' verbal behaviours during medical interviews. It can be used together with the VR-MICS/D (Verona-Medical Interview Classification System/Doctor; Saltini et al., 1998) to describe the medical interview, the quality of doctor-patient interview and can be used as a measure of patient centredness.

Humans↗

Are psychophysical functions derived from line bisection reliable?

Psychophysical functions are used to characterize both normal perception and altered perception among patients with neglect, yet the reliability of these functions is rarely examined. The present study examined two-week, test-retest reliability for power functions derived from line bisection data among 58 normal, young and old, male and female subjects. Power function exponents and constants were, at best, moderately reliable over time. The size of the exponent tended to decrease at retesting. Reliability coefficients varied by age and gender; they were highly significant for young men, marginally significant for older men, and non-significant for women. Race influenced reliability as coefficients were significant for Caucasian subjects but not for African American subjects. Age and gender effects in this study parallel those in the literature on pseudoneglect, and they may reflect hemispheric differences in visuo-spatial processing, magnitude estimation, or both.

Adolescent↗

Reliability and validity of the Japanese version of the Support Team Assessment Schedule (STAS-J).

OBJECTIVE: The aim of this project was to develop an appropriate and valid instrument for assessment by medical professionals in Japanese palliative care settings. METHODS: We developed a Japanese version of the Support Team Assessment Schedule (STAS-J), using a back translation method, and tested its reliability and validity. In the reliability study, 16 nurses and a physician who work in a palliative care unit evaluated 10 hypothetical cases twice at 3-month intervals. For the validity study, external researchers interviewed 50 patients with matignancy and their families and compared the results with ratings by the nurses in the palliative care unit. RESULTS: Our results with hypothetical cases were: interrater reliability weighted kappa = 0.53-0.77 and intrarater reliability weighted kappa = 0.64-0.85. In the validity study comparing nurse evaluations and the results of interviews with patients and families, complete agreement was 36-70%, and close agreement (+/-1) was 74-100%. As a whole, weighted kappa were low: between -0.07 and 0.51. Our results were similar to those in the United Kingdom and Canada. SIGNIFICANCE OF RESULTS: Although this research was conducted under methodologically limited conditions, we concluded that the STAS-J is a reliable tool and its validity is acceptable. The STAS-J should become a valuable tool, not only for daily clinical use, but also for research.

Adult↗

Reliability of data on smoking habit and coffee drinking collected by personal interview in a hospital-based case-control study.

A study on the reliability of information on smoking habits and coffee drinking collected via interview was conducted among 500 subjects enrolled in a case-control study on bladder cancer in Brescia, North Italy. A total of 215 cases (incident and prevalent) and 285 controls were interviewed personally in the hospital setting by a first interviewer, and then re-interviewed by telephone by either the same interviewer or another one. Agreement between the first and second interview was evaluated using the kappa statistic and the intra-class correlation coefficient and via multiple logistic regression modelling. No important differences in reliability were found according to sex, education or case/control status, while agreement was better among subjects below 65 than among older ones, and among incident than prevalent cases. A slightly better agreement was found among subjects interviewed twice by the same interviewer than those interviewed by two different individuals, which may reflect the presence of inter-observer reliability for the latter. Overall, these results show a very high reliability of data on smoking and a fairly high reliability regarding coffee drinking as collected through face-to-face interviews.

Adult↗

Test-retest reliability of the Spanish version of the Diagnostic Interview Schedule for Children (DISC-IV).

The test-retest reliability of the Spanish Diagnostic Interview Schedule for Children (DISC-IV) is presented. This version was developed in Puerto Rico in consultation with an international bilingual committee, sponsored by NIMH. The sample (N = 146) consisted of children recruited from outpatient mental health clinics and a drug residential treatment facility. Two different pairs of nonclinicians administered the DISC twice to the parent and child respondents. Results indicated fair to moderate agreement for parent reports on most diagnoses. Relatively similar agreement levels were observed for last month and last year time frames. Surprisingly, the inclusion of impairment as a criterion for diagnosis did not substantially change the pattern of results for specific disorders. Parents were more reliable when reporting on diagnoses of younger (4-10) than older children. Children 11-17 years old were reliable informants on disruptive and substance abuse/dependence disorders, but unreliable for anxiety and depressive disorders. Hence, parents were more reliable when reporting about anxiety and depressive disorders whereas children were more reliable than their parents when reporting about disruptive and substance disorders.

Adolescent↗

Reliability of instruments in a cooperative, multisite study: employment intervention demonstration program.

Reliability of well-known instruments was examined in 202 people with severe mental illness participating in a multisite vocational study. We examined interrater reliability of the Positive and Negative Syndrome Scale (PANSS) and the internal consistency and test-retest reliability of the PANSS, the Rosenberg Self-Esteem Scale, the Medical Outcomes Study Short Form-36 (SF-36), and the Quality of Life Interview. Most scales had good levels of reliability, with intraclass correlation coefficients (ICCs) and coefficient alphas above .70. However, the SF-36 scales were generally less stable over time, particularly Social Functioning (ICC = .55). Test-retest reliability was lower among less educated respondents and among ethnic minorities. We recommend close monitoring of psychometric issues in future multisite studies.

Adult↗

Assessing the reliability of the EORTC QLQ-C30 in a sample of older African American and Caucasian adults.

The purpose of this study was to examine the structure and reliability of the EORTC QLQ-C30. This 30-item instrument has five functional scales (physical, role, cognitive, emotional, and social), three symptom scales (fatigue, pain, and nausea and vomiting) and a global health and quality of life scale. Confirmatory factor analysis and Cronbach's alpha estimates were used to assess the functioning of the EORTC QLQ-C30 in a sample of 489 African American (n = 255) and Caucasian (n = 234) adults aged 50 + years. Seven of the nine EORTC QLQ-C30 scales showed good reliability for both the African Americans and the Caucasians in the sample (Cronbach's alpha > 0.75). In contrast, the cognitive functioning scale had a reliability coefficient of only 0.69 for the African Americans and 0.40 for the Caucasians, and the nausea and vomiting scale had a reliability coefficient of only 0.49 for the African Americans and 0.51 for the Caucasians. In summary, although the overall reliabilities of seven of the scales showed good fit, many of the item-to-scale correlations did not. Researchers planning to use the EORTC QLQ-C30 might first consider conducting separate analyses on the different racial or ethnic subgroups in their study populations to determine whether a common set of factors or scales is available for further analysis.

Activities of Daily Living↗

Reliable and valid self-report outcome measures in sexual (dys)function: a systematic review.

This paper examines the published reliability and validity of non-disease specific, self-report measures of sexual function. Relevant papers were found in a search of the Embase electronic bibliographic database, for English language papers (published 1980-99) reporting on the psychometric testing of sexual function questionnaires. Existing published reviews or collections of such instruments were also searched, and the reference lists of all papers obtained were "back-searched" to identify other measures. Included measures were evaluated in a systematic manner using published standards concerning the validity, internal consistency, and reproducibility of health measurement scales and quality of life measures. Twenty-three self-report measures were identified for inclusion in this review. A further 2 measures were identified by reviewers of this paper after the main searches were undertaken. One measure was found not to be exclusively self-report. Eleven (46% of 24 included measures) did not meet minimum published standards for reliability, internal consistency, and validity. However one of these was reliable and valid in the female-version only. Of the 14 reliable and valid measures, or versions thereof (58% of 24), 2 (8% of 24) met "superior" psychometric standards. Many measures were developed for use with patients in sex or marital therapy, and are mainly suitable for administration to people with long-term sex partners. It is sensible to assume that instruments are only reliable and valid in the often specialized populations in which they were developed.

Female↗

Reliability and validity of the COOP/WONCA health status measure in patients with chronic obstructive pulmonary disease.

The objective of the study was to assess the reliability and validity of the Dartmouth Primary Care Cooperative Information Project/World Organization of National Colleges, Academies, and Academic Associations of General Practice/Family Physicians (COOP/WONCA) questionnaire in outpatients with chronic obstructive pulmonary disease (COPD). The test-retest reliability of individual items of the COOP/WONCA questionnaire was assessed using a weighted kappa-statistic, and construct validity was assessed by correlating items of the COOP/WONCA with the EQ-5D health status measure. Discriminant validity was assessed by comparing scores for known groups, at the same time comparing the results with those of a lung-specific health status questionnaire. The individual items of the COOP/WONCA had test-retest reliabilities of 0.67-0.78 (weighted kappa). Spearman's rank correlations between COOP/WONCA single-item scores and corresponding EQ-5D ranged 0.45-0.72, which were generally higher than associations between non-corresponding items. Four of the five COOP/WONCA items did not discriminate between patient groups divided according to forced expiratory flow in 1 sec (FEV1) in percent of predicted and 6-min walking distance, while four of five items of the lung-specific questionnaire discriminated well between these groups. The COOP/WONCA chart system was reliable and showed properties supporting the construct validity of the measure. The items, however, did not discriminate well between known groups, indicating that this questionnaire is not very sensitive in patients with COPD. The reliability of the COOP/WONCA items was acceptable for use at group level, but lower than current recommendations for use in individual patients.

Adult↗

Test-retest reliability of the Isernhagen Work Systems Functional Capacity Evaluation in patients with chronic low back pain.

The aim of this study was to investigate test-retest reliability of the Isernhagen Work System Functional Capacity Evaluation (IWS FCE) in a sample of patients (n = 30) suffering from Chronic Low Back Pain (CLBP) and selected for rehabilitation treatment. The IWS FCE consists of 28 tests that reflect work-related activities like lifting, carrying, bending, etc. In this study, a slightly modified IWS FCE was used. Patients were included in the study if they were still at work or were less than 1 year out of work because of CLBP. Participants' mean age was 40 years, the duration of low back pain ranged between 5 and 10 years. Fifteen patients (50%) were out of work for a mean of 17 weeks, and they all received financial compensation. Two FCE sessions were held with a 2-week interval in between. Means per session, 95% confidence intervals of the mean difference, one-way random Intra Class Correlations (ICC), limits of agreement, Cohen's kappa and percentage of absolute agreement were calculated where appropriate. An ICC of 0.75 or more, a kappa value of more than 0.60 and a percentage of absolute agreement of 80% were considered as an acceptable reliability. Tests of the IWC FCE were divided into tests with and tests without an acceptable test-retest reliability on the basis of the kappa values, the percentage of absolute agreement and the ICC values. Fifteen tests (79%) showed an acceptable test-retest reliability based on Kappa values and percentage of absolute agreement. Eleven tests (61%) showed an acceptable test-retest reliability based on ICC values.

Adult↗

The eczema area and severity index (EASI): assessment of reliability in atopic dermatitis. EASI Evaluator Group.

OBJECTIVE: To test the reliability of the eczema area and severity index (EASI) scoring system by assessing inter- and intra-observer consistency. DESIGN: Training of evaluators, application, and assessment over 2 consecutive days. SETTING: An academic center. PATIENTS: Twenty adults and children with atopic dermatitis (AD); cohort 1 (10 patients > or = 8 years) and cohort 2 (10 patients < 8 years). INTERVENTIONS: None. MAIN OUTCOME MEASURE: The EASI was used by 15 dermatologist evaluators to assess atopic dermatitis in cohort 1 and cohort 2 on 2 consecutive days. Inter- and intraobserver reliability were analyzed. RESULTS: Overall intra-evaluator reliability of the EASI was in the fair-to-good range. Inter-evaluator reliability analyses indicated that the evaluators assessed the patients consistently across both study days. CONCLUSIONS: This study demonstrated that the EASI can be learned quickly and utilized reliably in the assessment of severity and extent of AD. There was consistency among the evaluators between consecutive days of evaluation. These results support the use of the EASI in clinical trials of therapeutic agents for AD.

Adolescent↗

Coefficient alpha and related internal consistency reliability coefficients.

The author studied the conditions under which coefficient alpha and 10 related internal consistency reliability coefficients underestimate the reliability of a measure. Simulated data showed that alpha, though reasonably robust when computed on n components in moderately heterogeneous data, can under certain conditions seriously underestimate the reliability of a measure. Consequently, alpha, when used in corrections for attenuation, can result in nontrivial overestimation of the corrected correlation. Most of the coefficients studied, including lambda 2, did not improve the estimate to any great extent when the data were heterogeneous. The exceptions were stratified alpha and maximal reliability, which performed well when the components were grouped into two subsets, each measuring a different factor, and maximized lambda 4, which provided the most consistently accurate estimate of the reliability in all simulations studied.

Humans↗

The intra- and inter-instrument reliability of DXA based on ex vivo soft tissue measurements.

OBJECTIVE: Comparison of ex-vivo soft tissue measurements using the GE/Lunar pencil (DPX-L; GE/Lunar Co., Madison, WI) and fan beam (Prodigy dual-energy X-ray absorptiometers (DXA) GE/Lunar Co.). RESEARCH METHODS AND PROCEDURES: Intra-instrument reliability was assessed by repeatedly scanning soft tissue phantoms for lean tissue (water) and fat tissue (methanol) using one DPX-L and two identical Prodigy DXAs at fast, medium, and slow scan modes. For each machine, 10 scans of each phantom were performed at each scan speed. The number of scans per instrument totaled 60. Data were analyzed using ANOVA to ascertain whether scan speed affected the intra-instrument reliability and to test whether soft tissue measurements differed among instruments. Percentage fat (phantom density) was the outcome variable. RESULTS: Intra-instrument reliability, expressed as coefficient of variation, ranged between 0.7% and 5.2% for the DPX-L and 0.4% and 4.5% for the Prodigy, with the lowest coefficients of variation observed when scanning the fat tissue phantom. Scan speed also affected the intra-instrument reliability (p < 0.01). Furthermore, differences in the measurement of percentage body fat for both the lean and fat tissue phantoms were observed among all three absorptiometers (all p < 0.01). After adjusting for scan speed, differences persisted for all three instruments. DISCUSSION: Intra- and inter-instrument reliability of DXA machines, even those from the same manufacturer, remains unpredictable. Thus, when measuring body composition using DXA, it is important to consider that even in the absence of measurement bias, the use of different DXA machines, particularly when using a variety of speed settings, will increase the residual error around the true value.

Absorptiometry, Photon↗

Reliability and validity of the Wheelchair User's Shoulder Pain Index (WUSPI).

Many long term wheelchair users develop shoulder pain. The purpose of this study was to examine the reliability and validity of the Wheelchair User's Shoulder Pain Index (WUSPI), an instrument which measures shoulder pain associated with the functional activities of wheelchair users. This 15-item functional index was developed to access shoulder pain during transfers, self care, wheelchair mobility and general activities. To establish test-retest reliability, the index was administered twice in the same day to 16 long term wheelchair users and their scores for the two administrations were compared by intraclass correlation. To establish concurrent validity, the index was administered to 64 long term wheelchair users and index scores were compared to shoulder range of motion measurements. Results showed that intraclass correlation for test-retest reliability of the total index score was 0.99. There were statistically significant negative correlations of total index scores to range of motion measurements of shoulder abduction (r = -0.485), flexion (r = -0.479) and shoulder extension (r = -0.304), indicating that there is a significant relationship of total index score to loss of shoulder range of motion in this sample. The Wheelchair User's Shoulder Pain Index shows high levels of reliability and internal consistency, as well as concurrent validity with loss of shoulder range of motion. As a valid and reliable instrument, this tool may be useful to both clinicians and researchers in documenting baseline shoulder dysfunction and for periodic measurement in longitudinal studies of musculoskeletal complications in wheelchair users.

Activities of Daily Living↗

Validity and reliability of SCREEN II (Seniors in the community: risk evaluation for eating and nutrition, Version II).

BACKGROUND: Nutrition risk screening for community-living seniors is of great interest in the health arena. However, to be useful, nutrition risk indices need to be valid and reliable. The following three studies describe construct validation, test-retest and inter-rater reliability of SCREEN II. METHODS: Study (1) seniors were recruited from the general community and from a geriatrician's clinic to complete a nutritional assessment and SCREEN II. 193 older adults provided medical and nutritional history, 3 days of dietary recall and anthropometric measurements. A dietitian reviewed all information collected and ranked seniors on risk: 1 (low) to 10 (high risk). Receiver operating characteristic curves were completed. An abbreviated SCREEN II was developed through statistical analysis and expert ranking of the 17 items. Studies (2) and (3) seniors were recruited from the community to self-administer (n = 149) or be interviewed (n = 97) using SCREEN II twice within 2 weeks. For self-administration one index was completed via mail. Interviewer administration was completed via telephone with two interviewers. Intra-class correlations were calculated. RESULTS: (1) Total and abbreviated SCREEN II have increased sensitivity and specificity as compared to SCREEN I in identifying seniors at nutritional risk. (2) Test-retest reliability was adequate (intra-class correlation (ICC) = 0.83). (3) Inter-rater reliability was adequate (ICC = 0.83). CONCLUSIONS: SCREEN II appears to be a valid and reliable tool for the identification of risk for impaired nutritional states in community-living older adults, and is an improvement over SCREEN I.

Aged↗

Statistical strategies to assess reliability in ophthalmology.

Reliability of measurements and measurers is important so that we can trust the measurements we record. However, the statistical techniques used to assess reliability of measurements or measurers in the ophthalmic literature are often inappropriate, and not able to evaluate reliability between measurements/measurers. We review the techniques used in reliability studies for both continuous and categorical data, and describe appropriate statistical methods for particular study designs. We also highlight current techniques that are not appropriate in the analysis of reliability, but that are still commonly used in the ophthalmic literature. We hope that by highlighting these, we shall discourage their future use.

Data Interpretation, Statistical↗

Reliability of anthropometric measurements in overweight and lean subjects: consequences for correlations between anthropometric and other variables.

OBJECTIVE: To estimate the reliability of anthropometric measurements in overweight and lean subjects, and to examine the influence of this reliability on correlations to other variables, since low reliability leads to underestimation of correlations. DESIGN: Replicate measurements by two observers in 26 overweight and 25 lean subjects measured at two occasions. MEASUREMENTS: Sagittal abdominal diameter (SAD), waist circumference (waist), waist-to-hip ratio (W/H) and skinfold measurements. RESULTS: Intra-class correlation coefficients (ICCs) for SAD and waist were higher than for W/H (0.98 vs. 0.90, P<0.001, and 0.97 vs. 0.90, P = 0.001, respectively). For waist, the ICC was lower for overweight than for lean subjects (0.85 vs 0.95, P=0.030), but the ICC values were comparable for SAD and W/H (0.92 vs. 0.95 and 0.78 vs. 0.83, respectively). Intra-observer variations (IOV) for SAD and waist were lower than for W/H (coefficients of variation; 1.6%, 1.4% and 2.3%, respectively), as were intra-subject variations (ISV) (2.7%, 3.0% and 3.4%, respectively). ICC values ranged from 0.84 to 0.93 and were lower for overweight than for lean subjects for biceps, subscapular and umbilical skinfolds (P=0.031, P<0.001 and P=0.048, respectively). Coefficients of variations for skinfold measurements ranged between 7.3% and 16.0% for IOV and between 14.9% and 20.8% for ISV. CONCLUSIONS: The low ICC values imply that correlations can be underestimated in overweight groups. We propose that, because of their higher reliability, SAD and waist have a higher predictive capacity for cardiovascular risk than W/H. SAD is the only measurement with high reliability in both weight groups and its use is recommended.

Abdomen↗

The QA pressure measurement system: an accuracy and reliability study.

OBJECTIVE: The main purpose of this study was to determine the accuracy and reliability of the Queen Alexandra Pressure Measurement System (QA PMS). Furthermore, we examined whether there were significant differences in measured pressures of the buttock area during sitting between normal subjects and spinal cord injured (SCI) patients. DESIGN: Accuracy (calibration) and reliability (test-retest) study. SETTING: The spinal cord unit of Tertiary Care Centre 'De Hoogstraat' in Utrecht, The Netherlands. PATIENTS: A convenience sample of 16 SCI patients and 15 normal subjects. MAIN OUTCOME MEASURES: The accuracy was determined by using the Standard Error of the Mean (SEM, in mmHg). The Technical Error of Measurement (TEM, in mmHg) was calculated as measure for differences between two paired measurements. The reliability was determined by using an Intraclass Correlation Coefficient (ICC). Significant differences in measured pressures between both groups (P<0.05) were determined by using an unpaired (two sample) t-test. RESULTS: Accuracy (calibration): mean SEM=0.30 (+/-0.1) mmHg, indicating a high level of accuracy. Differences between two paired measurements: mean TEM calibration= 1.87 (+/-0.76) mmHg; mean TEM normal subjects=4.76 (+/-1.78) mmHg; mean TEM SCI patients=6.34 (+/-2.19) mmHg. Reliability: mean ICC(3,1) calibration=0.85 (95% CI=0.74 0.95); mean ICC(2.1) normal subjects=0.92 (95% CI=0.90 0.94); mean ICC(2.1) SCI patients=0.90 (95% CI=0.88 0.92). The normal subjects had significantly higher mean pressures (P=0.028) than the SCI patients (mean pressures 31.0 vs 28.5 mmHg), whilst the SCI patients had significantly higher peak-pressures (P=0.0000) than the normal subjects (mean peak-pressures: 134.1 vs 75.7 mmHg). CONCLUSIONS: The QA Pressure Measurement System has sufficient accuracy and good reliability as a measurement procedure. There are significant differences between the measured pressures of both groups: the significantly higher peak pressures of the SCI patients seem to be the most important.

Buttocks↗