Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results↗

Reliability of self-reported sexual behavior in human immunodeficiency virus (HIV) concordant and discordant heterosexual couples in northern Thailand.

A partner study was conducted in northern Thailand between March 1992 and June 1996 which included data that allowed an assessment of the reliability of self-reports of sexual behavior and contraceptive use among heterosexual couples. The authors enrolled 529 couples among whom all male subjects were human immunodeficiency virus (HIV) seropositive voluntary blood donors and their female sexual partners were either HIV infected (n=246) or HIV seronegative (n=283). The levels of agreement within couples were assessed for recency of last sexual intercourse, sexual activity in the prior year, and contraceptive practices. For HIV discordant couples, a prospective study was conducted to examine risk factors for HIV transmission, the primary goal of the study. This allowed assessment of reliability of inter-partner reports over 6-12 months. Overall, agreement among couples was good for common sexual practices, especially vaginal intercourse and time since last intercourse, but was lower for condom use. Anal and oral sex were infrequently reported by these couples and there was greater disagreement for the occurrence of these practices. Partner agreement for contraceptive histories was good to excellent. Prospective data showed less frequent intercourse and more condom use but reliability remained good. Common sexual practices may be reliable for both HIV concordant and discordant couples in studies estimating prevalent infection. Estimates of incident heterosexually transmitted HIV may be made with greater reliability by studies which include assessment of reports of risk behavior by each member of a couple than studies of individuals.

Adult↗

Inter-rater reliability of a paediatric outcome measure in Nepal.

The Hospital and Rehabilitation Centre for Disabled Children (HRDC) in Nepal identified the need to evaluate their in-hospital and community programmes. An instrument was developed to provide information regarding the functional level of the children treated by the HRDC as well as to provide information regarding care-givers' attitudes towards disability. Inter-rater reliability of the measure was tested in three regions of Nepal with 49 children. Six HRDC field workers travelled in pairs to the childrens' homes and alternated in roles as test administrator and observer. Correlations between the scores documented by the administrator and observer were used to estimate inter-rater reliability. Inter-rater reliability coefficients, calculated using a weighted kappa statistic, varied from 0.60 to 1.0. We conclude that the instrument demonstrated an acceptable level of inter-rater reliability in the field setting. Future studies to measure construct and concurrent validity, test-retest reliability and responsiveness of the instrument are recommended as well as testing the instrument in different cultures.

Activities of Daily Living↗

Clinical criteria for the diagnosis of vascular dementia: a multicenter study of comparability and interrater reliability.

BACKGROUND: Several clinical criteria have been developed to standardize the diagnosis of vascular dementia (VaD). Significant differences in patient classification have been reported, depending on the criteria used. Few studies have examined interrater reliability. OBJECTIVE: To assess the concordance in classification and interrater reliability for the following 4 clinical definitions of VaD: the Hachinski Ischemic Score (HIS), the Alzheimer Disease Diagnostic and Treatment Centers (ADDTC), National Institute of Neurological Disorders and Stroke-Association Internationale pour la Recherche et l'Enseignement en Neurosciences (NINDS-AIREN), and Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM-IV). METHODS: Structured diagnostic checklists were developed for 4 criteria for VaD, 2 criteria for Alzheimer disease (AD), and 4 criteria for dementia. Twenty-five case vignettes, representing a spectrum of cognitive impairment and subtypes of dementia, were prepared in a standardized clinical format. Concordance in case classification using different criteria and interrater reliability among 7 ADDTCs given a specific set of criteria was assessed using the kappa statistic. RESULTS: The frequency of a diagnosis of VaD was highest using the modified HIS or DSM-IV criteria, intermediate using the original HIS and ADDTC criteria, and lowest using the NINDS-AIREN criteria. Scores for interrater reliability ranged from kappa = 0.30 (ADDTC) to kappa = 0.61 (original HIS). CONCLUSIONS: Clinical criteria for VaD are not interchangeable. Depending on the criteria selected, the reported prevalence of VaD will vary significantly. The traditional HIS has higher interrater reliability than the newer criteria for VaD. Prospective longitudinal studies with clinical-pathological correlation are needed to compare validity.

Aged↗

Accuracy and reliability of remote retinopathy of prematurity diagnosis.

OBJECTIVE: To determine the accuracy and reliability of retinopathy of prematurity (ROP) diagnosis using remote review of digital images by 3 masked ophthalmologist readers. METHODS: An atlas was compiled of 410 retinal photographs from 163 eyes of 64 low-birth-weight infants taken using a wide-angle digital fundus camera. All the images were independently reviewed by 3 readers, and the diagnosis in each eye was classified into 1 of 4 ordinal categories: no ROP, mild ROP, type 2 prethreshold ROP, or ROP requiring treatment. Findings were compared with a reference standard of dilated indirect ophthalmoscopy with scleral depression performed by an experienced pediatric ophthalmologist. RESULTS: Sensitivities/specificities of the diagnosis of any ROP were 0.845/0.910 for the first reader, 0.816/0.955 for the second reader, and 0.864/0.493 for the third reader. Sensitivities/specificities of the diagnosis of ROP requiring treatment were 0.850/0.960 for the first reader, 0.850/0.973 for the second reader, and 0.900/0.953 for the third reader. When ROP was classified into ordinal categories, the overall weighted kappa for interreader reliability was 0.743. Intrareader reliability for detection of low-risk prethreshold ROP or worse was 100% for all readers. CONCLUSION: The accuracy, interreader reliability, and intrareader reliability of remote diagnosis of clinically relevant ROP based on digital imaging are substantial.

Birth Weight↗

Interexaminer reliability in physical examination of pediatric patients with abdominal pain.

OBJECTIVE: To test the interexaminer reliability of abdominal examinations performed by pediatric emergency medicine physicians and surgeons in an emergency department. METHODS: A prospective cross-sectional study in which 3 different types of physicians (pediatric emergency department residents, pediatric emergency department attending physicians, and pediatric surgeons in training) independently examined a convenience sample of children (aged 3-19 years) with initial complaint of abdominal pain. The interexaminer reliability of 6 components of the abdominal examination (the presence or absence of abdominal distension, abdominal tenderness to percussion, abdominal tenderness to palpation, abdominal guarding, rebound tenderness, and bowel sounds) and the clinical diagnosis of peritonitis was tested. RESULTS: Sixty-eight patients were examined by pediatric emergency department residents and pediatric emergency department attending physicians. All 3 physician types examined 46 of these 68 patients. When comparing residents and attending physicians, the components of the abdominal examination showed less than moderate chance-adjusted agreement (kappa range, -0.04 to 0.38). When comparing attending physicians and surgeons, the presence of rebound tenderness showed moderate agreement (kappa = 0.54). The rest of the components demonstrated less than moderate chance-adjusted agreement (kappa range, -0.04 to 0.34). CONCLUSIONS: The components of the abdominal examination are poorly reliable between physician types. Only the "rebound tenderness" component of the abdominal examination shows moderate agreement between the pediatric emergency department attending physicians and the surgeon. No component of the abdominal examination appears to be consistently reliable. Interexaminer agreement must be considered when developing management strategies for acute abdomen. Interventions to improve reliability should be developed.

Abdominal Pain↗

Reliability studies of psychiatric diagnosis. Theory and practice.

The existing literature on the reliability of psychiatric diagnosis falls into two periods, the earlier reporting low reliability and the latter reporting much higher figures. The reasons for this trend are examined in the context of a discussion of the design of diagnostic reliability studies. The problems of research design and execution in studies of diagnostic reliability are reviewed, and statistical problems are examined. Solutions to many of these problems ae suggested, including recommendations of appropriate reliability coefficients and data analyses.

Humans↗

Six-month test-retest reliability of MRI-defined PET measures of regional cerebral glucose metabolic rate in selected subcortical structures.

Test-retest reliability of resting regional cerebral metabolic rate of glucose (rCMR) was examined in selected subcortical structures: the amygdala, hippocampus, thalamus, and anterior caudate nucleus. Findings from previous studies examining reliability of rCMR suggest that rCMR in small subcortical structures may be more variable than in larger cortical regions. We chose to study these subcortical regions because of their particular interest to our laboratory in its investigations of the neurocircuitry of emotion and depression. Twelve normal subjects (seven female, mean age = 32.42 years, range 21-48 years) underwent two FDG-PET scans separated by approximately 6 months (mean = 25 weeks, range 17-35 weeks). A region-of-interest approach with PET-MRI coregistration was used for analysis of rCMR reliability. Good test-retest reliability was found in the left amygdala, right and left hippocampus, right and left thalamus, and right and left anterior caudate nucleus. However, rCMR in the right amygdala did not show good test-retest reliability. The implications of these data and their import for studies that include a repeat-test design are considered.

Adult↗

Reliability and validity of an observer-rated disfigurement scale for head and neck cancer patients.

BACKGROUND: Facial disfigurement is considered to be one of the most distressing aspects of head and neck cancer and its treatment, but it has been the focus of little systematic study. Existing studies have yielded conflicting results about the psychosocial impact of disfigurement. No studies to date have examined disfigurement using a valid and reliable observer-rated measure. The purpose of the current study was to examine the validity (convergent and discriminant) and the inter-rater reliability of a novel nine-point observer-rated disfigurement scale. METHODS: The sample consisted of 74 ambulatory head and neck cancer patients more than 6 months post treatment. Ratings of disfigurement were assigned independently by surgical and nonsurgical raters. Validity was assessed by comparing the association between disfigurement ratings and sociodemographic and illness treatment variables. Reliability was assessed by examining the concordance between the surgical and nonsurgical ratings. RESULTS: Disfigurement ratings were not associated with several sociodemographic variables, supporting the discriminant validity of the scale. Disfigurement was significantly related to a diagnosis of oral cancer, a history of adjunctive radiation, the type of surgical procedure performed, the degree of physical dysfunction, and the presence of postoperative complications. Observer ratings of disfigurement were significantly related to patient ratings of disfigurement. These findings support the convergent validity of the disfigurement scale. Inter-rater reliability of the scale was high (intraclass correlation coefficient =.91). CONCLUSION: The study provides preliminary evidence for the validity and inter-rater reliability of a novel nine point observer-rated disfigurement scale that may be useful in evaluating the impact of disfigurement on quality of life in head and neck cancer.

Adaptation, Psychological↗

Internal consistency reliability of the Personality Assessment Inventory with psychiatric inpatients.

Internal consistency reliability (Cronbach's alpha) estimates and standard errors of measurement were determined for the Personality Assessment Inventory (PAI; Morey, 1991) with a group of psychiatric inpatients (N = 111). Full-scale reliabilities were large and acceptable, averaging .82. Subscale reliabilities were lower, averaging .66. Reliability estimates for full scales and subscales were comparable to those reported for the clinical portion of the PAI standardization group. Of the 20 full scales examined, only 2 (Somatization and Paranoia) had reliabilities (.85 and .80, respectively) that were significantly lower than comparable values reported for the standardization group. Despite statistically significant differences, these values were still considered large and acceptable.

Adolescent↗

Evaluating a national assessment strategy for urinary incontinence in nursing home residents: reliability of the minimum data set and validity of the resident assessment protocol.

Evaluation of 1 million incontinent American nursing home residents is hampered by both failure to detect incontinence and logistical barriers to diagnostic testing. The nationally mandated Minimum Data Set (MDS) and Resident Assessment Protocol (RAP) were devised to address these deficiencies. Although both instruments are also used in at least 18 other countries, neither has been evaluated. Our goal was to determine the reliability of the MDS and the accuracy of the RAP in predicting the lower urinary tract cause of incontinence. We determined interrater reliability for the 13 MDS items related to urinary incontinence in 123 randomly selected residents of 13 nursing homes in 5 states; forms were completed blindly by 2 nurses from each facility who were trained for a day. The RAP was assessed in 102 representative institutionalized women by blinded evaluation of its diagnostic accuracy compared with the multichannel videourodynamic criterion standard. For the MDS, interrater reliability for incontinence of all grades was excellent (weighted kappa correlation coefficient = 0.90), although reliability was greater at the extremes of measurement than for incontinence of intermediate severity. With the exception of delirium, correlations for the 11 MDS items related to incontinence were 0.65-0.96; for 6 items, correlations were > or = 0.8. The diagnostic accuracy of the RAP, successfully administered to 80% of women, was 70%. The accuracy of the nearly identical algorithm that formed the basis for the RAP was 84%. Importantly, serious misclassifications were not observed for either the RAP or the algorithm. Although its definitions should be modified slightly, the MDS appears to be feasible and reliable when administered by trained staff. In women, the diagnostic accuracy and safety of the RAP are good-particularly when administered as instructed-but the original, sex-specific algorithm is preferable. Together, the MDS and modified RAP provide a useful, stepwise, and non-urodynamically based strategy to guide evaluation and therapy of incontinence in this setting.

Aged↗

The interrater reliability of the Psychotic Inpatient Profile.

The interrater reliabilities of the 12 scales of the Psychotic Inpatient Profile (PIP) were assessed in three independent samples. These reliabilities were found to be consistently low for the three samples. Several possible sources of low reliability were investigated, and evidence to support the hypothesis that consistent rater biases might have contributed to this low reliability was found. The authors make several suggestions of ways to improve PIP's reliability.

Adjustment Disorders↗

WISC-R factor reliability at 11 age levels.

The reliability coefficients of the normative WISC-R factors (Verbal Comprehension, Perceptual Organization, Freedom from Distractibility) were computed for each age level of the standardization sample from a formula provided by Tellengen and Briggs (1967). Results indicated that the Verbal Comprehension and Perceptual Organization factors possess generally equivalent reliability to their analogous intelligence quotients. Freedom from Distractibility possesses adequate reliability for routine clinical interpretation. Criticisms of the reliability of the third factor (Distractibility) as being too low for individual assessment appear unwarranted. A table of reliability coefficients is presented for each age level in the standardization sample.

Adolescent↗

Reliability of the WAIS-R with psychiatric inpatients.

WAIS-R subtest and composite scale reliabilities, standard errors of measurement, and standard errors of estimate were determined for a sample of psychiatric inpatients (N = 100). For Digit Span and Digit Symbol, test-retest stability coefficients were obtained; split-half reliability coefficients were calculated for all other subtests. With the exception of Object Assembly (rxx = .38), all subtest and composite scale reliability coefficients were large and acceptable. Based on the standard error of measure, the most reliable WAIS-R subtests were Digit Symbol (.77), Information (1.04), and Picture Completion (1.07). Reliability coefficients for the psychiatric inpatient sample were, in general, comparable to those values reported for the standardization group (Wechsler, 1981). Significant differences were obtained only on the Object Assembly and Vocabulary subtests.

Adjustment Disorders↗

Test-retest reliability of the eating disorder examination.

OBJECTIVE: The purpose of this investigation was to determine the test-retest reliability of the Eating Disorder Examination (EDE). METHOD: This study examined the test-retest and interrater reliability of the EDE in 20 adult women with a range of eating disorder symptoms. Trained assessors administered the EDE to participants on two separate occasions, ranging from 2 to 7 days apart. RESULTS: Test-retest correlations were.7 or greater for all subscales and measures of eating disorder behaviors except for subjective bulimic episodes and subjective bulimic days. Interrater reliability was uniformly high with correlations above.9. DISCUSSION: Results provide further support for the reliability of the EDE, but suggest that smaller binge episodes may not be reliable indicators of eating pathology.

Adult↗

Inter-observer reliability of the Spondylitis Functional Index Instrument for assessing spondylarthropathies.

OBJECTIVE: To study the Spondylitis Functional Index (SFI) by having two physical therapists observe patients with spondylitis perform various tasks listed on the instrument. The physical therapists' observations were compared with each other and with the self-reported abilities of the patients. METHODS: Subjects (n = 30) were recruited from a cross-section of patients participating in a prospective randomized, multicenter, double-blind, parallel clinical trial of the efficacy of sulfasalazine on ankylosing spondylitis (n = 13), psoriatic arthritis (n = 13), and Reiter's syndrome (n = 4) conducted at the Veterans Affairs Medical Center in Salt Lake City. Percents of agreement and Cohen's kappa analysis were used to assess the reliability of the observations of the therapists and patients. RESULTS: The overall percent of agreement between the observers on the SFI was 93%. The overall percent of agreement between observer 1 and patients on the SFI was 66% and between observer 2 and patients was 67%. The overall inter-observer reliability measured by the Pearson coefficient was 0.91 and by Cohen's kappa was 0.86. Between observer 1 and the patients the Pearson was r = 0.53 and kappa = 0.39. For observer 2 the Pearson was r = 0.52 and kappa 0.39. CONCLUSIONS: We consider the agreement and reliability between observers to be high. The agreement and inter-observer reliability was poor between observers and patients. The SFI, as enhanced for use in this study to assess change in functional ability of patients with spondylitis, demonstrated high reliability when used by trained observers.

Activities of Daily Living↗

Reliability and validity of the Arthritis Hand Function Test in adults with systemic sclerosis (scleroderma).

OBJECTIVE: To determine the interrater and test-retest reliability and validity of the Arthritis Hand Function Test (AHFT) in persons with systemic sclerosis. METHODS: Interrater reliability of the AHFT was established by two raters independently scoring the performances of 20 women with systemic sclerosis. The same group of subjects was tested again 7-10 days later to determine test-retest reliability. Concurrent validity was established by the subjects' self-reports of their abilities to perform activities of daily living as measured by the Health Assessment Questionnaire and the Arthritis Impact Measurement Scales 2 (AIMS2). RESULTS: All of the items had excellent interrater intraclass correlation coefficients (ICC = 0.99-1.00). The ICCs for test-retest reliability were in the excellent (ICC = 0.80-0.97) range for most of the items and moderate (ICC = 0.57-0.73) for the others. Most of the items were moderately correlated with items on the AIMS2 (r = 0.45-0.69). CONCLUSION: The results from this study suggest that the AHFT is a reliable and valid test to measure hand function in persons with systemic sclerosis.

Activities of Daily Living↗

Relative reliability of circumferences and skinfolds as measures of body fat distribution.

The question has arisen whether patterns of body fat distribution can be identified by body circumferences, a method which is said to be more reliable and simpler than skinfold thickness or like measures of subcutaneous fat (Ashwell et al., Int. J. Obes. 6: 143-152, 1982). Here we address the question of whether body circumferences are inherently more reliable than skinfold thicknesses in 77 intra- and 224 interexaminer replicates from the Health Examination Survey of 12 to 17-year-olds in the U.S.A. Reliability of six body circumferences (0.96) was significantly (P less than .01) higher than that of skinfold thicknesses at five sites (0.91), suggesting that circumferences are a more reliable method. However, the reliability of skinfolds is still high, and skinfolds may be used in studies which focus on preadults or other groups in which the validity of circumferences as measures of body fat distribution is unknown.

Adipose Tissue↗