Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Poor interobserver reliability of AO classification of fractures of the distal radius. Additional computed tomography is of minor value.

Interobserver reliability of the AO system of classification of fractures of the distal radius was assessed using plain radiographs and CT. Five observers classified 30 Colles'-type fractures using only plain radiographs; two months later they were reclassified using CT in addition. Interobserver reliability was poor in both series when detailed classification was used. By reducing the categories to five, interobserver reliability was slightly improved, but was still poor. When only two AO types were used, the reliability was moderate using plain radiographs and good to excellent with the addition of CT. The use of CT as well as plain radiographs brings interobserver reliability to a good level in assessment of the presence or absence of articular involvement, but is otherwise of minor value in improving the interobserver reliability of the AO system of classification of fractures of the distal radius.

Colles' Fracture↗

Gender differences in the reliability of the EPQ? A bootstrapping approach.

Reliability indicates the degree of stability or homogeneity of a measurement, but also places an upper limit on the degree of association with other variables. Various methods are available to estimate the reliability of a measurement scale. However, an issue that has rarely been examined is that the reliability of a measurement, as estimated by coefficient alpha, may differ between groups. If a measurement has a different reliability for groups within a sample, spurious moderator effects may occur. The present study examines the reliability of the four subscales of a widely used psychological measurement instrument, the Eysenck Personality Questionnaire-Revised (EPQ-R), across gender. A bootstrapping methodology is employed which allows empirically derived standard errors to be calculated, and therefore tests of significance of difference to be computed. No significant differences were found in the reliability of the EPQ-R across sexes.

Adult↗

Long-term test-retest reliability of personality disorder diagnoses in opiate dependent patients.

This investigation reports the two year test-retest reliability of DSM-III-R personality disorder (PD) diagnoses in a sample of 219 patients with opiate dependence admitted to methadone treatment. Different MA/PhD interviewers at each assessment used a semistructured diagnostic interview for PD, the Structured Interview for DSM-III-R Personality Disorders (SIDP-R), to make their diagnoses. The reliability of any PD diagnosis versus no PD was fair (kappa = .51). The reliability for any specific PD (weighted kappa = .31) was poor. Antisocial (kappa = .45) and sadistic (kappa = .42), were the only specific PDs for which at least fair reliability was achieved. At the cluster level, only Cluster B had fair reliability (kappa = .47). The intraclass correlation coefficients between number of criteria for the specific PDs at the two evaluation points were consistently higher (range .22 to .62.) than were the corresponding kappas for categorical diagnoses. In that the base rates for most of the PDs were low and agreement for the specific PDs typically exceeded 90%. Increasing the base rate by lowering the diagnostic threshold, or examining more severe cases by raising the diagnostic threshold, did not consistently effect reliability. Reasons for the low kappa coefficients and the implications for PD research are discussed.

Adult↗

The reliability of alcohol abusers' self-reports of drinking and life events that occurred in the distant past.

This study investigated the test-retest reliability of 69 alcohol abusers' current reports about their past (approximately 8 years prior to interview) drinking behavior and life events. Drinking behavior was assessed by the Lifetime Drinking History (LDH) questionnaire and life events were assessed using the Recent Life Changes Questionnaire (RLCQ). Reliability coefficients for LDH variables were generally moderate to high (r = .52 to .81). Using empirical criteria, the diagnostic power of the two LDH interviews to classify correctly subjects as either having had or not having had a drinking problem was quite high. The reliability coefficient for the RLCQ was r = .85 and 91.7% of the identified events were reported in both interviews. Similarly high test-retest reliabilities and individual event agreement rates were obtained for the six homogeneous subscales of the RLCQ. Subjects were also asked why they had given inconsistent answers to life events questions in the two interviews. Inconsistencies often resulted from errors in the temporal placement of events or from misunderstanding items, rather than from failure to recall an event; this suggests that some sources of error in recalling life events can be reduced. It is concluded that alcohol abusers' reports of drinking and life events occurring many years prior to the date of interview are generally reliable. This finding is consistent with previous studies showing high test-retest reliabilities for reports of recent drinking and related events.

Adult↗

The Spanish Alcohol Use Disorder and Associated Disabilities Interview Schedule (AUDADIS): reliability and concordance with clinical diagnoses in a Hispanic population.

OBJECTIVE: The study reports the process of translation into Spanish and adaptation to the Hispanic culture of the Alcohol Use Disorder and Associated Disabilities Schedule (AUDADIS). This instrument is a structured diagnostic interview schedule specifically developed for the assessment of substance-related disorders and their comorbid disorders and disabilities. METHOD: A random sample (N = 169) of adults from a primary health care clinic in Puerto Rico was selected. The test-retest reliability of the instrument was examined across time and across interviewers, and the validity was assessed by comparing computer-derived diagnoses obtained through the administration of lay interviewers with best estimate diagnoses given by board-certified psychiatrists. RESULTS: For most diagnoses and symptoms studied, as well as for most of the alcohol consumption measures, the test-retest reliability of the Spanish AUDADIS was consistent with results reported in other national and international studies using this instrument. Good to excellent test-retest reliability was obtained for the diagnoses of alcohol dependence and major depression. Similarly, good to excellent agreement was obtained between the lay administered AUDADIS and best estimate diagnoses for most diagnostic categories, with the exception of dysthymia. As in other studies, the reliability and validity of the substance abuse category was poor. When agreement for this category was estimated independent of lifetime dependence, both the reliability and validity coefficients were considerably improved. CONCLUSIONS: The Spanish AUDADIS generally demonstrates good to excellent levels of reliability and validity that are comparable to findings reported for this instrument in other national and international studies.

Adult↗

Reliability and validity of the Children's Health Survey for Asthma.

OBJECTIVE: Describe the psychometric properties of the Children's Health Survey for Asthma (CHSA)- a condition-specific, self-report, functional health measure for parents of children 5 to 12 years of age with chronic asthma. METHOD: Data from two cross-sectional and one longitudinal study were used to assess internal consistency reliability, test-retest reliability, and validity of the CHSA. Over 275 parents and guardians of children with asthma completed the CHSA in one of three studies. The combined samples included a heterogenous mix of respondents by child age and race/ethnicity and parental marital and socioeconomic status. Five domain scores were computed: physical health, activity (child), activity (family), emotional health (child), and emotional health (family). Raw scale scores were transformed from 0 to 100 with higher scores indicating better or more positive outcomes. RESULTS: Across the three samples, mean scale scores ranged from a low of 61.5 (emotional health of the child) to a high of 86.1 (activity [family]). Internal consistency reliability for each of the scales was high (Cronbach's alpha =.81-. 92), and test-retest reliability (correlation between forms) ranged from.62 to.86. Significant differences in mean scores for four of five scales were noted between those with low versus moderate to high recent symptom activity. CONCLUSION: In three tests, the CHSA displays strong reliability and validity. Descriptive statistics demonstrate a range of scale scores. Internal consistency is good to excellent and short-term test-retest reliability is good for each of the five scales. Construct validity is demonstrated by the ability of CHSA to distinguish levels of disease severity, defined by symptom activity.

Asthma↗

Reliability of the information about the history of diagnosis and treatment of hypertension. Differences in regard to sex, age, and educational level. The Pró-Saúde study.

OBJECTIVE: To assess the intraobserver reliability of the information about the history of diagnosis and treatment of hypertension. METHODS: A multidimensional health questionnaire, which was filled out by the interviewees, was applied twice with an interval of 2 weeks, in July '99, to 192 employees of the University of the State of Rio de Janeiro (UERJ), stratified by sex, age, and educational level. The intraobserver reliability of the answers provided was estimated by the kappa statistic and by the coefficient of intraclass correlation (CICC). RESULTS: The general kappa (k) statistic was 0.75 (95% CI=0.73-0.77). Reliability was higher among females (k=0.88, 95% CI=0.85-0.91) than among males (k=0.62, 95% CI=0.59-0.65). The reliability was higher among individuals 40 years of age or older (k=0.79; 95% CI=0.73-0.84) than those from 18 to 39 years (k=0.52; 95% CI=0.45-0.57). Finally, the kappa statistic was higher among individuals with a university educational level (k=0.86; 95% CI=0.81-0.91) than among those with high school educational level (k=0.61; 95% CI=0.53-0.70) or those with middle school educational level (k=0.68; 95% CI=0.64-0.72). The coefficient of intraclass correlation estimated by the intraobserver agreement in regard to age at the time of the diagnosis of hypertension was 0.74. A perfect agreement between the 2 answers (k=1.00) was observed for 22 interviewees who reported prior prescription of antihypertensive medication. CONCLUSION: In the population studied, estimates of the reliability of the history of medical diagnosis of hypertension and its treatment ranged from substantial to almost perfect reliability.

Adolescent↗

[Cross-cultural adaptation and reliability analyses of the Southampton Assessment of Mobility to assess mobility of Brazilian elderly with dementia].

The objective was to perform a cross-cultural adaptation of the Southampton Assessment of Mobility and test its intra- and inter-examiner reliability for Brazilian elderly living in the community and diagnosed with dementia, with severity classified according to the Clinical Dementia Rating. The instrument was applied to 107 elderly (76.26 years +/- 7.59; 27.1% males, 72.9% females) diagnosed with dementia by the geriatric clinic at the university hospital of the Federal University in Minas Gerais. From the initial group, a randomized sample of 39 elderly (76.85 years +/- 7.75; 23.1% males, 76.9% females) was selected for the reliability tests. The statistical tool was the kappa test. The respective reliability indices were: mild dementia - 0.89-0.86; moderate - 0.79-0.85; and severe - 0.53-0.49. After cross-cultural adaptation and reliability tests, the instrument proved adequate for the target population, with "near-perfect" reliability for mild and moderate dementia. For severe dementia, the reliability was moderate.

Aged↗

Three-dimensional analysis of maxillary dental casts using Fourier transform profilometry: precision and reliability of the measurement.

OBJECTIVE: Fourier transform profilometry was used for the three-dimensional measurement of maxillary dental casts to analyze the size and shape of the palate. The objective of this study was to test the accuracy of the measuring system and determine the precision and reliability of the measurement METHODS: Images of dental casts were analyzed using newly developed measuring software. Based on five landmarks located on the alveolar ridge, the measuring software constructed 10 transversal sections of the palate. In each section profile, the width, area, and 23 height variables were assessed. SUBJECTS: Maxillary dental stone casts of 25 healthy girls, 14.1 to 15.3 years of age, were studied. RESULTS: The technical error of measurement exceeded 5% of the size of the measurement only in variables with means less than approximately 3 mm. In fact, such small absolute dimensions were exhibited only by the palate height in anterior profile 2 and the palate height at the margins of other profiles. Reliability of the measurements was found to be very high for the width and area of the profiles. For height measurements, the coefficient of reliability was slightly lower at the profile margins than near the midline. Nevertheless, only three height variables showed a coefficient of reliability lower than 0.90. The coefficients of reliability of other height measurements of profiles 3 through 10 were only sporadically lower than 0.97. CONCLUSION: With regard to the accuracy of the measuring system as well as the precision and reliability of the measurement, this method proved to be a suitable tool for studying palatal morphology.

Adolescent↗

The intrajudge reliability of the perceptual rating of cleft palate speech before and after pharyngeal flap surgery: the effect of judges and speech samples.

OBJECTIVE: In this pilot study, the reliabilities of the perceptual ratings of four types of speech samples by six judges, with and without expertise in evaluating cleft palate speech, were studied. DESIGN: Pre- and postoperative tape recordings of 15 patients with cleft lip and palate who had undergone a superiorly based pharyngeal flap operation were selected. Five speech-language pathologists and one oral and maxillofacial surgeon perceptually rated the following variables on separate 100-mm visual analog scales: hypernasality, audible nasal emission, intelligibility, misarticulations associated with velopharyngeal insufficiency, voice quality, and the presence or absence of hyponasality. These six variables were rated in four types of speech samples: reading of three sentences, repeating after the speech pathologist of three sentences, 10 sentences containing the aforementioned material, and the same 10 sentences in paired comparison. All speech samples were rerated after 3 months by the same judges. RESULTS: Judges differed largely in the range they used in their rating. Intrajudge reliability of .56 to .78 was found for ratings of hypernasality. No significant differences in intrajudge reliability were found for the ratings with the different types of speech samples. The intrajudge reliability of a judge with expertise was not necessarily higher than of a judge without this expertise. CONCLUSIONS: The improvement in speech is most reliably assessed with speech samples in paired comparison. A speech-language pathologist with expertise in evaluating cleft palate speech does not guarantee a high intrajudge reliability of the rating.

Adolescent↗

Meta-analysis of the reliability and validity of Part B of the Index of Work Satisfaction across studies.

Nurses' job satisfaction is a crucial factor in health care organizations. This study uses meta-analysis for reliability generalization and synthesis of construct validity of Part B of the Index of Work Satisfaction (IWS), a measure of job satisfaction. Meta-analysis was performed including assessments of study quality and descriptive coding of studies. Rater reliability was assessed for all coding and extraction of data. The mean reliability of Part B scores of the IWS based on 14 studies was .78 (df = 13, p < .05). The mean score reliability was .77 for university settings, .73 for community/acute care hospitals, .77 for multi-site studies, and .90 for other settings. For studies rated high and low quality, the mean score reliability was .77 and .83, respectively. Scores on Part B of the IWS correlated -.38 with turnover intent, .60 with organizational commitment, and -.53 with job stress. Scores on Part B of the IWS are reliable for measuring job satisfaction of nurses across samples. Construct validity needs additional testing.

Attitude of Health Personnel↗

Observer reliability as a function of circumstances of assessment.

THREE FACTORS CHARACTERISTIC OF EXPERIMENTAL SETTINGS WERE HYPOTHESIZED TO INFLATE ARTIFACTUALLY THE RELIABILITY OF OBSERVATIONAL RECORDINGS: (a) knowledge by observers of when and by whom their reliability is being assessed, (b) the absence of the experimenter or a monitor to prevent cheating, and (c) computation of reliability within- (versus between-) observer group. Three groups of four observers used a standard nine-category observational code for disruptive behavior in recording from videotapes of a classroom for 22 days. Analyses revealed considerable increases in average occurrence reliability as a function of the main effects of each of the experimental factors. The specific increases in reliability associated with each of the 12 combinations of the experimental factors are presented for each category of behavior. The possible role of observer-training procedures and behavioral definitions as determiners of nonartifactual reliability is discussed.

Journal Article↗

Reliability and validity of the self-report Quality of Life Questionnaire for Japanese School-aged Children with Asthma (JSCA-QOL v.3).

BACKGROUND: Asthma is a chronic disease prevalent in children which threatens their quality of life (QOL) through unexpected asthma attacks and/or the burden of daily self-management. As some conditions of chronic illness make it difficult for a child to accomplish normal developmental tasks, there may be fewer opportunities for the child to obtain a sense of achievement. This study investigated the reliability and validity of the Quality of Life Questionnaire for Japanese School-aged Children with Asthma Version 3 (JSCA-QOL v.3). This questionnaire includes 25 items with a 5-point Likert Scale format over five domains: "asthma attack triggers", "change in daily life", "family support", "satisfaction with daily life" and "restriction in participating in daily activities", and one summary scale. METHODS: In the present study, 2,425 children with asthma aged from 10 to 18 years were investigated in Japan. The internal consistency reliability of each domain was investigated with Cronbach's alpha reliability coefficient, and test-retest reliability with Spearman's correlations coefficient. Factorial validity by factor analysis using maximum-likelihood extraction with promax rotation was performed. Data analysis was performed using SPSS 12.0J. RESULTS: The final number of effective replies was 2,097 (the rate of effective data was 86.5%). "Asthma attack triggers", "change in daily life", "family support", "satisfaction with daily life" and "restriction in participating in daily activities" showed a high internal consistency (Cronbach's alpha = 0.66-0.86) as well as good test-retest reliability (Spearman's rho = 0.60, p < 0.01). The factorial validity was appropriate (KMO value = 0.90), because it was conceivable that the five factors extracted from factor analysis would be the same as in our hypothesis and support constructive validity. In addition, there was good correlation between the summary scale and the total QOL score (Spearman's rho = 0.58, p < 0.01). CONCLUSIONS: The present study showed that the JSCA-QOL v.3 is a reliable and valid measurement tool that can be used to appropriately assess QOL in school-aged children with asthma. As the JSCA-QOL v.3 can be easily completed in about 10 minutes, it can contribute as an efficient evaluation tool of the outcome of medical treatment through continual utilization in the outpatient clinic. The JSCA-QOL v.3 allows a health provider to help school-aged children with asthma to achieve their developmental tasks.

Adolescent↗

Diabetologists' judgments of diabetic control: reliability and mathematical simulation.

In study 1, laboratory and supervised blood or urine test data from actual cases were used to develop patient profiles. Seven diabetologists from the same institution rated the diabetic control of 125 profiles on a four-point scale (1 = poor, 2 = fair, 3 = good, 4 = excellent). Six of the 7 diabetologists demonstrated adequate intra- and interrater reliability. Study 2 assessed the reliability of judgments of diabetic control made by diabetologists working in two different settings. There were 9 raters from institution 1 and 8 from institution 2. The impact of the amount and type of information on judgment reliability was evaluated by developing two types of profiles. The test form contained only laboratory and supervised blood or urine test data similar to that utilized in study 1. The history form contained this information as well as other descriptive data typically available to diabetologists. The 17 diabetologists rated 125 anonymous profiles on each of two separate occasions approximately 1 wk apart. On one occasion they rated profiles presented on the test form. On the other occasion they rated profiles presented on the history form. As in study 1, the diabetologist raters demonstrated adequate intra- and interrater reliability. Intrarater reliability was somewhat better when rating test form profiles compared with history form profiles. Reliability was not higher within than between institutions. An analysis of the relative contribution of different diabetes control indices to the diabetologists' judgments indicated that HbA1 influenced raters' judgments at both institutions more than any other single variable.(ABSTRACT TRUNCATED AT 250 WORDS)

Adolescent↗

Comparative clinical reliability of fasting plasma glucose and glycosylated hemoglobin in non-insulin-dependent diabetes mellitus.

Because accurate determination of glycosylated hemoglobin (GHb) is difficult and relatively expensive in comparison with the modest cost and ready availability for tests of fasting plasma glucose (FPG), we examined the reliability of repeated measurements of FPG and GHb in typical diabetic outpatients taken in the usual clinical setting. We determined FPG and GHb concurrently on three separate occasions spanning 4 wk in 41 patients with non-insulin-dependent diabetes mellitus (NIDDM) and, for contrast, 5 with insulin-dependent diabetes mellitus (IDDM). Most of the NIDDM subjects were obese, with initial FPG levels ranging from 93 to 355 mg/dl. The reliability of each test was estimated by calculating two measures: the intraclass correlation coefficient (rho I) and the coefficient of variation (CV) for the repeated test values. For NIDDM patients treated with diet or oral hypoglycemic agents (OHA), rho I for FPG, log(FPG), and GHb were very similar. For insulin-treated NIDDM patients, rho I for FPG was somewhat lower than the coefficient in other treatment groups, and the reliability of FPG by this measure did not match the reliability of GHb within the limits of statistical significance. By analyzing the CV of test values repeated within subject, the reliability of FPG did not differ from GHb in any of the NIDDM treatment groups. Although patients were recruited sequentially to minimize sample selection bias, caution must be exercised in the interpretation of the statistical analyses of reliability with either rho I or CV due to limitations imposed by small sample size.(ABSTRACT TRUNCATED AT 250 WORDS)

Analysis of Variance↗

Percent of agreement among raters and rater reliability of the copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition.

The purpose of this study was to investigate the interrater reliability of the visual-motor portion of the Copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition. Eight raters independently scored 11 protocols completed by children aged 5 through 10 years, using the scoring criteria and guidelines in the manual. The raters marked each of 10 items pass or fail and computed a total raw score for each protocol. Interrater reliability coefficients were obtained for each child's protocol, and the Kappa coefficient was computed for each item. Significant raters' reliability coefficients ranged from .82 to .91, which were low in comparison to test-retest reliability and Kuder-Richardson-20 coefficients for this and other subtests of the Stanford-Binet in the technical manual. Percent agreement among 8 raters also indicated weak reliability. Although the obtained results suggested some interrater reliability coefficients within acceptable levels, questions were raised about the scoring criteria for individual items. Caution is warranted in the use of cognitive measures which include subjective judgement of the examiner in applying scoring criteria.

Child↗

Development of the ego and discomfort anxiety inventory: initial validity and reliability.

This article reports on four studies regarding the development, reliability, and validity of scales to measure two forms of anxiety, ego anxiety and discomfort anxiety. In the first study 140 undergraduates completed fourteen items related to ego anxiety and discomfort anxiety, as well as the Self-esteem Scale, the IPAT Anxiety Scale, and the Hopelessness Scale. Principal component analyses produced two factors, each with five items that showed differentiation between ego anxiety and discomfort anxiety. Guttman scales were developed from the items in the two factors. The resulting Ego Anxiety Scale had a coefficient of reproducibility of .94, a coefficient of scalability of .67 and estimated scale reliability of .84. The Discomfort Anxiety Scale had a coefficient of reproducibility of .91, a coefficient of scalability of .65 and estimated scale reliability of .83. Significant relationships were found between the scores on the two anxiety scales and scores on the Self-esteem Scale and the IPAT. Correlations between scores on the new anxiety scales and scores on the Hopelessness Scale were not significant. In the second study, undergraduates completed the Ego and Discomfort Anxiety Scales, the Self-esteem Scale, the IPAT Anxiety Scale, and the Hopelessness Scale. The reliability of the Ego Anxiety Scale (0.77) and the Discomfort Anxiety Scale (0.85) was estimated using Cronbach's alpha measure. t tests for scores of independent samples for students in Studies I and II were completed for scores on the Ego and Discomfort Anxiety Scales, Self-esteem Scale, IPAT Anxiety Scale, and Hopelessness Scale. None of these test comparisons were significant. The data from Studies I and II were pooled to provide tentative normative data for the Ego and Discomfort Anxiety Scales. The third study explored the reliability and validity of the new scales, testing 79 undergraduates who completed the Ego and Discomfort Anxiety Scales, the Fear of Negative Evaluation Scale, and the Problem Solving Inventory. The reliability coefficients of the Ego Anxiety Scale and Discomfort Anxiety Scale were 0.75 and 0.82, respectively. The differences between the combined scores of subjects in Studies I and II and the scores of subjects in the third study on the Ego and the Discomfort Anxiety Scale were not significant. A significant positive correlation, however, was found between scores on the Ego Anxiety Scale and scores on the Fear of Negative Evaluation Scale.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent↗

Reliability of measuring forward head posture in a clinical setting.

We believe there is a need to identify a practical method for determining objective measurement of forward head posture. In our study, we determined the within-tester and between-tester reliabilities for clinical measurements of static, sitting, forward head posture using the cervical range of motion (CROM) instrument. Repeated measurements were made using a standardized protocol on 40 patients seated in a standardized position. The seven testers had from 1 to 8 years of clinical experience. All measurements were recorded by the same investigator. The intraclass correlation coefficient (ICC[1,1]) was used to quantitate within-tester and between-tester reliability. Measurements of forward head position performed by the same physical therapist had high reliability (ICC = 0.93). Good reliability (ICC = 0.83) was demonstrated when different physical therapists measured the forward head posture of the same patient. We concluded that measurements of forward head posture made by physical therapists trained in the correct use of the CROM instrument are reliable. This reliability is important for determining the effectiveness of treatment programs. On the basis of our data, the CROM instrument will assist clinicians in the objective evaluation and reassessment of the patient population demonstrating forward head posture.

Adult↗