Search PubMedSearch

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

How to evaluate intraexaminer reliability using an interexaminer reliability study design.

Examiner reliability is often investigated to make generalizations about a profession's performance with various diagnostic tests. Although the evaluation of interexaminer reliability is straightforward, assessment of intraexaminer reliability can be problematic. Estimations of intrarater reliability can be inflated due to correlated error and the difficulty in blinding the examiners. A statistical method is presented that permits the investigator to compute intraexaminer as well as interexaminer reliability and precision from the findings of an interexaminer reliability investigation. In this approach, intraclass correlation coefficients are constructed from variance components estimated from a simple repeated-measures design: each of two or more raters evaluates each subject one time. The method can be applied to nominal and ordinal as well as interval data. Not only can spurious results be avoided, but time and funds can be saved by reducing the required number of subject ratings.

Analysis of Variance

Reliability of reliability coefficients in the estimation of asymmetry.

Although promising to provide insight into the interaction between genotype and environment, investigations into fluctuating asymmetry suffer from a lack of standardization in the reporting of measurement error. In the present paper we show, using both anthropometric and odontometric data, that the use of the reliability coefficient calculated for a bilateral measurement provides no indication of the reliability of the corresponding asymmetry estimate, because reliability of asymmetry depends on the relationship between measurement error and the difference between sides. Thus, we suggest that future investigations either provide reliability coefficients for asymmetry estimates specifically, or use methods that account for measurement error.

Anthropometry

Assessing the likelihood of reliable workplace behavior: further contributions to the validation of the Employee Reliability Inventory.

This paper summarizes a number of studies in which the validity of the Employee Reliability Inventory, a preemployment screening instrument designed to assess the likelihood of reliable and productive workplace behavior, was examined. Criterion-related studies compared the scores of a broadly diverse group of job applicants with those obtained from an array of criterion and comparison groups, for whom there was documented evidence of reliable or unreliable behavior. Criterion-related evidence indicates that the six scales are effective in differentiating a variety of criterion groups with unreliable behavior from a number of different job applicant comparison groups. Construct-related evidence for the validity of an emotional adjustment scale is reported as well. The issue of response distortion in preemployment inventories is discussed, and data are reported which indicate that scores on all six scales appear to be functionally free from the potentially confounding effects of response distortion. These results are consistent with the original validation and cross-validation findings, which supported the validity of the six scales when assessing the likelihood of reliable behavior in a population of job applicants.

Adult

Reliability of psychiatric diagnosis. II. The test/retest reliability of diagnostic classification.

In a study of interrater diagnostic reliability, 101 psychiatric inpatients were independently interviewed by physicians using a structured interview. Newly admitted patients were randomly selected and examined by one of three psychiatrists. A second psychiatrist reexamined the same patient about 24 hours later. Interviews from the two examinations were evaluated independently and diagnoses were made on the basis of objective criteria. The degree of diagnostic agreement for the two examinations were calculated using the kappa statistic. Agreement was found to be high as compared to other studies in the psychiatric literature, despite the fact that in most previous investigations diagnoses were not made independently. The results were also compared to studies of reliability of medical judgments. Possible reasons for the high interrater reliability are discussed and include the use of a structured interview and objective diagnostic criteria.

Attitude of Health Personnel

Reliability of life event assessments: test-retest reliability and fall-off effects of the Munich Interview for the Assessment of Life Events and Conditions.

This paper presents the findings of two independent studies which examined the test-retest reliability and the fall-off effects of the Munich Life Event List (MEL). The MEL is a three-step interview procedure for assessing life incidents which focuses on recognition processes rather than free recall. In a reliability study, test-retest coefficients of the MEL, based on a sample of 42 subjects, were quite stable over a 6-week interval. Stability for severe incidents appeared to be higher than for the less severe ones. In the fall-off study, a total rate of 30% fall-off was noted for all incidents reported retrospectively over an 8-year period. A more detailed analysis revealed average monthly fall-off effects of 0.36%. The size of fall-off effects was higher for non-severe and positive incidents than for severe incidents. This was particularly evident for the symptomatic groups. Non-symptomatic males reported a higher overall number of life incidents than females. This was partly due to more frequent reporting of severe incidents. The findings of the fall-off study do not support the common belief that the reliability of life incident report is much worse when the assessment period is extended over a period of several years as compared to the traditional 6-month period.

Adult

Optimum split-half reliabilities for the Rorschach: projective techniques are more reliable than we think.

A new technique for optimizing split-half reliability estimates yielded substantial and comparable increases in indexes of internal reliability for some major Rorschach variables across two diverse samples. Practical suggestions for applying this method to other projective tests were advanced. The inadvisability of computing odd-even reliability coefficients without regard to split-half distributional anomalies is addressed.

Adult

The reliability of reliability.

Forty-five original articles addressing the subject of examiner reliability were reviewed to determine if the findings were adequately substantiated by the statistical analyses and experimental designs employed by the authors. Only 10 studies were determined to have properly supported conclusions, while an additional three studies contained correct conclusions by coincidence. Eight investigations had invalid designs and three contained claims that were contradicted by the author's findings. Half the studies were found to have conclusions that were based on inappropriate or inconclusive statistical analysis. To date, the research presented in the chiropractic literature cannot substantiate claims concerning the reliability of any diagnostic instrumentation or palpatory procedures commonly employed by chiropractic physicians.

Chiropractic

Endodontic recall radiographs: how reliable is our interpretation of endodontic success or failure and what factors affect our reliability?

Three hundred thirty cases were selected from an endodontic practice. Postoperative and recall radiographs of each case were examined by four endodontists for an interpretation of treatment success or failure. One hundred eighteen cases were examined a second time by each endodontist. Initial analysis showed substantial inconsistency in both inter- and intraobserver interpretation. The cases were then categorized by average radiographic density differences within radiograph sets, anatomic location of the treated tooth, technical compatibility within radiograph sets, and by length of time between postoperative and recall radiographs. It appears that these factors do not affect reliability of success/failure interpretation.

Dental Pulp Cavity

Reliability of the AMDP-system. A preliminary report on a multicentre exercise on the reliability of psychopathological assessment.

The AMDP-System is a documentation system for psychiatric data widely in use in the German-speaking countries. A summary of results of a multicentered study of interrater agreement of the Psychopathology Scale is presented. A new index of rater agreement was tested and the notion is discussed that the judgement of the presence and the absence of a symptom are two different processes with different reliability.

Germany, West

[Goal attainment scaling: reliability and practical experiences with 397 psychiatric treatment courses. Part 1: Evaluation of validity and reliability].

Standardized outcome measures are often criticized, individualized criteria being preferred. Goal attainment scaling (GAS) is such an individual evaluation tool. It aims at measuring, whether a patient attains, what is thought to be his potential. For 36 psychiatric patients outcome estimated by GAS was compared with traditional outcome-measures ("Brief psychiatric rating scale", BPRS, "Clinical global impression", CGI, outcome-scales of Strauss and Carpenter and a patient self-rating. Concurrent validity was sufficient, compared to BPRS, CGI and the scale "absence of symptoms". The other scales of Strauss and Carpenter and the patient self-rating did not correlate closely with GAS. Traditional interraterreliability was sufficient, but if two raters constructed separate scales for one patient, their GAS scores correlated weakly. GAS is recommended for further study in spite of its obvious shortcommings: there seems to be no other way of quantifying, to what degree a patient realized his individual potential.

Activities of Daily Living

Reliability and validity of the chronic respiratory questionnaire (CRQ).

BACKGROUND: The Chronic Respiratory Questionnaire (CRQ) is frequently applied to assess quality of life in patients with chronic obstructive pulmonary disease (COPD). However, the reliability and validity of this questionnaire have not yet been determined. This study investigates the reliability and validity of the four separate dimensions of the CRQ. METHODS: The CRQ was administered on two consecutive days to 40 patients with COPD (mean FEV1 44% predicted, FEV1/IVC 37% predicted). Internal consistency reliability of each dimension was investigated by Cronbach's alpha reliability coefficient, test retest reliability by the Spearman-Brown reliability coefficient (p), and content validity by Pearson's correlation coefficient between the CRQ and the symptom checklist (SCL-90). RESULTS: Items of the fatigue, emotion, and mastery dimensions showed a high internal consistency reliability (alpha = 0.71-0.88) as well as a high test retest reliability (p above 0.90). These three dimensions correlated with comparable dimensions of the SCL-90. Items of the dyspnoea dimension showed a low internal consistency reliability (alpha = 0.53) and a test retest reliability of p = 0.73. CONCLUSIONS: Items of the dimensions fatigue, emotion, and mastery of the CRQ are reliable and valid and can be used to assess quality of life in patients with severe airways obstruction. Items of the dyspnoea dimension are less reliable and should not be included in the overall score of the CRQ in comparative research. However, by scoring the items of dyspnoea separately they may be useful for the evaluation of the effects of intervention in a specific patient.

Aged

Interrater reliability of auscultation of breath sounds among physical therapists.

BACKGROUND AND PURPOSE: Although auscultation is routinely used in the assessment of respiratory status, the ability of the rater to accurately and consistently identify lung sounds has been questioned. The literature on this issue is sparse and has focused on reliability of auscultation of tape-recorded rather than in vivo lung sounds. The purposes of this study were to determine the interrater reliability of physical therapists in the direct auscultation of lung sounds based on their clinical experience in chest physical therapy and to determine whether the adoption of standardized nomenclature and education on proper technique and interpretation affects reliability. SUBJECTS AND METHODS: A group of 57 registered physical therapists were stratified by clinical experience into four groups. Sixteen therapists (ie, 4 in each stratum) were randomly chosen using a random number table. The following criteria were developed to delineate clinical experience. Group 1 subjects were senior chest physical therapists with at least 5 years of experience in this area of practice. Group 2 subjects were experienced therapists who had a minimum of 2 years of experience in chest physical therapy and were currently practicing in this area. Group 3 subjects were experienced physical therapists in other areas who were also practicing in chest physical therapy on occasional weekend service. Group 4 subjects were new graduates. Ten patients were evaluated by each group of 4 physical therapists using a teaching stethoscope with one diaphragm/bell and four pairs of earpieces. The education session consisted of discussion of the adoption of standardized nomenclature and education on proper technique and interpretation of auscultation. Interrater reliability was assessed before and after the education session using kappa (kappa) values. Comparisons were made between kappa values before and after the education session to determine the effect of education and between groups to determine the effect of clinical experience. RESULTS: The kappa values before the education session were low, indicating poor reliability in detecting specific abnormal sounds (kappa = -.02-.59). Group 1 (seniors in respiratory therapy) and group 4 (new graduates) demonstrated the greatest reliability levels. The lowest kappa values were observed for detecting and categorizing the quality of breath sounds (normal, absent, bronchial, or decreased) (kappa = -.02-.25). Following the education session, there was a general improvement in reliability (kappa = -.30-.77), especially for group 3 (specialists in other areas). The most improvement was noted for the detection of the quality of breath sounds (kappa = .08-.50). CONCLUSION AND DISCUSSION: Reliability of auscultation was poor to fair, in general, before the education session. There was a definite improvement in reliability after the education session. There was no clear effect of clinical experience on reliability, and the agreement among observers appeared to depend on the abnormal lung sound present. Limitations of this study and recommendations for future research are discussed. [Brooks D, Thomas J. Interrater reliability of auscultation of breath sounds among physical therapists.

Auscultation