Search PubMedSearch

SEARCH · Search PubMed

Results for “Observer Variation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

[Observer variation in clinical practice. A review].

Observer variation denotes the discrepancy between two consecutive observations of the same indicant. Principles and results of observer variation studies are outlined. Observer variation seems to be present wherever it is searched for, in history taking as well as in physical examination. It has been found unanimously to be smaller for intra- than for inter- observer variation while it is still controversial whether clinical experience lessens observer variation. Even fundamental clinical observations and particularly evaluations of borderline cases between normal and abnormal exhibit observer variation. The possibilities of reducing observer variation, and the importance of observer variation to clinicians and patients, and also the influence on research and society are referred to briefly. The future role of observer variation studies is discussed. Observer variation often is evaluated by the Kappa statistics whose statistical advantages and clinical disadvantages are outlined.

Observer Variation

Is pancreatogram interpretation reliable?--a study of observer variation and error.

Observer variation in the interpretation of endoscopic pancreatograms has been examined by asking four experienced observers to assess 40 sets of well-documented radiographs (from 20 patients with cancer and 20 with pancreatitis), both without ("blind") and with clinical details, each on three occasions. Individual consistency for "blind" diagnoses ranged from 61% to 78%, increasing significantly with clinical information. Overall diagnostic accuracy with clinical information varied from 52% to 83% for cancer, and from 87% to 95% for pancreatitis. However, unanimous and correct opinions were given by the four observers for only 53% of all cases, even when clinical details were provided. Clinical information changed the radiographic diagnosis in 43% of assessments, 83% of these changes leading to improved accuracy. ERCP gives direct information about the major pancreatic and biliary duct systems and often provides an accurate diagnosis. Caution must be exercised in relying upon radiological appearances alone.

Diagnosis, Differential

Observer variation in quantification of immunocytochemistry by image analysis.

This paper reports the findings of a study designed to examine observer variation as a source of inaccuracy inherent in the use of computer-assisted image analysis to measure areas of stained tissue. The rat pituitary immunostained for prolactin and galanin was used as an example to estimate patterns of immunoreactivity exhibited by different cell types. Six observers, with differing experience, selected grey level threshold values on 40 fields of images of stained tissue making three repeats of each field. The 40 fields consisted of 20 serial pairs of colocalized fields, one immunostained for prolactin, the other for galanin. The 20 pairs consisted of four pairs from each of five animals. Analysis of observer variation in the selection of threshold values showed large differences in the within- and between-observer variation. Analysis of the components of variance in the estimation of the ratios of stained tissues showed that the major source of variation was the within-observer component. An additional experiment using two observers, where half of the images were compared to the original microscope images before setting threshold levels, showed that the opportunity to make a comparison did not reduce observer variation. It is suggested that any study which uses semi-automatic methods to segment regions of a digital image can benefit from an analysis of this kind so that the sources of variation can be determined to enable maximum discriminating power in future studies.

Animals

Observer variation in recording clinical data from women presenting with breast lesions. Report from the Yorkshire Breast Cancer Group.

The degree of observer variation in recording 11 186 items of clinical data from 242 woman who presented complaining of a lump in the breast to a group of 10 surgeons was studied. Each women was interviewed and examined twice and the findings (of the two clinicians) compared. There was a wide range of variation among the observers. Variation in recording the presence or absence of axillary nodes was considerable (45%), as was that in sizing the primary lesion (55%). In 20% of cases the two assessments of primary lesion size differed by over 2 cm. In other respects the results were more encouraging; much variation could be eliminated by wording the proforma more clearly. Moreover, variation was not person-specific, so that these findings are probably reasonably representative. Any future trial of breast lesions should (a) design specific proformata, (b) define terminology, (c) make these definitions universally available, and (d) conduct observer variation studies before the start of the full trial.

Analysis of Variance

[Observer variation and accuracy in the clinical diagnosis of ascites].

Seventeen observers participated in an observer variation study of the clinical evaluation of ascites. In a blinded design, eight patients with diagnoses of liver disease were examined. Fourteen observers examined the patients twice in order also to estimate the intra-observer variation. The accuracy of the observers' statements was compared with ultrasound findings, by which mean ascites was demonstrated in two patients. Poor correspondance between the observers' gradings and volume estimations, and poor accuracy of the gradings, make qualitative and quantitative ascites estimations useless. The inter-observer agreement was found to be low although the intra-observer agreement was good. The observers' subjective certainty of correctness of their own findings, marked as certain/uncertain, did not reflect the chance of making a correct statement on each particular occasion. The individual patient's general ability of inducing certainty did relate the chance of forming a correct diagnosis. Ultrasonic investigation of the abdomen is recommended in all situations in which demonstration of ascites is essential to diagnosis or therapy.

Ascites

[Observer variations in computerized 201 thallium myocardial scintigraphy].

The observer variation (reproducibility) in computerized 201Tl scintigraphy was determined by describing the results of 20 investigations as being normal, suggestive of ischaemia or representing scar tissue. At an interval of two months, two physicians evaluated the results independently. As to normality, the interobserver variation was small, the level of agreement corrected for chance agreement, the kappa coefficient, being from 69-100%. For the diagnosis of ischaemia, kappa was also high, 58-100%. The intraobserver variations were of the same magnitude. In scar tissue, the kappa coefficients were somewhat lower. The present method has an acceptable observer variation in distinguishing between a diseased and a normal myocardium. The lower kappa coefficients for scar tissue and recent findings showing that persistent defects on the delayed images are viable tissue in 30% of the cases, makes it difficult to diagnose scar tissue.

Adult

Observer variation in the measurement of Breslow depth and Clark's level in thin cutaneous malignant melanoma.

We have assessed the degree of observer variation of both Breslow depth and Clark's level in a series of 50 thin malignant melanomas. Our findings are similar to those of previous international studies in the Breslow depth is the more reproducible measure. Significant intra- and inter-observer variation exists and in some cases it was up to +/- 0.86 mm. Even small differences will potentially affect patient management at our centre and this was analysed using kappa statistics. Good agreement was found between observers and this could be improved by comparing the mean of two or more measurements. This removes larger errors, but smaller observer errors and differences in subjective interpretation of the deepest malignant cell mean that agreement will never be more than 90 per cent. This is high compared with studies of observer variation in other pathological conditions, e.g., dysplasia of the cervix, but where surgical management is potentially disfiguring it is not high enough. We conclude that Breslow depth and Clark's level should not be the sole basis of wide excision protocols.

Biometry

[Assessment of micturition cystourethrography. Intra- and inter-observer variation].

The reliability of any method of investigation depends upon the accuracy and reproducibility of the results of the investigation. The accuracy of voiding cystoureterography (VCU) which is greatly dependent on the radiographic assessment cannot be assessed because no standard answers exist. The reliability of the method may be assessed in the form of intra- and inter-observer variations. VCU investigations from 24 women with incontinence were assessed by two independent radiologists. Fifteen of the radiological examinations had previously been described by one of the radiologists so that intra-observer variation could be assessed by these. Inter-observer variation was 70% (95% confidence limits 51-89%), calculated from the diagnoses anterior and posterior suspension defects, combined defects and normal conditions. The corresponding intra-observer variation was 53% (95% confidence limits 27-78%). The radiographic criteria for subdivision of suspension defects into anterior and posterior defects are, theoretically, very simple but appear to be difficult to attain in practice. The indications for employing a form of examination where assessment of the result of examination is so obscure should be very weighty.

Adult

Observer variation in detecting the radiologic features associated with bronchiolitis.

Chest radiographs are commonly obtained to assess children for bronchiolitis, both to corroborate the diagnosis and to exclude other diagnostic possibilities. Their utility in this setting has not previously been examined. Using a blinded, randomized study design, we examined the interobserver and intraobserver variation in the detection of the radiologic features of bronchiolitis from the chest radiograph using "weighted kappa" statistics. This observer variation was compared with that found by other authors for other diagnoses. We also determined the reported presence of these radiologic features in radiographs from patients with bronchiolitis as compared with normal controls. Our study showed acceptable interobserver (kappa = 0.40-0.66) and intraobserver agreement (kappa = 0.50-0.78) on the radiologic features of bronchiolitis relative to other diagnoses. We demonstrated a higher reported presence of these accepted radiologic features in patients with bronchiolitis as compared to controls. Although kappa statistics are widely used in studies of observer variation, "weighted kappa" has received little attention in the radiologic literature. This statistical analysis allows observers to equivocate on the presence or absence of a feature and therefore allows the format of observer variation studies to simulate more closely the normal clinical setting.

Bronchiolitis, Viral

Observer variation in the radiographic classification of ankle fractures.

We recorded inter- and intra-observer variations in the classification of ankle fractures by the Lauge Hansen and Weber systems. Radiographs of 94 patients were classified independently by four observers. The observer variation was calculated by kappa statistics, which corrects the obtained values for the agreement expected by chance. There was an acceptable level of agreement for the overall classification into both systems. For the staging of supination-adduction and supination-eversion fractures in the Lauge Hansen system the agreement was poor. The results indicate that future classification systems should be subject to reliability analysis before they are accepted.

Ankle Injuries

A comparison of three methods of assessing inter-observer variation applied to measurement of the symphysis-fundal height.

This study assesses the inter-observer variation of symphysis-fundal height measurements by three methods--the coefficient of variation, the correlation coefficient and the limits of agreement. The coefficient of variation was 4% and the correlation coefficient 0.959, yet the limits of agreement were very wide. Applying these limits to centile charts of symphysis-fundal height shows that the fundal height cannot be measured by different observers with sufficient agreement to separate small fundal heights from those which are not small, and this severely limits the usefulness of measurement of the symphysis-fundal height as a screening test for intrauterine growth retardation. We conclude that inter-observer variation should be assessed by the method of limits of agreement, and not by calculating the coefficient of variation or the correlation coefficient.

Female

Inter-observer variation in assessment of undescended testis. Analysis of kappa statistics as a coefficient of reliability.

In a prospective study the inter-observer variation in the diagnosis of undescended testis was analysed. Two physicians assessed independently the position and motility of the testes of 37 boys referred for undescended testis. The boys were examined in the supine and squatting positions. The observed agreement rate between the observers was 0.90 to 0.97. Using kappa (kappa) statistics, the values were adjusted for the expected chance agreement; kappa values between 0.47 and 0.81 were obtained, slightly higher values for patients in the supine position. Complete agreement on all observations was reached in 13.5% of the patients. Inter-observer variation may be a substantial source of bias in diagnosing the undescended testis and one of the reasons for the varying results in studies of hormonal treatment of this condition. It is also a fact that the number of orchiopexies in some countries exceeds the incidence of this condition.

Child

Observer variation and depressive phenomenology.

An investigation into observer variation and depressive phenomenology is described. A group of 6 clinicians independently completed a 43-item sheet for 20 depressed patients at the same clinical interview. The coefficient of agreements on the items are given. Patient variables and observer variables did not have any influence on the degree of agreement. There was also an agreement in the subtyping of depression.

Adjustment Disorders

Histologic features and observational variation in cerebellar gliomas in children.

Variation existed in the recognition of histologic features commonly used in the evaluation of cerebellar gliomas of childhood. Some histologic features (e.g., perivascular pseudorosettes, leptomeningeal deposits, and calcification) were more reliably observed than were others (e.g., Rosenthal fibers, cell density, and hypervascularity). Knowledge of which features tend to have greater observational variation may lead to improved definitions, less reliance of these features in clinical decisions, further studies of the potential sources of the variation, and guidelines for minimizing observational variation.

Arachnoid

Observer variation in the assessment of scintigraphy of the thyroid gland.

In order to determine observer variation in the assessment of thyroid scintigrams two specialists in nuclear medicine and two specialists in endocrinology independently evaluated 240 thyroid pertechnetate scintigrams twice, and assessed a number of variables concerning size and isotope uptake. The observed agreement between pairs of observers for the variables ranged from 0.70 to 0.98. By the use of the kappa coefficient the observed agreement was adjusted for change agreement. Kappa can variate from -1 (total disagreement) to +1 (perfect agreement). Kappa values between 0.29 and 0.86 were found. In the intraobserver study the observed agreement ranged from 0.83 to 0.99 resulting in kappa coefficients between 0.53 and 0.96. Thus the level of agreement in the present study was "fair to substantial" for agreement in the interobserver part and "substantial to almost perfect" for agreement in the intraobserver part. No difference was found in the level of agreement between the nuclear specialists and the endocrinologists. Although the treatment of patients is based on knowledge of the case histories and clinical and laboratory findings the high degree of observer variation may lead to misclassification of a number of patients with thyroid disease and subsequently a less optimal choice of treatment.

Adolescent

Observer variation in the clinical assessment of the thyroid gland.

In order to evaluate the reliability of clinical assessment of the thyroid gland, two specialists in endocrinology and two younger doctors independently examined 53 patients twice, and assessed whether they had a diffuse goitre, a multinodular goitre, a solitary nodule or a normal gland. In 30% of the patients all four observers were in agreement, whereas in 47% and 23% of the patients, two and three different diagnoses were given, respectively. Inter-observer variation was determined and kappa values between -0.04 and 0.54 were found. Intra-observer variation was smaller, revealing kappa values between 0.44 and 1.00. The present study suggests that clinical assessment of the thyroid gland may lead to misclassification of the type of thyroid disease, and thereby to a less than optimal choice of therapy.

Adult

Observer variation in interpreting radiographs of the pituitary fossa.

Observer variation in interpreting sellar radiographs in patients suspected or known to have a pituitary tumor has been examined. Two radiologists experienced in interpreting sellar radiographs examined independently, without clinical details, plain films and tomograms of the sella of 101 patients. In most, only minor changes were anticipated. Of the 93 female patients, 67 were under investigation for amenorrhea. Radiographs were examined four times, each radiologist examining each set twice. Appearances were classified as normal, doubtful or abnormal on each occasion. Overall intraobserver agreement was 76%--85%. Neither radiologist changes his opinion by more than one category, e.g. from normal to doubtful. Overall interobserver agreement was 63%--75%. Disagreement between observers concerning 11 (11%) of the patients resulted from differences of opinion about whether minor changes in sellar outline represented an abnormality or merely a normal variation. Kappa analysis suggested that much of the agreement may be ascribed to chance. Agreement rates resemble those for other clinical and radiological investigations.

Adolescent

Observer variation in equine abdominal auscultation.

The reliability of abdominal auscultation was investigated via an observer variation study. Clinicians listened to a variety of minute-long equine gut sound recordings. They evaluated the amount of gut sounds as 'absent', 'decreased', 'normal', or 'increased'. They subsequently evaluated the same recordings replayed in a different order. Intra- and inter-observer agreement was measured by the statistic kappa. There was significant intra-observer (kappa 0.57) agreement, but less agreement between observers (kappa 0.37). The best agreement was on the classification of sound tracks as 'absent' (intra-observer kappa 0.72 and inter-observer kappa 0.55). There was significant correlation between the clinicians' average assessment of the recordings and their acoustic energy levels. In this study abdominal noise was reliably assessed by auscultation. Standardised techniques and definitions would probably enhance the reliability of abdominal auscultation for the evaluation of gastrointestinal disease.

Abdominal Pain