Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

Accuracy and reliability of periodic sharp wave complexes in Creutzfeldt-Jakob disease.

OBJECTIVE: To assess the sensitivity, specificity, and interobserver reliability of periodic sharp wave complexes in the electroencephalograms of patients with Creutzfeldt-Jakob disease. DESIGN: Sixty-eight electroencephalograms in 29 patients who had been suspected of having Creutzfeldt-Jakob disease were reanalyzed by an investigator who was unaware of the clinical data. The incidence of periodic sharp wave complexes in neuropathologically confirmed Creutzfeldt-Jakob disease vs progressive dementia other than Creutzfeldt-Jakob disease was assessed. Blinded electroencephalogram analysis was performed by a second investigator. The interobserver reliability was assessed by the kappa value. SETTING: University hospital, base of the German National Creutzfeldt-Jakob Disease Surveillance Study. PATIENTS: Fifteen patients with neuropathologically confirmed Creutzfeldt-Jakob disease and 14 patients who had been suspected of having Creutzfeldt-Jakob disease because of rapidly progressive dementia but in whom other dementias were diagnosed by unblinded investigators based on clinical and electroencephalographic criteria. MAIN OUTCOME MEASURE: Sensitivity and specificity of periodic sharp wave complexes assessed by their incidence in Creutzfeldt-Jakob disease vs other dementias. Interobserver reliability of periodic sharp wave complexes was expressed by the kappa value. RESULTS: For periodic sharp wave complexes, blinded electroencephalographic analysis resulted in a sensitivity and a specificity of 67% and 86%, respectively. Interobserver reliability was excellent (kappa = 0.95). CONCLUSION: This blinded electroencephalographic study in Creutzfeldt-Jakob disease confirms the high diagnostic value of electroencephalography, as previously reported by open studies.

Aged↗

Reliability and validity of the Ocular Surface Disease Index.

OBJECTIVE: To evaluate the validity and reliability of the Ocular Surface Disease Index (OSDI) questionnaire. METHODS: Participants (109 patients with dry eye and 30 normal controls) completed the OSDI, the National Eye Institute Visual Functioning Questionnaire (NEI VFQ-25), the McMonnies Dry Eye Questionnaire, the Short Form-12 (SF-12) Health Status Questionnaire, and an ophthalmic examination including Schirmer tests, tear breakup time, and fluorescein and lissamine green staining. RESULTS: Factor analysis identified 3 subscales of the OSDI: vision-related function, ocular symptoms, and environmental triggers. Reliability (measured by Cronbach alpha) ranged from good to excellent for the overall instrument and each subscale, and test-retest reliability was good to excellent. The OSDI was valid, effectively discriminating between normal, mild to moderate, and severe dry eye disease as defined by both physician's assessment and a composite disease severity score. The OSDI also correlated significantly with the McMonnies questionnaire, the National Eye Institute Visual Functioning Questionnaire, the physical component summary score of the Short Form-12, patient perception of symptoms, and artificial tear usage. CONCLUSIONS: The OSDI is a valid and reliable instrument for measuring the severity of dry eye disease, and it possesses the necessary psychometric properties to be used as an end point in clinical trials.

Diagnostic Techniques, Ophthalmological↗

Reliability and validity of refractive error-specific quality-of-life instruments.

OBJECTIVE: To evaluate the reliability and validity of the National Eye Institute Refractive Error Quality of Life Instrument (NEI-RQL-42) and the Refractive Status and Vision Profile survey (RSVP). METHODS: Eighty-one participants with good visual acuity (better than 20/30 best-corrected acuity in each eye) completed the NEI-RQL-42 and RSVP on 2 occasions. Noncycloplegic, subjective refractions and high-contrast visual acuity assessments were also performed. Statistical analyses addressed internal consistency, test-retest reliability, and validity (ie, concurrent and construct validity) of the 2 instruments. OUTCOME MEASURES: The NEI-RQL-42, RSVP survey, subjective refraction, and visual acuity. RESULTS: The internal consistency for the overall NEI-RQL-42 was excellent (Cronbach alpha = 0.91); and for the overall RSVP, good (Cronbach alpha = 0.81). Likewise, the test-retest reliability for the overall NEI-RQL-42 was excellent (intraclass correlation coefficient [ICC], 0.91; 95% limits of agreement, -9.1 to 10.1); and for the RSVP, fair (ICC, 0.76; 95% limits of agreement, -12.1 to 12.5). The NEI-RQL-42 overall score showed good concurrent validity as it correlated significantly with subjective refraction, whereas the RSVP overall score did not. The NEI-RQL-42 and RSVP showed similar construct validity in terms of refractive error discrimination, but the NEI-RQL-42 showed better construct validity when discriminating by the type of refractive correction used by patients. Between-instrument convergent and divergent validity was good. CONCLUSIONS: The NEI-RQL-42 and RSVP generally have good reliability and validity in this sample of patients with refractive error. However, other factors such as content should be considered in choosing 1 of these instruments for studies of refractive error correction.

Adult↗

Reliability of psychiatric diagnosis. I. A methodological review.

This article reviews some methodological aspects of studies of diagnostic reliability in psychiatry. We define and discuss the concept of interrater reliability and review some of the ways in which the design of the reliability study can influence the results. Three basic methodological issues are raised, including: importance of structured interviews and objective diagnostic criteria, the importance of a test/retest vs an interviewer/observer design, and the calculation of reliability in a way that takes chance agreement into account.

Attitude of Health Personnel↗

Reliability of the Group for Advancement of Psychiatry Diagnostic Categories in Child Psychiatry.

A total of 403 multiple diagnoses were independently assigned to 41 patient protocols by 73 psychiatrists, psychologists, and social workers to determine the levels of interrater reliability of the Group for the Advancement of Psychiatry (GAP) diagnostic categories. With the exception of the psychotic disorders category, these diagnostic categories were found to have low levels of interdiagnostician reliability. Differences in the reliabilities across disciplines and levels of training were found. It is noted, however, that neither years of experience, kind of training, nor direct contact with the patient can be regarded as a substitute for improvements in the classification system itself. The importance of a reliable classification system for child psychiatry is emphasized and suggestions for improvements in the present GAP system are made.

Child↗

Renard diagnostic interview. Its reliability and procedural validity with physicians and lay interviewers.

A psychiatric diagnostic interview that can be reliably and validly administered by nonpsychiatric physicians and lay interviewers has both research and clinical applications. We examined the interrater reliability and procedural validity of the Renard Diagnostic Interview (RDI), an instrument designed for these purposes. Randomly selected psychiatric inpatients were interviewed once by a psychiatrist using our standard departmental research interview and were then given RDIs by two psychiatrists, two lay interviewers, or one of each. The reliability of the RDI is estimated by examining diagnostic concordance for the two RDI interviews. Procedural validity is estimated by examining diagnostic concordance between the RDI and the traditional departmental interview. Both reliability and procedural validity were found to be high, and the study demonstrates that lay interviewers using the RDI after a brief period of training can obtain accurate diagnostic information.

Humans↗

Reliability of depression and associated clinical symptoms.

Interrater reliability assessments were undertaken for the Hamilton Depression Rating Scale, the Raskin Depression Rating Scale, and the Degree of Mental Illness Scale. Levels of reliability ranged from "poor" to "excellent" and varied as a function of (1) temporality (assessments made at termination of clinical trial more reliable than those made at randomization into treatment) and (2) unit of scoring (factor or total scores more reliable than single-item assessments). The implications of these results may be considered in the context of further studies evaluating the efficacy of treatment interventions on reduction of symptoms of clinical depression.

Adolescent↗

Long-term reliability of diagnosing lifetime major depression in a community sample.

Limited information is available on the reliability of diagnostic assessments in community populations. This study analyzed the 18-month test-retest stability of lifetime major depression determined from the Schedule for Affective Disorders and Schizophrenia-Lifetime Version using the Research Diagnostic Criteria. Overall, the reliability among the 391 female subjects was poor. Clinical status during the 18-month interval influenced reliability, while demographic, psychosocial, and interviewer characteristics were unrelated. The women who reliably reported lifetime episodes of depression were consistent about details such as medication use, but were inconsistent about other features, eg, number of episodes, length of longest episode, and age at first episode. The results suggest the need for caution in analyzing data on the lifetime prevalence of depression in community samples.

Adult↗

The reliability of the family history method for psychiatric diagnoses.

We evaluated the test-retest interrater reliability of the Family History Research Diagnostic Criteria (FH-RDC) in 58 depressed patients who described 341 first-degree relatives. Reliability was examined as a function of the threshold to determine caseness. In general, diagnostic reliability was good-excellent for specific FH-RDC disorders, but not for the residual category of other psychiatric disorder. A higher diagnostic threshold was associated with greater reliability, especially for the diagnosis of depression. Patient variance accounted for a greater percentage of the disagreements between the interviewers than did rater variance.

Data Collection↗

The lifetime history of major depression in women. Reliability of diagnosis and heritability.

BACKGROUND: In epidemiologic samples, the assessment of lifetime history (LTH) of major depression (MD) is not highly reliable. In female twins, we previously found that LTH of MD, as assessed at a single personal interview, was moderately heritable (approximately 40%). In that analysis, errors of measurement could not be discriminated from true environmental effects. METHODS: In 1721 female twins from a population-based register, including both members of 742 pairs, LTH of MD, covering approximately the same time period, was obtained twice, once by self-administered questionnaire and once at personal interview. RESULTS: Reliability of LTH of MD was modest (kappa = +.34, tetrachoric r = +.56) and was predicted by the number of depressive symptoms, treatment seeking, number of episodes, and degree of impairment. Deriving an "index of caseness" from these predictors, the estimated heritability of LTH of MD was greater for more restrictive definitions. Incorporating error of measurement into a structural equation model including both occasions of measurement, the estimated heritability of the liability to LTH of MD increased substantially (approximately 70%). More than half of what was considered environmental effects when LTH of MD was analyzed on the basis of one assessment appeared, when two assessments were used, to reflect measurement error. CONCLUSIONS: Major depression, as assessed over the lifetime, may be a rather highly heritable disorder of moderate reliability rather than a moderately heritable disorder of high reliability.

Adult↗

Diagnostic reliability of bipolar II disorder.

BACKGROUND: Although the diagnostic reliability of major depression and mania has been well established, that of hypomania and bipolar II (BPII) disorder has not. This remains an important issue for clinicians, especially for those undertaking genetic studies of BP disorder since bipolar I (BPI) and BPII disorders often cluster in the same families. We have assessed our diagnostic reliability of BP disorders, recurrent unipolar disorder, and their constituent episodes (major depression, mania, and hypomania) using interview and best-estimate diagnostic procedures used in a genetic study of families with BPI disorder. METHODS: Reliability was assessed for (1) co-rated Schedule for Affective Disorders and Schizophrenia-Lifetime version interviews of 37 subjects including 15 with BP disorders; (2) test-retest Schedule for Affective Disorders and Schizophrenia-Lifetime version interviews of 26 subjects including 13 with BP disorders; and (3) best-estimate diagnoses made by 2 noninterviewing psychiatrists on 524 subjects in a genetic linkage study of BPI disorder. Diagnoses were based on Research Diagnostic Criteria for a Selected Group of Functional Disorders, except that recurrent major depression as well as hypomania was required for a diagnosis of BPII disorder. RESULTS: On co-rated interviews, we observed complete agreement between interviewers for diagnosing major depressive, manic, and hypomanic episodes. For test-retest interviews, the Cohen kappa coefficients were 0.83 for manic, 0.72 for hypomanic, and 1.0 for major depressive episodes. At the best-estimate level, the Cohen kappa coefficients were 0.99 for BPI, 0.99 for BPII, and 0.98 for recurrent unipolar disorder. CONCLUSION: Good interrater reliability for BPII can be achieved when the interviews and best-estimate diagnoses are done by experienced psychiatrists.

Adult↗

Delirium in mechanically ventilated patients: validity and reliability of the confusion assessment method for the intensive care unit (CAM-ICU).

CONTEXT: Delirium is a common problem in the intensive care unit (ICU). Accurate diagnosis is limited by the difficulty of communicating with mechanically ventilated patients and by lack of a validated delirium instrument for use in the ICU. OBJECTIVES: To validate a delirium assessment instrument that uses standardized nonverbal assessments for mechanically ventilated patients and to determine the occurrence rate of delirium in such patients. DESIGN AND SETTING: Prospective cohort study testing the Confusion Assessment Method for ICU Patients (CAM-ICU) in the adult medical and coronary ICUs of a US university-based medical center. PARTICIPANTS: A total of 111 consecutive patients who were mechanically ventilated were enrolled from February 1, 2000, to July 15, 2000, of whom 96 (86.5%) were evaluable for the development of delirium and 15 (13.5%) were excluded because they remained comatose throughout the investigation. MAIN OUTCOME MEASURES: Occurrence rate of delirium and sensitivity, specificity, and interrater reliability of delirium assessments using the CAM-ICU, made daily by 2 critical care study nurses, compared with assessments by delirium experts using Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition, criteria. RESULTS: A total of 471 daily paired evaluations were completed. Compared with the reference standard for diagnosing delirium, 2 study nurses using the CAM-ICU had sensitivities of 100% and 93%, specificities of 98% and 100%, and high interrater reliability (kappa = 0.96; 95% confidence interval, 0.92-0.99). Interrater reliability measures across subgroup comparisons showed kappa values of 0.92 for those aged 65 years or older, 0.99 for those with suspected dementia, or 0.94 for those with Acute Physiology and Chronic Health Evaluation II scores at or above the median value of 23 (all P<.001). Comparing sensitivity and specificity between patient subgroups according to age, suspected dementia, or severity of illness showed no significant differences. The mean (SD) CAM-ICU administration time was 2 (1) minutes. Reference standard diagnoses of delirium, stupor, and coma occurred in 25.2%, 21.3%, and 28.5% of all observations, respectively. Delirium occurred in 80 (83.3%) patients during their ICU stay for a mean (SD) of 2.4 (1.6) days. Delirium was even present in 39.5% of alert or easily aroused patient observations by the reference standard and persisted in 10.4% of patients at hospital discharge. CONCLUSIONS: Delirium, a complication not currently monitored in the ICU setting, is extremely common in mechanically ventilated patients. The CAM-ICU appears to be rapid, valid, and reliable for diagnosing delirium in the ICU setting and may be a useful instrument for both clinical and research purposes.

APACHE↗

Reliability of computer image analysis of pigmented skin lesions of Australian adolescents.

BACKGROUND: The diagnosis of melanomas at an early stage is associated with improved survival, so the recognition of changes in pigmented skin lesions over time is important. We have developed a computer imaging system with the aim of assisting clinicians in differentiating early melanomas from benign pigmented skin lesions. The objective of this study was to investigate the system's reliability over time in measuring diagnostic characteristics of pigmented skin lesions, including their color, size, shape, and distinctness of boundary. METHODS: We captured video images of 5 lesions, all larger than 2 mm in greatest dimension, on each of 66 Australian adolescents on 2 occasions approximately 1 month apart. Features extracted by computer image analysis included area, perimeter, and regularity of outline of the lesions, the mean and standard deviation of reflectance at red, green, and blue wavelengths, and the mean and standard deviation of the gradients of red, green, and blue reflectance at the lesion boundary. RESULTS: All measurements showed moderate to high reliability (intraclass correlation coefficients 0.66-0.94), except for the standard deviations of the color gradients, whose reliability improved to moderate levels (0.68-0.71) when the mean of 5 lesions was considered. For most outcomes, reasonable within subject reliability was achieved when five lesions per subject were measured. CONCLUSIONS: These results, in combination with previous work demonstrating the reasonable ability of this computer imaging system to discriminate between malignant melanomas and other pigmented lesions, indicates that the system has the potential to become a useful tool for clinicians in following people with pigmented lesions over time to detect early malignant changes.

Adolescent↗

Multiple genetic diagnoses from single cells using multiplex PCR: reliability and allele dropout.

We used a multiplex fluorescent PCR system containing seven primer sets on single cells from three different cell types (buccal, corneal and blastomere cells) and more than 3500 heterozygous alleles to investigate reliability and extent of allele dropout in multiplex PCRs at the single cell level. All three cell types gave similarly high reliability, accuracy and allele dropout rates, with similar reliability between singleplex and multiplex PCRs. Allele dropout was also consistent between the three cell types and did not significantly increase as allele size increased. These results indicate that multiplex fluorescent PCR is a reliable and accurate method of obtaining multiple diagnosis (eight chromosomes simultaneously) from single cells and maximizes the information available from single cell analysis.

Alleles↗

A matrix of kappa-type coefficients to assess the reliability of nominal scales.

To investigate the reliability of nominal scales, Kraemer proposed a measurement model from which kappa coefficients could be derived. More recently she suggested a matrix of coefficients as a comprehensive summary of reliability, contrasting this with use of a single summary kappa statistic. The main diagonal of the matrix consists of binary kappa coefficients for each category which measure the reliability of each category relative to all others, while the off-diagonal elements are correlation coefficients for pairs of categories. The off-diagonal elements were suggested as measures of confusion between categories. Schouten also suggested coefficients to assess confusion between pairs of categories, which might be used as alternative off-diagonal elements in a summary matrix. The two types of off-diagonal element will be compared. It will be shown that Schouten's coefficients can be expressed in terms of the parameters of Kraemer's measurement model and that they are more easily interpreted as measures of confusion. First, the maximum value for Schouten's coefficient is one. Secondly, for any pair of categories, Schouten's coefficient equals the proportionate reduction in the probability of classifying a subject in one category of the pair having previously classified them in the other. Thirdly, where the coefficient for a pair of categories is less than the summary kappa statistic, it will be shown that combining these two categories will increase the value of the summary kappa statistic. The methods of analysis are applied to data from a study of the reliability of psychiatric diagnosis and used to identify pairs of classifications between which there is substantial confusion.

Humans↗

Reliability and base rates of interpersonal themes in narratives from psychotherapy sessions.

We present an analysis of the reliability and base rates of the interpersonal contents of narratives told by patients in psychotherapy. Trained judges rated two samples, including 60 opiate-dependent patients in cognitive or psychodynamic therapy and 72 depressed patients in cognitive or interpersonal therapy. Using a comprehensive system based upon a circumplex model and involving 104 separate categories, we found that most categories of interpersonal behavior could be rated reliably. Potential problem categories were identified and strategies for increasing reliability are discussed. In particular, categories related to the concept of the introject (what the self does to the self) had low reliability. An analysis of the base rates of interpersonal themes revealed that issues related to autonomy/ assertion were most prevalent, although some differences between the two samples were evident. The implications of the results for research on narratives and models of psychotherapy are discussed.

Adult↗

Interrater reliability of a Danish version of the Morgan Russell scale for assessment of anorexia nervosa.

OBJECTIVE: To study the interrater reliability of a Danish version of the Morgan Russell scale for assessment of patients with anorexia nervosa, and subsequently to clarify the existing rating instructions. METHOD: Ten patients undergoing treatment for anorexia nervosa at a regional center participated and had their interview videotaped. Two interviews were reserved for a training phase only. The group of raters comprised eight clinicians, and measures of interrater reliability were computed using intraclass correlation coefficient (ICC). RESULTS: The ICC for the total score was good (0.79), while reliability for the single items varied from poor to excellent (0.14-0.99). Internal consistency as expressed by Cronbach's coefficient alpha was acceptable (0.74). DISCUSSION: The Morgan Russell scale stands out as an easily applied and reliable measure of severity of anorexia nervosa, though the rating instructions need clarification in some items.

Adolescent↗

Interrater reliability of the short-memory questionnaire in a variety of health professional representatives.

A study was performed to assess interrater reliability of the Japanese version of the Short-Memory Questionnaire (SMQ), which is an easily-administered, informant-based scale of cognitive function. The subjects were 18 consecutive patients with Alzheimer's disease who were outpatients of Department of Neuropsychiatry in Ehime University School of Medicine and their principal caregivers. One neuropsychiatrist (NP) administered the SMQ, and all sessions were videotaped. Then one nurse (Ns), one clinical psychologist (CP), one occupational therapist (OT), and one neurologist (NL) from another institution viewed the videotape and performed reassessments independently. Interrater reliability between the NP and Ns, CP, OT, or NL were all extremely good. Interrater reliability between the Ns and CP, between the Ns and OT, between the Ns and NL, between the CP and OT, between the CP and NL, and between the OT and NL were also extremely good. The SMQ is a convenient, quantitative scale, and in this study it showed good interrater reliability between personnel from different fields. Therefore, it is a very useful test for everyday medical consultations and for clinical research.

Aged↗