Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Improved reliability of the Standardized Alzheimer's Disease Assessment Scale (SADAS) compared with the Alzheimer's Disease Assessment Scale (ADAS).

OBJECTIVES: To compare the interrater and intrarater reliability of the Alzheimer's Disease Assessment Scale (ADAS) with the Standardized Alzheimer's Disease Assessment Scale (SADAS). DESIGN: A randomized, double blind trial. Sixteen university students were randomized to administer either version of the instrument. Subjects were randomized to three assessments, at 2-week intervals, using the ADAS or the SADAS. Each subject's first and third tests were administered by the same rater, the second by a different rater. SETTING: A geriatric outpatient clinic in a university teaching hospital. PARTICIPANTS: Fifty-four patients with possible or probable Alzheimer's disease living in the community or in a long-term care facility. MEASUREMENTS: The primary outcome was the interrater reliability of total ADAS and SADAS scores. Secondary outcomes were ADAS and SADAS cognitive scores, noncognitive scores, duration of testing, and sample size estimates. RESULTS: The interrater reliability of the SADAS total score was significantly better than that of the ADAS (interrater ICC 0.93 SADAS vs 0.83 ADAS), and the interrater standard deviation of the total SADAS score was lower than that of the ADAS (38%, P < .05). The SADAS cognitive subscale inter and intrarater reliability, although higher than the ADAS, was not significantly different when used by different raters (interrater ICC 0.91 SADAS vs 0.90 ADAS; intrarater ICC 0.88 SADAS vs 0.86 ADAS). The SADAS noncognitive subscale was significantly more reliable than the ADAS (interrater ICC 0.89 SADAS vs 0.42 ADAS; intrarater ICC 0.87 SADAS vs 0.70 ADAS; P < or = .05) and had a lower standard deviation between raters (59%; P < .01) and within raters (40%; P < .05) compared with the ADAS. CONCLUSION: The improved reliability of the SADAS total score means that investigators can now use this score as a primary outcome measure, and important behavioral symptomatology can be included as a marker for treatment efficacy in AD. The smaller standard deviation of the SADAS means that clinical trials using the SADAS as a primary outcome will demonstrate differences, if present, with smaller sample sizes than with the ADAS.

Aged↗

Reliability of the visual analog scale for measurement of acute pain.

OBJECTIVE: Reliable and valid measures of pain are needed to advance research initiatives on appropriate and effective use of analgesia in the emergency department (ED). The reliability of visual analog scale (VAS) scores has not been demonstrated in the acute setting where pain fluctuation might be greater than for chronic pain. The objective of the study was to assess the reliability of the VAS for measurement of acute pain. METHODS: This was a prospective convenience sample of adults with acute pain presenting to two EDs. Intraclass correlation coefficients (ICCs) with 95% confidence intervals (95% CIs) and a Bland-Altman analysis were used to assess reliability of paired VAS measurements obtained 1 minute apart every 30 minutes over two hours. RESULTS: The summary ICC for all paired VAS scores was 0.97 [95% CI = 0.96 to 0.98]. The Bland-Altman analysis showed that 50% of the paired measurements were within 2 mm of one another, 90% were within 9 mm, and 95% were within 16 mm. The paired measurements were more reproducible at the extremes of pain intensity than at moderate levels of pain. CONCLUSIONS: Reliability of the VAS for acute pain measurement as assessed by the ICC appears to be high. Ninety percent of the pain ratings were reproducible within 9 mm. These data suggest that the VAS is sufficiently reliable to be used to assess acute pain.

Acute Disease↗

Interobserver and intraobserver reliability of the Rome II criteria in children.

BACKGROUND: Functional gastrointestinal disorders are common in children. It has been suggested that the diagnosis of these conditions should be based on the "pediatric Rome II" criteria. The interobserver reliability for the DSM-IV, another symptom-based criteria is considered almost perfect in multiple studies. There are no studies assessing the reliability of the Rome II criteria in children. OBJECTIVES: To evaluate the reliability of the pediatric Rome II criteria. METHODS: Interobserver reliability-Ten pediatric gastroenterologists and 10 fellows in pediatric gastroenterology were provided with 20 clinical vignettes, the Rome II criteria, and a list of 15 possible diagnoses. Each of the raters was instructed to select one or more diagnoses for each vignette. Intraobserver reliability-The specialists were provided with the same set of vignettes 4 months later. RESULTS: Average percentage of agreement coefficient: 45% (specialists), 47% (fellows). In order to correct for possible agreement by chance, we calculated the kappa coefficient, a measure of pairwise agreement corrected for chance. Specialists: k = 0.37 (p < 0.0001), trainees: k = 0.41, (p < 0.0001). Physicians with a special interest in functional gastrointestinal disorders (k = 0.37, p < 0.0001), and other specialists (k = 0.38, p < 0.0001). Analysis of data in pain and constipation diagnosis subgroups revealed even lower kappa (constipation: k = 0.2, p < 0.0001; pain: k = 0.3, p < 0.0001). Intraobserver agreement: k = 0.63 (p < 0.0001). CONCLUSION: The interobserver reliability of the Rome II criteria among pediatric gastroenterologists and fellows is low. Further validation of the criteria is necessary.

Child↗

Interrater reliability of the Structured Clinical Interview for the Positive and Negative Syndrome Scale for schizophrenia.

The Swedish version of the Positive and Negative Syndrome Scale for schizophrenia (PANSS) has been tested and construct validity, internal reliability and interrater reliability have been demonstrated to be quite satisfactory. However, the interrater reliability of the negative symptoms was unsatisfactory low. In this study, the Swedish version of the Structured Clinical Interview for the PANSS has been tested. The interrater reliability is increased as compared with the inter-rater reliability achieved by means of the PANSS. As concerns the positive scale, the intraclass coefficients increased to 0.98-0.99 with the SCID-PANSS. For the negative scale, the intraclass coefficients increased to 0.83-0.90 with the SCID-PANSS, and for the general scale the increase was to 0.95-0.98 with the SCID-PANSS. It was also demonstrated that the interrater reliability is higher for the positive and the negative factors derived from the PANSS than for the positive and the negative scales.

Anxiety Disorders↗

Reliability of linear alveolar bone loss measurements of mandibular posterior teeth from digitized bitewing radiographs.

Observer reliability in performing linear measurements between the cementoenamel junction and alveolar crest was determined for mandibular posterior teeth from digitized clinical bitewing radiographs acquired during recall examinations. 6 measurements (corresponding to traditional probing measurements) were made per tooth by 3 observers. Mesial and distal measurements made to the most coronal aspects of the alveolar crest were the most reliable and least biased. As was anticipated, intra-observer reliability was better than inter-observer reliability although the 3 observers of our study were able to detect a significant mean change (0.1 mm, p<0.0001) in alveolar bone height over a 1-year period for 10 patients. For our most reliable and unbiased measurements (mesial measurements to the alveolar crest), a change of 0.54 mm (90th percentile) would be required to indicate change at a site from one time to the next. Based on the reliability of our digital radiographic measurements, with the alpha error rate set at 0.05 and beta at 0.20, a difference in alveolar bone height of 0.3 mm could be detected with a patient sample size of between 13 (best case) and 54 (worst case).

Alveolar Bone Loss↗

Reliability of anthropometric measurements in the WHO Multicentre Growth Reference Study.

AIM: To describe how reliability assessment data in the WHO Multicentre Growth Reference Study (MGRS) were collected and analysed, and to present the results thereof. METHODS: There were two sources of anthropometric data (length, head and arm circumferences, triceps and subscapular skinfolds, and height) for these analyses. Data for constructing the WHO Child Growth Standards, collected in duplicate by observer pairs, were used to calculate inter-observer technical error of measurement (TEM) and the coefficient of reliability. The second source was the anthropometry standardization sessions conducted throughout the data collection period with the aim of identifying and correcting measurement problems. An anthropometry expert visited each site annually to participate in standardization sessions and provide remedial training as required. Inter- and intra-observer TEM, and average bias relative to the expert, were calculated for the standardization data. RESULTS: TEM estimates for teams compared well with the anthropometry expert. Overall, average bias was within acceptable limits of deviation from the expert, with head circumference having both lowest bias and lowest TEM. Teams tended to underestimate length, height and arm circumference, and to overestimate skinfold measurements. This was likely due to difficulties associated with keeping children fully stretched out and still for length/height measurements and in manipulating soft tissues for the other measurements. Intra- and inter-observer TEMs were comparable, and newborns, infants and older children were measured with equal reliability. The coefficient of reliability was above 95% for all measurements except skinfolds whose R coefficient was 75-93%. CONCLUSION: Reliability of the MGRS teams compared well with the study's anthropometry expert and published reliability statistics.

Anthropometry↗

The validity of reliability assessments.

This paper focuses on reliability and evaluation of health education programs in school settings. Reliability is a concept that guides researchers in selecting or developing instruments, and is used as a standard, with validity and acceptability, for judging the credibility of research findings and inferences. Reliability is defined within the context of research design, and methods for estimating the reliability of cognitive measures are reviewed. Using data gathered in a school health education curriculum evaluation as an example, possible errors in hypotheses testing that may occur when estimating internal consistency of cognitive test scores obtained in quasi-experimental designs are examined. The appropriateness of internal consistency as a measure of reliability of cognitive measures is discussed and suggestions for reliability assessment and related issues such as power analysis are presented.

Attitude to Health↗

Instrument validity and reliability in three health education journals, 1980-1987.

Investigators examined how often validity and reliability measures were reported for research articles in three health education journals: Health Education, Health Education Quarterly, and the Journal of School Health. Articles published from 1980 to 1987 were considered in the analysis. Of the 611 articles published by Health Education during the period used for analysis, 128 (21%) met the criteria of a research article. Reliability was reported for 22 (17%) articles, and validity was reported for 78 (61%) articles. Health Education Quarterly published 212 articles; 74 (35%) were research articles. Reliability was reported for 16 (21%) articles and validity was reported for 40 (54%) articles. The Journal of School Health published 778 articles, of which 243 (31%) were research articles. Reliability was reported for 62 (25%), and validity was reported for 164 (67%) of the research articles. A chi-square test found a significant difference among the number of research articles published by the journals. Chi-square tests also found significant differences among the journals in the proportion of research articles that reported reliability information and the proportion that reported validity. A significant trend was noted for Health Education Quarterly and the Journal of School Health; the proportion of research articles that reported validity and reliability increased over time for both publications.

Health Education↗

The oral health assessment tool--validity and reliability.

BACKGROUND: The Oral Health Assessment Tool (OHAT) was a component of the Best Practice Oral Health Model for Australian Residential Care study. The OHAT provided institutional carers with a simple, eight category screening tool to assess residents' oral health, including those with dementia. This analysis presents OHAT reliability and validity results. METHODS: A convenience sample of 21 residential care facilities (RCFs) in urban and rural Victoria, NSW and South Australia used the OHAT at baseline, three-months and six-months to assess intra- and inter-carer reliability and concurrent validity. RESULTS: Four hundred and fifty five residents completed all study phases. Intra-carer reliability for OHAT categories: percent agreement ranged from 74.4 per cent for oral cleanliness, to 93.9 per cent for dental pain; Kappa statistics were in moderate range (0.51-0.60) for lips, saliva, oral cleanliness, and for all other categories in range of 0.61-0.80 (substantial agreement) (p < 0.05). Inter-carer reliability for OHAT categories: percent agreement ranged from 72.6 per cent for oral cleanliness to 92.6 per cent for dental pain; Kappa statistics were in moderate range (0.48-0.60) for lips, tongue, gums, saliva, oral cleanliness, and for all other categories in range of 0.61-0.80 (substantial agreement) (p < 0.05). Intraclass correlation coefficients for OHAT total scores were 0.78 for intra-carer and 0.74 for inter-carer reliability. Validity analyses of the OHAT categories and examination findings showed complete agreement for the lips category, with the natural teeth, dentures, and tongue categories having high significant correlations and percent agreements. The gums category had significant moderate correlation and percent agreement. Non-significant and low correlations and percent agreements were evident for the saliva, oral cleanliness and dental pain categories. CONCLUSION: The Oral Health Assessment Tool was evaluated as being a reliable and valid screening assessment tool for use in residential care facilities, including those with cognitively impaired residents.

Aged↗

How reliably do rheumatologists measure shoulder movement?

OBJECTIVE: To assess the intrarater and interrater reliability among rheumatologists of a standardised protocol for measurement of shoulder movements using a gravity inclinometer. METHODS: After instruction, six rheumatologists independently assessed eight movements of the shoulder, including total and glenohumeral flexion, total and glenohumeral abduction, external rotation in neutral and in abduction, internal rotation in abduction and hand behind back, in random order in six patients with shoulder pain and stiffness according to a 6x6 Latin square design using a standardised protocol. These assessments were then repeated. Analysis of variance was used to partition total variability into components of variance in order to calculate intraclass correlation coefficients (ICCs). RESULTS: The intrarater and interrater reliability of different shoulder movements varied widely. The movement of hand behind back and total shoulder flexion yielded the highest ICC scores for both intrarater reliability (0.91 and 0.83, respectively) and interrater reliability (0.80 and 0.72, respectively). Low ICC scores were found for the movements of glenohumeral abduction, external rotation in abduction, and internal rotation in abduction (intrarater ICCs 0.35, 0.43, and 0.32, respectively), and external rotation in neutral, external rotation in abduction, and internal rotation in abduction (interrater ICCs 0.29, 0.11, and 0.06, respectively). CONCLUSIONS: The measurement of shoulder movements using a standardised protocol by rheumatologists produced variable intrarater and interrater reliability. Reasonable reliability was obtained only for the movement of hand behind back and total shoulder flexion.

Aged↗

Precision and reliability for measurement of change in MRI lesion volume in multiple sclerosis: a comparison of two computer assisted techniques.

OBJECTIVE: The serial quantification of MRI lesion load in multiple sclerosis provides an effective tool for monitoring disease progression and this has led to its increasing use as an outcome measure in treatment trials. Segmentation techniques must display a high degree of precision and reliability if they are to be responsive to small changes over time. This study has evaluated the performance of two such techniques, the manual outlining and contour methods, in serial lesion load quantification. METHODS: Sixteen patients with clinically definite multiple sclerosis were scanned at baseline and after two years. Scan analysis was performed twice, independently by three observers using each technique. RESULTS: For the absolute lesion volumes the median intrarater coefficient of variation (CV) was 3.2% for the contour technique and 7.6% for the manual outlining method (p < 0.005), the interrater CVs were 3.8% and 6.1% respectively (p < 0.01) and the reliability of both techniques was very high. For the change in lesion volume the intrarater and interrater repeatability coefficients were respectively 2.6 cm3 and 2.8 cm3 for the contour technique, and 3.3 cm3 and 3.7 cm3 for the manual outlining method (lower values reflect higher precision). The values for intrarater and interrater reliability for measuring change in lesion volume were respectively, 0.945 and 0.944 for the contour technique, and 0.939 and 0.921 for the manual outline method (perfect reliability = 1.0). CONCLUSIONS: With such high values for reliability, the impact of measurement error in lesion segmentation on sample size requirements in multiple sclerosis treatment trials is minor. This study shows that a change in lesion volume can be measured with a higher level of precision and reliability with the contour technique and this supports its further application in serial studies.

Adjuvants, Immunologic↗

Inter- and intra-rater reliability for classification of medication related events in paediatric inpatients.

BACKGROUND: In medication safety research studies medication related events are often classified by type, seriousness, and degree of preventability, but there is currently no universally reliable "gold standard" approach. The reliability (reproducibility) of this process is important as the targeting of prevention strategies is often based on specific categories of event. The aim of this study was to determine the reliability of reviewer judgements regarding classification of paediatric inpatient medication related events. METHODS: Three health professionals independently reviewed suspected medication related events and classified them by type (adverse drug event (ADE), potential ADE, medication error, rule violation, or other event). ADEs and potential ADEs were then rated according to seriousness of patient injury using a seven point scale and preventability using a decision algorithm and a six point scale. Inter- and intra-rater reliabilities were calculated using the kappa (kappa) statistic. RESULTS: Agreement between all three reviewers regarding event type ranged from "slight" for potential ADEs (kappa = 0.20, 95% CI 0.00 to 0.40) to "substantial" agreement for the presence of an ADE (kappa = 0.73, 95% CI 0.69 to 0.77). Agreement ranged from "slight" (kappa = 0.06, 95% CI 0.02 to 0.10) to "fair" (kappa = 0.34, 95% CI 0.30 to 0.38) for seriousness classifications but, by collapsing the seven categories into serious versus not serious, "moderate" agreement was found (kappa = 0.50, 95% CI 0.46 to 0.54). For preventability decision, overall agreement was "fair" (kappa = 0.37, 95% CI 0.33 to 0.41) but "moderate" for not preventable events (kappa = 0.47, 95% CI 0.43 to 0.51). CONCLUSION: Trained reviewers can reliably assess paediatric inpatient medication related events for the presence of an ADE and for its seriousness. Assessments of preventability appeared to be a more difficult judgement in children and approaches that improve reliability would be useful.

Adverse Drug Reaction Reporting Systems↗

Reliability of seizure diaries in adult epileptic patients.

Daily diaries are used widely in neurologic research and clinical practice to assess alterations in seizure frequency among patients with epilepsy. However, no formal tests of the reliability of this data collection method have been performed. We investigated the reliability of seizure recall in adult patients participating in a longitudinal study of stress, mood and seizure frequency. Patients maintained daily diaries for 10-36 weeks. The reliability study entailed completion of a single additional diary on the evening of a randomly selected day with reference to the preceding day. This design produced two diaries, completed 1 day apart, for the same 24-hour period. Measuring reliability with the Pearson correlation coefficient, overall reliability of seizure recall was 0.95 and was not markedly influenced by the subjects' sociodemographic characteristics, neurological or psychological status. In sum, the assumption in the literature that the daily diary is a reliable method for securing data on seizure counts appears warranted.

Adolescent↗

Reliability of assessing percentage of diffusion-perfusion mismatch.

BACKGROUND AND PURPOSE: Emergent neurovascular imaging holds promise in identifying new and optimum target populations for thrombolysis in stroke. Recent research has focused on patients with diffusion-weighted MRI (DWI)-perfusion-weighted MRI (PWI) mismatch as a marker of tissue at risk of infarction and a means to select the most suitable candidates for thrombolysis. The present study sought to estimate the reliability of assessing the percentage of DWI-PWI mismatch. METHODS: Thirteen patients with acute strokes had DWI and PWI within 7 hours of symptom onset. Six raters independently created relative mean transit time (rMTT) maps and then compared them with DWI images to assess the percentage of mismatch (PWI>DWI) in 10% increments. The MR scans were reassessed by 4 raters, tracing around the lesions to calculate the volume percentage of mismatch. RESULTS: Visual assessment had an interrater reliability of 0.68 (95% CI, 0.52 to 1.0; SEM=21.6%) and an intrarater reliability of 0.80 (95% CI, 0.47 to 1.0; SEM=16.9%). Hand-drawn assessment had an interrater reliability of 0.66 (95% CI, 0.45 to 1.0; SEM=26.2%) and an intrarater reliability of 0.94 (95% CI, 0.81 to 1.0; SEM=10.9%). CONCLUSIONS: Results from the present study suggest that quantifying mismatch by the human eye is reproducible but not reliable among observers. This raises doubts about using mismatch for clinical decision making and clinical trial enrollment.

Acute Disease↗

Improving the reliability of stroke subgroup classification using the Trial of ORG 10172 in Acute Stroke Treatment (TOAST) criteria.

BACKGROUND AND PURPOSE: We sought to improve the reliability of the Trial of ORG 10172 in Acute Stroke Treatment (TOAST) classification of stroke subtype for retrospective use in clinical, health services, and quality of care outcome studies. The TOAST investigators devised a series of 11 definitions to classify patients with ischemic stroke into 5 major etiologic/pathophysiological groupings. Interrater agreement was reported to be substantial in a series of patients who were independently assessed by pairs of physicians. However, the investigators cautioned that disagreements in subtype assignment remain despite the use of these explicit criteria and that trials should include measures to ensure the most uniform diagnosis possible. METHODS: In preparation for a study of outcomes and management practices for patients with ischemic stroke within Department of Veterans Affairs hospitals, 2 neurologists and 2 internists first retrospectively classified a series of 14 randomly selected stroke patients on the basis of the TOAST definitions to provide a baseline assessment of interrater agreement. A 2-phase process was then used to improve the reliability of subtype assignment. In the first phase, a computerized algorithm was developed to assign the TOAST diagnostic category. The reliability of the computerized algorithm was tested with a series of synthetic cases designed to provide data fitting each of the 11 definitions. In the second phase, critical disagreements in the data abstraction process were identified and remaining variability was reduced by the development of standardized procedures for retrieving relevant information from the medical record. RESULTS: The 4 physicians agreed in subtype diagnosis for only 2 of the 14 baseline cases (14%) using all 11 TOAST definitions and for 4 of the 14 cases (29%) when the classifications were collapsed into the 5 major etiologic/pathophysiological groupings (kappa=0.42; 95% CI, 0.32 to 0.53). There was 100% agreement between classifications generated by the computerized algorithm and the intended diagnostic groups for the 11 synthetic cases. The algorithm was then applied to the original 14 cases, and the diagnostic categorization was compared with each of the 4 physicians' baseline assignments. For the 5 collapsed subtypes, the algorithm-based and physician-assigned diagnoses disagreed for 29% to 50% of the cases, reflecting variation in the abstracted data and/or its interpretation. The use of an operations manual designed to guide data abstraction improved the reliability subtype assignment (kappa=0.54; 95% CI, 0.26 to 0.82). Critical disagreements in the abstracted data were identified, and the manual was revised accordingly. Reliability with the use of the 5 collapsed groupings then improved for both interrater (kappa=0.68; 95% CI, 0.44 to 0.91) and intrarater (kappa=0.74; 95% CI, 0.61 to 0.87) agreement. Examining each remaining disagreement revealed that half were due to ambiguities in the medical record and half were related to otherwise unexplained errors in data abstraction. CONCLUSIONS: Ischemic stroke subtype based on published TOAST classification criteria can be reliably assigned with the use of a computerized algorithm with data obtained through standardized medical record abstraction procedures. Some variability in stroke subtype classification will remain because of inconsistencies in the medical record and errors in data abstraction. This residual variability can be addressed by having 2 raters classify each case and then identifying and resolving the reason(s) for the disagreement.

Acute Disease↗

A multisite investigation of the reliability of the Scale for the Assessment of Negative Symptoms.

OBJECTIVE: The Scale for the Assessment of Negative Symptoms is a widely used instrument for measuring negative symptoms in schizophrenia, but few studies have examined its reliability. This study examined the interrater, internal, and test-retest reliabilities of the scale and its factor structure in the context of a multisite study. METHOD: Two hundred seven patients with schizophrenia who were participating in the Treatment Strategies in Schizophrenia study were assessed with the Scale for the Assessment of Negative Symptoms following a symptom exacerbation and again 3-6 months later. All assessments were performed by trained psychiatrists who were treating the patients. RESULTS: Interrater reliabilities ranged from low to high for the items on the Scale for the Assessment of Negative Symptoms but were statistically significant in most cases. Most correlations between individual items and subscale total scores were moderate to high, as were coefficient alphas for each subscale, indicating adequate internal consistency. Test-retest correlations were of moderate magnitude. Few differences in reliability statistics between sites were found, although differences in mean scale ratings between sites were present. A factor analysis indicated three factors corresponding to the Affective Flattening or Blunting subscale, the Avolition-Apathy and Anhedonia-Asociality subscales, and the Alogia and Inattention subscales. CONCLUSIONS: The results suggest that the Scale for the Assessment of Negative Symptoms has good reliability and is a useful instrument for the measurement of negative symptoms in multisite clinical studies. The internal reliability of the Alogia, Avolition-Apathy, and Inattention subscales could be improved by replacing some items and including additional items.

Adult↗

Reliability of diagnostic reporting for children aged 6-11 years: a test-retest study of the Diagnostic Interview Schedule for Children-Revised.

OBJECTIVE: This study examined the reliability of symptom reporting by community children of elementary school age and their parents on a version of the Diagnostic Interview Schedule for Children-Revised (DISC-R). METHOD: A sample of 109 children aged 6-11 years from an ongoing epidemiologic study were recruited for retest DISC-R interviews after completion of the study protocol. Retest interviews took place 7-18 days after the first interview and were conducted by interviewers who had no prior information about the subjects. Test-retest reliability for five common childhood psychiatric diagnoses was evaluated with the kappa statistic; the intraclass correlation coefficient was used to evaluate test-retest reliability of symptom scales. RESULTS: The reliability of the parents' reports on the DISC-R was good to excellent for attention deficit hyperactivity disorder and separation anxiety disorder; it was fair for overanxious disorder, oppositional defiant disorder, and conduct disorder. The children reported many fewer symptoms than the parents except for separation anxiety disorder; reliability was fair for separation anxiety disorder and poor for attention deficit hyperactivity disorder. The children were particularly unreliable in reporting about time factors, such as duration and onset of symptoms. When symptoms were considered without duration and onset, children's reports reached fair reliability for separation anxiety disorder, overanxious disorder, and attention deficit hyperactivity disorder but remained poor for oppositional defiant disorder. CONCLUSIONS: The results suggest that highly structured diagnostic interviews such as the DISC-R may not be appropriate for use with younger children of elementary school age in community-based studies.

Adolescent↗

Psychiatric Research Interview for Substance and Mental Disorders (PRISM): reliability for substance abusers.

OBJECTIVE: The purpose of this study was to investigate the reliability of a new semistructured diagnostic interview, the Psychiatric Research Interview for Substance and Mental Disorders (PRISM), for substance-abusing patients. The reliability of psychiatric diagnoses for individuals who drink heavily or use drugs has been shown to be problematic. The PRISM was designed to improve the reliability for such individuals. METHOD: A test-retest reliability study of the PRISM was conducted with 172 patients being treated in dual-diagnosis or substance abuse settings. RESULTS: Good to excellent reliability was shown for many diagnoses, including affective disorders, substance use disorders, eating disorders, some anxiety disorders, and psychotic symptoms. The interview has recently been updated for DSM-IV diagnoses. CONCLUSIONS: The PRISM offers a method of producing psychiatric diagnoses with improved reliability for patients and other research subjects who have problems with alcohol or drugs.

Adult↗