Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Reliability”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

Intra and inter-observer reliability of determining degree of pelvic incidence in high-grade spondylolisthesis using a computer assisted method.

Pelvic incidence was described as a fundamental parameter to describe spino-pelvic balance. In high-grade spondylolisthesis, severe dystrophic changes of the upper sacral endplate may be responsible for technical difficulties in pelvic incidence measurement. We propose to evaluate the reliability of PI measurement in high-grade spondylolisthesis patients and to compare the manual method with a computer-assisted method. In 30 high-grade spondylolisthesis patients, pelvic incidence was measured by manual and computer-assisted technique by the Spineview software package. We statistically assessed agreement between the manual and the computer-assisted technique, the intra-observer and the inter-observer reliability of the computer-assisted technique. Significant correlation was found (Spearman's rank R = 0.921 with P<0.001) between manual and computer-assisted results. The paired t test (t = 0.979 with P<0.001) and the intraclass correlation coefficient (ICC) were also significant. Intra- and inter-observer reliability of the computer-assisted technique were excellent with Spearman's rank correlation from 0.964 to 0.985 with P<0.001, a paired t test from 0.978 to 0.983 with P<0.001) and an ICC from 0.986 to 0.992. Intra- and inter-observer repeatability were better with the computer-assisted method than with the manual technique. We proved the reliability and repetability of a computer-assisted angular measurement method in high-grade spondylolisthesis patients. This validated measurement technique could be now used to measure the main parameters of the sagittal balance of the spine in further studies on spondylolisthesis patients.

Humans↗

The Camino intracranial pressure device in clinical practice: reliability, handling characteristics and complications.

Intracranial pressure monitoring has a key role in the management of patients developing increased intracranial pressure (ICP). We adopted the Camino fiberoptic system for intracranial pressure measurement in 1993 in our neurosurgical department. The aim of this study was to investigate reliability, handling characteristics and complication rate of the Camino intracranial pressure device. In an eighteen month period, we prospectively investigated 118 patients with intracranial pathology undergoing Camino fiberoptic intraparenchymal or intraventricular ICP monitoring. The assessment of reliability of ICP monitoring according to patients clinical condition, to cranial computed tomography (CCT) findings and ICP waveform was carried out. Position of the probe and intracranial bleeding complications related to probe insertion were confirmed by CCT. Technical complications, as well as infections due to the device, were documented. In vivo recalibration was performed in 22 patients. At the end of the measuring period the drift of the probe was evaluated and the accuracy of the fiberoptic device was measured by performing a two point calibration. Recordings of intracranial pressure were carried out with 136 Camino devices (104 parenchymal, 32 ventricular) in 118 patients with an average measuring time of 94.1 +/- 79.1 hrs. One hundred and fifteen Camino intracranial pressure devices (85.2%) demonstrated reliability according to the predetermined clinical parameters. The actual mean drift after removal of the devices was 3.4 mmHg +/- 3.2 with an actual daily drift of 3.2 +/- 17.2 mmHg. Recorded complications included infection (0.7%), intraparenchymal haematoma (5.1%), and a high complication rate (23.5%) with regard to technical aspects. The Camino intracranial pressure system offers reliable ICP measurements in an acceptable percentage of devices, and the advantage of in vivo recalibration. The high incidence of technical complications identifies a need for improvement in the fiberoptic cable and the fixation system.

Adolescent↗

The Children's Global Assessment Scale (CGAS) and Global Assessment of Psychosocial Disability (GAPD) in clinical practice--substance and reliability as judged by intraclass correlations.

Studies on the inter-rater reliability on the Children's Global Assessment Scale (CGAS) and the Global Assessment of Psychosocial Disability (GAPD) involving different subgroups of 145 outpatients from 4 to 16 years of age showed fair to substantial intraclass correlations of 0.59 to 0.90. Raters of different training levels participated. Interrater reliability was dependent on number of ratings per rater, training, available data sources and experience. A more detailed description of anchor points resulted in higher inter-rater agreement by psychiatrists training in child and adolescent psychiatry, but did not influence the inter-rater reliability among more (widely) experienced raters. Both the CGAS and the GAPD seem to be sufficiently reliable tools in clinical practice. The CGAS seems to be more sensitive to inter-rater variation than the GAPD.

Adolescent↗

Adaptation of the Bath Ankylosing Spondylitis Functional Index to the Turkish population, its reliability and validity: functional assessment in AS.

The aim of this study was to adapt the Bath Ankylosing Spondylitis Functional Index (BASFI) to the Turkish population and investigate the reliability and the validity of the Turkish version. Seventy-six patients with ankylosing spondylitis (AS) were included in the study. The functional status of the patients was assessed by using the adapted Turkish version of the BASFI twice, at recruitment and 24 h later. For validity analysis, patients were also assessed by the Bath Ankylosing Spondylitis Disease Activity Index (BASDAI) evaluating disease activity, the Bath Ankylosing Spondylitis Global Score (BAS-G) indicating effect of the disease on patient's well-being, physician's assessment of the disease activity and pain intensity. Spinal mobility was assessed by the Bath Ankylosing Spondylitis Metrology Index (BASMI). Erythrocyte sedimentation rate (ESR) and serum C-reactive protein (CRP) levels of the patients were also recorded. The lumbar region and the sacroiliac joints were assessed by Stoke Ankylosing Spondylitis Spine Score (SASSS) and the hip joints were assessed by Bath Ankylosing Spondylitis Radiology Index hip (BASRI-h). The internal consistency was 0.89 (Cronbach's alpha), which showed a high reliability for the Turkish version of the BASFI. Test-retest reliability was good, with a high intraclass correlation coefficient between the two time points (ICC=0.93). Significant correlations were detected between the BASFI and the BASDAI, BAS-G, doctor's global assessment, and general pain intensity (r=0.62, p<0.001; r=0.47, p<0.001; r=0.55, p<0.001; r=0.47, p<0.001, respectively). The adaptation of the BASFI to the Turkish population was successful and it was found to be reliable and valid among Turkish patients. Thus, studies using the Turkish BASFI can be compared with international studies.

Adult↗

The Turkish version of the Bath Ankylosing Spondylitis Functional Index: reliability and validity.

The purpose of this study was to investigate the reliability and validity of the Turkish version of the Bath Ankylosing Spondylitis (AS) Functional Index (BASFI). The Turkish version of the BASFI was obtained after a process of translation and back-translation. Eighty-one consecutive patients meeting the 1984 New York criteria for AS were enrolled. Patients were evaluated and requested to complete the questionnaire at days 1 and 2 and on a third occasion between days 15-90. Reliability, reproducibility, validity and sensitivity to change of the Turkish version of the index were assessed. Each score correlated closely with the index score, with coefficients between 0.727 and 0.844. Reliability analysis showed a Cronbach's alpha score of 0.926. Correlations were found between all items of the BASFI and Schober's test (r=-0.258 to -0.531, p<0.001-0.05), occiput-to-wall distance (r=0.284 and 0.589, p<0.001-0.05), and finger-to-floor distance (r=0.334 to 0.613, p<0.001-0.01). The total index score was correlated with the number of nocturnal awakenings (r=0.515, p<0.001), Schober's test (r=-0.444, p<0.001), finger-to-floor distance (r=0.567, p<0.001), occiput-to-wall distance (r=0.535, p<0.001), chest expansion (r=-0.403, p<0.001), and the Dougados articular index (r=0.371, p<0.01). A good correlation was found between day 0 and 1 BASFI indices (r=0.765-0.917, p<0.001), showing good reproducibility of the index. The Turkish version of the BASFI showed reliability, reproducibility, and validity, confirming its utility in the research of AS in Turkey. However, sensitivity to changes due to drug therapy and/or rehabilitation remains to be determined.

Activities of Daily Living↗

Validity and reliability of the Modified Manchester Health Questionnaire in assessing patients with fecal incontinence.

PURPOSE: To date, no measures of fecal incontinence severity or its impact on quality of life have been validated for telephone interview. This study was designed to 1) compare responses of a self-administered and a telephone-administered Fecal Incontinence Severity Index; 2) compare a self-administered Fecal Incontinence Quality of Life Scale to the Manchester Health Questionnaire after modifying the latter for telephone administration and American English (Modified Manchester Health Questionnaire); 3) assess test-retest reliability of the telephone-administered Modified Manchester Health Questionnaire; and 4) assess the internal consistency of the Modified Manchester Health Questionnaire subscales. METHODS: Consecutive, English-speaking, nonpregnant females known to have fecal incontinence were invited to participate. Two validated paper questionnaires accompanied the letter informing them of the study: Fecal Incontinence Severity Index and Fecal Incontinence Quality of Life Scale. Consenting patients were contacted for the initial telephone administration of the Modified Manchester Health Questionnaire, and patients who agreed to continue the study were contacted for a repeat telephone administration of the Modified Manchester Health Questionnaire two to four weeks after completing the first interview. RESULTS: Fifty-one females were invited to participate in the study; however, 13 declined or were ineligible. Thirty females, aged 49.3 +/- 10.3 years, returned self-administered questionnaires and completed the first telephone interview, and 21 completed a second telephone interview after an average interval of 23 days. The telephone-administered Fecal Incontinence Severity Index scores were significantly lower than those yielded by the self-administered Fecal Incontinence Severity Index, (6.19 vs. 9.85; P < 0.001), but the telephone and written administrations were significantly correlated (r = 0.5; P < 0.02). Correlations between the Modified Manchester Health Questionnaire quality of life subscales and the paper Fecal Incontinence Quality of Life subscales ranged from 0.6 to 0.9 (median, r = 0.81). The correlation between the total score for the Fecal Incontinence Quality of Life and the total score for the Modified Manchester Health Questionnaire quality of life scales was 0.93 (P < 0.001). Test-retest reliability for the eight Modified Manchester Health Questionnaire subscales ranged from 0.55 to 0.98 (median, r = 0.83), and test-retest reliability for the two telephone administrations of the Fecal Incontinence Severity Index was r = 0.75. Cronbach's alpha for the eight Modified Manchester Health Questionnaire subscales ranged from 0.79 to 0.92 (median, alpha = 0.85). CONCLUSIONS: Telephone-administered versions of the Modified Manchester Health Questionnaire showed good-to-excellent validity, internal consistency, and test-retest reliability. The telephone-administered Fecal Incontinence Severity Index yielded lower severity scores than the written Fecal Incontinence Severity Index; however, the difference (3.66 units) was not clinically significant.

Adult↗

Item analysis to improve reliability for an internal medicine undergraduate OSCE.

Utilization of objective structured clinical examinations (OSCEs) for final assessment of medical students in Internal Medicine requires a representative sample of OSCE stations. The reliability and generalizability of OSCE scores provides validity evidence for OSCE scores and supports its contribution to the final clinical grade of medical students. The objective of this study was to perform item analysis using OSCE stations as the unit of analysis and evaluate the extent to which OSCE score reliability can be improved using item analysis data. OSCE scores from eight cohorts of fourth-year medical students (n = 435) in a 6-year undergraduate program were analyzed. Generalizability (G) coefficients of OSCE scores were computed for each cohort. Item analysis was performed by considering each OSCE station as an item and computing the corrected item-total correlation. OSCE stations which negatively impacted the reliability were deleted and the G-coefficient was recalculated. The G-coefficients of OSCE scores from the eight cohorts ranged from 0.48 to 0.80 (median 0.62). The median number of OSCE stations that negatively impacted the G-coefficient was 3.5 (out of a median of 25 total stations). When the ''problem stations'' were deleted, the median G-coefficient across eight cohorts increased to 0.62--0.72. In conclusion, item analysis of OSCE stations is useful and should be performed to improve the reliability of total OSCE scores. Problem stations can then be identified and improved.

Clinical Competence↗

Can concept sorting provide a reliable, valid and sensitive measure of medical knowledge structure?

CONTEXT: Evolution from novice to expert is associated with the development of expert-type knowledge structure. The objectives of this study were to examine reliability and validity of concept sorting (ConSort) as a measure of static knowledge structure and to determine the relationship between concepts in static knowledge structure and concepts used during diagnostic reasoning. METHOD: ConSort was used to identify static knowledge concepts and analysis of think-aloud protocols was used to identify dynamic knowledge concepts (used during diagnostic reasoning). Intra- and inter-rater reliability, and correlation across cases, were evaluated. Construct validity was evaluated by comparing proportions of nephrologists and students with expert-type knowledge structure. Sensitivity and specificity of static knowledge concepts as a predictor of dynamic knowledge concepts were estimated. RESULTS: Thirteen first-year medical students and 19 nephrologists participated. Intra- and inter-rater agreement for determination of static knowledge concepts were 1.0 and 0.90, respectively. Reliability across cases was 0.45. The proportions of nephrologists and students identified as having expert-type knowledge structure were 82.9% and 55.8%, respectively (p=0.001). Sensitivity and specificity of ConSort((c)) in predicting concepts that were used during diagnostic reasoning were 96.8% and 27.8% for nephrologists and 87.2% and 55.1% for students. CONCLUSIONS: ConSort is a reliable, valid and sensitive tool for studying static knowledge structure. The applicability of tools that evaluate static knowledge structure should be explored as an addition to existing tools that evaluate dynamic tasks such as diagnostic reasoning.

Alberta↗

Test-retest reliability of the multifocal electroretinogram and humphrey visual fields in patients with retinitis pigmentosa.

We examined the reliability of Humphrey visual field thresholds and multifocal electroretinogram (mfERG) amplitudes and timing in a group of patients with Retinitis Pigmentosa (RP). Eight patients with RP and seven control subjects were tested five times: at baseline (visit #0), at three weekly follow-up visits (visits #1 - #3), and at three months (visit #4). For the Humphrey thresholds, differences between dB values on repeat visits were obtained. Differences between log values on repeat visits were calculated for mfERG amplitude and implicit time. We used the standard deviations of these difference scores as a measure of reliability and the means of the difference scores as a measure of progression. We found that the majority of the patients' repeat data were more variable than that of the control subjects for both the Humphrey and mfERG. We found no single factor that predicted the magnitude, or the variance, of the SD of differences scores for the patients. We recommend that each patient's reliability be assessed individually. Ultimately, the choice of an outcome measure must be guided by its reliability, as well as its ability to assess the visual function of interest.

Adult↗

Transductive machine learning for reliable medical diagnostics.

In the past decades Machine Learning tools have been successfully used in several medical diagnostic problems. While they often significantly outperform expert physicians (in terms of diagnostic accuracy, sensitivity, and specificity), they are mostly not being used in practice. One reason for this is that it is difficult to obtain an unbiased estimation of diagnose's reliability. We discuss how reliability of diagnoses is assessed in medical decision making and propose a general framework for reliability estimation in Machine Learning, based on transductive inference. We compare our approach with a usual (Machine Learning) probabilistic approach as well as with classical stepwise diagnostic process where reliability of diagnose is presented as its posttest probability. The proposed transductive approach is evaluated on several medical data sets from the UCI (University of California, Irvine) repository as well as on a practical problem of clinical diagnosis of the coronary artery disease. In all cases significant improvements over existing techniques are achieved.

Adult↗

The Personality Disorders Institute/Borderline Personality Disorder Research Foundation Randomized Control Trial for Borderline Personality Disorder: reliability of Axis I and II diagnoses.

The Personality Disorder Institute/Borderline Personality Disorder Research Foundation randomized control trial (PDI/BPDRF RCT) is a randomized control trial comparing three treatments for borderline personality disorder (BPD). An important issue for any RCT is diagnostic reliability, demonstration of which is necessary to evaluate claims of a treatment's efficacy for a given population. The present paper examines the interrater reliability of Axis I and II disorders in the context of a high base rate of BPD features for participants referred for inclusion in the RCT. Our results indicate good to excellent levels of interrater reliability for all Axis I and II disorders in this context. Assessors were able to reliably diagnose BPD, exclusionary criteria, and comorbid diagnoses. This data is important for comparing findings and sample composition across different studies using similar sampling strategies, especially as treatments are increasingly being developed and tested for BPD.

Adult↗

The Dartmouth COOP Charts: a simple, reliable, valid and responsive quality of life tool for chronic obstructive pulmonary disease.

The negative impact of chronic obstructive pulmonary disease (COPD) on health-related quality of life (HRQL) is substantial. Measurement of HRQL is increasingly advocated in clinical practice; traditional outcome measures such as lung function are poorly responsive. However many HRQL tools are not user-friendly in the clinic setting. Hence HRQL is often neglected. The Dartmouth Cooperative Functional Assessment Charts (COOP) have the requisite attributes of a tool suitable for routine clinical practice: they are simple, reliable, quick and easy to perform and score and well accepted. We aimed to determine the reliability, validity and responsiveness of the COOP in patients with significant COPD. HRQL was assessed during a prospective, randomised, placebo-controlled, double-blind, 12 week cross-over interventional study of ambulatory oxygen in patients (n = 50) with COPD. Test-retest reliability of the COOP domains was only modest however it was measured over a 2 month period. Significant correlations ranging between 0.4 and 0.8 were observed between all comparable domains of the COOP and the Medical Outcomes Study 36-item Short-form Health Survey, Chronic Respiratory Questionnaire (CRQ) and Hospital Anxiety and Depression (HAD) scale. Following ambulatory oxygen significant improvements were noted in all CRQ and HAD domains. Several domains of the generic SF-36 (role emotional, social functioning, role-physical) showed significant improvements. Comparable domains of the COOP (social activities, feelings) also showed significant improvements. The COOP change in health domain improved very significantly. The COOP is a simple, reliable HRQL tool which proved valid and responsive in our study population of COPD patients and may have a valuable role in routine clinical practice.

Activities of Daily Living↗

The validity and reliability of the functional impairment checklist (FIC) in the evaluation of functional consequences of severe acute respiratory distress syndrome (SARS).

Severe acute respiratory distress syndrome (SARS) contributed to significant mortality and morbidity worldwide. We aimed to establish the validity, reliability and responsiveness of the functional impairment checklist (FIC) as a measurement tool for physical dysfunction in SARS survivors. One hundred and sixteeen (65 females and 51 males, mean age 45.6) patients who joined the SARS rehabilitation programme were analysed. The factor analysis yielded two latent factors. The mean FIC-symptom and FIC-disability score were 24.12 (SD +/- 20.2) and 26.11 (SD +/- 27.32), respectively. Based on the item-scale correlation coefficients, the Cronbach's alpha coefficients reflecting the internal consistency reliability of scale score were 0.75 for FIC-symptom and 0.86 for FIC-disability. Test-retest reliability in 23 patients showed no statistical significant difference in the FIC scores between tests with intraclass correlation coefficient (ICC) 0.49-0.57. The FIC scales correlated both with 6 munute walking test (6MWT) distance (-0.26 and -0.38) and handgrip strength (HGS) (-0.20 and -0.27). Moreover, the FIC scales correlated with St. George's respiratory questionnaire (SGRQ) (0.19 to 0.52) and short form 36 Hong Kong (SF-36) domains (-0.19 to -0.59). Both FIC scales correlated stronger with physical component summary (PCS) (-0.41 and -0.55) than with mental component summary (MCS) (-0.30 and -0.23). FIC reduced significantly at 6 months while the SF-36 PCS and MCS did not show any change. In conclusion, the study results indicate the FIC is reliable, valid and responsive to change in symptom and disability as a consequence of SARS, suggesting it may provide a means of assessing health related quality of life (HRQOL) outcomes in a longitudinal follow up.

Adult↗

Reliability, validity and responsiveness of the German Short Musculoskeletal Function Assessment Questionnaire in patients undergoing surgical or conservative inpatient treatment.

OBJECTIVE: The patient-based evaluation of outcome is gaining increased importance. The aim of the study was to demonstrate the reliability, validity and responsiveness of the German version of the Short Musculoskeletal Function Assessment Questionnaire (SMFA-D) in patients undergoing surgical or conservative treatment. METHODS: Three hundred and thirty-two patients suffering from osteoarthritis of the hip or knee, rheumatoid arthritis or rotator cuff tear undergoing surgical or medical inpatient treatment were followed up for 12 month. Patients underwent both SMFA-D and other assessments and clinical as well as radiological examinations. Reliability, validity and responsiveness of the SMFA-D were evaluated. RESULTS: Values of the SMFA-D subscales, Function index (M 22-49, SD 12-20, range 0-96) and Bother index (M 29-52, SD 15-23, range 0-100), showed a normal distribution. Internal consistency (0.88-0.97) and retest reliability (0.71-0.96) coefficients were satisfactory to excellent. In most cases, the SMFA-D correlated significantly with function tests, physicians' function ratings, patients' pain ratings and other quality-of-life questionnaires in all patient subgroups. The results support both the construct and criterion validity of the measure. Different patient groups and subgroups could be discriminated with the SMFA-D scales. The standardized response means of SMFA-D subscales were in surgical patients better than in conservatively treated patients and comparable to those of the SF-36 Physical Component Summary scale. CONCLUSIONS: The German version of SMFA is a reliable, valid and responsive questionnaire in patients with osteoarthritis of the hip or knee, rheumatoid arthritis or rotator cuff tear undergoing surgical or medical inpatient treatment. Thus, the use of the SMFA-D in these patients can be recommended.

Activities of Daily Living↗

Validity and reliability of reported dietary intake data.

OBJECTIVE: To compare two training techniques for validity and reliability of dietary instruments and the measurement of total energy expenditure (TEE) to determine whether technique could influence the accuracy of food portion estimates. DESIGN: Adult women were randomized into a control group and an experimental group for comparison of training technique. SETTING: University and research center. SUBJECTS: Five hundred women were screened using the Three-Factor Eating Questionnaire to identify restrained eaters or disinhibitors. Other criteria for selection included good health; absence of thyroid, respiratory, or other diseases; normal menstrual cycles; between the ages of 18 and 50 years. Forty-nine were recruited, with an attrition rate of 10% for a total sample of 44 subjects. INTERVENTION: The control group (n = 26) was trained with food models and the experimental group (n = 18) was trained with a combination of food models and life-sized food photographs. All subjects completed two 24-hour recalls and 14 consecutive days of food records. TEE was measured by the doubly-labeled water method. MAIN OUTCOME MEASURES: Training would improve the accuracy of food portion estimates. STATISTICAL ANALYSES PERFORMED: Analysis of variance, the paired t test, Pearson's correlation coefficient, and Wilcoxon's ranking test. RESULTS: The mean reported intake between instruments was found to be reliable; however, the comparison with TEE was underreported by 21.4% and was thus nonvalid. Training technique made no difference in validity or reliability. Both training techniques improved the accuracy of food portion estimates; however, improvement was enhanced with food photographs. APPLICATIONS/CONCLUSION: The findings indicate that training can improve food portion estimates, and dietary instruments may provide reliable but nonvalid results.

Adult↗

Reliability and validity of clinical assessments of malocclusion.

The purpose of this research was to determine the reliability and validity of selected clinical judgments of malocclusion, including general evaluations of occlusal status and more specific aspects of dentofacial malrelations. Study casts of twenty-one adolescents planning orthodontic treatment and twenty-nine not planning treatment were examined and rated. The examiners were five dentists in an orthodontic specialty-training program. They completed ratings on six dimensions: (1) need for treatment, (2) degree of malocclusion, (3) potential for tissue loss, (4) negative effect on occlusal stability, (5) negative effect on dental-facial attractiveness, and (6) negative effect on masticatory function. Six weeks later the same five rates scored the fifty casts, using the standardized Treatment Priority Index (TPI). Three weeks later, or 9 weeks after the initial ratings, the casts were again rated on the two general dimensions: need for treatment and degree of malocclusion. Correlations among all the measures were examined. Inter-rater reliability was highest for the ratings of impact on dental-facial attractiveness (r = 0.88). The two general assessments also yielded relatively high rater reliabilities, and the second rating yielded stability coefficients of 0.84 for both of these ratings. Correlations with total TPI scores were 0.70 for the dental-facial attractiveness measure and 0.65 and 0.64, respectively, for assessments of need for treatment and degree of malocclusion. The data indicate that clinical evaluations of the severity of malocclusions are comparable to objective measures in terms of inter-rater reliability. Clinical evaluations are also relatively stable over time. Correlations with the TPI scores also provide evidence of the concurrent validity of clinical judgments.

Adolescent↗

Reliability of different grading systems used in evaluating surgical students.

Inter-rater agreement in assigning grades using five different grading systems was determined. The performance of 16 students in a surgery clerkship was rated by 21 faculty raters using a pass-fail grading system, a pass-fail-honors system, a letter grade system, a number grade scale from 1 to 10, and a number grade scale from 1 to 100. Inter-rater agreement coefficients were used to assess relative and absolute reliabilities, respectively. Both the letter grade and 1 to 10 number grade systems provided good discrimination, had high to moderate reliability, and required only five raters to achieve a mean rating with the commonly recommended reliability of 0.80. Using the letter grade system, however, a majority of raters agreed on a specific grade assignment for 14 of 16 students, in contrast to the 1 to 10 scale, for which this was true for only 4 of 16 students. The results of this reliability study favor the use of a letter grading system.

Clinical Clerkship↗

The Clinician-Administered Rating Scale for Mania (CARS-M): development, reliability, and validity.

There are currently seven rating scales available to assess manic symptomatology. All, however, have some limitations that could restrict their clinical and research utility. To resolve these deficiencies the Clinician-Administered Rating Scale for Mania (CARS-M) was developed and normed on 96 patients with mixed diagnoses during baseline and following treatment. Interrater reliability was established across multiple raters viewing 14 videotaped interviews and comparing agreement among individual items and total scores. Test-retest reliability was assessed on 36 patients twice during baseline. The mean intraclass correlation coefficient among five raters across items for each of the 14 patients was 0.81, and for total scores 0.93. Principal components analysis of items revealed two factors: mania, and psychosis. Test-retest reliability was significant for both factors (range = 0.78 to 0.95). Internal validity, comparing each item with its respective total factor score, revealed significant correlations for all items. Correlation of CARS-M total scores with mania rating scale (MRS) total scores was 0.94. Results indicate the CARS-M is both a reliable and valid measure of the severity of manic symptomatology, which incorporates a number of methodological improvements leading to greater precision and clinical utility.

Adolescent↗