Search PubMed⌕ Search

Biomedical subjects

G Regehr

Publications and source records attributed to G Regehr.

At least 37 records · Page 2Linked to original sources

A virtual reality module for intravenous catheter placement.

BACKGROUND: Virtual reality (VR) is a potential tool for technical skills training. We tested the validity and instructional effectiveness of a prototype VR module for learning intravenous (i.v.) catheter placement. METHODS: First-year medical students (n = 37), third-year medical students (n = 14), and surgical residents (n = 9) attempted two pretest i.v.s into each other, used the VR module for 12 minutes, and subsequently attempted two posttest i.v.s. Success or failure were recorded for each attempt. For each successful attempt, time and global rating of i.v. insertion were also recorded. RESULTS: The pretest success rate was significantly different between groups (chi square = 28.71, P <0.01). VR success rate was not significantly different between groups (F(2,57) = 1.47, ns). Although there was improvement in all groups during VR training (F(2,114) = 44.16, P <0.01), this did not result in improvement in posttest performance. CONCLUSIONS: Significant differences between groups were observed in performance of i.v. insertion in physical reality. However, no significant difference was observed in performance in VR. Thus, performance in VR demonstrated neither construct nor concurrent validity. While performance improved in VR, transfer of skill from VR to physical reality was not observed. Additional development and testing of VR as a training tool is warranted before its widespread use can be recommended.

Catheterization↗

OSCE checklists do not capture increasing levels of expertise.

PURPOSE: To evaluate the effectiveness of binary content checklists in measuring increasing levels of clinical competence. METHOD: Fourteen clinical clerks, 14 family practice residents, and 14 family physicians participated in two 15-minute standardized patient interviews. An examiner rated each participant's performance using a binary content checklist and a global process rating. The participants provided a diagnosis two minutes into and at the end of the interview. RESULTS: On global scales, the experienced clinicians scored significantly better than did the residents and clerks, but on checklists, the experienced clinicians scored significantly worse than did the residents and clerks. Diagnostic accuracy increased for all groups between the two-minute and 15-minute marks without significant differences between the groups. CONCLUSION: These findings are consistent with the hypothesis that binary checklists may not be valid measures of increasing clinical competence.

Analysis of Variance↗

Assessing the generalizability of OSCE measures across content domains.

PURPOSE: To assess the degree to which OSCE measures generalize across multiple administrations to the same students. METHODS: Students' scores from three OSCEs at one institution were correlated to determine the generalizability of the scoring systems across course domains. RESULTS: Analysis revealed that while checklist scores showed quite low correlations across examinations from different domains (ranging from 0.14 to 0.25), global process scores showed quite reasonable correlations (ranging from 0.30 to 0.44), with the correlations for global scores being significantly higher than those for checklist scores in all three comparisons. CONCLUSION: These data would seem to confirm the intuitions about each of these measures: the checklist scores are highly content-specific, while the global scores are evaluating a more broadly based set of skills. Implications for the use of these scales are discussed.

Clinical Clerkship↗

Computer-assisted learning versus a lecture and feedback seminar for teaching a basic surgical technical skill.

BACKGROUND: Rapid improvements in computer technology allow us to consider the use of computer-assisted learning (CAL) for teaching technical skills in surgical training. The objective of this study was to compare in a prospective, randomized fashion, CAL with a lecture and feedback seminar (LFS) for the purpose of teaching a basic surgical skill. METHODS: Freshman medical students were randomly assigned to spend 1 hour in either a CAL or LFS session. Both sessions were designed to teach them to tie a two-handed square knot. Students in both groups were given knot tying boards and those in the CAL group were asked to interact with the CAL program. Students in the LFS group were given a slide presentation and were given individualized feedback as they practiced this skill. At the end of the session the students were videotaped tying two complete knots. The tapes were independently analyzed, in a blinded fashion, by three surgeons. The total time for the task was recorded, the knots were evaluated for squareness, and each subject was scored for the quality of performance. RESULTS: Data from 82 subjects were available for the final analysis. Comparison of the two groups demonstrated no significant difference between the proportion of subjects who were able to tie a square knot. There was no difference between the average time required to perform the task. The CAL group had significantly lower quality of performance (t = 5.37, P <0.0001). CONCLUSIONS: CAL and LFS were equally effective in conveying the cognitive information associated with this skill. However, the significantly lower performance score demonstrates that the students in the CAL group did not attain a proficiency in this skill equal to the students in the LFS group. Comments by the students suggest that the lack of feedback in this model of CAL was the significant difference between these two educational methods.

Computer-Assisted Instruction↗

Validation of an objective structured clinical examination in psychiatry.

PURPOSE: To examine the validity of a psychiatry clerkship's objective structured clinical examination (OSCE). METHOD: In 1996, 33 clinical clerks and 17 psychiatry residents at the University of Toronto participated in an eight-station OSCE evaluated by psychiatrist-examiners using binary checklists and global ratings. Prior to the OSCE, communication course instructors were asked to rank the clerks on interviewing ability, and faculty supervisors were asked to identify the OSCE stations on which the clerks were likely to do well or poorly. RESULTS: Mean OSCE scores were significantly higher for the residents than for the clerks on global ratings but not on checklists. The communication instructors accurately predicted the clerks' rankings on the global scores but not their scores on the checklists. The faculty supervisors predicted with moderate accuracy the clerks' success on the OSCE stations as measured by the checklists but not by the global ratings. The residents rated the OSCE scenarios as highly realistic. CONCLUSIONS: The evidence of construct and concurrent validity together with high ratings of realism suggest that a psychiatry OSCE can be a valid assessment of clerks' clinical competence.

Clinical Clerkship↗

Comparing the psychometric properties of checklists and global rating scales for assessing performance on an OSCE-format examination.

PURPOSE: To compare the psychometric properties of checklists, global rating scales preceded by a checklist, and global rating scales alone in assessing surgery residents' performances on an OSCE-like technical skills examination. METHOD: In 1996, 53 general surgery residents with one to six years of postgraduate training participated in a performance-based examination of technical skills consisting of eight 15-minute stations (bench-model simulations of operative procedures in general surgery). Two qualified surgeons marked at each station, one using a task-specific checklist (C) and a subsequent global rating scale (Gc), the other using a global rating scale only (G). RESULTS: Interstation reliabilities measured by Cronbach's alpha were .79 for C, .89 for Gc, and .85 for G. A series of multiple regressions predicting level of training from test scores revealed an R2 of .584 for C alone, which increased to .711 when Gc was entered after (p < .001), and increased to .704 when G was entered after C (p < .001). However, R2 for Gc alone was .711, and for G alone was .704, neither of which changed when C was entered into the prediction (p > .10). The R2 for Gc and G predicting level of training (.725) was not significantly greater than that of either Gc or G alone. A very similar pattern of results was seen when C, Gc, and G were used to predict independent evaluations of the operative outcomes. CONCLUSIONS: Global rating scales scored by experts showed higher inter-station reliability, better construct validity, and better concurrent validity than did checklists. Further, the presence of the checklists did not improve the reliability or validity of the global rating scale over that of the global rating scale alone. These results suggest that global rating scales administered by experts are a more appropriate summative measure when assessing candidates on performance-based examinations.

Educational Measurement↗

Using videotaped benchmarks to improve the self-assessment ability of family practice residents.

PURPOSE: To address methodologic and statistical problems of previous studies of self-assessment by exposing participants to relevant standards, anchoring rating scales, and providing practice in the use of the assessment tool. METHOD: Fifty first- and second-year family practice residents performed a ten-minute patient interview with a difficult communication problem. Following each interview, the resident and two experts independently evaluated the resident's communication skills. The resident was then shown a videotape of four performances (ranging in quality from poor to good) of the same scenario. The resident evaluated the communication skills displayed in each performance and then reevaluated his or her own performance. RESULTS: The correlation between experts' evaluations and residents' self-evaluations was moderate immediately after the interview (r = 0.38) but increased significantly after the residents viewed the videotape (r = 0.52). This effect was more pronounced for first-year residents (0.22 to 0.45) than for second-year residents (0.53 to 0.65), although the difference was not significant. Post-hoc analysis revealed that neither initial nor post-benchmark self-assessment ability was related to the ability to accurately evaluate the benchmarks in a manner consistent with the experts. CONCLUSIONS: The ability to self-assess does not seem strongly tied to the ability to assess the performances of others on the same task. Nonetheless, providing a set of benchmarks against which trainees can compare their own performances improves their ability to self-evaluate even if the qualities of the benchmarks are not explicitly identified.

Adult↗

The integration of child psychiatry into a psychiatry clerkship OSCE.

OBJECTIVE: To integrate child psychiatry into a psychiatry clerkship OBJECTIVE Structured Clinical Examination (OSCE). METHOD: Child psychiatry OSCE stations were designed to evaluate clerk' skills in the identification of 4 common conditions. Child psychiatrists wrote case scenarios and checklists and supported standardized patient (SP) training for the stations. A bank of 4 child psychiatry OSCE stations is now available for use in the psychiatry OSCE. Child psychiatry faculty have been trained as examiners for ongoing administration of his OSCE. RESULTS: This bank of child psychiatry OSCE stations has examined 402 clerks. Mean student scores for content were 68% to 86% and for process were 69% to 76%. Station reliability and examiner feedback were acceptable. CONCLUSIONS: Child psychiatry has been successfully integrated into a psychiatry clerkship OSCE. Although the commitment in terms of monetary and faculty costs has been considerable, the accompanying educational benefits of such integration warranted this expense.

Child↗

A model for predicting depression in victims of rape.

This article proposes a model for understanding the factors contributing to long-standing depression in women who have been raped. A path analysis of data obtained from 71 women who had been raped revealed that women with generalized beliefs that they could not control events in their lives were more likely to attribute responsibility for their rape to permanent intrapsychic factors and were more likely to be depressed. Women who perceived that they had higher levels of internal control tended to have higher levels of education, were more likely to be employed, and were less likely to be depressed more than one year after having been raped. Childhood sexual abuse was not associated with internal control or attributions of causality or depression in this analysis. Implications for the determination of prognosis and treatment recommendations in civil litigation assessments are discussed.

Adolescent↗

Testing technical skill via an innovative "bench station" examination.

BACKGROUND: A new approach to testing operative technical skills, the Objective Structured Assessment of Technical Skill (OSATS), formally assesses discrete segments of surgical tasks using bench model simulations. This study examines the interstation reliability and construct validity of a large-scale administration of the OSATS. METHODS: A 2-hour, eight-station OSATS was administered to 48 general surgery residents. Residents were assessed at each station by one of 48 surgeons who evaluated the resident using two methods of scoring: task-specific checklists and global rating scales. RESULTS: Interstation reliability was 0.78 for the checklist score, and 0.85 for the global score. Analysis of variance revealed a significant effect of training for both the checklist score, F(3,44) = 20.08, P <0.001, and the global score, F(3,44) = 24.63, P <0.001. CONCLUSIONS: The OSATS demonstrates high reliability and construct validity, suggesting that we can effectively measure residents' technical ability outside the operating room using bench model simulations.

Clinical Competence↗

A new assessment tool: the patient assessment and management examination.

BACKGROUND: The major goal of certification is to assure the public that the candidate is competent in all facets required of the position. The patient assessment and management examination (PAME) was developed to enable a more comprehensive assessment of competence in the practice of surgery. METHODS: A six-station, 3-hour, standardized-patient-based evaluation was developed. Each station was scored using a set of five-point global rating scales. PAME results were compared to the last two in training evaluation reports (ITER), the clinical knowledge component of the ITER (ITER-CK), an in-house oral examination (OE), and the Canadian Association of General Surgeons' multiple-choice examination (CAGS). RESULTS: Eighteen senior general surgery residents were evaluated. Overall reliability was 0.70 (Cronbach's alpha). Fifth-year residents scored significantly better than fourth-year residents (t = 3.062; p = 0.0074), with 1 year of training accounting for 37% of the variance in scores. Correlations between the PAME and each of the other measures were ITER, 0.24; ITER-CK, 0.38; OE, -0.13; and CAGS, 0.061, with the PAME demonstrating better reliability and stronger evidence of validity than any other. CONCLUSIONS: The PAME had better psychometric properties than other measures and assessed areas often not evaluated. This type of evaluation may be useful for feedback, remediation, or certification decisions.

Adult↗

Methodological problems in the retrospective computation of responsiveness to change: the lesson of Cronbach.

OBJECTIVE: To examine the relation between responsiveness coefficients derived directly from a calculation of average change resulting from a treatment intervention (Responsiveness-Treatment or RT) and those derived from retrospective analysis of changed and unchanged groups (Responsiveness Retrospective or RR) based on a global measure of change. METHOD: Two approaches were used. First, we used simulation methods to examine the analytical relationship between the RT and RR coefficients. We then located eight studies where it was possible to compute both RT and RR coefficients. As anticipated from theoretical arguments, the RR coefficients were larger than the RT coefficients (1.50 versus 0.41, p < .0001). Within study there was no predictable relationship between the two indices. Across studies, the magnitude of the RR coefficient was strongly related to the correlation with the retrospective global scale, and unrelated to the magnitude of the RT coefficient. The simulated curves fit well with the observed data, and substantiated the observation that the relation between RT and RR coefficients is complex and only weakly related to the size of the treatment effect. CONCLUSION: Retrospective methods of computing responsiveness yield little information about the ability of an instrument to detect treatment effects, and should not be used as a basis for choice of an instrument for applications to clinical trials.

Computer Simulation↗

Applying a relative ranking model to the self-assessment of extended performances.

BACKGROUND: Accurate self-assessment is an important but underdeveloped skill in medicine that, in the past, has received little formal attention from educators. METHOD: Following an orthopedic rotation, twenty-five orthopedic surgery residents performed a self-assessment task for ten skills using a new relative ranking method, in which an individual's skills are ranked relative to each other rather than being compared to the individual's peers. Supervising faculty assessed residents using the same instrument. Faculty inter-rater reliability was measured and comparisons were made between each resident's self-assessment and the faculty assessments using Spearman rank order correlation coefficients. RESULTS: The mean correlation between faculty rating the same resident was 0.27 (sd = 0.49). The mean correlation between resident and faculty rankings was 0.20 (sd = 0.38), but was higher for junior residents (0.33) than for senior residents (0.12), apparently because senior residents do not alter their self-assessments while faculty change their assessments of senior residents. CONCLUSIONS: Consistent with the literature in other fields, we find that self-assessment is poor among surgical trainees when they are asked to assess their own performance over an extended time period.

Journal Article↗

Objective structured assessment of technical skill (OSATS) for surgical residents.

BACKGROUND: The technical skill of surgical trainees is not well assessed. This study aimed (1) to compare the reliability of three scoring systems, (2) to compare live and bench formats and (3) to assess construct validity of a test of operative skill. METHODS: Parallel examinations of operative skill, one using live animals and one using simulations, were developed. Performance was graded using operation-specific checklists, detailed global rating forms and pass/fail judgements. Twenty surgical residents each took both formats. RESULTS: Disattenuated correlations between live and bench scores were high (0.69-0.72). Mean interrater reliability across stations ranged from 0.64 to 0.72. Internal consistency was moderate to high (alpha: 0.61-0.74) for the live format using the checklist and for live and bench formats using global ratings. Global ratings discriminated between resident levels for both formats (bench: F(2,17) = 4.45, P < 0.05; live: F(2,17) = 3.55, P < 0.05), checklists did not. CONCLUSION: This preliminary study suggests that the Objective Structured Assessment of Technical Skill can reliably and validly assess surgical skills. Global ratings are a better method of assessment than task-specific checklists. Bench model simulation gives equivalent results to use of live animals for this test format.

Clinical Competence↗

An objective structured clinical examination for evaluating psychiatric clinical clerks.

PURPOSE: To assess the feasibility, reliability, and validity of an objective structured clinical examination (OSCE) for psychiatric clinical clerks. METHOD: In 1995 two parallel forms of a ten-station OSCE (eight clinical stations, two writing stations) were developed at the University of Toronto Faculty of Medicine Each 12-minute performance-based clinical station was assessed by a faculty psychiatrist using both a checklist for each student's performance content and a global-rating scale of the performance process. The students' clinical-station scores were calculated as the average of their content and process scores (expressed as percentages). Examiners also recorded an overall judgment of each students' performance (pass, borderline, or fail) and wrote [in collaboration with the standardized patient (SP) at that station] comments on each student's performance. There were two criteria for a passing grade: a total mark of 60% or higher across all ten stations and a "pass" or "borderline" mark in at least five of the eight clinical stations. Each OSCE form was administered three times. RESULTS: The first form was used to examine 94 clerks, the second form to examine 98 clerks. The students' mean scores for the two forms were 70.47% (SD, 6.33%) and 67.66% (SD, 7.05%), respectively. In addition to the standard evaluation information collected on the students, several critical incidents occurred (e.g., a student's loss of control of emotions) that may identify potential problems in professional conduct. The direct cost for one administration of the examination was approximately Can$3,300: the largest portion of this was for the SPs' time spent in training and performing their roles. CONCLUSION: Preliminary evidence suggests that a psychiatry OSCE is feasible for assessing complex psychiatric skills. However, careful attention must be paid to SP training, examination monitoring, detection of critical incidents, and provision of feedback to students, faculty, and SPs. The university's previous system of oral examinations required approximately 600 faculty hours per year. The OSCE requires approximately 450 faculty hours, and the 150 hours saved almost cover the Can$20,000 that the examination costs each year. In all, the OSCE is an evaluation system that has demonstrable reliability and is more enjoyable for both the faculty and the students.

Clinical Clerkship↗