Search PubMed⌕ Search

Biomedical subjects

Kevin W Eva

Publications and source records attributed to Kevin W Eva.

35 records · Page 2Linked to original sources

Do clinical clerks provide candidates with adequate formative assessment during Objective Structured Clinical Examinations?

CONTEXT: Various research studies have examined the question of whether expert or non-expert raters, faculty or students, evaluators or standardized patients, give more reliable and valid summative assessments of performance on Objective Structured Clinical Examinations (OSCEs). Less studied has been the question of whether or not non-faculty raters can provide formative feedback that allows students to take advantage of the educational opportunity that OSCEs provide. This question is becoming increasingly important, however, as the strain on faculty resources increases. METHODS: A questionnaire was developed to assess the quality of feedback that medical examiners provide during OSCEs. It was pilot tested for reliability using video recordings of OSCE performances. The questionnaires were then used to evaluate the feedback given during an actual OSCE in which clinical clerks, residents, and faculty were used as examiners on two randomly selected test stations. RESULTS: The inter-rater reliability of the 19-item feedback questionnaire was 0.69 during the pilot test. The internal consistency was found to be 0.90 during pilot testing and 0.95 in the real OSCE. Using this form, the feedback ratings assigned to clinical clerks were significantly greater than those assigned to faculty evaluators. Furthermore, performance on the same OSCE stations eight months later was not impaired by having been evaluated by student examiners. DISCUSSION: While evidence of mark inflation within the clinical clerk examiners should be addressed with examiner training, the current results suggest that clerks are capable of giving adequate formative feedback to more junior colleagues.

Adult↗

How can I know what I don't know? Poor self assessment in a well-defined domain.

As the rapidity with which medical knowledge is generated and disseminated becomes amplified, an increasing emphasis has been placed on the need for physicians to develop the skills necessary for life-long learning. One such skill is the ability to evaluate one's own deficiencies. A ubiquitous finding in the study of self-assessment, however, is that self-ratings are poorly correlated with other performance measures. Still, many educators view the ability to recognize and communicate one's deficiencies as an important component of adult learning. As a result, two studies have been performed in an attempt to improve upon this status quo. First, we tried to re-define the limits within which self-assessments should be used, using Rosenblit and Keil's argument that calibration between perceived and actual performance will be better within taxonomies that are regularly tested (e.g., factual knowledge) compared to those that are not (e.g., conceptual knowledge). Second, we tried to norm reference individuals based on both the performance of their colleagues and their own historical performance on McMaster's Personal Progress Inventory (a multiple choice question test of medical knowledge). While it appears that students are able to (a) make macro-level self-assessments (i.e., to recognize that third year students typically outperform first year students), and (b) judge their performance relatively accurately after the fact, students were unable to predict the percentage of questions they would answer correctly with a testing procedure in which they have had a substantial amount of feedback. Previous test score was a much better predictor of current test performance than were individuals' expectations.

Adult↗

An admissions OSCE: the multiple mini-interview.

CONTEXT: Although health sciences programmes continue to value non-cognitive variables such as interpersonal skills and professionalism, it is not clear that current admissions tools like the personal interview are capable of assessing ability in these domains. Hypothesising that many of the problems with the personal interview might be explained, at least in part, by it being yet another measurement tool that is plagued by context specificity, we have attempted to develop a multiple sample approach to the personal interview. METHODS: A group of 117 applicants to the undergraduate MD programme at McMaster University participated in a multiple mini-interview (MMI), consisting of 10 short objective structured clinical examination (OSCE)-style stations, in which they were presented with scenarios that required them to discuss a health-related issue (e.g. the use of placebos) with an interviewer, interact with a standardised confederate while an examiner observed the interpersonal skills displayed, or answer traditional interview questions. RESULTS: The reliability of the MMI was observed to be 0.65. Furthermore, the hypothesis that context specificity might reduce the validity of traditional interviews was supported by the finding that the variance component attributable to candidate-station interaction was greater than that attributable to candidate. Both applicants and examiners were positive about the experience and the potential for this protocol. DISCUSSION: The principles used in developing this new admissions instrument, the flexibility inherent in the multiple mini-interview, and its feasibility and cost-effectiveness are discussed.

Adult↗

The relationship between interviewers' characteristics and ratings assigned during a multiple mini-interview.

PURPOSE: To assess the consistency of ratings assigned by health sciences faculty members relative to community members during an innovative admissions protocol called the Multiple Mini-Interview (MMI). METHOD: A nine-station MMI was created and 54 candidates to an undergraduate MD program participated in the exercise in Spring 2003. Three stations were staffed with a pair of faculty members, three with a pair of community members, and three with one member of each group. Raters completed a four-item evaluation form. All participants completed post-MMI questionnaires. Generalizability Theory was used to examine the consistency of the ratings provided within each of these three subgroups. RESULTS: The overall test reliability was found to be .78 and a Decision Study suggested that admissions committees should distribute their resources by increasing the number of interviews to which candidates are exposed rather than increasing the number of interviewers within each interview. Divergence of ratings was greater within the pairing of community member to faculty member and least for pairings of community members. Participants responded positively to the MMI. CONCLUSION: The MMI provides a reliable protocol for assessing the personal qualities of candidates by accounting for context specificity with a multiple sampling approach. Increasing the heterogeneity of interviewers may increase the heterogeneity of the accepted group of candidates. Further work will determine the extent to which different groups of raters provide equally valid (albeit different) judgments.

Adult↗

The ability of the multiple mini-interview to predict preclerkship performance in medical school.

PROBLEM STATEMENT AND BACKGROUND: One of the greatest challenges continuing to face medical educators is the development of an admissions protocol that provides valid information pertaining to the noncognitive qualities candidates possess. An innovative protocol, the Multiple Mini-Interview, has recently been shown to be feasible, acceptable, and reliable. This article presents a first assessment of the technique's validity. METHOD: Forty five candidates to the Undergraduate MD program at McMaster University participated in an MMI in Spring 2002 and enrolled in the program the following autumn. Performance on this tool and on the traditional protocol was compared to performance on preclerkship evaluation exercises. RESULTS: The MMI was the best predictor of objective structured clinical examination performance and grade point average was the most consistent predictor of performance on multiple-choice question examinations of medical knowledge. CONCLUSIONS: While further validity testing is required, the MMI appears better able to predict preclerkship performance relative to traditional tools designed to assess the noncognitive qualities of applicants.

Clinical Clerkship↗

Issues to consider when planning and conducting educational research.

This article is intended to provide students and clinicians aspiring to perform educational research with some background information pertaining to many of the issues inherent in performing research within this domain. It is not intended to provide a comprehensive review of the quantitative methods one might adopt, nor will it fully reflect all of the debate that currently exists within the educational research community. Rather, it is intended to offer an overview of issues and controversies within the field that will hopefully provide a starting point from which interested individuals can begin to engage in the study of educational effectiveness. Using investigations of the efficacy of problem-based learning as background, the article represents an attempt to guide new researchers through the process of generating and refining scientific research questions, identifying appropriate outcome measures, and selecting or adapting the optimal research design for the questions to be addressed. The article focuses on quantitative methods in general with particular attention paid to experimental designs.

Controlled Clinical Trials as Topic↗

Stemming the tide: cognitive aging theories and their implications for continuing education in the health professions.

As demographic drift among health care providers mimics that of the larger population, it becomes increasingly clear that theory pertaining to the impact of aging on cognitive processing should inform the continuing education efforts designed for health care professionals. The purpose of this article is to offer a critical review of the major theories in this area and outline a sample of the implications that can be derived from these views. Research articles examining the relationship between age and physician performance were identified using MEDLINE, PsychLit, and ERIC. In addition, the psychology literature on age-related changes in cognitive processing was reviewed. Evidence from the medical education literature and psychological theory suggest the importance of increased environmental supports, decreased time demands, and peer review programs as barriers against the impact of aging. The implications of these findings include the potential to tailor continuing education (and physician remediation) efforts toward the age-related abilities/deficiencies of individual physicians.

Aging↗

Can the strength of candidates be discriminated based on ability to circumvent the biasing effect of prose? Implications for evaluation and education.

PURPOSE: Residents have greater confidence in diagnoses when indicative features are presented in medical terminology. The current study examines the implications of this result by assessing its relationship to clinical ability. METHOD: Candidates writing the Medical Council of Canada's Qualifying Examination completed six questions in which the terminology used was manipulated. The influence of aptitude was examined by contrasting groups based on performance on the medicine section of Part I. RESULTS: The difference between the candidates was greatest in the mixed conditions in which the features consistent with one diagnosis were presented in medicalese and those consistent with a second diagnosis were presented using lay terminology; weaker candidates were more biased by language than stronger candidates. CONCLUSIONS: The results suggest that the language used in presenting case histories will influence the reliability of medical examinations. Furthermore, they suggest that weaker candidates might benefit from practice in making the translation between lay terminology and medicalese.

Age Factors↗

The privileged status of prestigious terminology: impact of "medicalese" on clinical judgments.

PURPOSE: Health professionals frequently use medical terminology like dyspnea or nasopharyngitis. These two studies examine how the use of medical terms affects the judgments of seriousness, prevalence, and disease; and diagnostic judgments. METHOD: In study 1, a survey containing the names of 22 diseases with either a medical or lay description was completed by 47 undergraduate psychology students and 25 medical students, who were asked to judge seriousness, prevalence, and how "disease-like" it was. In study 2, undergraduate students learned four "pseudopsychiatry" conditions, each with four associated features. Features were presented in lay or medical versions. They were then tested with 18 new cases with two medical features from one condition and two lay terms from the other. RESULTS: In study 1, the medical students rated conditions as more disease-like, more serious, and less prevalent than did the psychology students. Medical descriptions were seen as significantly less common and somewhat more serious and more disease-like. In study 2, the participants rated the condition with medical features consistently more likely than the alternative, regardless of training condition. CONCLUSIONS: The specific words used to describe a feature or condition can have an impact on judgments of likelihood of disease, and, to a lesser extent, judgments of seriousness.

Diagnosis↗

Self and peer assessment in tutorials: application of a relative-ranking model.

PURPOSE: While self assessment continues to be touted as being of paramount importance for continuing professional competence, problem-based learning curricula, and adult learning theory, techniques for ensuring valid judgments have proven elusive. This study tested the applicability of an innovative relative-ranking procedure to problem-based learning tutorials. METHOD: A total of 36 students in the McMaster University Faculty of Health Sciences' MD program were provided relative-ranking forms listing seven domains of competence along with their definitions. The student, two of the student's peers, and the student's tutor were asked to complete the ranking exercise after their second, fourth, and sixth tutorials. RESULTS: Combining each level of the time and rater variables generated 66 correlation coefficients, none of which was significantly different from zero. Re-performing the analysis on only the extreme domains did not improve this result. CONCLUSION: The relative-ranking instrument developed did not prove to be a reliable measure of tutorial performance. Ratings were inconsistent from one week to the next as well as across raters within a week.

Attitude of Health Personnel↗

Expert-novice differences in memory: a reformulation.

BACKGROUND: One of the most discriminating measures of expertise in multiple domains has been performance on memory tasks. In medicine, however, the relation between expertise and memory is more equivocal. PURPOSE: To compare and contrast the sufficiency of multiple explanations of this finding by using three probes of memory rather than the traditional free recall task alone. METHODS: Students, residents, and internists were asked to read case histories and assign diagnoses before undertaking free recall, cued recall, and recognition tests. RESULTS: Students consistently outperformed internists. Resident performance was more variable. CONCLUSIONS: Our data appear to rule out (a) the notion that expert memory for cases takes on an encapsulated form, (b) the idea that experts simply say less than students in response to a free recall task, and (c) the possibility that experts attend differentially to highly diagnostic features. The results can best be explained by the idea that students process the featural details of a case history more elaborately than do expert diagnosticians who, instead, read medical cases more holistically.

Clinical Competence↗