Search PubMed⌕ Search

Biomedical subjects

Brian E Clauser

Publications and source records attributed to Brian E Clauser.

14 recordsLinked to original sources

A multilevel analysis of the relationships between selected examinee characteristics and United States Medical Licensing Examination Step 2 Clinical Knowledge performance: revisiting old findings and asking new questions.

BACKGROUND: This study examines: (1) the relationships between examinee characteristics and United States Medical Licensing Examination Step 2 Clinical Knowledge (CK) performance; (2) the effect of gender and examination timing (time per item) on the relationship between Steps 1 and 2 CK; and (3) the effect of school characteristics on the relationships between examinee characteristics and Step 2 CK performance. METHOD: A series of hierarchical linear models (examinees-nested-in-schools) predicting Step 2 CK scores was fit to the data set. The sample included 54,487 examinees from 114 U.S. Liaison Committee on Medical Education-accredited medical schools. RESULTS: Consistent with past examinee-level research, women generally outperformed men on Step 2 CK, and examinees who received more time per item generally outperformed examinees who received less time per item. Step 1 score was generally more strongly associated with Step 2 CK performance for men and for examinees who received less time per item. School-level characteristics (size, average Step 1 performance) influenced the relationship between Steps 1 and 2 CK. CONCLUSION: Both examinee-level and school-level characteristics are important for understanding Step 2 CK performance.

Female↗

Relationships among subcomponents of the USMLE Step 2 Clinical Skills Examination, the Step 1, and the Step 2 Clinical Knowledge Examinations.

BACKGROUND: This research examined relationships between and among scores from the United States Medical Licensing Examination (USMLE) Step 1, Step 2 Clinical Knowledge (CK), and subcomponents of the Step 2 Clinical Skills (CS) examination. METHOD: Correlations and failure rates were produced for first-time takers who tested during the first year of Step 2 CS Examination administration (June 2004 to July 2005). RESULTS: True-score correlations were high between patient note (PN) and data gathering (DG), moderate between communication and interpersonal skills and DG, and low between the remaining score pairs. There was little overlap between examinees failing Step 2 CK and the different components of Step 2 CS. CONCLUSION: Results suggest that combining DG and PN scores into a single composite score is reasonable and that relatively little redundancy exists between Step 2 CK and CS scores.

Clinical Competence↗

Use of the mini-clinical evaluation exercise to rate examinee performance on a multiple-station clinical skills examination: a validity study.

BACKGROUND: Multivariate generalizability analysis was used to investigate the performance of a commonly used clinical evaluation tool. METHOD: Practicing physicians were trained to use the mini-Clinical Skills Examination (CEX) rating form to rate performances from the United States Medical Licensing Examination Step 2 Clinical Skills examination. RESULTS: Differences in rater stringency made the greatest contribution to measurement error; more raters rating each examinee, even on fewer occasions, could enhance score stability. Substantial correlated error across the competencies suggests that decisions about one scale unduly influence those on others. CONCLUSIONS: Given the appearance of a halo effect across competencies, score interpretations that assume assessment of distinct dimensions of clinical performance should be made with caution. If the intention is to produce a single composite score by combining results across competencies, the presence of these effects may be less critical.

Analysis of Variance↗

Psychometric characteristics and response times for content-parallel extended-matching and one-best-answer items in relation to number of options.

BACKGROUND: This study investigated the impact of item format and number of options on the psychometric characteristics (p values and biserials) and response times for multiple-choice questions (MCQs) appearing on Step 2 of the United States Medical Licensing Examination. METHOD: In all, 192 MCQ items were used in the study. Each item was presented in two formats: in a two-item extended-matching set and as an independent item. For the extended matching format, there were two versions: a base version that included all options (10 to 26) and an 8-option version. For the independent-item format, there were three versions: a base version that included all options, and 8-option and 5-option versions created by a group of physicians that selected options without information about examinee performance. All items were embedded in unscored sections of the 2005-06 Step 2 test forms. RESULTS: Versions of items with more options were harder and required more testing time; no differences in item discrimination were observed. Mean response times for items presented in the extended-matching format were lower than for those presented as independent items, primarily because of shorter response times for the second item presented in a set. CONCLUSION: Use of the extended-matching format and smaller numbers of options per item (and more items) should result in more efficient use of testing time and greater score precision per unit of testing time.

Licensure, Medical↗

Psychometric characteristics and response times for one-best-answer questions in relation to number and source of options.

BACKGROUND: The research reported here investigated the impact of number and source of response options on the psychometric characteristics and response times for one-best-answer MCQs. METHOD: Ninety sets of MCQs were used in two studies; numbers of options in base versions of items ranged from 11 to 25. For each set, a United States Medical Licensing Examination Step 2 item-writing committee selected the five options viewed as most appropriate. For 40 used sets, two NBME staff constructed five- and eight-option versions to maximize item discrimination. All versions of items were embedded unscored on 2003-04 Step 2 test forms. RESULTS: Versions of items with more options were harder and required more testing time; no differences in item discrimination were observed in either study, but previous versions of the items in extended matching format were more discriminating than those used in the study. CONCLUSION: Use of smaller numbers of options (and more items) results in more efficient use of testing time.

Analysis of Variance↗

Assessing the validity of the USMLE step 2 clinical knowledge examination through an evaluation of its clinical relevance.

PURPOSE: To assess the validity of the USMLE Step 2 Clinical Knowledge (CK) examination by addressing the degree to which experts view item content as clinically relevant and appropriate for Step 2 CK. METHOD: Twenty-seven experts were asked to complete three survey questions related to the clinical relevance and appropriateness of 150 Step 2 CK multiple-choice questions. Percentages, reliability estimates, and correlation coefficients were calculated and ordinary least squares regression was used. RESULTS: Results showed that 92% of expert judgments indicated the item content was clinically relevant, 90% indicated the content was appropriate for Step 2 CK, and 85% indicated the content was used in clinical practice. The regression indicated that more difficult items and more frequently used items are considered more appropriate for Step 2 CK. CONCLUSIONS: Results suggest that the majority of item content is clinically relevant and appropriate, thus providing validation support for Step 2 CK.

Clinical Competence↗

The impact of timing changes on examinee pacing on the USMLE Step 2 exam.

PURPOSE: To examine the impact of a timing change on pacing behavior and perceptions in a high-stakes multiple-choice examination. METHOD: Two samples of standard-time examinees were analyzed: (1) 29,796 examinees that completed the examination prior to the timing change, and (2) 28,373 examinees that completed the examination after the change. Subgroups of examinees were identified and compared within and across samples with respect to perceptions, accuracy, and pacing. RESULTS: After the timing change, more examinees reported having sufficient time to complete examination sections; a small improvement in overall accuracy was observed, and there was a shift in the time-per-item strategy, though examinees continued to use more than the average amount of time available per item at the beginning of sections. CONCLUSIONS: Examinees are more satisfied with the new timing constraints, although an effect due to the time limit continues to impact performance at the end of test sections.

Attitude of Health Personnel↗

Scoring the computer-based case simulation component of USMLE Step 3: a comparison of preoperational and operational data.

PURPOSE: Operational USMLE(TM) computer-based case simulation results were examined to determine the extent to which rater reliability and regression model performance met expectations based on preoperational data. METHOD: Operational data resulted from Step 3 examinations given between 1999 and 2004. Plots were produced using reliability and multiple correlation coefficients. RESULTS: Operational testing reliabilities increased over the four years but were lower than the preoperational reliability. Multiple correlation coefficient results are somewhat superior to the results reported during the preoperational period and suggest that the operational scoring algorithms have been relatively consistent. CONCLUSIONS: Changes in the rater population, changes in the rating task, and enhancements to the training procedures are several factors that can explain the identified differences between preoperational and operational results. The present findings have important implications for test development and test validity.

Algorithms↗

Analysis of the relationship between score components on a standardized patient clinical skills examination.

PURPOSE: This work investigated the reliability of and relationships between individual case and composite scores on a standardized patient clinical skills examination. METHOD: Four hundred ninety two fourth-year U.S. medical students received three scores [data gathering (DG), interpersonal skills (IPS), and written communication (WC)] for each of 10 standardized patient cases. mGENOVA software was used for all analyses. RESULTS: Estimated generalizability coefficients were 0.69, 0.80, and 0.70 for the DG, IPS, and WC scores, respectively. The universe-score correlation between DG and WC was high (.83); those for DG/IPS and IPS/WC were not as strong (0.51 and 0.37, respectively). Task difficulty appears to be modestly but positively related across the three scores. Correlations between the person-by-task effects for DG/IPS and DG/WC were positive yet modest. The estimated generalizability coefficient for a ten-case test using an equally weighted composite DG/WC score was 0.78. CONCLUSIONS: This work allows for interpretation of correlations between (1) proficiencies measured by multiple scores and (2) sources of error that affect those scores as well as for estimation of the reliability of composite scores. Results have important implications for test construction and test validity.

Clinical Competence↗

A demonstration of the impact of response bias on the results of patient satisfaction surveys.

OBJECTIVES: The purposes of the present study were to examine patient satisfaction survey data for evidence of response bias, and to demonstrate, using simulated data, how response bias may impact interpretation of results. DATA SOURCES: Patient satisfaction ratings of primary care providers (family practitioners and general internists) practicing in the context of a group-model health maintenance organization and simulated data generated to be comparable to the actual data. STUDY DESIGN: Correlational analysis of actual patient satisfaction data, followed by a simulation study where response bias was modeled, with comparison of results from biased and unbiased samples. PRINCIPAL FINDINGS: A positive correlation was found between mean patient satisfaction rating and response rate in the actual patient satisfaction data. Simulation results suggest response bias could lead to overestimation of patient satisfaction overall, with this effect greatest for physicians with the lowest satisfaction scores. CONCLUSIONS: Findings suggest that response bias may significantly impact the results of patient satisfaction surveys, leading to overestimation of the level of satisfaction in the patient population overall. Estimates of satisfaction may be most inflated for providers with the least satisfied patients, thereby threatening the validity of provider-level comparisons.

Bias↗