Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Reliability and validity of the objective structured clinical examination in assessing surgical residents.

The purpose of this research was to assess reliability and construct validity of the objective structured clinical examination (OSCE) for evaluating the clinical skills of surgical residents. Reliability refers to precision of the examination and construct validity to the degree to which the examination can discriminate between different levels of training. Twenty-seven second postgraduate year surgical residents took a 38-station OSCE representing seven surgical specialties and that tested history-taking, physical examination, problem-solving, technical skills, and attitudes. A couplet methodology was used wherein a patient encounter was followed by written questions aimed at testing problem-solving and patient management capabilities. Thirty-six standardized patients were trained and 36 surgeons served as examiners marking from structured checklists. Overall reliability, Cronbach's alpha, was 0.89. Construct validity was examined by comparing the scores of the residents with those of a group of graduates of foreign medical schools applying for a "pre-internship" program. For 17 of 19 stations that both groups took, the residents performed significantly better (p less than 0.01). Individual station validity was significant for 32 of 38 stations (r = 0.36 to 0.82, p less than 0.05). The examinations took 3.83 hours at a cost of $5,293 (Canadian dollars). The OSCE has been shown to be a reliable method of assessing clinical skills of surgical residents, construct validity has been established, and inter-item validity confirmed. Reliabilities achieved exceed those traditionally required for both acceptance and promotion decisions.

Analysis of Variance↗

Development and validation of a self-report symptom inventory to assess the severity of oral-pharyngeal dysphagia.

BACKGROUND & AIMS: The aim of this study was to develop and evaluate the validity and reliability of a self-report inventory to measure symptomatic severity of oral-pharyngeal dysphagia. METHODS: Test-retest reliability and face, content, and construct validity of a prototype visual analogue scale inventory were assessed in 45 patients who had stable, neuromyogenic dysphagia. RESULTS: Normalized scores varied over time by -0.5% +/- 17.6% (95% confidence interval, -9.2% to 8.2%). Factor analysis identified a single factor (dysphagia), to which 18 of 19 questions contributed significantly, that accounted for 56% of total variance (P < 0.0001). After deletion of 2 questions with poor face validity and patient compliance, this proportion increased to 59%; mean test-retest change was -2% (95% confidence interval, -11% to 7%); and total score correlated highly with an independent global assessment severity score (r = 0.7; P < 0.0001). A mean 70% reduction in score (P < 0.0001) was observed after surgery in patients with Zenker's diverticulum (discriminant validity). CONCLUSIONS: Applied to patients with neuromyogenic dysphagia, the 17-question inventory shows strong test-retest reliability over 2 weeks as well as face, content, and construct validity. Discriminant validity (responsiveness) has been demonstrated in a population with a correctable, structural cricopharyngeal disorder. Responsiveness of the instrument to treatment in neuromyogenic dysphagia remains to be quantified.

Adult↗

Validating an instrument for clinical supervision using an expert panel.

Use of research instruments standardised in English-speaking countries is commonplace in non-English speaking countries. This article describes technical, linguistic and conceptual issued raised in the translation and validation of an acceptable country-specific instrument. The Manchester Clinical Supervision Scale was validated for use in clinical supervision in Finland using quantitative and qualitative methods. The approach applied was triangulation. The focus of this paper, is to describe the item validation process of the scale using an expert panel (n=11) by means of the content validity index and panel interview. The process resulted in a country-specific 33-item instrument. The study indicated that content validity index and panel discussion are easy, but reliable ways of demonstrating the instrument's content validity.

Cross-Cultural Comparison↗

Factor analysis and validity of the Transplant Evaluation Rating Scale in a large bone marrow transplant sample.

OBJECTIVE: Clinical and methodological challenges are involved in screening bone marrow transplant (BMT) recipients for pretransplant psychosocial adjustment in an attempt to anticipate and prevent behavioral difficulties. Validity of the Transplant Evaluation Rating Scale (TERS), which quantifies the disparate salient elements in a structured clinical assessment, has not been adequately established. This study comprehensively investigated three questions about convergent, internal-structural, and predictive validity of the TERS: how indicative the TERS is of psychosocial difficulties; whether the TERS is uni- or multidimensional; to what degree the TERS predicts long-range adjustment during recovery posttransplant. METHODS: Pre-BMT, 345 consecutive patients were prospectively assessed and completed the MMPI. TERS ratings were assigned retrospectively by two raters (interrater reliability r=.89). RESULTS: The TERS showed good convergent validity relative to MMPI subscales, and a clear, simple, two-factor structure accounting for 47% of the variance. On a subset of our sample (n=29), the factor subscales, "Defiance" and "Emotional Sensitivity," exhibited differential predictive validity to functional status at 1 year posttransplant. CONCLUSIONS: This study, the first large-scale statistical investigation of TERS validity, provided evidence for the validity of the TERS on all three questions. The TERS is indeed indicative of psychosocial risk indexed by MMPI behavioral pathology. It has an understandable, clinically useful factor structure. Its subordinate constructs, Defiance and Emotional Sensitivity, can and should be distinguished conceptually and measured separately. The TERS has clinical utility for specifying behavioral concerns before and guiding proactive intervention after BMT.

Adaptation, Psychological↗

The National Institutes of Health chronic prostatitis symptom index: development and validation of a new outcome measure. Chronic Prostatitis Collaborative Research Network.

PURPOSE: Chronic abacterial prostatitis is a syndrome characterized by pelvic pain and voiding symptoms, which is poorly defined, poorly understood, poorly treated and bothersome. Research and clinical efforts to help men with this syndrome have been hampered by the absence of a widely accepted, reliable and valid instrument to measure symptoms and quality of life impact. We developed a psychometrically valid index of symptoms and quality of life impact for men with chronic prostatitis. MATERIALS AND METHODS: We conducted a structured literature review of previous work to provide a foundation for the new instrument. We then conducted a series of focus groups comprising chronic prostatitis patients at 4 centers in North America, in which we identified the most important symptoms and effects of the condition. The results were used to create an initial draft of 55 questions that were used for formal cognitive testing on chronic prostatitis patients at the same centers. After expert panel review formal validation testing of a revised 21-item draft was performed in a diverse group of chronic prostatitis patients and 2 control groups of benign prostatic hyperplasia patients and healthy men. Based on this validation study, the index was finalized. RESULTS: Analysis yielded an index of 9 items that address 3 different aspects of the chronic prostatitis experience. The primary component was pain, which we captured in 4 items focused on location, severity and frequency. Urinary function, another important component of symptoms, was captured in 2 items (1 irritative and 1 obstructive). Quality of life impact was captured with 3 items about the effect of symptoms on daily activities. The 9 items had high test-retest reliability (r = 0.83 to 0.93) and internal consistency (alpha = 0.86 to 0.91). All but the urinary items discriminated well between men with and without chronic prostatitis. CONCLUSIONS: The National Institutes of Health chronic prostatitis symptom index provides a valid outcome measure for men with chronic prostatitis. The index is psychometrically robust, easily self-administered and highly discriminative. It was formally developed and psychometrically validated, and may be useful in clinical practice as well as research protocols.

Chronic Disease↗

Prediction of outcome in acute lower-gastrointestinal haemorrhage based on an artificial neural network: internal and external validation of a predictive model.

BACKGROUND: Models based on artificial neural networks (ANN) are useful in predicting outcome of various disorders. There is currently no useful predictive model for risk assessment in acute lower-gastrointestinal haemorrhage. We investigated whether ANN models using information available during triage could predict clinical outcome in patients with this disorder. METHODS: ANN and multiple-logistic-regression (MLR) models were constructed from non-endoscopic data of patients admitted with acute lower-gastrointestinal haemorrhage. The performance of ANN in classifying patients into high-risk and low-risk groups was compared with that of another validated scoring system (BLEED), with the outcome variables recurrent bleeding, death, and therapeutic interventions for control of haemorrhage. The ANN models were trained with data from patients admitted to the primary institution during the first 12 months (n=120) and then internally validated with data from patients admitted to the same institution during the next 6 months (n=70). The ANN models were then externally validated and direct comparison made with MLR in patients admitted to an independent institution in another US state (n=142). FINDINGS: Clinical features were similar for training and validation groups. The predictive accuracy of ANN was significantly better than that of BLEED (predictive accuracy in internal validation group for death 87% vs 21%; for recurrent bleeding 89% vs 41%; and for intervention 96% vs 46%) and similar to MLR. During external validation, ANN performed well in predicting death (97%), recurrent bleeding (93%), and need for intervention (94%), and it was superior to MLR (70%, 73%, and 70%, respectively). INTERPRETATION: ANN can accurately predict the outcome for patients presenting with acute lower-gastrointestinal haemorrhage and may be generally useful for the risk stratification of these patients.

Acute Disease↗

Development and validation of the insulin treatment satisfaction questionnaire.

BACKGROUND: Treatment of diabetes mellitus (DM) is complex, requiring multifaceted lifestyle change or regulation and, for many, self-regulation of insulin levels in the blood. Historically, daily insulin treatment has been viewed as burdensome to patients, prompting newer formulations and improved delivery methods. OBJECTIVE: This multicenter, clinical study was designed to develop a conceptually sound, clinically meaningful, and psychometrically valid measure of insulin treatment satisfaction, applicable to a wide range of insulin therapies. METHODS: A 3-phase iterative process was employed to develop and validate the Insulin Treatment Satisfaction Questionnaire (ITSQ): (1) conceptual development of items, (2) preliminary validation among patients with DM, and (3) confirmatory validation among patients with DM. RESULTS: The ITSQ was validated with 170 patients in phase 2 and 402 patients in phase 3. Confirmatory factor analysis produced a 5-factor, 22-item instrument assessing regimen inconvenience, lifestyle flexibility, glycemic control, hypoglycemic control, and satisfaction with the insulin delivery device. Results for reliability and construct validity of the final version were consistent in both samples of patients treated with insulin, with different data collection methods. Internal consistency (using Cronbach alpha coefficient) of the subscales ranged from 0.79 to 0.91. Test-retest reliability (using Spearman rank correlation coefficients) ranged from 0.63 to 0.94. ITSQ scores showed moderate to high correlation with related measures of treatment burden. The ITSQ differentiated among insulin delivery methods, glycosylated hemoglobin values, the number of times the patient required assistance administering insulin, and insulin adherence. CONCLUSION: In our study samples, the ITSQ appeared to be conceptually and psychometrically sound and applicable to a wide range of insulin therapies.

Diabetes Mellitus↗

Validation of a poultry biosecurity survey.

A questionnaire for farm managers was designed, to obtain information regarding biosecurity on Ontario commercial broiler chicken and turkey operations, and then pre-tested. The questions that could be validated were verifiable by seeing the facility, by using farm records or by interviewing technical personnel other than the survey respondent. The survey was validated using a convenience sample of 24 farms from two companies. For 15 questions with dichotomous responses, the sensitivity ranged from 16.7 to 100%; the specificity ranged from 0 to 100%. For example, fences and gates seen during the farm visit were not accurately reported on the survey (poor sensitivity). Chance-corrected agreement was low (kappa < 0.4) for 34 questions, fair to good (0.4 < kappa < 0.8) for 25 questions, and excellent (kappa > 0.8) for seven questions. The percent agreement for questions where only one of the possible options was observed on validation ranged from 60.9 to 100%. Five questions with continuous numeric variables were analysed. A difference was observed (P < 0.1) between the survey and validation data for three questions regarding the number of birds, the bird sources and the downtime between flocks. In spite of pre-testing, the lack of clear wording and the absence of definitions for technical terms appeared to reduce validity. Response bias seems to be an issue with biosecurity surveys. The value of validating questionnaires before their use in epidemiologic research is confirmed.

Animal Husbandry↗

[Validation of ESCAP-CD as an instrument of measure for the evaluation of the quality of life in prostatic cancer. Part 2a].

UNLABELLED: In order to make a measure of quality of life related to health (QLRH) useful in the investigation, it must fulfill the psychometric properties (validity, reliability and sensibility). The selection of an instrument is a job for the clinic that must choose the most effective for each proposed objective. We set out the objectives to validate the ESCAP-CDV in a multicentric study in Andalusia. We studied 88 patients who were submitted to the instrument presented to validation and two more tests recognized already: the QLQ-C30 from EORTC gold standard in Europe in the valuation of the neoplastic patients' quality of life and the KARNOFSKY the most clinic utility index in neoplastic patients, used to correlate the items. RESULTS: Questionnaire acceptance analysis: The difficulty of understanding was greater for QLQ C30 items (6.81%) than ESCAP items (1.98%). The lapse of time needed to carry out the test was shorter in the ESCAP test (9.84 min) than in the QLQ C30 (13.13), test. Structural analysis or internal validity analysis: The homogeneity index of the items is high (alfa of Cronbach = 0.93). The dimensionality proposed is not accepted, due to the existence of some modifications pund in the factorial analysis. Finally, the established dimensions: Physical and Emotional Capacity (PEC), 5 items; General Symptoms (GS), 4 items; Pain (P), 3 items; Ligh Functional Capcity (LFC), 4 items; Serious Functional Capacity (SFC), 2 items; Economic State (ES), 3 items; Social and Family State (SFE), 5 items; Capacity Sexual (CSX), 2 items; Isolated Variables (IV), 2 items; and Specific Questionnaire (P), 6 items. The ESCAP is a scale with a normal distribution. Approach or external validity analysis: The ESCAP test is well correlated with the other two scales. Reliability test retest: The interclass correlation coefficient is 0.94 in the ESCAP, not so in the KARNOFSKY that is 0.77. CONCLUSIONS: The ESCAP-CDV is a new instrument of valuation of the QLRH composed of a general questionnaire and other specific test of prostate cancer. It has turned out to be a very homogeneous scale due to its internal consistence (alfa of Cronbach of 0.93), showing that it has a normal distribution, that correlates correctly with the scales compared and it is a valid scale to measure the prostate cancer patients' quality of life. The ESCAP-CDV has shown to be a scale with a high reliability (0.94), setting up as an instrument not only useful for investigation, but to clinical use, as well.

Humans↗

[Validity of the occupational history assessed by interview].

The study was carried out on 114 workers of a large firm in the city of Barcelona. Its objectives were to measure the validity of a questionnaire on professional history, achieved by telephone interviews, using the firm's employee register and to identify the socio-professional characteristics associated with the firm. The validity of the main and secondary occupation (name and length of the occupation) was assessed by company reported information with that in the firm's registry, which was considered as a reference criterion. The validity of the name of the main and secondary occupation was 88% and 74% respectively and the validity of their length was 57% and 64.4% respectively. The validity of professional history, in as much as the name of the main and secondary occupation was high. The validity of their length was low.

Epidemiologic Methods↗

[Construction and validation of a severity of illness index of patients hospitalized in clinical areas].

Severity of illness indexes for hospitalized patients are frequently used for the assessment of hospital performance. The purpose here was to develop and validate a quantitative index for clinical areas which could measure severity of illness during hospitalization period and would be easy to obtain from the information contained in a regular clinical chart. The construction included item selection and search for items weights. Experts and literature were consulted and 74 clinical records provided empirical information. The result was the proposal of an index with two alternatives: one quantitative and the other ordinal with four levels of severity. Validation included four aspects of validity, general reliability, interrater agreement and internal consistency. Sixteen specialized physicians assessed content and face validity. One hundred clinical records of discharged patients were used to assess criterion, construct validity and reliability. Results show satisfactory validity in almost all aspects. The reliability coefficient was 0.95, Kappa coefficient was 0.4 for the ordinal index and correlation coefficients over 0.93 for the quantitative one. The index is ready for current applications in this context although some suggestions for future improvements were also included.

Adolescent↗

Logistic regression model to predict outcome after in-hospital cardiac arrest: validation, accuracy, sensitivity and specificity.

OBJECTIVE: To develop and validate a logistic regression model to identify predictors of death before hospital discharge after in-hospital cardiac arrest. DESIGN: Retrospective derivation and validation cohorts over two 1 year periods. Data from all in-hospital cardiac arrests in 1986-87 were used to derive a logistic regression model in which the estimated probability of death before hospital discharge was a function of patient and arrest descriptors, major underlying diagnosis, initial cardiac rhythm, and time of year. This model was validated in a separate data set from 1989-90 in the same hospital. Calculated for each case was 95% confidence limits (C.L.) about the estimated probability of death. In addition, accuracy, sensitivity, and specificity of estimated probability of death and lower 95% C.L. of the estimated probability of death in the derivation and validation data sets were calculated. SETTING: 560-bed university teaching hospital. PATIENTS: The derivation data set described 270 cardiac arrests in 197 inpatients. The validation data set described 158 cardiac arrests in 120 inpatients. INTERVENTIONS: none. MEASUREMENTS AND RESULTS: Death before hospital discharge was the main outcome measure. Age, female gender, number of previous cardiac arrests, and electrical mechanical dissociation were significant variables associated with a higher probability of death. Underlying coronary artery disease or valvular heart disease, ventricular tachycardia, and cardiac arrest during the period July-September were significant variables associated with a lower probability of death. Optimal sensitivity and specificity in the validation set were achieved at a cut-off probability of 0.85. CONCLUSIONS: Performance of this logistic regression model depends on the cut-off probability chosen to discriminate between predicted survival and predicted death and on whether the estimated probability or the lower 95% C.L. of the estimated probability is used. This model may inform the development of clinical practice guidelines for patients who are at risk of or who experience in-hospital cardiac arrest.

Confidence Intervals↗

The Chronic Pain Grade questionnaire: validation and reliability in postal research.

The Chronic Pain Grade questionnaire has been proposed as an interview-administered, multi-dimensional measure of chronic pain severity in selected populations with chronic pain in the United States of America. It has not previously been tested in the United Kingdom, in self-completion form or in an unselected general population. We undertook a postal survey to assess its reliability, validity and acceptability in these circumstances, using a general practice population in Scotland, with a practice population of 11202 patients. A random sample of 400 patients aged over 18 was drawn, stratified for age, gender and receipt or non-receipt of regular prescriptions for pain-relieving medication. The dimensions and sub-scales of the Chronic Pain Grade were compared with the SF-36 general health questionnaire and questions relating to duration of any pain and attempts to seek treatment for this. The methodological approach proposed by Streiner and Norman (1989) was used to assess validity and reliability. A response rate of 76% was achieved. Cronbach's alpha was > 0.9 and item-total correlations were all high, indicating good internal consistency and reliability. Validity was confirmed by psychometric testing, including confirmatory factor analysis. Good correlations with comparable dimensions of the SF-36 general health questionnaire confirmed convergent validity. Construct validity was confirmed by testing scores against duration of pain and treatment sought for pain. We concluded that the Chronic Pain Grade questionnaire is a useful, reliable and valid measure of severity of chronic pain. It translates well into UK English and is acceptable in general population postal research.

Chronic Disease↗

Some empirical evidence regarding the validity of the Spanish version of the McGill Pain Questionnaire (MPQ-SV).

Despite the fact that the McGill Pain Questionnaire (MPQ) is a useful pain assessment tool with widespread acceptance, empirical analyses have questioned its validity because they have not consistently supported the three a priori factors that guided its construction. The Spanish version that has followed the most systematic and rigorous reconstruction process (Lázaro C, Bosch F, Torrubia R, Banos JE. The development of a Spanish Questionnaire for assessing pain: preliminary data concerning reliability and validity. Eur J Psychol Assess, 1994;10:145-151) lacks evidence to support its construct validity. In the present study, the internal structure of the Spanish version of the McGill Pain Questionnaire (Lázaro C, Bosch F, Torrubia R, Banos JE. The development of a Spanish Questionnaire for assessing pain: preliminary data concerning reliability and validity. Eur J Psychol Assess, 1994;10:145-151) was examined in a sample of 202 acute pain patients and 207 chronic pain patients. Confirmatory factor analyses were carried out to compare alternative models postulating different internal structures (one-factor model, the classic three-factor model, and the semantic model inspired by the alternative structure found by Donaldson in 1995 (Donaldson GW. The factorial structure and stability of the McGill Pain Questionnaire in patients experiencing oral mucositis following bone marrow transplantation. Pain 1995;62:101-109)). Results from the LISREL CFA analysis indicated that the semantic model fitted better than the other models. On the other hand, intercorrelations between scales were smaller than the reliability indexes. In relation to concurrent evidence, significant correlations (0.001) were found between each subscale and the criteria measurements of every pain dimension. Only the affective subscale presented discriminant validity. Evidence supports the validity of the affective and sensory subscales but not the evaluative scale.

Adolescent↗

Validation and acceptance of new and revised tests: a flexible but transparent process.

Scientific validation is a necessary prerequisite for the regulatory acceptance of new toxicological tests, and for acceptance as Organisation for Economic Co-operation and Development Test Guidelines. Validation is defined as the determination of the reliability and relevance of a test method for a particular purpose. Reliability refers to the reproducibility of the test within and among laboratories. The relevance of a test addresses how well it measures or predicts what it is supposed to measure or predict. There are no fixed values for how reliable or relevant a test should be. Acceptable values depend on the type of test, the inherent variability in the biological systems used, and what regulatory decisions the test will be asked to support. The validation of in vitro and in vivo tests should follow the same general principles. The validation process should be flexible, according to the type of test being evaluated and its proposed uses. If a test is designed to replace a test currently in use, its reliability and relevance should be measured against that reference test or against the effect that the test is designed to predict. For the evaluation of health effects where adequate human data are available, the performance of the new test should be measured against the relevant effect in humans, rather than a surrogate species. All proceedings from the validation process and the subsequent peer-review of the method should be transparent, i.e., publicly available. It is also recommended that the peer-review of the validation study be accessible to the interested public.

Guidelines as Topic↗

Validation of liquid chromatographic and gas chromatographic methods. Applications to pharmacokinetics.

Validations of analytical methods are important for the generation of data for bioavailability, bioequivalence and pharmacokinetic studies. It is essential to use well defined and fully validated analytical methods to obtain reliable results that can be satisfactorily interpreted. This manuscript is intended to provide guiding principles for the evaluation of a method's overall performance. For this purpose, all of the variables of the method are considered, including sampling procedure, sample preparation, chromatographic separation, detection and data evaluation. The criteria considered are as follows: stability, selectivity, limits of quantification and of detection, accuracy, precision, linearity, recovery and ruggedness. Models used for analytical calibration curves are explained in term of validity and limitations, along with a presentation of the most common statistical considerations used to validate the model. Appropriate means of testing precision and accuracy, the most important factors in assessing method quality, are presented. Other issues, such as re-validation, cross-validation, partial sample volume, endogenous drugs and biological matrix of limited availability, are also discussed.

Calibration↗

Validation of immunoassays for bioanalysis: a pharmaceutical industry perspective.

Immunoassays are bioanalytical methods in which quantitation of the analyte depends on the reaction of an antigen (analyte) and an antibody. Although applicable to the analysis of both low molecular weight xenobiotic and macromolecular drugs, these procedures currently find most consistent application in the pharmaceutical industry to the quantitation of protein molecules. Immunoassays are also frequently applied in such important areas as the quantitation of biomarker molecules which indicate disease progression or regression, and antibodies elicited in response to treatment with macromolecular therapeutic drug candidates. Currently available guidance documents dealing with the validation of bioanalytical methods address immunoassays in only a limited way. This review highlights some of the differences between immunoassays and chromatographic assays, and presents some recommendations for specific aspects of immunoassay validation. Immunoassay calibration curves are inherently nonlinear, and require nonlinear curve fitting algorithms for best description of experimental data. Demonstration of specificity of the immunoassay for the analyte of interest is critical because most immunoassays are not preceded by extraction of the analyte from the matrix of interest. Since the core of the assay is an antigen-antibody reaction, immunoassays may be less precise than chromatographic assays; thus, criteria for accuracy (mean bias) and precision, both in pre-study validation experiments and in the analysis of in-study quality control samples, should be more lenient than for chromatographic assays. Application of the SFSTP (Societe Francaise Sciences et Techniques Pharmaceutiques) confidence interval approach for evaluating the total error (including both accuracy and precision) of results from validation samples is recommended in considering the acceptance/rejection of an immunoassay procedure resulting from validation experiments. These recommendations for immunoassay validation are presented in the hope that their consideration may result in the production of consistently higher quality data from the application of these methods.

Algorithms↗

Internal validation of predictive models: efficiency of some procedures for logistic regression analysis.

The performance of a predictive model is overestimated when simply determined on the sample of subjects that was used to construct the model. Several internal validation methods are available that aim to provide a more accurate estimate of model performance in new subjects. We evaluated several variants of split-sample, cross-validation and bootstrapping methods with a logistic regression model that included eight predictors for 30-day mortality after an acute myocardial infarction. Random samples with a size between n = 572 and n = 9165 were drawn from a large data set (GUSTO-I; n = 40,830; 2851 deaths) to reflect modeling in data sets with between 5 and 80 events per variable. Independent performance was determined on the remaining subjects. Performance measures included discriminative ability, calibration and overall accuracy. We found that split-sample analyses gave overly pessimistic estimates of performance, with large variability. Cross-validation on 10% of the sample had low bias and low variability, but was not suitable for all performance measures. Internal validity could best be estimated with bootstrapping, which provided stable estimates with low bias. We conclude that split-sample validation is inefficient, and recommend bootstrapping for estimation of internal validity of a predictive logistic regression model.

Aged↗