Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Patient management problems. Issues of validity.

Patient management problems (PMP) are being used in medical examinations with increasing frequency despite evidence which throws doubt on their validity as measures of clinical competence. This study investigated the construct validity of a PMP constructed in both written and interview formats. Each test was administered to groups of students of different seniorities and to two groups of doctors, interns and post-interns. The pattern of scores for the different groups was not that expected of a valid test of competence. The most competent groups (the post-interns) generally scored less well on the calculated indices than the senior students and interns. These findings were similar for both formats of the test so cueing was not thought to be the major factor. It appears that the scoring system is at fault. A comparison of performance on the written and interview (uncued) formats showed that many more options were chosen by all groups tested on the written PMP. It was concluded that written PMPs cannot yet be regarded as a valid simulation of clinical performance. Although content validity is high this does not appear to be so for construct validity or concurrent validity.

Australia↗

What is the validity evidence for assessments of clinical teaching?

BACKGROUND: Although a variety of validity evidence should be utilized when evaluating assessment tools, a review of teaching assessments suggested that authors pursue a limited range of validity evidence. OBJECTIVES: To develop a method for rating validity evidence and to quantify the evidence supporting scores from existing clinical teaching assessment instruments. DESIGN: A comprehensive search yielded 22 articles on clinical teaching assessments. Using standards outlined by the American Psychological and Education Research Associations, we developed a method for rating the 5 categories of validity evidence reported in each article. We then quantified the validity evidence by summing the ratings for each category. We also calculated weighted kappa coefficients to determine interrater reliabilities for each category of validity evidence. MAIN RESULTS: Content and Internal Structure evidence received the highest ratings (27 and 32, respectively, of 44 possible). Relation to Other Variables, Consequences, and Response Process received the lowest ratings (9, 2, and 2, respectively). Interrater reliability was good for Content, Internal Structure, and Relation to Other Variables (kappa range 0.52 to 0.96, all P values < .01), but poor for Consequences and Response Process. CONCLUSIONS: Content and Internal Structure evidence is well represented among published assessments of clinical teaching. Evidence for Relation to Other Variables, Consequences, and Response Process receive little attention, and future research should emphasize these categories. The low interrater reliability for Response Process and Consequences likely reflects the scarcity of reported evidence. With further development, our method for rating the validity evidence should prove useful in various settings.

Education, Medical↗

Validation of two prognostic models predicting outcome at two years after diagnosis in a new cohort of children with epilepsy: the Dutch Study of Epilepsy in Childhood.

PURPOSE: To validate two prognostic models for childhood-onset epilepsy designed to predict a terminal remission of <6 months at 2 years after diagnosis in children referred to the hospital. METHODS: A hospital-based cohort of children with newly diagnosed epilepsy was recruited and followed up for 2 years to validate previously developed models. One model was based on variables collected at intake, and the other was based on intake variables plus variables collected during the first 6 months of follow-up. The accuracy of both models was estimated by measuring the area under the receiver-operant-characteristic curves (ROC area). RESULTS: The ROC area of the model developed with intake variables was 0.69 [95% confidence interval (CI), 0.64-0.74] for the original cohort and 0.62 (95% CI, 0.55-0.69) for the validation cohort. The best combination of sensitivity and specificity for the original cohort was 61.6% and 69.1%, whereas it was 60.0% and 61.4% for the validation cohort. For the model with intake and 6-month variables combined, the ROC area was 0.78 (95% CI, 0.73-0.82) for the original cohort and 0.71 (95% CI, 0.64-0.78) for the validation cohort. The sensitivity and specificity were 72.6% and 73.1%, respectively, for the original cohort and 67.4% and 60.2%, respectively, for the validation cohort. CONCLUSIONS: Although both models predict outcome better than chance, they are insufficiently accurate to be of practical value. Both models performed marginally less well with the validation cohort than with the original cohort, but in both instances, the model based on intake and 6-month variables was more accurate.

Adolescent↗

A comparison of individual genotyping and pooled DNA analysis for polymorphism validation prior to large-scale genetic studies.

Polymorphism validation is an important issue in genetic studies because only polymorphic markers provide useful information. We analyzed genetic data for 180 SNPs in the human major histocompatibility complex region in Caucasian and Taiwanese populations, and evaluated ethnic heterogeneity between these populations to illustrate the importance of polymorphism validation. An initial individual genotyping experiment (IGE) with 95 samples was compared with a DNA pooling allele-typing experiment (PAE) of 630 individuals for polymorphism validation based on authentic data sets. Afterwards, all samples were genotyped individually in a confirmation study. Under narrow (broad) polymorphism criteria, 24 (41) polymorphic SNPs in Caucasians could not be validated in the Taiwanese population, suggesting a 13% (23%) inconsistency rate and revealing a strong discrepancy between genetic backgrounds, probably due to ethnic heterogeneity. IGE yielded high sensitivity and specificity for polymorphism validation, but may be sensitive to sampling variation. PAE showed high sensitivity (97%) and specificity (100%) using a narrow polymorphism criterion, but reduced specificity (83%) using a broad criterion. Public domain polymorphism databases should therefore be used with caution and polymorphism validation should be performed routinely prior to conducting large-scale genetic studies. PAE is a cost-saving, reliable alternative to IGE for polymorphism validation, especially for a stringent polymorphism criterion.

Asian People↗

Statistical methodology: II. Reliability and validity assessment in study design, Part B.

Validity measures the correspondence between a test and other purported measures of the same or similar qualities. When a reference standard exists, a criterion-based validity coefficient can be calculated. If no such standard is available, the concepts of content and construct validity may be used, but quantitative analysis may not be possible. The Pearson and Spearman tests of correlation are often used to assess the correspondence between tests, but do not account for measurement biases and may yield misleading results. Techniques that measure interest differences may be more meaningful in validity assessment, and the kappa statistic is useful for analyzing categorical variables. Questionnaires often can be designed to allow quantitative assessment of reliability and validity, although this may be difficult. Inclusion of homogeneous questions is necessary to assess reliability. Analysis is enhanced by using Likert scales or similar techniques that yield ordinal data. Validity assessment of questionnaires requires careful definition of the scope of the test and comparison with previously validated tools.

Humans↗

Development and content validity testing of a comprehensive classification of diagnoses for pediatric nurse practitioners.

Pediatric nurse practitioners (PNPs) need an integrated, comprehensive classification that includes nursing, disease, and developmental diagnoses to effectively describe their practice. No such classification exists. Further, methodologic studies to help evaluate the content validity of any nursing taxonomy are unavailable. A conceptual framework was derived. Then 178 diagnoses from the North American Nursing Diagnosis Association (NANDA) 1986 list, selected diagnoses from the International Classification of Diseases, the Diagnostic and Statistical Manual, Third Revision, and others were selected. This framework identified and listed, with definitions, three domains of diagnoses: Developmental Problems, Diseases, and Daily Living Problems. The diagnoses were ranked using a 4-point scale (4 = highly related to 1 = not related) and were placed into the three domains. The rating scale was assigned by a panel of eight expert pediatric nurses. Diagnoses that were assigned to the Daily Living Problems domain were then sorted into the 11 Functional Health patterns described by Gordon (1987). Reliability was measured using proportions of agreement and Kappas. Content validity of the groups created was measured using indices of content validity and average congruency percentages. The experts used a new method to sort the diagnoses in a new way that decreased overlaps among the domains. The Developmental and Disease domains were judged reliable and valid. The Daily Living domain of nursing diagnoses showed marginally acceptable validity with acceptable reliability. Six Functional Health Patterns were judged reliable and valid, mixed results were determined for four categories, and the Coping/Stress Tolerance category was judged reliable but not valid using either test. There were considerable differences between the panel's, Gordon's (1987), and NANDA's clustering of NANDA diagnoses. This study defines the diagnostic practice of nurses from a holistic, patient-centered perspective. It is the first study to use quantitative methods to test a diagnostic classification system for nursing. The classification model could also be adapted for other nurse specialties.

Humans↗

Greek brief pain inventory: validation and utility in cancer pain.

OBJECTIVE: The Brief Pain Inventory (BPI) is a pain assessment tool. It has been translated into and validated in several languages. The purpose of this study was the translation into and validation of the BPI in Greek. Moreover, we wanted to detect cultural and social differences, if any, of pain interference in patients' lives. METHODS: The translation and validation of the inventory took place at the Areteion Hospital. The final validation sample consisted of 220 cancer patients (123 males, 97 females, age range 21-87 years, mean age 61.3). Primary cancer locations were lung 25.6%, gastrointestinal tract 25.6%, breast 11.5%, prostate 7.07%, gynecological cancers 9.6% and others 20.57%. The patients themselves completed the majority of the Greek BPI (G-BPI) papers. The pain management index (PMI) was also calculated in order to assess the adequacy of pain treatment. Assessing the reliability and the validity made the actual validation of the G-BPI. RESULTS: Pain severity and pain management: 147 patients reported severe pain, 48 patients moderate, and 25 patients mild pain (mean average pain 6.22). From these patients only 21 were found on strong and 33 on weak opioid treatment, while 166 patients were found on no opioid analgesic treatment. In agreement with these data is the PMI which was positive only for 9 patients, while 44 patients had PMI = 0 and all the others had negative PMI scores. Reliability and Validity of the G-BPI: Coefficient alphas were 0.849 for the interference items and 0.887 for the severity items. Additionally, the factor analysis of the G-BPI items results in a two-factor solution, that satisfies the criteria of reproducibility, interpretability and confirmatory setting. CONCLUSION: This study shows the efficacy of the G-BPI for the assessment of pain severity as well as the pain management in Greece, and therefore its utility in improving the analgesic treatment outcome in Greek patients.

Activities of Daily Living↗

Unified Neurological Stroke Scale is valid in ischemic and hemorrhagic stroke.

BACKGROUND AND PURPOSE: The growing interest in testing new therapeutic agents for acute brain injury has lead to increased use of stroke scales. The reliability and validity of these measures need to be examined more completely. We used structural equation modeling, a technique that merges the analytic procedures of factor analysis and multiple regression, to examine the reliability and construct validity of the Middle Cerebral Artery Neurological Scale and the Scandinavian Neurological Stroke Scale used together as the Unified Neurological Stroke Scale. We also analyzed the predictive validity, sensitivity, and specificity of the scales in predicting mortality and functional outcome. METHODS: We prospectively studied 84 consecutive patients admitted to a neurology/neurosurgery intensive care unit with intracerebral hemorrhage (n = 30), subarachnoid hemorrhage (n = 15), ischemic stroke (n = 15), and traumatic brain injury (n = 24). Patients were evaluated within 24 hours of admission and at 48-hour intervals until intensive care unit discharge. A total of 386 assessments were obtained. The Functional Independence Measure was administered by telephone 3 months after hospital discharge. RESULTS: High levels of reliability and construct validity were observed for the majority of the Unified Stroke Scale items. Facial palsy and eye movement items had the lowest reliability and validity. Both the Middle Cerebral Artery and Scandinavian Scales were significant predictors of outcome. Sensitivity and specificity varied by diagnosis. Predictive validity of functional outcome was best in groups with ischemic and hemorrhagic stroke rather than traumatic brain injury and subarachnoid hemorrhage. CONCLUSIONS: The Unified Stroke Scale demonstrates reliability and construct and predictive validity, and its use is supported in ischemic and hemorrhagic stroke. Structural equation modeling is an appropriate technique for use with scales of this type.

Activities of Daily Living↗

Reducing MMPI-2 defensiveness: the effect of specialized instructions on retest validity in a job applicant sample.

The MMPI-2 is often used for screening job applicants when public safety or security are at risk. Inherent in such applications is concern for profile validity and test defensiveness. In this study, we examine the impact of revised instructions on profile validity for a group of job applicants who initially produced invalid profiles. Participants were 271 male applicants for airline pilot positions. Of these, 72 produced invalid defensive MMPI-2 profiles during preemployment screening. The MMPI-2 was readministered to these applicants with instructions informing them of validity scales and instructing them to respond in a more open, honest manner. Comparisons were made between valid and invalid profiles for initial administrations and between valid and invalid profiles at readministration. Some clinical scales were more elevated for valid, nondefensive profiles. Most content scales showed more elevation for valid profiles, and 12% of the applicants who were retested produced significant elevations (T>or=65) on the content scales. Profiles were similar to those produced by employed pilots of a previous study.

Journal Article↗

Reliability and validity of the perioperative opioid-related symptom distress scale.

A reduction in opioid use may reduce the incidence and severity of opioid-related side effects. However, no published studies have demonstrated this relationship. In a prospective, placebo-controlled, randomized trial of analgesia for laparoscopic cholecystectomy, we validated an opioid-related symptom distress scale (SDS) questionnaire and clinically meaningful events (CMEs). A total of 193 patients completed the SDS questionnaire every 24 h after discharge for 7 days. This analysis was based on data from Day 1 only. The SDS assessed 12 common opioid-related symptoms, including nausea, vomiting, and difficulty passing urine, by 3 ordinal measures: frequency, severity, and bothersomeness. Patients with responses of "frequently" to "almost constantly," "moderate" to "very severe," or "quite a bit" to "very much bothered" were considered to have a CME. A detailed postoperative recovery survey of patient functional status and experience of adverse effects was used to validate the SDS. Validation measures in the recovery survey were categorized as nonspecific (e.g., level of normal activities) and specific (e.g., number of times vomited in 24 h, minutes of nausea in 24 h, and ability to void normally). SDS scores and CMEs for nausea, vomiting, and difficulty passing urine were strongly associated with three related validation measures from the recovery survey: minutes of nausea within 24 h, number of times vomited within 24 h, and ability to void normally, respectively (P < 0.0001). There was also a strong association between SDS scores and CMEs for nausea, vomiting, and voiding and general recovery validation measures, although the association was significantly weaker than that for symptom-specific validation measures. CMEs for nausea, vomiting, and voiding showed a high specificity and lower sensitivity with directly assessed responses. The SDS questionnaire and CMEs are valid tools for assessing postoperative opioid-related symptoms after laparoscopic cholecystectomy. Symptoms defined as CMEs through the SDS may be more sensitive than those identified by direct assessment.

Adult↗

Validation of alternative methods for toxicity testing.

Before nonanimal toxicity tests may be officially accepted by regulatory agencies, it is generally agreed that the validity of the new methods must be demonstrated in an independent, scientifically sound validation program. Validation has been defined as the demonstration of the reliability and relevance of a test method for a particular purpose. This paper provides a brief review of the development of the theoretical aspects of the validation process and updates current thinking about objectively testing the performance of an alternative method in a validation study. Validation of alternative methods for eye irritation testing is a specific example illustrating important concepts. Although discussion focuses on the validation of alternative methods intended to replace current in vivo toxicity tests, the procedures can be used to assess the performance of alternative methods intended for other uses.

Animal Testing Alternatives↗

Reliability, validity, and responsiveness of the Lysholm knee score and Tegner activity scale for patients with meniscal injury of the knee.

BACKGROUND: A torn meniscus is one of the most common indications for knee surgery. The purpose of this study was to determine the psychometric properties of the Lysholm knee score and the Tegner activity scale when used for patients with a meniscal injury of the knee. METHODS: Test-retest reliability, content validity, criterion validity, construct validity, and responsiveness to change were determined for the Lysholm score and the Tegner activity scale. Test-retest reliability was measured in a group of 122 patients at least two years after they had undergone surgery for a meniscal lesion. This group completed a follow-up form and then completed it again within four weeks. The other tests were performed in a group of 191 patients who had only a meniscal lesion at the time of the surgery and a group of 477 patients who had a meniscal lesion and other intra-articular lesions. RESULTS: The overall Lysholm score showed acceptable test-retest reliability, floor and ceiling effects, criterion validity, construct validity, and responsiveness to change. There were unacceptable ceiling effects (>30%) for the Lysholm domains of limp, instability, support, and locking. The Tegner activity scale showed acceptable test-retest reliability, floor and ceiling effects, criterion validity, construct validity, and responsiveness to change. CONCLUSIONS: Overall, the Lysholm knee score and the Tegner activity scale demonstrated acceptable psychometric performances as outcome measures for patients with a meniscal injury of the knee. Some domains of the Lysholm score showed suboptimal performance, and the Tegner scale had only a moderate effect size. Psychometric testing of other condition-specific knee instruments for patients with a meniscal lesion of the knee would be helpful to allow comparison of the properties of the various knee instruments.

Humans↗

The validity and reproducibility of a work productivity and activity impairment instrument.

The construct validity of a quantitative work productivity and activity impairment (WPAI) measure of health outcomes was tested for use in clinical trials, along with its reproducibility when administered by 2 different methods. 106 employed individuals affected by a health problem were randomised to receive either 2 self-administered questionnaires (self administration) or one self-administered questionnaire followed by a telephone interview (interviewer administration). Construct validity of the WPAI measures of time missed from work, impairment of work and regular activities due to overall health and symptoms, were assessed relative to measures of general health perceptions, role (physical), role (emotional), pain, symptom severity and global measures of work and interference with regular activity. Multivariate linear regression models were used to explain the variance in work productivity and regular activity by validation measures. Data generated by interviewer-administration of the WPAI had higher construct validity and fewer omissions than that obtained by self-administration of the instrument. All measures of work productivity and activity impairment were positively correlated with measures which had proven construct validity. These validation measures explained 54 to 64% of variance (p less than 0.0001) in productivity and activity impairment variables of the WPAI. Overall work productivity (health and symptom) was significantly related to general health perceptions and the global measures of interference with regular activity. The self-administered questionnaire had adequate reproducibility but less construct validity than interviewer administration. Both administration methods of the WPAI warrant further evaluation as a measure of morbidity.

Absenteeism↗

[Validation of a measurement scale: example of a French Adverse Drug Reactions Preventability Scale].

Adverse drug reactions (ADRs) have been recognised as an important cause of hospital admission. Most of these drug-related admissions were expected ADRs and, thus, partly preventable. However, as far as we know, the assessment of the preventability of ADRs was addressed in only two studies performed in France. In contrast, several other studies have been performed, mainly in the USA, and using different methods of assessing preventability. None of these methods were clearly evaluated with regard to reproducibility, validity or relevance. The purpose of this study was to initiate the validation of a French preventability scale. Here, we propose the first two phases of validation: the content validity and reliability of the scale. A working group of pharmacovigilance experts has been specifically established for this purpose. The content validity was assessed by collecting items representative of preventability. The choice and the formulation of items and a proposal of a score (global and for each item) were adopted after the consensus of the experts. A definitive version of the ADR preventability scale was used for the assessment of reliability. During the second phase, experts independently tested the new scale from observations of ADRs (49 central nervous system haemorrhages with antivitamine K). The concordance of the experts' judgements was calculated using two statistical methods (Kappa statistic and correlation coefficient). The content validity phase was performed during several workshops where experts discussed the choice and formulation of the best items. We decided to construct a scale with a small number of items, allowing a rapid evaluation of the preventability of ADRs. On the basis of a global score, four categories of preventability of ADRs ("preventable", "potentially preventable", "unclassable", "not preventable" ADRs) were proposed. The agreement of experts regarding the global score was low, with a poor correlation coefficient value (coefficient interclass = 0.491). Classification of ADRs in the four categories by the experts showed discrepancies (Kappa = 0.1136). The preventability assessment using this scale was feasible, although poor concordance between the judges has raised some questions. Several experts found use of this scale difficult in terms of a clear understanding of the items, and found that two of them were redundant. We have oversimplified some items and revision of their formulation will be necessary. Moreover, most of ADR notifications were poorly documented, resulting in a frequent choice of an "unevaluable" item. This represented an important bias in the calculation of the global score. This experience suggests the need for further studies to improve this French ADR preventability scale and validate it in differing circumstances, in order to provide a useful tool to enhance the rational use of drugs.

Adverse Drug Reaction Reporting Systems↗

Validation of commercial DNA tests for quantitative beef quality traits.

Associations between 3 commercially available genetic marker panels (GeneSTAR Quality Grade, GeneSTAR Tenderness, and Igenity Tender-GENE) and quantitative beef traits were validated by the US National Beef Cattle Evaluation Consortium. Validation was interpreted to be the independent confirmation of the associations between genetic tests and phenotypes, as claimed by the commercial genotyping companies. Validation of the quality grade test (GeneSTAR Quality Grade) was carried out on 400 Charolais x Angus crossbred cattle, and validation of the tenderness tests (GeneSTAR Tenderness and Igenity Tender-GENE) was carried out on over 1,000 Bos taurus and Bos indicus cattle. The GeneSTAR Quality Grade marker panel is composed of 2 markers (TG5, a SNP upstream from the start of the first exon of thyroglobulin, and QG2, an anonymous SNP) and is being marketed as a test associated with marbling and quality grade. In this validation study, the genotype results from this test were not associated with marbling score; however, the association of substituting favorable alleles of the marker panel with increased quality grade (percentage of cattle grading Choice or Prime) approached significance (P < or = 0.06), mainly due to the effect of 1 of the 2 markers. The GeneSTAR Tenderness and Igenity TenderGENE marker panels are being marketed as tests associated with meat tenderness, as assessed by Warner-Bratzler shear force. These marker panels share 2 common mu-calpain SNP, but each has a different calpastatin SNP. In both panels, there were highly significant (P < 0.001) associations of the calpastatin marker and the mu-calpain haplotype with tenderness. The genotypic effects of the 2 tenderness panels were similar to each other, with a 1 kg difference in Warner-Bratzler shear force being observed between the most and least tender genotypes. Unbiased and independent validation studies are important to help build confidence in marker technology and also as a potential source of data required to enable the integration of marker data into genetic evaluations. As DNA tests associated with more beef production traits enter the marketplace, it will become increasingly important, and likely more difficult, to find independent populations with suitable phenotypes for validation studies.

Alleles↗

Effects of three interview factors on the validity of alcohol abusers' self-reports.

Using 54 outpatient male court-referred alcohol abusers as subjects, this study investigated the effects of three different interview factors--interview setting (group vs individual), method of interview administration (self vs other), and question type (alcohol vs nonalcohol vs demographic)--on the validity of alcohol abusers' self-reports of verifiable life events. Overall, subjects gave relatively valid self-reports, and when answers were invalid they were more often overreported than underreported. Of the three question types, demographic questions were answered the most validly. The validity of subjects' answers was not differentially affected by whether they answered the questions themselves or were interviewed by an experimenter. While subjects who were interviewed individually gave significantly more valid responses to questions than subjects interviewed in a group setting, the difference (5%) was not great. Given that the overall validity rate was quite high for both groups, consideration must be given to whether it is worth the added time of interviewed subjects individually as compared to interviewing subjects in groups and settling for a slightly lower rate of validity.

Alcohol Drinking↗

Robustness of validation criteria in the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology.

CONTEXT: Field validation of slides used in gynecologic cytology proficiency testing has surfaced as an important issue. Although the precision of diagnoses in peer-reviewed educational programs has been examined, the robustness of the validation criteria for specific types of interpretations used in proficiency testing has not been previously studied. OBJECTIVE: To evaluate the robustness of validation criteria for slides entering an educational slide program. DESIGN: We reviewed the results of the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology and compared the robustness of validation criteria for different reference diagnoses, using a total of 16,948 circulating slides. RESULTS: Validation criteria could be divided into 2 significantly different groups. The criteria for herpes, Trichomonas, squamous cell carcinoma, and adenocarcinoma were significantly more robust than the diagnoses of unsatisfactory; negative for intraepithelial lesion and malignancy, not otherwise specified; low-grade squamous intraepithelial lesion; and high-grade squamous intraepithelial lesion (P < .001). CONCLUSIONS: The validation criteria used in the College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology show 2 different levels of robustness or redundancy. These results have implications for the design of fair proficiency tests. Proficiency testing can be designed with the necessary number of reviews needed for slide validation.

Clinical Competence↗

Development of a new prognostic system and validation of APACHE II for surgical ICU mortality: a multicenter study in Taiwan.

BACKGROUND: To develop and to validate a new prognostic prediction system for patients admitted to the surgical intensive care unit (ICU), and to compare its performance with the Acute Physiology and Chronic Health Evaluation (APACHE) II system. METHODS: The database was derived from three surgical ICUs in three hospitals. For each patient, demographic data, diagnosis, APACHE II score and hospital survival data were collected. The accuracy in outcome prediction of the APACHE II was assessed by means of receiver operating characteristic (ROC) analysis. The new prognostic system was developed by using a multiple logistic regression in the developmental data set and validated with the validation data set. RESULTS: A total of 1,248 patients were included from three ICUs. The area under the ROC curve was 0.74 for the APACHE II score. The new prognostic system includes 18 variables. Goodness-of-fit tests indicated that the model performed well in the developmental and validation samples (p = 0.235 in the developmental data set and p = 0.297 in the validation set). The area under the ROC curve was 0.84 in the developmental sample and 0.77 in the validation sample for the new prognostic score. The area under the ROC curve was 0.71 in the validation sample for the APACHE II score. CONCLUSIONS: Although APACHE II correlates with mortality for surgical ICU patients in Taiwan, its accuracy is not as good as in the original study. Mortality prediction performance improved with the use of the new, local scoring system.

Adult↗