Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Independent prospective multicenter validation of biochemical markers (fibrotest-actitest) for the prediction of liver fibrosis and activity in patients with chronic hepatitis C: the fibropaca study.

OBJECTIVES: Fibrotest (FT) and Actitest (AT) are biochemical markers of fibrosis and activity for use as a non-invasive alternative to liver biopsy in patients with chronic hepatitis C virus (HCV). The aim of this study was to perform an external validation of FT and AT and to study the discordances between FT/AT and liver biopsy in patients with chronic hepatitis C. METHODS: A total of 519 consecutive patients with chronic HCV were prospectively included in five centers, with liver biopsy and biochemical markers taken at the same day. Fifteen patients were excluded because their biopsies could not be interpreted. Diagnostic accuracies were assessed by receiver operating characteristic (ROC) curve analysis. RESULTS: Median biopsy size was 15 mm (range: 2-58), with 9 portal tracts (1-37) and 1 fragment (1-12). 46% (230/504) were classified F2-F4 in fibrosis and 39% A2-A3 in activity. FT area under ROC curve for diagnosis of activity (A2-A3), significant fibrosis (F2-F4), and severe fibrosis (F3-F4) were 0.73 [0.69-0.77], 0.79 [0.75-0.82], and 0.80 [0.76-0.83], respectively. Among the 92 patients (18%) with 2 fibrosis stages of discordance between FT and biopsy, the discordance was attributable to FT in 5% of cases, to biopsy in 4%, and undetermined in 9%. CONCLUSIONS: This prospective independent and multicenter study confirms the diagnostic value of FT and AT found in the princeps study and suggests that FT and AT can be an alternative to biopsy in most patients with chronic HCV.

Adolescent↗

Beyond frequency and severity: development and validation of the brief coercion and conflict scales.

Responding to calls for improved measurement in the field of domestic violence, this paper reports the development and initial validation of the Brief Coercion and Conflict Scales in a sample of incarcerated women. Confirmatory factor analyses tested the scales hypothesized structure and supported coercion and conflict as moderately and positively related but distinct constructs. Although women reported experiencing both conflict and coercion in their most recent relationship before incarceration, they reported that their experiences were more often marked by interpersonal conflict than by coercion. Further, the coercion and conflict scales differentially predicted women's behavioral and psychological responses to abuse. Only coercion consistently predicted strategic responses and posttraumatic stress symptoms. Overall, findings support the instrument as a viable option, but further psychometric evaluation of internal and external validity with additional samples is warranted.

Adult↗

Metrics for external model evaluation with an application to the population pharmacokinetics of gliclazide.

PURPOSE: The aim of this study is to define and illustrate metrics for the external evaluation of a population model. MATERIALS AND METHODS: In this paper, several types of metrics are defined: based on observations (standardized prediction error with or without simulation and normalized prediction distribution error); based on hyperparameters (with or without simulation); based on the likelihood of the model. All the metrics described above are applied to evaluate a model built from two phase II studies of gliclazide. A real phase I dataset and two datasets simulated with the real dataset design are used as external validation datasets to show and compare how metrics are able to detect and explain potential adequacies or inadequacies of the model. RESULTS: Normalized prediction errors calculated without any approximation, and metrics based on hyperparameters or on objective function have good theoretical properties to be used for external model evaluation and showed satisfactory behaviour in the simulation study. CONCLUSIONS: For external model evaluation, prediction distribution errors are recommended when the aim is to use the model to simulate data. Metrics through hyperparameters should be preferred when the aim is to compare two populations and metrics based on the objective function are useful during the model building process.

Algorithms↗

Prospective validation of the CLIP score: a new prognostic system for patients with cirrhosis and hepatocellular carcinoma. The Cancer of the Liver Italian Program (CLIP) Investigators.

Prognosis of patients with cirrhosis and hepatocellular carcinoma (HCC) depends on both residual liver function and tumor extension. The CLIP score includes Child-Pugh stage, tumor morphology and extension, serum alfa-fetoprotein (AFP) levels, and portal vein thrombosis. We externally validated the CLIP score and compared its discriminatory ability and predictive power with that of the Okuda staging system in 196 patients with cirrhosis and HCC prospectively enrolled in a randomized trial. No significant associations were found between the CLIP score and the age, sex, and pattern of viral infection. There was a strong correlation between the CLIP score and the Okuda stage. As of June 1999, 150 patients (76.5%) had died. Median survival time was 11 months, overall, and it was 36, 22, 9, 7, and 3 months for CLIP categories 0, 1, 2, 3, and 4 to 6, respectively. In multivariate analysis, the CLIP score had additional explanatory power above that of the Okuda stage. This was true for both patients treated with locoregional therapy or not. A quantitative estimation of 2-year survival predictive power showed that the CLIP score explained 37% of survival variability, compared with 21% explained by Okuda stage. In conclusion, the CLIP score, compared with the Okuda staging system, gives more accurate prognostic information, is statistically more efficient, and has a greater survival predictive power. It could be useful in treatment planning by improving baseline prognostic evaluation of patients with HCC, and could be used in prospective therapeutic trials as a stratification variable, reducing the variability of results owing to patient selection.

Adult↗

Prospective validation of Wells Criteria in the evaluation of patients with suspected pulmonary embolism.

STUDY OBJECTIVE: The literature suggests that the d -dimer is useful in patients suspected of having pulmonary embolism and who have a low pretest probability of disease. A previously defined clinical decision rule, the Wells Criteria, may provide a reliable and reproducible means of determining this pretest probability. We evaluate the interrater agreement and external validity of Wells Criteria in determining pretest probability in patients suspected of having pulmonary embolism. METHODS: This was a prospective observational study. Trained research assistants enrolled patients during 120 random 8-hour shifts. Patients who underwent imaging for pulmonary embolism after a medical history, physical examination, and chest radiograph were enrolled. Treating providers and research assistants determined pretest probability according to Wells Criteria in a blinded fashion. Two d -dimer assays were run. Three-month follow-up for the diagnosis of pulmonary embolism was performed. Interrater agreement tables were created. kappa Values, sensitivities, and specificities were determined. RESULTS: Of the 153 eligible patients, 3 patients were missed, 16 patients declined, and 134 (88%) patients were enrolled. Sixteen (12%) patients were diagnosed with pulmonary embolism. The kappa values for Wells Criteria were 0.54 and 0.72 for the trichotomized and dichotomized scorings, respectively. When Wells Criteria were trichotomized into low pretest probability (n=59, 44%), moderate pretest probability (n=61, 46%), or high pretest probability (n=14, 10%), the pulmonary embolism prevalence was 2%, 15%, and 43%, respectively. When Wells Criteria were dichotomized into pulmonary embolism-unlikely (n=88, 66%) or pulmonary embolism-likely (n=46, 34%), the prevalence was 3% and 28%, respectively. The immunoturbidimetric and rapid enzyme-linked immunosorbent assay d -dimer assays had similar sensitivities (94%) and specificities (45% versus 46%). CONCLUSION: Wells Criteria have a moderate to substantial interrater agreement and reliably risk stratify pretest probability in patients with suspected pulmonary embolism.

Adult↗

Modeling hepatic fibrosis in African American and Caucasian American patients with chronic hepatitis C virus infection.

Assessment of histological stage is an integral part of disease management in patients infected with the hepatitis C virus (HCV). The aim of this study was to develop a model incorporating objective clinical and laboratory parameters to estimate the probability of severe fibrosis (i.e., Ishak fibrosis > or = 3) in previously untreated African American (AA) and Caucasian American (CA) patients with HCV genotype 1. The Ishak fibrosis scores of 205 CA and 194 AA patients enrolled in the Viral Resistance to Antiviral Therapy of Chronic Hepatitis C study (Virahep-C) were modeled using simple and multiple logistic regression. The model was then validated in an independent cohort of 461 previously untreated patients with HCV. The distribution of fibrosis scores was similar in the AA and CA patients as was the proportion of patients with severe fibrosis (35% vs. 39%, P = .47). After accounting for the number of portal areas in the biopsy, patient age, serum aspartate aminotransferase, alkaline phosphatase, and platelet count were independently associated with severe fibrosis in the overall cohort, and the relationship with fibrosis was similar in both the AA and CA subgroups. The area under the receiver operating characteristic curve (AUROC) of the Virahep-C model (0.837) was significantly better than in other published models (P = .0003). The AUROC of the Virahep-C model was 0.851 in the validation population. In conclusion, a model consisting of widely available clinical and laboratory features predicted severe hepatic fibrosis equally well in AA and CA patients with HCV genotype 1 and was superior to other published models. The excellent performance of the Virahep-C model in an external validation cohort suggests the findings are replicable and potentially generalizable.

Black or African American↗

Inhibition and substrate recognition--a computational approach applied to HIV protease.

We have developed a computational approach in which an inhibitor's strength is determined from its interaction energy with a limited set of amino acid residues of the inhibited protein. We applied this method to HIV protease. The method uses a consensus structure built from X-ray crystallographic data. All inhibitors are docked into the consensus structure. Given that not every ligand-protein interaction causes inhibition, we implemented a genetic algorithm to determine the relevant set of residues. The algorithm optimizes the q2 between the sum of interaction energies and the observed inhibition constants. The best possible predictive model resulting has a q2 of 0.63. External validation by examining the predictivity for compounds not used in derivation of the model leads to a prediction accuracy between 0.9 and 1.5 log10 unit. Out of 198 residues in the whole protein, the best internally predictive model defines a subset of 20 residues and the best externally predictive model one of 9 residues. These residues are distributed over the subsites of the enzyme. This approach provides insight in which interactions are important for inhibiting HIV protease and it allows for quantitative prediction of inhibitor strength.

Amino Acids↗

[Evaluating the reliability, validity and responsiveness of the german short musculoskeletal function assessment questionnaire, SMFA-D, in inpatient rehabilitation of patients with conservative treatment for hip osteoarthritis].

BACKGROUND: Modern patient based outcome measures like the SMFA-D (German Short Musculoskeletal Function Assessment Questionnaire) are able to detect the impairment and functional capacity of patients with musculoskeletal extremity disorders. The SMFA-D was successfully evaluated in several cohorts treated operatively for osteoarthritis of the knee and hip, rotator cuff tears and rheumatoid arthritis. The aim of the present study was the evaluation of the SMFA-D in patients with conservative treatment for hip osteoarthritis. PATIENTS AND METHODS: 69 patients with osteoarthritis of the hip were enrolled in a prospective controlled clinical trial. All patients completed the SMFA-D, SF-36, WOMAC, FFbH-OA. A standardized test of walking speed and the functional status of the patient as judged by the physician were recorded. Statistical analysis were done for the following: re-test reliability (ICC), internal consistency (Cronbach's alpha), validity and responsiveness. RESULTS: Internal consistency (Cronbach's alpha) was alpha = 0.89 and alpha = 0.97 for the SMFA-D scales. The retest reliability (ICC, unjust, mixed effect) was 0.91 (p < 0.001) for the function index and 0.73 (p < 0.001) for the bother index. Both indices correlated significantly with the FFbH-OA (r = 0.66 to r = 0.84), the WOMAC (r = 0.55 to r = 0.86) and the scales of the SF-36 (r = - 0.34 to r = - 0.85) on all three time points, which supports construct validity. There was mainly a significant correlation between the SMFA-D scales and the functional status of the patient (r = 0.21 to r = 0.44), pain reported by the patient (r = 0.43 to = 0.54) and the self selected walking speed (r = 0.28 to r = 0.51), which supports external validity. We were able to differentiate operatively and conservatively treated patients (discriminant construct validity). At the end of the rehabilitation program we were able to demonstrate small to medium treatment effects in SMFA-D and SF-36. The WOMAC and FFbH-OA were not able to demonstrate these treatment effects. CONCLUSION: Even in patients with conservative treatment of hip osteoarthritis the SMFA-D represents a reliable, valid and responsive measure. The use of the SMFA-D can be recommended as a patient based outcome measure.

Activities of Daily Living↗

A cross-sectional validation study of Self-Evaluation of Communication Experiences after Laryngeal Cancer--a questionnaire for use in the voice rehabilitation of laryngeal cancer patients.

A psychometric evaluation of the questionnaire 'Self-Evaluation of Communication Experiences after Laryngeal Cancer' (S-SECEL) addressing communication dysfunction in patients with laryngeal cancer was carried out. Ninety-three patients with laryngeal cancer were studied. For comparison of response patterns and external validation, 21 patients with non-small cell lung cancer (NSCLC) and 26 patients with hoarseness, caused by benign laryngeal disease, were included in the analysis. The patients completed three questionnaires; the S-SECEL, the Sickness Impact Profile (SIP) and the Hospital Anxiety and Depression scale (HAD). The S-SECEL questionnaire was well-accepted by the patients, compliance was satisfactory, and missing value rates were low. The reliability of the S-SECEL was satisfactory for the Environment and Attitude subscales, whereas the General subscale did not reach the reliability levels recommended for group comparisons. In general, the response pattern in the three diagnostic groups and the pattern of correlations between the S-SECEL scores and the SIP- and HAD-subscales and dimensions lent support to the construct validity of the S-SECEL.

Aged↗

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans↗

The Italian SF-36 Health Survey: translation, validation and norming.

This article reports on the development and validation of the Italian SF-36 Health Survey using data from seven studies in which an Italian version of the SF-36 was administered to more than 7000 subjects between 1991 and 1995. Empirical findings from a wide array of studies and diseases indicate that the performance of the questionnaire improved as the Italian translation was revised and that it met the standards suggested by the literature in terms of feasibility, psychometric tests, and interpretability. This generally satisfactory picture strengthens the idea that the Italian SF-36 is as valid and reliable as the original instrument and applicable and valid across age, gender, and disease. Empirical evidence from a cross-sectional survey carried out to norm the final version in a representative sample of 2031 individuals confirms the questionnaire's characteristics in terms of hypothesized constructs and psychometric behavior and gives a better picture of its external validity (i.e., robustness and generalizability) when administered in settings that are very close to real world.

Cross-Cultural Comparison↗

Validation of Mayo Clinic risk adjustment model for in-hospital complications after percutaneous coronary interventions, using the National Heart, Lung, and Blood Institute dynamic registry.

OBJECTIVES: We sought to validate the recently proposed Mayo Clinic risk score model for complications after percutaneous coronary interventions (PCI), using an independent data set. BACKGROUND: The Mayo Clinic risk score has eight simple clinical and angiographic variables for the prediction of complications defined as either death, Q-wave myocardial infarction, emergent or urgent coronary artery bypass graft surgery, or cerebrovascular accident after PCI. External validation using an independent data set is lacking. METHODS: A total of 3,264 patients undergoing PCI at each of the 17 sites in the National Heart, Lung, and Blood Institute's Dynamic Registry during two enrollment periods (July 1997 to February 1998 and February to June 1999) were studied. Logistic regression was used to model the calculated risk score and major procedural complications. The expected number of complications, with 95% confidence bounds (CBs), was also calculated. RESULTS: There were 96 (2.94%) observed procedural complications, and the Mayo Clinic risk score predicted 93.5 events (2.86%; 95% CB 2.32% to 3.41%; p = NS). The Hosmer-Lemeshow goodness-of-fit p value was 0.28, and the area under the receiver operating curve was 0.76, indicating excellent overall discrimination. There were no statistical differences between observed and predicted procedural complications using the Mayo Clinic risk score among the most selected high- and low-risk subgroups. CONCLUSIONS: Eight variables were combined into a convenient risk scoring system that accurately predicts cardiovascular complications after PCI. The Mayo clinic predictive model for procedural complications yielded excellent results when applied to a multi-center external data set.

Aged↗

Evidence of the validity of a model of determinants of quality of restorative dental care.

A model describing the relationship between self-reported quality of restorative dentistry and dentist characteristics for 119 Montana general dentists is presented. The best predictors formed a significant model explaining 22% of the variance of the quality measure. Results are contrasted with a previous estimation of the model for 102 Washington general practitioners. Evidence for the external validity of the model is presented.

Adult↗

Using and interpreting analogue designs.

Researchers in rehabilitation counseling and disability studies sometimes use analogue research, which involves materials that approximate or describe reality (e.g., written vignettes, videotaped exemplars) rather than investigating phenomena in real-world settings. Analogue research often utilizes experimental designs, and it thereby frequently possesses a high degree of internal validity. Analogue research allows investigators to exercise tight control over the implementation of the independent or treatment variable and over potentially confounding variables, which enables them to isolate the effects of those treatment variables on selected outcome measures. However, the simulated nature of analogue research presents an important threat to external validity. As such, the generalizability of analogue research to real-life settings and situations may be problematic. These and other issues germane to analogue research in vocational rehabilitation are discussed in this article, illustrated with examples from the contemporary literature.

Humans↗

A validation of two preoperative nomograms predicting recurrence following radical prostatectomy in a cohort of European men.

Kattan et al. at Baylor College of Medicine and D'Amico et al. at Harvard Medical School have each developed preoperative nomograms for prostate cancer recurrence after radical prostatectomy based on readily available clinical variables. Calibration and validation of those tools was achieved using North American patient cohorts, and their validity has not yet been shown in patients from other continents. We investigated the predictive accuracy of these nomograms when applied to European men with localized prostate cancer. Clinical data from patients who underwent radical prostatectomy at the University-Hospital Hamburg and fitted the respective derivation criteria were used for external validation (n = 1003 for the Kattan-Nomogram, n = 932 men for the D'Amico-Nomogram). Nomogram predictions of the probability for 2-years and 5-years freedom from recurrence predicted by the D'Amico-Nomogram and the Kattan-Nomogram respectively were compared with actual follow-up. The predictive accuracy of the nomograms was tested using areas under the receiver-operating-characteristic curves (AUC). The D'Amico-Nomogram AUC predicting 2-years probability of freedom from PSA recurrence was 0.80 vs. Kattan-Nomogram 5-years prediction with an AUC of 0.83. Using the 932 patients who exactly fit the derivation criteria of both nomograms, the predictive accuracy of the Kattan-Nomogram was 0.81. The superiority in predictive accuracy of the Kattan-Nomogram was statistically significant (p = 0.0274) but of unclear clinical significance. The two nomograms predicted recurrence with similar accuracy when applied to men diagnosed with localized prostate cancer in Germany. The high predictive accuracy of both nomograms demonstrates that these predictive tools derived in the U.S. can be applied to non-U.S. patients.

Cohort Studies↗

A downscaled practical measure of mood lability as a screening tool for bipolar II.

BACKGROUND: Current data indicate a strong association between Cyclothymic temperament (and its more ultradian counterpart of mood lability) and Bipolar II (BPII). Administration of elaborate measures of temperament are cumbersome in routine practice. Accordingly, the aim of the present analyses was to test if a practical measure of mood lability was unique to BPII, in comparison with major depressive disorder (MDD). METHODS: Using the Structured Clinical Interview for DSM-IV Axis I Disorders, Clinician Version as modified by us [J. Affect. Disord. 73 (2003) 33; Curr. Opin. Psychiatry 16 (2003) S71], we interviewed 62 consecutive BPII outpatients, as well as their 59 MDD counterparts during a major depressive episode (MDE). Hypomanic symptoms during MDE were systematically assessed: three or more such symptoms defined depressive mixed state (DMX3) on the basis of previous work by us [J. Affect. Disord. 73 (2003) 113]. A downscaled definition of trait mood lability was adapted from Akiskal et al. [Arch. Gen. Psychiatry 52 (1995) 114] and Angst et al. [J. Affect. Disord. 73 (2003) 133], requiring a positive response to one of two queries on whether one is a person with frequent "ups and downs" in mood, and whether such mood swings occur for no reason. The patients selected for inclusion had not received neuroleptics and antidepressants for at least 2 weeks prior to the index episode, they were free of substance and alcohol abuse, and did not meet the DSM-IV criteria for borderline personality disorder (BPD). Associations between mood swings and clinical variables were tested by logistic regression (STATA 7). RESULTS: Mood swings were endorsed by 50.4% of the entire sample: 62.9% of BPII and 37.2% of MDD (p = 0.0047). This practical measure of mood lability was significantly associated with BPII, lower age at onset, high depressive recurrences, atypical features, and DMX3. When controlled for number of major affective episodes, mood swings were still significantly associated with BP-II. Sensitivity and specificity of mood swings for predicting BPII were 62.9% and 62.7%, respectively. LIMITATION: The low specificity of trait mood lability for BPII diagnosis is probably due to the fact that we used a downscaled simplified measure of this trait. CONCLUSIONS: On the other hand, the relatively high sensitivity of our downscaled measure of mood lability for predicting BPII supports its usefulness as a screening tool for this diagnosis. The lack of association between self-reported mood lability and number of major mood episodes indicates that such lability does not reflect the perception of history of frequent episodes, and that it has some validity as a trait indicator. Given that our sample excluded patients meeting the DSM-IV criteria for BPD, contradicts the opinion of the latter manual that such mood lability represents its pathognomonic characteristic that distinguishes it from BPII. The bipolar nature of mood lability is further supported by significant associations with external validating criteria for bipolarity. Overall, these data indicate that in the differential diagnosis between MDD and BPII, trait mood lability favors the latter at a significant statistical level.

Adult↗

Ethnic specification, validation prospects, and the future of drug use research.

Interest in drug use among America's major ethnic minority groups is rapidly increasing. Despite the welcomed interest, researchers tend to use broad ethnic labels to identify their samples. Such labels as "ethnic glosses" provide little information concerning the heterogeneity of each ethnic group and, in most instances, violate the guidelines concerning appropriate descriptions of sample characteristics. Use of broad ethnic descriptors particularly in drug use research creates external validity problems and prevents replications. Researchers are encouraged to obtain detailed information on the sociocultural characteristics of their samples by obtaining measures on ethnic identification, situated identity, and acculturative status. Use of ethnic identity and acculturation measures, however, can create problems in defining appropriate sample frames.

Acculturation↗

A clinical index for predicting visual acuity after cataract surgery.

We developed a clinical index for predicting postoperative visual acuity of cataract patients and cross-validated it using data from 182 patients aged 70 years and older. The index consisted of four statistically combined indicators: age, preoperative visual acuity, frequency of reading, and comorbidity. Validation of the index included comparisons to two standard technical instruments for measurement of retinal visual acuity. For the clinical index, 72% of predictions were accurate within one Snellen line of postoperative visual acuity compared to 37% using a laser interferometer and 33% using a potential acuity meter. Testing of the clinical index's external validity using data from 111 patients in a different ophthalmology clinic disclosed 61% of predictions accurate within one Snellen line.

Aged↗