Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Randomized controlled trials in psychiatry. Part II: their relationship to clinical practice.

OBJECTIVE: To discuss the extent to which the results of randomized controlled trials (RCTs) in psychiatry can be generalized to clinical practice. METHOD: Threats to internal and external validity in psychiatric RCTs are reviewed. RESULTS: Threats to internal validity increase the possibility of bias. Psychiatric RCTs have problems with small samples, arbitrary definitions of caseness, disparate definitions of outcome and high spontaneous recovery rates. Particular issues arise in psychotherapy RCTs. Threats to external validity reduce the extent to which the results of a RCT produce a correct basis for generalization to other circumstances. These include high rates of comorbidity and sub syndromal pathology in normal clinical practice, manual-based treatment protocols and varying definitions of successful treatment. CONCLUSIONS: Randomized controlled trials remain the most robust design to investigate the effectiveness of treatments. They should be applied to important clinical questions; and carried out, as far as possible, with typical patients in the clinical conditions in which the treatment is likely to be used.

Evidence-Based Medicine↗

The reliability of upper limb anthropometry in older Chinese people.

OBJECTIVE: To evaluate the validity of the Durnin-Womersley equations and to derive our local predictive equations for body fat from upper limb skinfold thicknesses in older Chinese people in Hong Kong. To evaluate the validity of mid-arm circumference and corrected arm muscle area in predicting lean tissue mass in the same population. DESIGN: Comparison of fat percentages predicted by Durnin-Womersley (D-W) equations with those estimated by Dual energy X ray absorptiometry (DXA). Predictive equations derived from regression between upper limb skinfold thicknesses and fat percentages estimated by DXA were similarly evaluated in internal and external validation groups. Mid-arm circumference (MAC) and corrected arm muscle area (CAMA) were correlated with the limb lean tissue mass, body lean tissue mass and fat percentage. SUBJECTS: 354 female and 263 male, apparently well, community dwelling subjects, aged 69-82 y; of which 40 subjects of each sex were randomly selected from the study population for internal validation of the local predictive equations; 60 female and 33 male hospital medical outpatients, aged 61-87 y, were recruited for external validation. MEASUREMENTS: Triceps and biceps skinfold thicknesses, mid-arm circumference, body mass index, fat percentages, limb and whole body lean tissue masses estimated by Hologic QDR-2000 bone densitometer. RESULTS: Fat percentages calculated by D-W equations were significantly different from those estimated by DXA (average difference -2.4 (s.d. 4.8)% and +2.1 (5.2)% in females and males respectively). The corresponding differences for our local predictive equations were not significant (-0.9 (4.7)% and -0.5 (5.0)% in females and males respectively). There was a trend of under-estimation of body fat with increasing fatness. In the hospital medical outpatients, there was a significant difference between fat percentages predicted by our equation and those by DXA in female (-2.9(5.3)%), but not in male (+0.3(4.3)%) subjects. In males, MAC correlated with limb and body lean tissue masses as well as with fat percentage (r = 0.60, 0.68, 0.65 respectively). CAMA correlated similarly well with lean tissue masses but was more independent of fat percentage (r = 0.61, 0.65, 0.44 respectively). In females, both MAC and CAMA correlated poorly with limb and body lean tissue masses. Moreover, MAC correlated well with fat percentage (r = 0.80). CONCLUSION: Upper limb skinfold thicknesses measurement is a valid means of predicting body fat in older Chinese people. Local predictive equations were more reliable that D-W equations. They were, however, subject to errors at the extreme ends of body fatness and in the presence of disease. In older females, MAC and CAMA were not reliable in predicting lean tissue mass, but MAC could be used to predict fat percentages. In older males, CAMA was more reliable than MAC in predicting lean tissue mass.

Absorptiometry, Photon↗

Detection of depressive symptomatology in elderly people: a short version of the CES-D scale.

This study aims to test a short form of the Center for Epidemiological Studies--Depression Scale (CES-D) which can be a useful screening tool for depressive symptomatology in epidemiological studies of elderly patients. The study was conducted on 2792 subjects from the PAQUID (Personnes Agées QUID?) cohort, an epidemiological survey of community dwellers living in South-West France. CES-D items with high sensitivity and good specificity were selected for the short form, then the best cut-off scores were determined with Receiver Operating Characteristics (ROC) curves. The external validity of the 5-item scale was then assessed against the full scale at different PAQUID follow-ups. Sensitivity was 99% and specificity 81% for detecting depressive symptomatology when compared to the 20-item scale. The external validity on the different follow-ups was good, yielding a sensitivity varying from 95 to 100%, and a specificity from 83 to 89%. In conclusion, the 5-item CES-D is a simple, rapid and reliable tool which could be useful for screening depressive symptoms in epidemiological studies of the elderly.

Aged↗

Five-factor model of schizophrenic psychopathology: how valid is it?

Aim of the study was to examine the consistency of the five-factor model of schizophrenic symptoms, assess its validity and evaluate its dimensional factor structure using confirmatory factor (CFA) analysis. A sample of 258 randomly assigned DSM-III R patients with schizophrenic disorders were studied by means of the structured clinical interview for the Greek validated Positive and Negative Syndrome Scale (PANSS) and were rated on its 30 items. Patients' scores were subjected to principal component analysis (PCA) with varimax rotation. Internal consistency for each of the components was determined by the use of Cronbach's alpha. External validity of the model derived was investigated by searching for possible relationships between the components and sociodemographic characteristics with the aid of canonical correlation analysis. Confirmatory factor analysis (CFA) was also performed. Using the scree plot criterion PCA revealed a five-factor model. These factors were interpreted as representing--in a decreasing order of relative importance--the following dimensions of schizophrenic psychopathology: negative, excitement, depression, positive and cognitive impairment. The model was comparable with six previous factor analytic studies. Internal consistency was quite satisfactory whereas external validity was found to be not so powerful. CFA did not show that the proposed model yields an adequate factor structure.

Adult↗

Catecholaminergic polymorphic ventricular tachycardia mediated by ryanodine receptor 2: a validated risk stratification.

BACKGROUND AND AIMS: Patients with catecholaminergic polymorphic ventricular tachycardia (CPVT) are at risk for potentially life-threatening arrhythmic events (AEs) even while treated with β-blockers. The aim was to develop a model for individualized prediction of AEs in patients with RYR2-mediated CPVT on β-blocker monotherapy. METHODS: The derivation and independent validation cohorts included 743 and 129 patients, respectively. AEs were defined as arrhythmic syncope, appropriate implantable cardioverter-defibrillator shock, sudden cardiac arrest (SCA), and sudden cardiac death. Near-fatal or fatal AEs (nf/fAEs) included all AEs except for arrhythmic syncope. Prediction models using Cox regression were developed and internally and externally validated. RESULTS: A total of 102 (13.7%) patients in the derivation cohort and 24 (18.6%) patients in the validation cohort experienced ≥1 AE over a median follow-up of 5.1 [interquartile range (IQR), 7.7] and 2.4 (IQR, 4.4) years, respectively. Predictors of AE were arrhythmic syncope or SCA prior to diagnosis and age at β-blocker initiation. In the derivation and validation cohorts, the optimism-corrected C-indices of the models for AE were 0.67 [95% confidence interval (CI) 0.62-0.72] and 0.59 (95% CI 0.48-0.71), respectively. For nf/fAEs, ventricular arrhythmia severity before β-blocker initiation was a fourth independent predictor, and C-indices of the models in the derivation and validation cohorts were 0.74 (95% CI 0.68-0.80) and 0.60 (95% CI 0.47-0.72), respectively. In the derivation cohort, calibration slopes were 1.00 (95% CI 0.59-1.41) for AE and 1.00 (95% CI 0.69-1.32) for nf/fAE. CONCLUSIONS: These externally validated risk prediction models using clinical parameters accurately distinguished CPVT patients on β-blocker monotherapy at low and high risk for future AEs while treated with β-blockers. These models provide guidance for implementation of clinical management therapies to prevent AEs in patients with CPVT.

Humans↗

Choice and validation of a near infrared spectroscopic application for the identity control of starting materials. practical experience with the EU draft Note for Guidance on the use of near infrared spectroscopy by the pharmaceutical industry and the data to be forwarded in part II of the dossier for a marketing authorization.

Recently the CPMP/CVMP sent out for consultation the draft Note for Guidance (dNfG) on the use of near infrared spectroscopy (NIRS) by the pharmaceutical industry and the data to be forwarded in part II of the dossier for a marketing authorization. We explored the practicability of this dNfG with respect to the verification of the correct identity of starting materials in a generic tablet-manufacturing site. Within the boundaries of the dNfG, a release procedure was developed for 12 substances containing structurally related compounds and substances differing only in particle size. For the method development literature data were also taken into consideration. Good results were obtained with wavelength correlation (WC), applied on raw spectra or second derivative spectra both without smoothing. The defined threshold of 0.98 for raw spectra differentiated between all molecular structures. Both methods were found to be robust over a period of 1 year. For the differentiation between the different particle sizes a subsequent second chemometric technique had to be used. Soft independent modelling of class analogy (SIMCA) with a probability level of 0.01 proved suitable. Internal and external validation I according to the dNfG showed no incorrect rejections or false acceptances. External validation II according to the dNfG was carried out with 95 potentially interfering substances from which 46 were tested experimentally. Macrogol 400 was not distinguished from macrogol 300. For the complete verification of the identity of macrogol 300 test A of the European Pharmacopoeia is needed in addition to the NIRS application. A release procedure developed with WC applied on raw spectra and SIMCA as a second method, which is different from the preferred method of the dNfG, was tested in practice with good results. We conclude that the dNfG has good practicability and that deviations from the preferred methods of the dNfG can also give good differentiation.

Drug Industry↗

Near-infrared reflectance spectroscopy as a fast and non-destructive tool to predict foliar organic constituents of several woody species.

Near-infrared reflectance spectroscopy (NIRS) was used to estimate N, neutral detergent fibre (NDF), acid detergent fibre (ADF), lignin and cellulose contents in leaves of a heterogeneous group of 17 woody species from the Central Western region of the Iberian Peninsula. The sample set consisted of 182 samples of leaves of deciduous and evergreen species, showing a wide range of concentrations determined by reference methods: 6.60-35.2 g kg-1 (N), 15.5-66.0% (NDF), 10.2-57.3% (ADF), 3.45-27.4% (lignin) and 5.79-31.3% (cellulose). Reflectance spectra, obtained for samples of dried and ground leaves, were recorded as log1/R (R=reflectance) from 1,100 to 2,500 nm. NIRS calibrations were developed using multiple linear (MLR) and partial least-squares (PLSR) regressions, and tested by external validation. Spectral data were transformed to the first and second derivative (1D, 2D). The PLSR method and derivative transformations provided the best statistics and showed lower standard errors of calibration (SEC) and higher coefficients of multiple determination (R2). In the external validation the standard errors of prediction (SEP) were 0.76 g kg-1 (N), 2.11% (NDF), 1.47% (ADF), 0.85% (lignin) and 0.86% (cellulose). The results obtained show that NIRS is very effective for the estimation of these organic constituents in leaf tissue of woody species. This technique can be used in ecological or ecophysiological studies as an alternative to the more time-consuming standard methods.

Cellulose↗

Die Vision einer pragmatischen klinischen Forschung oder das Ende der Diskussion über < > und < >

The Concept of Pragmatic Clinical Research or the End of Discussion about 'Placebo' and 'Specific Effects&rsquoThe reorientation of clinical research towards the questions of treatment benefit (beyond the question of treatment efficacy) and of how much clinical trials represent actual practice (external validity) is the timely path to clinical research questions of real interest and importance. Postmodern 'anything goes' makes it possible to also consider thus far looked down on placebo effects as valuable, however, it requires the precise documentation of the external validity of such effects. Not disease as such, but the disease context, not therapy as such, but the therapy context, not the patient as such, but the patient context, not a test as such, but the test situation have become the important focuses of clinical research. In respect to test results current medicine has to recognize its illiterate mystification of allegedly 'objective' and 'hard' data. The patient context can determine whether an 'efficacious' therapy is beneficial or harmful, and thus, it is the proper definition of the patient context which makes medicine scientific, no matter how 'objective' or 'subjective' the effect of therapy is. The consideration of the therapy context leads to the important distinction between efficacy and effectiveness (or benefit), and it becomes intelligible that the randomised controlled trial in its traditional design as the placebo-controlled double-blind trial is limited to the evaluation of an agent theory. The evaluation of treatment effectiveness requires more pragmatic trials which study treatment operations and not isolated components and which may even compare entire treatment strategies. Pragmatic clinical trials, in future, will not only allow the study of 'pathogenesis blockers'., but also the study of 'salutogenetic' interventions working with the formation of the host. The focus of attention and research in the new school of evidencebased medicine with clinical epidemiology as its basic science (if not superficially understood as mere literature medicine) has long ago been identified as illness as the product of host, disease and environment. The dispute about 'placebo' and 'specific effects', in the meantime, has become obsolete.

Journal Article↗

Development of a COPD severity score.

OBJECTIVE: Treatment of chronic obstructive pulmonary disease (COPD) is based on symptom control. This suggests that COPD severity can be determined by analyzing treatment intensity. The objective of this analysis was to develop and validate a severity score for adult COPD based on treatments. RESEARCH DESIGN: Using principal components analysis, a COPD severity score was developed using data based on treatments extracted from an employer claims database (development group). Variables included were identified from literature review and clinical expert opinion. External validity was tested in a separate group of adult chronic bronchitis patients in whom principal components analysis was re-conducted and factor loadings were compared to the development group. Construct validity was tested by comparing the incidence of acute exacerbations of chronic bronchitis (AECB) in patients with high and lower severity scores. To illustrate the use of the COPD severity score, effectiveness of alternative AECB antibiotic treatments was compared in a separate patient sample categorized by severe versus mild/moderate COPD. RESULTS: In the development group (n = 2068), principal components analysis produced a single main factor for severity scoring. Of the 12 variables contributing to this factor, the 6 with the highest factor loadings were treatment related. The factor performed similarly in the external validity group (n = 9127) as it did in the development group. In construct validity testing, severe COPD patients were 4 times more likely to have AECB episodes than mild/moderate patients. Patients with severe COPD and an AECB were more likely to fail treatment with antibiotics than those with mild/moderate COPD. Based on the COPD severity score developed, we found that treatment of patients with severe COPD and an AECB with fluoroquinolones was more likely to result in treatment failure than treatment with macrolides (OR = 2.01; p = 0.03). CONCLUSIONS: The analysis was successful in developing and validating a method to score COPD severity based on treatments. This method may prove useful in providing insights about the benefits of COPD treatments.

Anti-Bacterial Agents↗

Developing an empirical typology for regular exercise.

BACKGROUND: Tailored interventions require the identification of distinct homogenous subgroups that will benefit from different intervention materials. One way to identify such subgroups is to use cluster analysis to identify an empirical typology. METHODS: A sample of 346 adults completed surveys through a telephone interview that included questions related to participating in regular exercise. The three variables used in the cluster analysis were the Pros of Exercise, the Cons of Exercise, and Exercise Self-Efficacy. RESULTS: Six resulting clusters were labeled Disengaged, Immotive, Relapse Risk, Early Action, Maintainers, and Habituated. A series of analyses tested the internal and external validity of the typology. The internal validity test revealed that four of the clusters demonstrated high stability and replicability, while the Relapse Risk and Early Action clusters were less stable. Differences among clusters on self-reported exercise behavior and a strong association with stage of change for regular exercise provided external validity evidence of the typology. CONCLUSIONS: The resulting typology reflects a range of motivational patterns that are likely to be responsive to different types of messages and strategies regarding adoption and maintenance of regular exercise. The typology also generates a number of hypotheses about the identified clusters that can be empirically tested in further studies.

Adult↗

Meaning in life depth in the active married elderly.

To discover if elderly people have developed a deeper meaning in life than younger individuals, a sample of active married elderly people was compared to a group of younger adults. Two dimensions of meaning in life depth were investigated. The first was a self-suitability measure indicating comfort with one's own meaning, measured by Crumbaugh and Maholick's (1969) Purpose in Life Test. The second was an external validation measure derived from a statement about their own strongest meaning in life, written by the participants and rated for depth by two outside judges. The older group scored significantly higher than the younger adults on the self-suitability measure and significantly lower on the external validation measure. Such results could mean that toward the end of life we are better able to appreciate life's beauty though less able to communicate our depth of appreciation to others. An alternative interpretation of the results is that the elderly participants were engaging in self-deception.

Aged↗

[Cardiovascular risk assessment for informed decision making. Validity of prediction tools].

BACKGROUND AND PURPOSE: Patient involvement in health care decisions is increasingly requested. The authors investigated whether currently available assessment tools for prediction of cardiovascular risk can be used for individual risk prediction as a basis of informed decision making. METHODS: The authors searched for risk assessment tools and respective validation studies in Medline (until August 16, 2004) and the Cochrane Library (issue 2/2004). The following criteria were used for evaluation of prognostic studies: (1) discrimination between risk groups; (2) predictive values; (3) prognostic agreement; (4) transferability across populations. RESULTS: A total of twelve assessment tools were identified. The Framingham function, Sheffield Tables, Canadian Tables, Framingham Categorial, New Zealand, Joint British, and European Charts (1994 and 1998) are based on the Framingham Study; PROCAM Risk Score, UKPDS Risk Engine, and SCORE Risk Charts use different source data. Framingham-based instruments overestimate cardiovascular risk of Central-European populations by at least 30%, with substantial regional variation even within a country (between 30% and 100%, British Regional Heart Study). Therefore, prior to application the assessment tools would need recalibration using regional data of cardiovascular mortality and adjustment for social class differences. Published sensitivity, specificity, and C-statistics for external validation (area under the curve [AUC] approximately 0.6) are clearly inferior to internal validation (AUC approximately 0.8). Agreement between instruments beyond chance is moderate (kappa approximately 0.5). No studies on external validation could be identified for the new European SCORE Risk Charts and UKPDS Risk Engine. CONCLUSION: Validation of currently available assessment tools for cardiovascular risk prediction is inadequate. Uncritical use may lead to substantial under- or overestimation of individual cardiovascular risk and inappropriate treatment decisions.

Adult↗

An Exosomal Signature for Preoperative Detection of Occult Liver Metastasis in Pancreatic Cancer.

IMPORTANCE: Early liver metastasis (early-LiM) after pancreatectomy represents an aggressive biological phenotype of pancreatic ductal adenocarcinoma (PDAC) and is associated with markedly poor survival. Reliable preoperative biomarkers to identify occult hepatic micrometastasis remain lacking. OBJECTIVE: To develop and externally validate a circulating exosomal microRNA (exo-miRNA)-based machine learning model for preoperative detection of occult early-LiM in PDAC. DESIGN, SETTING, AND PARTICIPANTS: This multicenter retrospective case-control study included 3 phases: genome-wide discovery using exo-miRNA sequencing (discovery cohort), model development (training cohort), and independent external validation (2 validation cohorts). The study took place at 4 medical centers in China, Japan, and South Korea. A total of 372 patients were enrolled between 2011 and 2024. Data were analyzed from July 2024 to November 2025. EXPOSURES: Circulating plasma-derived exosomal miRNA expression profiles. MAIN OUTCOMES AND MEASURES: The primary outcome was early-LiM, defined as liver recurrence within 6 months after curative-intent resection. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC) and survival outcomes were assessed using Kaplan-Meier analysis. RESULTS: Among 372 patients with PDAC (median [IQR] age, 67 [59-73] years; 229 [61.6%] male and 143 [38.4%] female; median follow-up among survivors, 969 days),early-LiM was associated with significantly worse overall survival compared with other recurrence patterns (median OS, 9.1 months vs 26.6-31.8 months; log-rank P&#x2009;<&#x2009;.001). A 7-exo-miRNA extreme gradient boosting model demonstrated discrimination in the training cohort (AUC, 0.899; 95% CI, 0.822-0.976) and maintained performance in external testing cohorts (AUC, 0.876; 95% CI, 0.846-0.951 and AUC, 0.862; 95% CI, 0.744-0.981). The exo-miRNA panel score remained an independent identifier of early-LiM in multivariable analysis (odds ratio, 26.49; 95% CI, 18.45-55.28; P&#x2009;<&#x2009;.001) and stratified overall survival (log-rank P&#x2009;<&#x2009;.001). Decision curve analysis suggested improved net clinical benefit compared with conventional clinicopathologic variables. CONCLUSION AND RELEVANCE: In this multicenter study, a circulating exo-miRNA-based machine learning model enabled preoperative detection of occult early liver metastasis risk in PDAC. These findings support the potential of exosomal biomarkers to inform biology-guided treatment sequencing and warrant prospective validation.

Journal Article↗

Threats to the validity of clinical trials employing enrichment strategies for sample selection.

Subject selection and exclusion criteria employed in typical clinical effectiveness trials of investigational new drugs have two fundamental aims: (1) to ensure that patients entering a study are truly suffering from the condition the drug is intended to treat and (2) to maximize the likelihood that the study will detect an effect of the drug if, in fact, one exists. Typical protocol selection criteria not only specify exacting procedures for establishing and documenting the diagnosis of those recruited for a study but also seek to increase, relative to the prevalence in the general population, the proportion of individuals in the sample likely to respond to pharmacological treatment. Because it is ordinarily impossible to learn prior to extensive clinical experience with a new drug which, if any, patient characteristics reliably predict a consistent treatment response, strategies for sample "enrichment" typically operate by excluding patients (for example, those with very advanced and/or complicated illness, those with serious concomitant illness, those at the extremes of age, those with very mild illness, and so forth) in whom a dependable response to treatment seems unlikely on logical and/or generic grounds. Some studies use positive strategies for sample "enrichment." In studies evaluating drugs intended to treat recurrent episodes of psychiatric illnesses, many protocols recommend selective recruitment of patients with a history of meaningful positive responses to antipsychotic treatment during prior episodes. Sample selection procedures of these kinds impose limits on the generalizability of a study's results (i.e., external validity), but the use of nonrandom patient samples is ordinarily held to have no effect on the internal validity of the results. In short, studies employing highly selected patient samples are, despite their limited external validity, regularly accepted as valid sources of evidence bearing on a drug's effectiveness. There are exceptions, however; this paper describes one in which the use of a seemingly innocuous sample enrichment maneuver proved highly damaging to the ultimate credibility of an important multicenter trial. In particular, exposure to an experimental treatment during an open qualification phase may invalidate drug-placebo comparisons made during a later randomized, blinded, controlled phase. Our review of the trial also reveals that the enrichment maneuver employed probably failed to accomplish its intended aims, selecting patients whose improvements on the outcome variable may be as reasonably ascribed to chance as to drug effect. This is all the more surprising because the method of sample enrichment employed has much in common with those long recommended in the clinical trial literature.

Aged↗

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans↗

Risk Factors and Predictive Model for Postoperative High Myopia in Children Undergoing Congenital Cataract Surgery With Intraocular Lens Implantation.

PURPOSE: To identify risk factors associated with the development of high myopia following congenital cataract surgery and to establish a robust predictive model. DESIGN: Retrospective clinical cohort study. SUBJECTS: This retrospective study included 106 pediatric patients who underwent congenital cataract surgery with primary IOL implantation (mean follow-up 8.19 years). The model was externally validated in an independent cohort of 72 patients with a mean follow-up of 7.83 years. METHODS: Preoperative and postoperative ocular biometric parameters were collected. Risk factors for postoperative high myopia were analyzed using Cox proportional hazards regression, which served as the basis for model construction. The predictive performance of the model was rigorously evaluated for discrimination and calibration. Discriminative ability was quantified using Harrell's C-index and the area under the receiver operating characteristic curve (AUC). Model calibration was assessed via calibration plots by comparing predicted probabilities with actual observed outcomes. Internal validation was performed using a bootstrapping method (500 iterations) to ensure model stability and adjust for potential overfitting. RESULTS: An initial postoperative refraction of <+0.75D, and a higher IOL Power to Axial length Ratio (IOL/AL ratio) were identified as significant risk factors for the development of postoperative high myopia. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. The predictive model demonstrated robust performance, achieving a C-index of 0.711 (internal validation C-index: 0.713). The area under the receiver operating characteristic curve (AUC) values for predicting high myopia at 5 and 10 years were 0.858 and 0.745, respectively. Furthermore, calibration curves demonstrated excellent agreement between the predicted and observed outcomes throughout the follow-up period. In external validation, the model achieved a C-index of 0.825, 5-year AUC of 0.833, and 10-year AUC of 0.713. CONCLUSIONS: Our analysis established that initial postoperative refraction <+0.75D, and an elevated IOL/AL ratio are key determinants of high myopia risk following surgery. Shorter preoperative axial length was associated with a greater magnitude of postoperative myopic shift. This predictive framework provides clinicians with a practical tool to optimize preoperative IOL selection and identify high-risk infants who require vigilant myopia prevention and balanced amblyopia management.

Humans↗

The Major Depression Rating Scale (MDS). Inter-rater reliability and validity across different settings in randomized moclobemide trials. Danish University Antidepressant Group.

The Major Depression Rating Scale (MDS) has been derived from the Hamilton Depression Scale and the Melancholia Scale. The MDS contains the nine DSM-IV items for major depression which all have anchoring scores from 0 to 4; hence, the theoretical score range is up to 36. The Major Depression Rating Scale has in this study been psychometrically analysed in randomized moclobemide trials. The results showed that the MDS had higher internal validity than the Hamilton Depression Scale. Thus, the homogeneity of the items was higher; factor analysis identified only one general depression factor (after 4 weeks of treatment explaining more than 50% of the variance). The inter-rater reliability of the two scales was of the same high level. The ability to measure changes (external validity) was tested in randomized clinical trials with moclobemide versus tricyclics (clomipramine and notriptyline) performed in Denmark in the psychiatric setting as well as in the general practice. The results showed that in the psychiatric setting tricyclics were superior to moclobemide with effect sizes ranging between 0.43 and 0.53. The highest effect size was obtained with the Melancholia Scale and the Major Depression Rating Scale, while the Hamilton Depression Scale was below 0.50. In the general practice setting no difference was found between moclobemide and clomipramine. In conclusion, the Major Depression Rating Scale has been found to have a more homogeneous factor structure than the Hamilton Depression Scale, but still with the same level of reliability and external validity. However, studies are needed to standardize the scale, especially in the general practice setting.

Antidepressive Agents↗

Binary classification of dyslipidemia from the waist-to-hip ratio and body mass index: a comparison of linear, logistic, and CART models.

BACKGROUND: We sought to improve upon previously published statistical modeling strategies for binary classification of dyslipidemia for general population screening purposes based on the waist-to-hip circumference ratio and body mass index anthropometric measurements. METHODS: Study subjects were participants in WHO-MONICA population-based surveys conducted in two Swiss regions. Outcome variables were based on the total serum cholesterol to high density lipoprotein cholesterol ratio. The other potential predictor variables were gender, age, current cigarette smoking, and hypertension. The models investigated were: (i) linear regression; (ii) logistic classification; (iii) regression trees; (iv) classification trees (iii and iv are collectively known as "CART"). Binary classification performance of the region-specific models was externally validated by classifying the subjects from the other region. RESULTS: Waist-to-hip circumference ratio and body mass index remained modest predictors of dyslipidemia. Correct classification rates for all models were 60-80%, with marked gender differences. Gender-specific models provided only small gains in classification. The external validations provided assurance about the stability of the models. CONCLUSIONS: There were no striking differences between either the algebraic (i, ii) vs. non-algebraic (iii, iv), or the regression (i, iii) vs. classification (ii, iv) modeling approaches. Anticipated advantages of the CART vs. simple additive linear and logistic models were less than expected in this particular application with a relatively small set of predictor variables. CART models may be more useful when considering main effects and interactions between larger sets of predictor variables.

Adult↗