Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Psychotherapy for patients with complex disorders and chronic symptoms. The need for a new research paradigm.

BACKGROUND: A clear distinction has been made between efficacy and effectiveness in relation to the methods of evaluation of new psychological treatments in psychiatry. Efficacy trials target patients with relatively pure conditions, who may not be representative of the patients who are usually referred for psychological treatment in a clinical setting. Few studies have explored the benefits of psychotherapy in patients with complex disorders and enduring symptoms. AIMS: To explore the rationale for the distinction between efficacy and effectiveness, particularly in relation to outcome studies of patients with complex and enduring disorders. METHOD: A narrative review with examples drawn from the literature, and an illustration of a recent naturalistic outcome study which combines features of both efficacy and effectiveness. RESULTS: Studies of patients with complex and mixed disorders can be designed so that they retain internal validity, but also have external validity and are relevant to clinical practice. CONCLUSION: Studies which evaluate psychological interventions should be carried out in populations of patients clinically representative of those who are likely to receive the intervention, should it be shown to be of benefit.

Adult↗

Noninvasive detection and differentiation of gastric malignancy using cell-free DNA biomarkers.

INTRODUCTION: Gastric cancer remains a major global health burden, with high mortality driven by late-stage diagnoses that limit treatment options and reduce survival. Current diagnostic methods such as endoscopy and biopsy are invasive, resource-intensive, and impractical for large-scale early detection. OBJECTIVES: This study aimed to develop and validate an ensemble machine learning model integrating four cell-free DNA (cfDNA) fragmentomic feature classes derived from 5 × whole genome sequencing (WGS) data to non-invasively differentiate malignant gastric cancer from benign gastric lesions in high-risk or symptomatic patients. METHODS: A total of 681 plasma samples were prospectively collected, comprising 329 from patients with gastric cancer or high-grade intraepithelial neoplasia (HGIN) and 352 from individuals with benign gastric conditions. The dataset was divided into a training cohort (n = 333) and a temporally independent validation cohort (n = 348). An external validation cohort of 305 participants was also included. RESULTS: The ensemble model achieved an AUROC of 0.920 in cross-validation testing on the training cohort, 0.912 in the independent validation cohort, and 0.896 (95% CI 0.860-0.932) in the external cohort. At a pre-specified prediction threshold of 0.402, the model demonstrated 93.3% sensitivity and 71.9% specificity in the validation cohort, yielding a PPV of 71.3% and an NPV of 93.5%. In the external cohort, sensitivity and specificity were 91.7% and 69.1%, respectively (PPV 75.7%, NPV 88.8%). Model scores correlated with clinical stage, tumor grade, and histopathological subtype. Approximately 71% of non-cancer patients could have been spared unnecessary endoscopy. CONCLUSIONS: The cfDNA fragmentomics-based ensemble model enables accurate, non-invasive differentiation between gastric cancer and benign gastric lesions in high-risk or symptomatic patients. This approach demonstrates strong potential as a pre-endoscopy triage tool, supporting earlier detection and more efficient use of diagnostic resources.

Humans↗

A principal-components analysis of the Narcissistic Personality Inventory and further evidence of its construct validity.

We examined the internal and external validity of the Narcissistic Personality Inventory (NPI). Study 1 explored the internal structure of the NPI responses of 1,018 subjects. Using principal-components analysis, we analyzed the tetrachoric correlations among the NPI item responses and found evidence for a general construct of narcissism as well as seven first-order components, identified as Authority, Exhibitionism, Superiority, Vanity, Exploitativeness, Entitlement, and Self-Sufficiency. Study 2 explored the NPI's construct validity with respect to a variety of indexes derived from observational and self-report data in a sample of 57 subjects. Study 3 investigated the NPI's construct validity with respect to 128 subject's self and ideal self-descriptions, and their congruency, on the Leary Interpersonal Check List. The results from Studies 2 and 3 tend to support the construct validity of the full-scale NPI and its component scales.

Adolescent↗

How far is the preoperative Kattan nomogram applicable for the prediction of recurrence after prostatectomy in patients presenting with PSA levels of more than 20 ng/ml? A validation study.

OBJECTIVE: We present an external validation study investigating the applicability of the preoperative Kattan nomogram for predicting recurrence after prostatectomy in a population of patients with serum prostate-specific antigen (PSA) levels exceeding 20 ng/ml. MATERIALS: In the evaluation of clinical parameters pooled from a total of 191 patients presenting with PSA levels ranging between 20.1 and 100 ng/ml, the PSA-free survival rate 60 months after surgery was calculated according to Kattan nomograms. Subsequently, the results were statistically compared with the corresponding actual survival rates obtained from Kaplan-Meier analysis. For this purpose, the patients were assigned to one of four different risk groups according to predictions derived from the Kattan nomograms, enabling a direct comparison of expected (as predicted by Kattan nomogram) versus actual survival of each patient investigated in our study. RESULTS: Predicted PSA-free survival rates were determined to be as follows: 83% (low risk group); 66% (intermediate risk group); 39% (intermediate-high risk group), and 10% (high risk group) in comparison with the actual survival rates determined to be 63, 62, 40 and 21%, respectively. For PSA levels ranging between 20.1 and 30 ng/ml, 30.1 and 50 ng/ml, and 50.1 and 100 ng/dl, PSA-free survival rates were found to be 57, 37, and 27% (p=0.0017), respectively, during a 5-year post-prostatectomy follow-up. CONCLUSIONS: The Kattan nomogram shows good statistical concordance with actual survival rates in the mean risk quadrants, but considerable differences were demonstrated concerning individuals with either a high or with a low risk of cancer progression.

Aged↗

Validity and reliability: part I.

Internal and external validity of a study determine the generalizability of its findings. Measurement reliability and validity, which also have an impact on the overall validity of the study, will be discussed in the next column. Although the concepts have been pulled apart for this discussion, all of these factors influence the overall confidence with which researchers can say that they have measured what was intended and that the findings can be applied to a larger population of individuals.

Humans↗

Modeling blood-brain barrier partitioning using the electrotopological state.

The challenging problem of modeling blood-brain barrier partitioning is approached through topological representation of molecular structure. A QSAR model is developed for in vivo blood-brain partitioning data treated as the logarithm of the blood-brain concentration ratio. The model consists of three structure descriptors: the hydrogen E-State index for hydrogen bond donors, HS(T)(HBd); the hydrogen E-State index for aromatic CHs, HS(T)(arom); and the second order difference valence molecular connectivity index, d(2)chi(v) (q(2) = 0.62.) The model for the set of 106 compounds is validated through use of an external validation test set (20 compounds of the 106, MAE = 0.33, rms = 0.38), 5-fold cross-validation (MAE = 0.38, rms = 0.47), prediction of +/- values for an external test set (27/28 correct), and estimation of logBB values for a large data set of 20 039 drugs and drug-like compounds. Because no 3D structure information is used, computation of logBB by the model is very fast. The quality of the validation statistics supports the claim that the model may be used for estimation of logBB values for drug and drug-like molecules. Detailed structure interpretation is given for the structure indices in the model. The model indicates that molecules that penetrate the blood-brain barrier have large HS(T)(arom) values (presence of aromatic groups) but small values of HS(T)(HBd) (fewer or weaker H-Bond donors) and smaller d(2)chi(v) values (less branched molecules with fewer electronegative atoms). These three structure descriptors encode influence of molecular context of groups as well as counts of those groups.

Blood-Brain Barrier↗

Predictive ability of level A in vitro-in vivo correlation for ringcap controlled-release acetaminophen tablets.

PURPOSE: The goal of this study was to establish and validate an in vitro-in vivo correlation (IVIVC) for two sustained-release formulations (i.e., a matrix tablet and a RingCap banded matrix tablet) containing 750 mg of acetaminophen. METHODS: The in vitro dissolution and in vivo disposition of these formulations were examined by using a USP type III dissolution apparatus and a single-dose, three-way, crossover study that included an immediate-release acetaminophen dosage form, respectively. An IVIVC was established by using the mean fraction dissolved (FD) and mean fraction absorbed (FA) and used to simulate the plasma concentration-time profile of acetaminophen after administration of the matrix tablet (i.e., internal validation) and RingCap banded matrix tablet (i.e., external validation). RESULTS: A statistically significant relationship (r2 = 0.997, P < 0.001) existed between the FD and FA for matrix tablets and was best described by the equation (FA) = 0.984 x (FD) + 0.0133. The percent predictions errors in CMAX and AUCL were <10% when predicting the plasma concentration-time profiles for the two formulations, validating the internal and external predictability of the IVIVC. CONCLUSIONS: The data (i) show that in vitro dissolution data are a good predictor of in vivo fraction absorbed for acetaminophen, (ii) support the general use of in vitro dissolution data for readily soluble and readily absorbed drugs, (iii) suggest that acetaminophen may serve as a model drug for evaluating novel sustained-release delivery systems, and (iv) provide a tangible example of the limitations of current methods for predicting and validating IVIVC.

Acetaminophen↗

Validation of a comprehensive classification tool for treatment-related problems.

OBJECTIVE: Several drug-related problem classification systems can be found in the literature. However, it is generally agreed that a comprehensive, well constructed and validated instrument is currently lacking. The aim of this study is the development and validation of a comprehensive treatment-related problems assessment and classification tool for use in teaching, practicing and researching pharmaceutical care and to improve identification, resolving and preventing of treatment-related problems. METHOD: The development and validation involved five steps starting with literature search to define a treatment related problem and also to form a database of treatment-related problems identified in the literature. In the next step, all problems that were identified in the first step and passed the evaluation of the three authors were pooled together and then divided into groups according to their common or shared construct, in the third step a suitable assessment method was developed according to the construct of the different problems, in the next step the developed instrument was validated for content, internal and external validity. Finally the tool was finalized and tested for reproducibility and inter-rater agreement. RESULTS: The final validated version included six main categories for treatment-related problems (Indication, Effectiveness, Safety, Knowledge, Adherence and Miscellaneous). These categories include a total of nine subcategories and a total of 29 treatment related problems. CONCLUSION: The treatment-related problems assessment and classification tool introduced in this paper was applied to actual patient cases and proved to be valid. This tool also has several features that are new.

Adverse Drug Reaction Reporting Systems↗

New molecular surface-based 3D-QSAR method using Kohonen neural network and 3-way PLS.

Comparative molecular field analysis (CoMFA) has been widely used as a standard three dimensional quantitative structure-activity relationship (3D-QSAR) method. Although CoMFA is a useful technique, it does not always reflect real ligand-receptor interaction. Molecular interactions between the ligand and receptor are mainly occurred near the van der Waals surface of ligand. All grid points surrounding whole molecule in CoMFA are not important as molecular descriptors. If each molecule is represented by physico-chemical parameters on molecular surface, more precise and realistic 3D-QSAR is possible. We developed a new surface-based 3D-QSAR method using Kohonen neural network (KNN) and three-way partial least squares (3-way PLS). This method was applied to 25 dopamine 2 (D2) receptor antagonists for validation. First, the 3D coordinates of all sampling points on the van der Waals surface were projected into the 2D map by KNN. Each node in the map was coded by the associated molecular electrostatic potential (MEP) value of the original sampling point. Then, the correlation between the MEP values of all 2D maps and D2 receptor antagonist activities was analyzed by 3-way PLS. The statistics of the 3-way PLS model was excellent and the coefficients back-projected on the van der Waals surface had reasonable 3D distribution. Lastly, all data was divided into the calibration and validation sets by D-optimal designs and the activities of validation set were predicted. The external validation suggested that 3-way PLS is better than standard (2-way) PLS for prediction.

Dopamine Antagonists↗

[Registries of morbimortality in cardiology: methods].

The effectiveness of diagnostic, preventive and therapeutic procedures, whose efficacy has been assessed in clinical trials, should be tested in a real treatment scenario. The procedures used in acute myocardial infarction (AMI) management can be evaluated by means of cohort studies that include all consecutive patients admitted to one or several hospitals. Such studies are called hospital registries. They are simpler to organize and cheaper than clinical trials. On the other hand, the AMI population-based registries allow the establishment of the incidence and mortality rates, as well as case-fatality as they include those patients who die before reaching hospital facilities. In both types of registries a set of variables on co-morbidity, age, sex, severity, and the utilization of procedures along with the course of the disease are systematically recorded in each patient using standard definitions to warrant the internal validity. In hospital registries, the external validity of the results will depend on whether the sample of hospitals represents the population where it was obtained. A good registry should include patients with a wide age range, allow the analysis of specific subgroups of patients such as non-Q wave or first AMI to allow for comparison with other registries. In addition, it should also permit a mid-term follow-up, respect ethical issues, receive appropriate funding and keep a multidisciplinary team involved in its design and development.

Cardiology↗

Beware of q2!

Validation is a crucial aspect of any quantitative structure-activity relationship (QSAR) modeling. This paper examines one of the most popular validation criteria, leave-one-out cross-validated R2 (LOO q2). Often, a high value of this statistical characteristic (q2 > 0.5) is considered as a proof of the high predictive ability of the model. In this paper, we show that this assumption is generally incorrect. In the case of 3D QSAR, the lack of the correlation between the high LOO q2 and the high predictive ability of a QSAR model has been established earlier [Pharm. Acta Helv. 70 (1995) 149; J. Chemomet. 10(1996)95; J. Med. Chem. 41 (1998) 2553]. In this paper, we use two-dimensional (2D) molecular descriptors and k nearest neighbors (kNN) QSAR method for the analysis of several datasets. No correlation between the values of q2 for the training set and predictive ability for the test set was found for any of the datasets. Thus, the high value of LOO q2 appears to be the necessary but not the sufficient condition for the model to have a high predictive power. We argue that this is the general property of QSAR models developed using LOO cross-validation. We emphasize that the external validation is the only way to establish a reliable QSAR model. We formulate a set of criteria for evaluation of predictive ability of QSAR models.

Data Interpretation, Statistical↗

QSAR prediction of estrogen activity for a large set of diverse chemicals under the guidance of OECD principles.

A large number of environmental chemicals, known as endocrine-disrupting chemicals, are suspected of disrupting endocrine functions by mimicking or antagonizing natural hormones, and such chemicals may pose a serious threat to the health of humans and wildlife. They are thought to act through a variety of mechanisms, mainly estrogen-receptor-mediated mechanisms of toxicity. However, it is practically impossible to perform thorough toxicological tests on all potential xenoestrogens, and thus, the quantitative structure--activity relationship (QSAR) provides a promising method for the estimation of a compound's estrogenic activity. Here, QSAR models of the estrogen receptor binding affinity of a large data set of heterogeneous chemicals have been built using theoretical molecular descriptors, giving full consideration to the new OECD principles in regulation for QSAR acceptability, during model construction and assessment. An unambiguous multiple linear regression (MLR) algorithm was used to build the models, and model predictive ability was validated by both internal and external validation. The applicability domain was checked by the leverage approach to verify prediction reliability. The results obtained using several validation paths indicate that the proposed QSAR model is robust and satisfactory, and can provide a feasible and practical tool for the rapid screening of the estrogen activity of organic compounds.

Algorithms↗

Improving the evidence-base in surgery: evaluating surgical effectiveness.

This second of two articles about clinical epidemiology reviews the generation and synthesis of evidence for the effectiveness of surgical procedures. While well-designed randomized controlled trials of surgical procedures are considered the 'gold standard' of evaluation design, they may achieve high internal validity at the expense of external validity (generalizability). Improving the -evidence-base in surgery likely will require a comprehensive approach to surgical outcomes assessment, involving both improvements in the quality and quantity of randomized controlled trials as well as recognition of the complementary role of alternate study designs.

Algorithms↗

Prediction of aromatic amines mutagenicity from theoretical molecular descriptors.

In the present research the mutagenicity data (Ames tests TA98 and TA100) for various aromatic and heteroaromatic amines, a data set extensively studied by other quantitative structure-activity relationship (QSAR)-authors, have been modeled by a wide set of theoretical molecular descriptors using linear multivariate regression (MLR) and genetic algorithm-variable subset selection (GA-VSS). The models have been calculated on a subset of compounds selected by a D-optimal experimental design. Moreover, they have been validated by both internal and external validation procedures showing satisfactory predictive performance. The models proposed here can be useful in predicting data and setting a testing priority for those compounds for which experimental data are not available or are not yet synthesized.

Algorithms↗

Is orthotopic bladder replacement the new gold standard? Evidence from a systematic review.

PURPOSE: In this systematic review we determined whether the outcome of orthotopic bladder replacement is superior to that of continent and incontinent urinary diversion. MATERIALS AND METHODS: We searched MEDLINE, PubMed, EMBASE, CINAHL and the Cochrane Library from January 1990 to January 2003. A total of 3,370 abstracts were reviewed, including all types of studies from prospective, randomized, controlled studies to small, retrospective series. All relevant articles with at least 10 patients and a mean followup of at least 1 year were retrieved. There were no language restrictions. NonEnglish articles were translated. Comparisons were made between the major surgery types, including ileal conduit, continent diversion, bladder reconstruction and bladder replacement. All studies were scored using a predetermined quality assessment checklist to assess internal validity (bias and confounding) and external validity. RESULTS: A total of 405 studies met inclusion criteria. There were 32 prospective and 373 retrospective studies describing a total of 32,795 patients. The majority of studies were incompletely or poorly described and outcomes were often not defined. When they were defined, definitions varied. In clinical outcomes ileal conduit diversions had the lowest operative complications rate but highest reported postoperative morbidity. They also had a higher reported incidence of symptomatic urinary tract infections. The rates of postoperative morbidity, mortality and need for reoperation varied widely among studies even for the same procedure. Of physiological outcomes metabolic acidosis was the most commonly reported metabolic complication in patients with various urinary diversions. The quality of the reported literature was poor. There were no studies of the health economic implications of performing 1 type of surgery vs another type. CONCLUSIONS: While enthusiasts regard orthotopic bladder replacement as the new gold standard when lower urinary tract function must be replaced, the level and quality of current evidence are poor. The immediate concern must be to rectify this paucity of evidence with well designed and well reported prospective studies, ideally in a randomized setting, comparing the various major forms of urinary diversion and bladder replacement surgery.

Humans↗

Risk Prognostication After Hypomethylating Agents Combined With Venetoclax in AML: The PRISM Risk Model.

PURPOSE: As risk stratification for patients with AML treated with lower-intensity venetoclax-based therapy remains suboptimal, we developed and validated a prognostic model integrating clinical, cytogenetic, and molecular features. METHODS: We assembled a multinational data set comprising 2,092 adults with newly diagnosed AML treated with hypomethylating agents plus venetoclax (HMA + VEN). One thousand nine hundred eighteen patients with complete data were randomly divided into training (70%) and internal validation (30%) cohorts. Two independent external validation cohorts were assembled (n = 500 and n = 222). Modeling overall survival (OS), Elastic Net regression was applied in 1,000 bootstrap samples from the training cohort to select variables for a Ridge regression, which generated a continuous Prognostic Risk Integration for Survival Modeling (PRISM) score and risk categories based on tertiles (PRISM-3: low, moderate, high). These PRISM indices were then computed for the validation cohorts and compared with the 4-gene classifier (based on mutations in FLT3-ITD, N/KRAS, and TP53). RESULTS: PRISM integrated 17 clinical and genomic variables and demonstrated a linear association with OS. PRISM-3 stratified survival consistently across all cohorts (median OS: 25.1-28.8 months for low risk, 12.5-14.7 months for moderate risk, and 5.8-6.7 months for high risk; P < .001). Compared with the 4-gene classifier, PRISM-3 reassigned approximately 40% of patients (and >50% of those with favorable risk) and demonstrated significantly better discrimination in validation cohorts (C-index 0.63-0.65 v 0.59-0.61; P < .05). CONCLUSION: PRISM is a validated prognostic model for patients with AML receiving HMA + VEN that improves survival risk stratification beyond current standard tools and supports individualized, risk-adapted clinical decision making. The model, the PRISM-AML Risk Calculator, is publicly available.

Humans↗

Evaluation and measurement: some dilemmas for health education.

Seven dilemmas of evaluation and measurement posed by the nature of health education are presented, together with suggestions for their resolution. These include the dilemmas of : 1) rigor of experimental design vs significance or program adaptability; 2) internal validity or "true" effectiveness vs external validity or feasibility; 3) experimental vs placebo effectsl 4) effectiveness vs economy of scale; 5) risk vs payoff; 6) measurement of long-term vs short-term out-comon. Emphasis is placed on the need to develop a more cumulative data base through standardization of measures, replication of experiments in different settings, and better documentation, reporting, and diffusion of experiences in practice.

Cost-Benefit Analysis↗