Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Who enrolls in prevention trials? Discordance in perception of risk by professionals and participants.

Internal and external validity problems permeate all intervention studies but are accentuated in primary preventive intervention research, particularly when studies target or recruit individuals based on their risk for psychopathology. Since many people who are at risk do not yet experience distress, they may not perceive the need for intervention. Recruitment tactics based on explaining extent of risk are unlikely to be persuasive and may have negative consequences. If respondents are not motivated to participate, a small or biased subset of the target population will participate in the intervention. Bias is of special concern when those enrolled represent only part of the continuum of risk. Selective enrollment may compromise both internal validity (the interpretation of the research results) and external validity (the generalizability of the findings) of intervention trials in primary prevention. This article discusses the effects of partial enrollment and the resultant bias. It suggests several strategies for increasing the enrollment of the target population and examines some of their ethical ramifications. It also stresses the importance of collecting systematic data documenting how the participants in the intervention differ from the target group as a whole.

Bias↗

Short form of a situational temptation scale for heavy, episodic drinking.

PURPOSE: A short form for situational temptations to drink scale was developed from an original 21-item inventory by Migneault. METHODS: The form measured four hypothesized subscales of temptations on a sample of 348 college drinkers (66% female). Peer pressure, social anxiety, negative affect, and positive/social situations subscales were replicated and reduced. RESULTS: Strong empirical support was found for a hierarchical model, indicating that the four subscales can be summed to provide a global measure of situational temptations. Confirmatory factor results, internal and external validity, and high correlations with the original measures indicate that the short form was as psychometrically valid as the original measure. IMPLICATIONS: Measures of external validity demonstrated the applicability of this measure to heavy drinking prevention programs.

Adult↗

Changes in obsessive/compulsive patients as measured by the Leyton Inventory before and after treatment with clomipramine.

The Leyton Obsessional Inventory has been found to be a useful measure in assessing patients before and after treatment with clomipramine. Mean scores for symptoms and interference altered significantly during the course of treatment. The Leyton Obsessional Inventory, however, lacks external validation owing to the absence of some valid alternative quantification. In the absence of such external validation it seems justifiable to use the mean Leyton score diagnostically but not as a sole indication of severity or response to treatment.

Clinical Trials as Topic↗

Model validation for external doses due to environmental contaminations by the Chernobyl accident.

The objective of the present paper is to validate the deterministic JSP5 model for external exposures to population groups living in the areas contaminated with radionuclides after the Chernobyl accident. For this purpose inhabitants of contaminated areas wore TL-dosimeters for about 1 mo in the spring/summer periods of the years 1989 to 1994. External doses due to the Chernobyl accident were determined from the dosimeter readings by subtracting the natural background. 2,342 results for rural inhabitants and 420 results for inhabitants of the town Novozybkov passed reliability checks. These data show that the average dose in inhabitants of a rural settlement predicted by the model is in the range 0.69-1.55 of the measured values with a confidence level of 95%. Differences are attributed to settlement specific location factors, which are supported by the very good agreement of model and measurements in Novozybkov. In this case location factors of the model were obtained from Novozybkov directly.

Adult↗

Machine Learning-Driven Prediction of Coronary Artery Disease Risk Based on UK Biobank Plasma Proteomics.

BACKGROUND: Coronary artery disease (CAD) is a leading global cause of mortality, yet the predictive accuracy of conventional risk models is limited. Here, we integrate conventional risk factors, polygenic risk scores, and large-scale proteomics to develop a unified model for enhanced CAD risk prediction. METHODS: Using data from UK Biobank, participants with plasma proteomics and genetic risk data were included after excluding prevalent CAD. Participants from England were split into training (n=32 330) and internal validation (n=13 857) sets, and Scotland/Wales participants formed an external validation set (n=5775). Incident CAD was ascertained from linked health records. A 202-protein proteomic risk score was derived by least absolute shrinkage and selection operator Cox regression, and CatBoost models were trained using conventional risk factors alone and with incremental addition of polygenic risk scores and protein proteomic risk scores; Shapley Additive Explanations-guided forward selection identified a compact protein panel. RESULTS: Across cohorts, the median age was 58 years and ∼45% were men. Protein proteomic risk score was dose-dependently associated with CAD risk. Compared with conventional risk factors alone, integrating polygenic risk scores and protein proteomic risk scores improved discrimination, with the area under the curve increasing from 0.750 (95% CI, 0.732-0.767) to 0.789 (95% CI, 0.772-0.805) in internal validation and from 0.717 (95% CI, 0.683-0.750) to 0.762 (95% CI, 0.732-0.791) in external validation. A 9-protein panel (GDF15 [growth differentiation factor 15], MMP12 [matrix metalloproteinase 12], NPPB [natriuretic peptide B], PGF [placental growth factor], REN [renin], ADGRG2 [adhesion G-protein coupled receptor], ACE2 [angiotensin-converting enzyme 2], CDCP1 [CUB domain-containing protein 1], CXCL17 [C-X-C motif chemokine ligand 17)]) captured most proteomic predictive information. CONCLUSIONS: Our findings demonstrate that integrating conventional risk factors, polygenic risk scores, and proteomic data improves CAD risk prediction. This study highlights the utility of proteomics in precision cardiovascular medicine and simplified risk stratification tools.

Humans↗

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n = 26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette ≈ 0.16) that remained unassociated with overall survival (log-rank p = 0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) = 0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p = 0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans↗

Quantification of communication processes, is it possible?

The PAS system (Problem-Analysis-Solution-system) is developed to quantify oral communication processes during counselling in pharmacy practice. The pharmacist translates the patient's drug-related questions into a P-code, the analysis of the question into an A-code and finally the given solution upon the question into a S-code. The PAS system has been developed for two goals. First, for the registation of drug-related questions from patients which gives the pharmacist insight in the most common issues addressed by patients. Second, it might help the pharmacist to structure the communication with the patient during the consultation. Forty-one pharmacists participated in the evaluation of the PAS system. The validation of the PAS system consisted of two phases: the external validation and the internal validation. Kappa values were calculated as a measure of agreement in the coding by the pharmacists. The kappa-value of the external validation for the P-, A- and S-codes for the total set of questions indicate a moderate to poor agreement. This means that pharmacists categorize drug-related questions from patients in a different way. Therefore we conclude that the PAS system is less reliable for research purpose. The kappa-value of the internal validation for the P-code varies from 0.42 to 0.91. For the A-code it varies from 0.07 to 0.35 and for the S-code from zero to 0.68. Internal reproducibility is good for P-code but not for the A-code and S-code. This implies that the pharmacist can use the P-codes for registration of patients' questions in his own pharmacy. Moreover, the usage of the PAS system during counselling in pharmacy practice can structure the consultation.

Communication↗

Treatment satisfaction of patients with lower urinary tract symptoms: randomised controlled trials vs. real life practice.

Randomised controlled trials (RCTs) are an important scientific tool to determine the efficacy and tolerability of a given treatment relative to placebo or other treatment forms. However, due to strict inclusion and exclusion criteria the patient populations in RCTs may not be fully representative for those routinely consulting the physician. Moreover, participation in a formal study puts physician and patient in a situation where they may react different than in real life. In contrast real life practice (RLP) studies cannot determine treatment efficacy or tolerability in absolute terms since they typically do not include a control group and are purely observational. On the other hand, they tend to be more representative for real treatment outcomes. Thus, RCTs have high internal but less external validity whereas RLP studies have less internal and greater external validity. Hence, RCTs and RLP studies should not be considered as mutually exclusive but rather as complementing each other. Specific advantages and disadvantages of RCTs and RLP studies will be discussed using published evidence for the treatment of lower urinary tract symptoms suggestive of benign prostatic obstruction with alpha1-adrenoceptor antagonists and other treatments.

Humans↗

Cost estimates for hospital inpatient care in Australia: evaluation of alternative sources.

OBJECTIVE: This paper presents a framework for evaluation of alternative sources of estimates of the costs of hospital inpatient care in Australia. It argues that the choice of costing methods depends on the decision-context and the sensitivity of the decision to estimation errors. METHOD: Five criteria are proposed for evaluation of sources of hospital cost data, with detailed consideration of the way estimates are derived in two computerised approaches which use accounting data. Three broad approaches to cost estimation are evaluated against these criteria. RESULTS: Choosing an estimation method entails an optimisation analysis for each decision context. 'Microcosting' techniques remains the most valid approach to cost estimation, but are costly and this may, in turn, limit the sample of patients or institutions. Protocol-based cost estimates vary widely in their validity, depending on source data, but there is little justification for continued use of crude per diem cost estimates in such protocols. When precision and resolution are important objectives, clinical costing approaches provide the most valid inpatient cost estimates at a reasonable data cost. When external validity is important, or where standardisation of hospital costs is desired, use of published national cost weights may be preferred. CONCLUSION: Both primary and secondary sources of cost data must withstand challenges to internal and external validity. The 'resolution' (or precision) of cost estimates and the relative costs of collection must also be considered. IMPLICATIONS: Studies using estimates of the costs of hospital care should defend the appropriateness of the costing approach and data source for the decision context.

Accounting↗

Treatment research at the crossroads: the scientific interface of clinical trials and effectiveness research.

OBJECTIVE: Policy and clinical management decisions depend on data on the health and cost impacts of psychiatric treatments under usual care, i.e., effectiveness. Clinical trials, however, provide information on treatment efficacy under best-practice conditions. An understanding of the design, analysis, and conventions of both efficacy and effectiveness studies can lead to research that better informs clinical and societal questions. METHOD: This paper contrasts the strengths and limitations of clinical trials and effectiveness studies for addressing policy and clinical decisions. These research approaches are assessed in terms of outcomes, treatments, service delivery context, implementation conventions, and validity. RESULTS: Clinical trials and effectiveness research share problems of internal and external validity despite more attention to internal validity in clinical trials (e.g., randomization, blinding, standardized protocols) and to external validity in effectiveness studies (e.g., community-based treatments, representative samples). CONCLUSIONS: To develop research at the interface of clinical trials and effectiveness studies, research goals must be redefined, and methods, such as cost-utility and econometric analyses, must be shared and developed. Development of hybrid designs that combine features of efficacy and effectiveness research will require separation of conventions such as frequency of follow-up, intensity of measurement, and sample size from the central scientific issues of aims and validity.

Clinical Protocols↗

Evaluation and application of models for the prediction of ready biodegradability in the MITI-I test.

Three existing models and one newly developed model for the prediction of ready biodegradability of organic compounds are evaluated by comparing the descriptors they use, and the consistency of the models when applied to the set of High Production Volume Chemicals (HPVC) in the European Union. Linear regression models developed for the OECD showed the best performance in the external validation (84.7% correct), although comparison with the other three models is flawed because of the class specificity of these models. With these models 567 of the 894 compounds could be predicted in the validation. The multivariate statistical model showed the best performance in the external validation (82.7% correct) combined with the broadest applicability of the model. The evaluation of the predictions of the models for the HPVC shows that all models are highly consistent in their prediction of not-ready biodegradability, but much less consistency is seen in the prediction of ready biodegradability. This complies with the observation that all 4 models show better performance in their predictions of not-ready biodegradability.

Analysis of Variance↗

Randomized and non-randomized patients in clinical trials: experiences with comprehensive cohort studies.

In clinical research, randomized trials are widely accepted as the definitive method of evaluating the efficacy of therapies. Random assignment of patients to treatment ensures internal validity of the comparison of new treatments with controls. An assessment of external validity can best be achieved by comparing the randomized study sample to the population of patients who met the eligibility criteria but did not consent to randomization. The Comprehensive Cohort Study (CCS) is designed to recruit all patients fulfilling the clinical eligibility criteria regardless of their consent to randomization. The CCS concept was adopted in the major clinical trials of the German Breast Cancer Study Group (GBSG) conducted between 1983 and 1989. In this period 124 centres recruited 2084 patients in three clinical trials. 734 (35 per cent) of these patients accepted being randomized, while 1350 (65 per cent) chose one of the treatments under study; the randomization rates differed remarkably between trials. In this paper we examine the representativeness of the randomized patients in the three trials. Based on a median follow-up of about 5 years we present results on the external validity of the treatment effects estimated in the randomized patients by means of Cox's proportional hazards model and compare them between trials. We discuss advantages and disadvantages of the CCS design and conclude that its use is only justified under extraordinary circumstances.

Breast Neoplasms↗

Pros and cons of permutation tests in clinical trials.

Hypothesis testing, in which the null hypothesis specifies no difference between treatment groups, is an important tool in the assessment of new medical interventions. For randomized clinical trials, permutation tests that reflect the actual randomization are design-based analyses for such hypotheses. This means that only such design-based permutation tests can ensure internal validity, without which external validity is irrelevant. However, because of the conservatism of permutation tests, the virtues of permutation tests continue to be debated in the literature, and conclusions are generally of the type that permutation tests should always be used or permutation tests should never be used. A better conclusion might be that there are situations in which permutation tests should be used, and other situations in which permutation tests should not be used. This approach opens the door to broader agreement, but begs the obvious question of when to use permutation tests. We consider this issue from a variety of perspectives, and conclude that permutation tests are ideal to study efficacy in a randomized clinical trial which compares, in a heterogeneous patient population, two or more treatments, each of which may be most effective in some patients, when the primary analysis does not adjust for covariates. We propose the p-value interval as a novel measure of the conservatism of a permutation test that can be defined independently of the significance level. This p-value interval can be used to ensure that the permutation test have both good global power and an acceptable degree of conservatism.

Humans↗

Factor structure and clinical validity of competing models of positive symptoms in schizophrenia.

BACKGROUND: The factor structure of four competing models of positive symptoms and their clinical validity was studied in a sample of 253 schizophrenia inpatients. METHODS: The following models were tested using confirmatory factor analysis: a one-dimension severity model, a two-dimension model comprising a psychosis factor and a disorganization factor, a four-dimension model based on the Scale for the Assessment of Positive Symptoms (SAPS) structure in subscales, and a five-dimension model derived from the previous one by further differentiating Schneiderian delusions from non-Schneiderian ones. RESULTS: More complex multifactorial models fit the data better than simpler models. The five-dimension model was the best adjusted (goodness of fit index = .844, nonnormed fit index = .812, normed fit index = .728). Whereas the one-dimension model did not display significant association with the clinical variables, multidimensional models were related to age at onset and illness severity. The two-dimension model captured well the clinical correlates of the more complex models. CONCLUSION: None of the tested models showed good fit to the data. The one-dimension model displayed both poor factor validity and poor external validity; therefore, research relying on the SAPS total score may reach misleading conclusions.

Adult↗

The future of pharmacoeconomics: bridging science and practice.

In the context of new challenges, issues facing the science, practice, and future of pharmacoeconomics will be discussed. Certain methodologic weaknesses have been observed in published pharmacoeconomic studies, and compromises need to be made between developing an "ideal" method and allowing a study to remain practicable. The objective is to reach a balance between clinical trial-based studies and projective models; trials have high internal validity but low external validity, while models can help explore relevance to real-life settings. Cross-national differences also have an important impact on pharmacoeconomic data; however, using some basic standardized guidelines results from pharmacoeconomic studies may be generalized to other settings. The use of pharmacoeconomic results by decision makers in the United Kingdom has been restrained by unclear priorities within their authority and by the limited availability of credible studies. The future of pharmacoeconomics lies in developing both trial-based and modeling studies, improving their credibility, and meeting the needs of decision makers.

Clinical Trials as Topic↗

Empirically supported psychosocial interventions for children: an overview.

Discusses issues related to the identification of psychosocial interventions for children that have demonstrated efficacy. Recent debate concerning differences between clinical trials research and clinical practice is summarized, including the tradeoff between interpretability (internal validity) and generalizability (external validity) of outcome studies. This article serves as an introduction to the special issue containing articles that have as their focus the identification of empirically supported psychosocial interventions for children as part of a task force. The article provides an overview of the history, agenda, and methodology used by the task force to define and identify specific empirically supported interventions for children with specific disorders. Whereas a number of well-established or probably efficacious interventions are identified within the series, more work directed at closing the gap between research and practice is needed.

Adolescent↗

Analysis of randomized and nonrandomized patients in clinical trials using the comprehensive cohort follow-up study design.

In clinical research, randomized trials are widely accepted as the definitive method of evaluating the efficacy of therapies. The random assignment of patients to their treatment ensures the internal validity of the comparison of new treatments with controls. An assessment of the external validity of trial results can best be achieved by comparing the study population to the population of patients who met the eligibility criteria but did not consent to randomization. A part of the data of the Coronary Artery Surgery Study (CASS), in which coronary artery bypass surgery is compared to conventional medical therapy in patients with coronary artery disease, is used to illustrate a strategy of multivariate analysis of randomized and nonrandomized patients which allows an investigation of both internal and external validity. The method used Cox's proportional hazards regression model with inclusion of covariates for randomization status and corresponding interactions in addition to the usual covariates for treatment and the important prognostic factors.

Cohort Studies↗

[Validity of the clinical prediction rule for the diagnosis of renal arterial stenosis in hypertensive patients resistant to treatment].

PURPOSE: To perform an external validation of the clinical prediction rule established by Krijnen et al. (Ann Intern Med 1998; 129: 705-11) designed to identify renal artery stenoses (RAS) in hypertensive patients. METHODS: We included 102 patients with a refractory hypertension treated with at least two antihypertensive drugs. All subjects had the research of RAS by renal angiography, or angio-computed tomography, or doppler ultrasound. Probability to detect RAS was calculated with Krijnen's algorithm (Pre-test probability) from the following parameters: age, smoking status, diffuse atherosclerosis, recent hypertension (< 2 y), obesity (BMI > 25), abdominal bruit, hypercholesterolemia (> 6.5 mmol/L), creatinine. ROC curves were plotted for each pre-test probability value. A "post-test probability" was obtained from the likelihood ratio calculated at each pre-test probability level. RESULTS: RAS prevalence in this population was 49%. Area under the ROC curve was 0.79 and Youden index was maximal for a pre-test probability of 15%. Maximal likelihood ratio was obtained for a pre-test probability of 46%. Table shows post-test probability as a function of pre-test probability obtained with Krijnen's algorithm. [table: see text] CONCLUSION: Krijnen's algorithm is valid in a population of resistant hypertensives treated with a bi-therapy. This external validation obtained on a population with a high prevalence of RAS should also be tested on a population with a lower prevalence of SAR.

Age Factors↗