Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Reliability of logP predictions based on calculated molecular descriptors: a critical review.

Correct QSAR analysis requires reliable measured or calculated logP values, being logP the most frequently utilized and most important physico-chemical parameter in such studies. Since the publication of theoretical fundamentals of logP prediction, many commercial software solutions are available. These programs are all based on experimental data of huge databases therefore the predicted logP values are mostly acceptable - especially for known structures and their derivatives. In this study we critically reviewed the published methods and compared the predictive power of commercial softwares (CLOGP, KOWWIN, SciLogP/ULTRA) to each other and to our recently developed automatic QS(P)AR program. We have selected a very diverse set of 625 known drugs (98%) and drug-like molecules with experimentally validated logP values. We have collected 78 reported "outliers" as well, which could not be predicted by the "traditional" methods. We used these data in the model building and validation. Finally, we used an external validation set of compounds missing from public databases. We emphasized the importance of data quality, descriptor calculation and selection, and presented a general, reliable descriptor selection and validation technique for such kind of studies. Our method is based on the strictest mathematical and statistical rules, fully automatic and after the initial settings there is no option for user intervention. Three approaches were applied: multiple linear regression, partial least squares analysis and artificial neural network. LogP predictions with a multiple linear regression model showed acceptable accuracy for new compounds therefore it can be used for "in-silico-screening" and/or planning virtual/combinatorial libraries.

Combinatorial Chemistry Techniques↗

[Advances of the investigation of "Ambulatory Care Sensitive Conditions" in primary care in Spain].

Hospitalization due to ambulatory care sensitive conditions (ACSC) is an indicator of hospital activity that has demonstrated its usefulness as an indirect measurement of primary care effectiveness. Since this indicator was recently introduced in Spain, a collaborative effort between the different research groups could facilitate and promote its development and progress. The objective of this paper is to propose a working agenda that, starting from the most recent information, enhances the advance in this research field. The agenda includes the following sections: 1) To draw up specific ACSC lists for adult and pediatric population, as well as to look in greater depth into the concepts of, and differences in avoidable hospitalization and ACSC. 2) To complete the indicator validation process by assessing the external validity. 3) To propose, for future studies, the municipality as the unit of analysis, as well as to individualize the analysis of health conditions allowing for the differences between acute and chronic ones. 4) To adjust the indicators of hospital activity by hospital use index, when data from some hospitals are lacking and comparisons are wanted 5) To include a new variable, provider of primary health care services, in the Minimum Basic Data Set of Hospital Discharges. 6) To use this indicator as a measure of both the distribution of functions between levels of care and the coordination among them.

Ambulatory Care↗

Molecular staging of lymph nodes from patients with esophageal adenocarcinoma.

PURPOSE: This study was designed to evaluate molecular markers for the detection of micrometastasis in esophageal adenocarcinoma, define algorithms to distinguish positive from benign lymph nodes and to validate these findings in an independent tissue set and in patients with p(N0) esophageal adenocarcinoma. EXPERIMENTAL DESIGN: Potential markers were identified through literature and database searches. All markers were analyzed by quantitative reverse transcription (QRT)-PCR on a limited set of primary tumors and benign lymph nodes. Selected markers were further evaluated on a larger tissue set and classification algorithms were generated for individual markers and combinations. Algorithms were statistically validated internally as well as externally on an independent set of lymph nodes. Selected markers were then used to identify occult disease in lymph nodes from 34 patients with p(N0) esophageal adenocarcinoma. RESULTS: Thirty-nine markers were evaluated, six underwent further analysis and five were analyzed in the external validation study. Two markers provided perfect classification in both the screening and validation sets, although parametric bootstrap analysis estimated 2% to 3% optimism in the observed classification accuracy. Several marker combinations also gave perfect classification in the observed data sets, and estimates of optimism were lower, implying more robust classification than with individual markers alone. Five of thirty-four patients with esophageal adenocarcinoma had positive nodes by multimarker QRT-PCR analysis and disease-free survival was significantly worse in these patients (P = 0.0023). CONCLUSIONS: We have identified novel QRT-PCR markers for the detection of occult lymph node disease in patients with esophageal adenocarcinoma. The objective nature of QRT-PCR results, and the ability to detect occult metastases, make this an attractive alternative to routine pathology.

Adenocarcinoma↗

[Is the randomized controlled trial overvalued as a basis for clinical decision-making? A review with comments].

The randomized controlled trial (RCT) may have considerable limitations in clinical research. Lacking the possibility of blinding impairs the internal validity of the trials. The external validity is often impaired, as results of RCTs obtained in an ideal situation, may be difficult to generalize to a clinical routine situation. Pragmatic randomized trials move from ideal situations towards routine situations, and by modifying the design it is possible to reduce selection bias due to patient and physician preferences. Quasi-experimental studies have varying degrees of problems with internal validity but are necessary contributions to our knowledge of the effect of treatment in clinical routine situations. Limitations of the usefulness of RCTs as well as pragmatic and quasi-experimental studies in clinical research make it necessary to recognise that different methods complement one another. Research in development of RCTs and new methods in clinical research should be encouraged.

Decision Making↗

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n = 38, 74%). Hierarchical clustering (n = 20) and K-means clustering (n = 14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans↗

A topological substructural approach for the prediction of P-glycoprotein substrates.

A topological substructural molecular design approach (TOPS-MODE) has been used to predict whether a given compound is a P-glycoprotein (P-gp) substrate or not. A linear discriminant model was developed to classify a data set of 163 compounds as substrates or nonsubstrates (91 substrates and 72 nonsubstrates). The final model fit the data with sensitivity of 82.42% and specificity of 79.17%, for a final accuracy of 80.98%. The model was validated through the use of an external validation set (40 compounds, 22 substrates and 18 nonsubstrates) with a 77.50% of prediction accuracy; fivefold full cross-validation (removing 40 compounds in each cycle, 80.50% of good prediction) and the prediction of an external test set of marketed drugs (35 compounds, 71.43% of good prediction). This methodology evidenced that the standard bond distance, the polarizability and the Gasteiger-Marsilli atomic charge affect the interaction with the P-gp; suggesting the capacity of the TOPS-MODE descriptors to estimate the P-gp substrates for new drug candidates. The potentiality of the TOPS-MODE approach was assessed with a family of compounds not covered by the original training set (6-fluoroquinolones), and the final prediction had a 77.7% of accuracy. Finally, the positive and negative substructural contributions to the classification of 6-fluoroquinolones, as P-gp substrates, were identified; evidencing the possibilities of the present approach in the lead generation and optimization processes.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Development and validation of a nomogram for predicting outcome of patients with vulvar cancer.

OBJECTIVE: To construct and validate a nomogram to predict relapse-free survival of patients treated for vulvar cancer. METHODS: Data from 244 patients treated for vulvar cancer at a single institution (Creteil, France) were used as a training set to develop and calibrate a nomogram for predicting relapse-free survival and local relapse-free survival. We used bootstrap resampling for the internal validation and we tested the nomogram on an independent validation set of patients (Torino, Italy) for the external validation. RESULTS: The nomograms were based on a Cox proportional hazards regression model. Covariates for the relapse-free survival model included age, T stage, number of metastatic nodes, bilateral lymph node involvement, omission of the lymphadenectomy, margin status, lymphovascular space invasion, and depth of invasion. The concordance indices were 0.85 and 0.83 in the training set before and after bootstrapping, respectively, and 0.83 in the validation set. The predictions of our nomogram discriminated better than did the International Federation of Gynecology and Obstetrics stage (0.83 compared with 0.78, P = .01). The calibration of our nomogram was good. In the validation set, 2-year and 5-year relapse-free survival were well predicted with less than 5% difference between the predicted and observed survivals for each quartile. A nomogram for predicting local relapse was also developed. CONCLUSION: We have developed nomograms for predicting distant and local relapse of vulvar cancer at 2 and 5 years and validated them both internally and externally. These nomograms will be freely available on the International Society for the Study of Vulvovaginal Disease Web site. LEVEL OF EVIDENCE: III.

Adult↗

TOPS-MODE approach for the prediction of blood-brain barrier permeation.

The blood-brain barrier permeation has been investigated by using a topological substructural molecular design approach (TOPS-MODE). A linear regression model was developed to predict the in vivo blood-brain partitioning coefficient on a data set of 119 compounds, treated as the logarithm of the blood-brain concentration ratio. The final model explained the 70% of the variance and it was validated through the use of an external validation set (33 compounds of the 119, MAE = 0.33), a leave-one-out crossvalidation (q(2) = 0.65, S(press) = 0.43), fivefold full crossvalidation (removing 28 compounds in each cycle, MAE = 33, RMSE = 0.43) and the prediction of +/- values for an external test set (85.7% of good prediction). This methodology evidenced that the hydrophobicity increase the blood-brain barrier permeation, while the polar surface and its interaction with the atomic mass of compounds decrease it; suggesting the capacity of the TOPS-MODE descriptors to estimate brain penetration potential of new drug candidates. Finally, by the present approach, positive and negative substructural contributions to the brain permeation were identified, and their possibilities in the lead generation and optimization processes were evaluated.

Blood-Brain Barrier↗

Integrated Genomic and Tumor Microenvironment Subtyping Improved Risk Stratification in Primary Central Nervous System Lymphoma.

Current prognostic models fail to capture the biological complexity of primary central nervous system lymphoma (PCNSL). We integrated whole-genome sequencing and multiplex immunofluorescence in 68 treatment-na&#xef;ve patients to define four genomic subtypes (C1, C2, C3, and C4) with divergent survival (C4 worst: median overall survival [OS], 26&#x2009;months). In parallel, a novel tumor microenvironment (TME) classification based on CD8+T/M2 macrophage ratio stratified patients into High (>&#x2009;1.5), Intermediate (0.8-1.5), and Low (<&#x2009;0.8) groups. Unexpectedly, the Intermediate TME group showed the poorest outcomes (5-year OS: 10%). Integration revealed a lethal subgroup (C4&#x2009;+&#x2009;Intermediate TME; 9.8% of cohort) with a median OS of 3.0&#x2009;months (hazard ratio&#x2009;=&#x2009;7.24, p&#x2009;=&#x2009;0.006). Prognostic nomograms incorporating these subtypes showed promising discriminative performance in internal validation (C-index >&#x2009;0.78), but external validation is needed. Together, these findings identify a high-risk biological subset and provide a hypothesis-generating framework for future biomarker-driven risk stratification and therapeutic discovery in PCNSL.

Humans↗

Prediction of chemical carcinogenicity from molecular structure.

Carcinogens represent a serious threat to human health. In vivo determination of carcinogenicity is time-consuming and expensive, thus in silico models to predict chemical carcinogenicity are highly desirable for virtual screening of compound libraries of both pharmaceutically and other commercially interesting molecules. In the present study, a PLS-DA (partial least squares discriminant analysis) model was developed to predict carcinogenicities in each of four rodent models: male mouse (MM), female mouse (FM), male rat (MR), and female rat (FR). The data set that was used contained over 520 compounds from both the NTP and the FDA databases. All the models were built from the same molecular descriptor system, which is based on atom typing [Sun, H. J. Chem. Inf. Comput. Sci. 2004, 44, 748-757], enabling the comparison of atomic contributions to carcinogenicity with respect to species and gender. Using four components, the models were able to achieve excellent fitting and prediction, with r(2) = 0.987 and q(2) = 0.944 for MM, r(2) = 0.985 and q(2) = 0.950 for FM, r(2) = 0.989 and q(2) = 0.962 for MR, and r(2) = 0.990 and q(2) = 0.965 for FR. The models were further validated by response permutation testing and external validation, and the results indicated that the models were both statistically significant and predictive. Variable influence on projection (VIP) analysis identified the key atom types and fragments that contributed to carcinogenicities and response differences across species and gender.

Animals↗

Age-based construct validation using structural equation modeling.

In this paper we describe some mathematical and statistical models based on structural equation modeling (SEM) using computer programs like LISREL. We focus on SEM methodology for the simultaneous examination of the internal validity of psychological constructs and the external validity represented by age relations. To illustrate these ideas we use a latent variable path model to examine the organization of intellectual abilities measured by the WAIS-R in the standardization sample. We also examine different ways in which age can be used to structure this organization. This is primarily a methodological paper, but we try to integrate conceptual principles of modeling with some substantive issues of research on the psychology of aging.

Aging↗

Confidence intervals versus p-values for interpretation of clinical trial results: introduction.

The following three papers summarize the presentations at a Society for Clinical Trials annual meeting session on the relative merits of estimation versus testing for analysis of randomized clinical trials. By design, randomized clinical trials have internal validity. Whether they also possess quantitative external validity--generalizability of effect size to some population represented by the trial subjects--is one of the main points of disagreement among the three authors. It may be unrealistic to expect a resolution that applies across the wide variety of therapeutic areas and clinical trial goals. Extrapolation from clinical trial to clinical practice is often endorsed in connection with large trials having loose entry criteria and focusing on an objective, clearly meaningful clinical endpoint. By contrast, the relevance of estimates of effect size is less clear in the case of many clinical trials conducted in the course of drug development.

Clinical Trials as Topic↗

Measuring quality of life in children with adenotonsillar disease with the Child Health Questionnaire: a first U.K. study.

OBJECTIVE: To validate the Child Health Questionnaire (CHQ) and assess the quality of life of inner-city British children with adenotonsillar disease. METHODS: The primary caregiver of a consecutive series of 43 patients referred for adenotonsillar disease to a pediatric otolaryngology clinic completed the Child Health Questionnaire. Questionnaires were analyzed for data quality and completeness, items/scale correlation, internal consistency and discriminant validity, interscale correlation, reliability estimates and external validity. RESULTS: CHQ demonstrated excellent measuring characteristics in our population. In a comparison with healthy children, 11 out of 15 measures of quality of life were significantly depressed in our sample. Compared with children with rheumatoid arthritis, scores were equivalent in most areas, with the exception of the global health subscale and overall physical score, where our sample scored significantly lower. CONCLUSION: The CHQ (PF 28 version) is an accurate and reliable way of assessing the impact of adenotonsillar disease on the quality of life in children in Britain. This appears to be quite significant in most aspects of a child's life.

Adolescent↗

A stroke-adapted 30-item version of the Sickness Impact Profile to assess quality of life (SA-SIP30).

BACKGROUND AND PURPOSE: In view of the growing therapeutic options in stroke, measurement of quality of life has become increasingly relevant as an outcome parameters. The Sickness Impact Profile (SIP) is one of the most widely used measures to assess quality of life. To overcome the major disadvantage of the SIP, its length, we constructed a short stroke adapted 30-item SIP version (SA-SIP30). METHODS: Data on the original SIP version were collected for 319 communicative patients at 6 months after stroke. The 12 subscales and the 136 items of the original SIP were reduced to 8 subscales with 30 items in a three step procedure, on the basis of relevancy and homogeneity. Reliability of the SA-SIP30 was evaluated by means of an analysis of homogeneity (Cronbach's alpha coefficient). Different types of validity were assessed: construct, clinical, and external validities. RESULTS: Homogeneity of the SA-SIP30 was demonstrated by a high Cronbach's alpha (0.85). Principal component analyses revealed the same two dimensions as in the original SIP (a physical and a psychosocial dimension). The SA-SIP30 could explain 91% of the variation in scores of the original SIP in the same cohort of patients, and 89% in a different cohort. Furthermore, the SA-SIP30 was related to other functional health measures similar to how the original SIP was. We could demonstrate that the SA-SIP30 was able to distinguish patients with lacunar infarctions from patients with cortical or subcortical lesions. CONCLUSIONS: We conclude that the SA-SIP30 is a feasible and clinimetrically sound measure to assess quality of life after stroke.

Aged↗

An evaluation of the quantity and quality of empirical research in three pastoral care and counseling journals, 1990-1999: has anything changed?

This article summarizes a review of all articles published in Pastoral Psychology, The Journal of Rleigion and Health, and The Journal of Pastoral Care between 1900 and 1999, identifying a total of 737 scholarly articles, of which 165 (22.4%) were research studies. The proportion of research studies, especially quantitative studies, increased significantly between the first and second half of the study period (p < .05). There was a significant positive correlation between compliance with three out of four criteria of internal validity. Three of five criteria of external validity were also positively related to one another. Compared to previous research using identical criteria to assess quantitative studies in the same journals in 1980-1989, the 1990-1999 sample showed improved compliance with respect to specifying the sampling method (p < .001), reporting the response rate (p < .05), and discussing the limitations of research studies (p < .001). However, the overall findings suggest that many researchers in the field do not have a sophisticated knowledge of statistical sampling, statistical analysis, or research design. Several recommendations for increasing the quality of quantitative research are offered.

Bibliometrics↗

SERPINE1-centric inflammatory signature associates with treatment resistance and survival in laryngeal squamous cell carcinoma.

BACKGROUND: Laryngeal squamous cell carcinoma (LSCC) prognosis remains poor despite treatment advances. More accurate prognostic assessment models can help guide individualized treatment and improve prognosis. Chronic inflammation contributes to tumorigenesis, yet inflammatory response-related genes (IRGs) in LSCC prognosis are underexplored. This study aimed to construct an IRG prognostic signature for LSCC and further dissect core IRG-mediated mechanisms of immune escape and chemoresistance. METHODS: Transcriptional profiles and clinical data from LSCC patients were retrieved from The Cancer Genome Atlas (TCGA). IRGs were sourced from Gene Set Enrichment Analysis (GSEA) hallmark gene set. We identified differentially expressed IRGs linked to survival outcomes in LSCC. Key IRGs were subsequently selected using least absolute shrinkage and selection operator (LASSO) Cox regression analysis to establish an inflammatory risk score model. This model underwent internal validation within the TCGA cohort and external validation using independent Gene Expression Omnibus (GEO) datasets. We further assessed the model's association with the tumor immune microenvironment and the impact of IRGs on chemotherapy response. Finally, the functional roles of interested signature IRG were experimentally validated in LSCC cell lines. RESULTS: Four significant IRGs (AQP9, ITGA5, LCK, SERPINE1) were identified to build the risk score model. The model stratified LSCC patients into distinct prognostic groups: TCGA cohort: 5-year area under the curve (AUC) =0.836, P<0.001; GSE25727 cohort: 5-year AUC =0.706, P=0.02; GSE27020 cohort: 5-year AUC =0.798, P<0.01. Multivariate analysis confirmed the risk score as an independent prognostic factor (P<0.05). High-risk patients showed reduced immune cell infiltration (CD8+ T cells, dendritic cells) and suppressed immune pathways. Multi-algorithm immune analysis further revealed defective antigen presentation and reduced anti-tumor immune infiltration in high-risk LSCC, promoting tumor immune escape. GSEA/Gene Ontology (GO) enrichment combined with drug sensitivity prediction further revealed that high-risk tumors activate invasive signaling and acquire broad chemoresistance alongside impaired anti-tumor immunity. SERPINE1 might be associated with chemotherapy resistance and exhibited the highest alteration frequency (predominantly amplification) and overexpression in LSCC tissues. Its knockdown significantly suppressed proliferation, migration, invasion and chemoresistance in LSCC cells. Immunohistochemistry (IHC) confirmed tumor SERPINE1 overexpression (P=0.002 vs. normal tissues), correlating with poor survival (P<0.001). CONCLUSIONS: The 4-IRG risk signature is a reliable prognostic indicator reflecting immune dysfunction in LSCC. SERPINE1 is validated as a therapeutic target and biomarker, enriching our understanding of gene regulation dynamics in LSCC.

Laryngeal cancer↗

Can we use contingent valuation to assess the demand for childhood immunisation in developing countries?: a systematic review of the literature.

Childhood immunisation is one of the most cost-effective public health interventions, yet its population coverage in low- and middle-income countries is severely limited by the fiscal constraints that health services face. A recent proposal suggested that commitments to purchase vaccines and make them available to developing countries for modest co-payments could solve the problem. However, this is dependent on communities being willing and able to share the cost in this way, which is difficult to assess. One possible method to assess this demand is contingent valuation (CV). This article evaluates the usefulness of using CV in this way, by reviewing applications of CV in developing countries against current 'standards' for CV of immunisation in the literature. A structured review was adopted with reference to the standard frameworks for methodological evaluation. A set of five criteria were developed for evaluating an 'acceptable' CV study: (i) response rate; (ii) association between willingness to pay (WTP) and socioeconomic status (SES); (iii) sensitivity of WTP to benefit scale/scope; (iv) predictive validity; and (v) reliability in elicitation formats. Two strands of literature search were conducted using electronic databases (MEDLINE, EMBASE, HEALTHSTAR and Econlit) from 1966 to 2003, one for CV studies of immunisation and one for CV studies in developing countries. Twelve CV studies of vaccination and 13 CV studies undertaken within developing countries were identified and reviewed. The quality of existing CV studies conducted in developing countries exceeded the benchmark standard set by studies of immunisation in the developed world in four of the five criteria. WTP estimates appeared both internally valid (i.e. associations with SES) and externally valid (i.e. predictive validity), reliability in developing countries was no less than that of the benchmark level in the existing literature, and the high response rates suggested that CV can be administered to a rural, and perhaps less literate, population. Only sensitivity to scale/scope was not well demonstrated. Our assessment indicated that the CV technique offers a promising tool to estimate the demand for childhood immunisation in low- and middle-income countries. International agencies are therefore encouraged to devote resources to such an application when designing their support to the immunisation programmes.

Benchmarking↗

Clinical trials of primary care treatments for major depression: issues in design, recruitment and treatment.

The objective of this article is to consider whether randomized clinical trials (RCTs) are able to determine the validity of transferring treatments for major depression from the psychiatric to the primary care sector. This clinical issue is of growing concern in the United States since both governmental and professional bodies are establishing guidelines for the treatment of medical patients with the affective disorder. The article's method involves analysis of how the competing aims of rigorous scientific methodology (internal validity) and generalization of study findings (external validity) are best balanced within the RCT. Experiences in recruiting medical patients with major depression and providing pharmacologic, psychotherapeutic, and usual care interventions compatible with the sociotechnical characteristics of ambulatory medical centers are described to illustrate the complexities of investigating transferability of treatments for major depression with RCT methodology.

Antidepressive Agents↗