Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

[Validity of an ELISA test for CD4+ T lymphocyte count and validity of total lymphocyte count in the assessment of immunodeficiency status in HIV infection].

A newly available commercial ELISA (TRAx CD4, T Cell Diagnostics USA) for enumerating CD4+ T lymphocytes has been evaluated with blood samples of 105 HIV seropositive and 6 seronegative subjects. Results from the flow cytometric analysis were used as reference. The sensitivity and specificity of the ELISA to identify HIV seropositive subjects having less than 200 CD4+ T lymphocytes/microliters were assessed and studied using the ROC curve. The reproducibility of the ELISA test was analyzed on 40 samples. The results of the ELISA correlated well with these of the flow cytometric analysis (r = 0.79, p < 0.001). However, the ELISA test tends to overestimate the true CD4 count in HIV seropositives. This overestimation could not be explained by the aspecific contribution of monocytic CD4. The threshold for identifying HIV seropositive subjects with less than 200 CD4+ T lymphocytes with a maximum sensitivity and specificity was determined with ROC curve and equalled 400 cell equivalents with the ELISA (sensitivity and specificity were equal to 80%) and 1,450 lymphocytes/microliters with the total absolute lymphocyte count (sensitivity and specificity were equal to 75%). Using this curve, a threshold of 300 cell equivalents for the ELISA test and of 1,100 lymphocytes/microliters for the absolute lymphocyte count was shown to maximize the specificity (> 95%) without a significant loss of sensitivity.

CD4-Positive T-Lymphocytes↗

Lessons learned from validation of in vitro toxicity test: from failure to acceptance into regulatory practice.

As no scientific approach or regulatory guidelines existed for the experimental validation of in vitro toxicity tests, in 1990 a US/European validation workshop agreed in Amden (Switzerland) on a simple definition of the validation process. Several international validation studies failed, although they were conducted according to these recommendations. Taking into account the lessons learned from this experience, a second validation workshop was held by ECVAM in Amden in 1994 to develop a more precisely defined validation concept. Prevalidation and the development of biostatistically defined prediction models were added as essential elements to the validation process. In 1995/1996 the ECVAM validation procedure was officially accepted by EU member countries and at the international level by the US regulatory agencies and the OECD. The improved validation concept was immediately introduced into ongoing validation studies. In 1996 the ECVAM/COLIPA validation study of the in vitro phototoxicity test, which was conducted according to the ECVAM/OECD validation concept, was finished successfully and in 1998 a supporting study on UV-filter chemicals was undertaken. In 1998 the 3T3 NRU PT in vitro phototoxicity test was the first experimentally validated in vitro toxicity test that was recommended for regulatory purposes by ESAC, the ECVAM Scientific Advisory Committee, and by the DG ENV of the EU Commission. Meanwhile, two in vitro skin corrosivity tests have successfully been validated by ECVAM. Finally, in June 2000 the three experimentally validated tests were accepted by EU member states for regulatory purposes as the first in vitro toxicity tests. In addition, ECVAM has funded a successful validation study of three in vitro embryotoxicity tests, which was conducted in 12 European laboratories and finished in July 2000. The three tests validated in this study were the whole embryo culture (WEC) test applied to rat embryos, the micromass (MM) test employing primary cultures of dissociated mouse limb bud cells and the mouse embryonic stem cell test (EST). Examples will be given of successful validation studies during the past decade with particular reference to in vitro toxicity tests that were evaluated for regulatory purposes either by the US validation centre ICCVAM or ECVAM in the fields of sensitisation, phototoxicity and embryotoxicity

Animal Testing Alternatives↗

Measurement validity in physical therapy research.

This article considers the role of measurement validity within physical therapy research. The concept of measurement validity is identified as a component of internal validity, and it is differentiated from the notion of reliability; these concepts are related to systematic and random sources of error, respectively. Using examples from physical therapy and rehabilitation, four main types of validity are reviewed: face validity, criterion-related validity, content validity, and construct validity. The differing implications of these types of validity for quantitative and qualitative research are discussed. Three principal areas of concern are then addressed, based on a critical discussion of selected examples from the literature. First, it is argued that validity is often poorly distinguished from the allied concept of reliability and that purported claims for validity often only demonstrate reliability. Second, it is claimed that validity is too often neglected in favor of reliability, and specific examples relating to gait analysis are put forward to support this argument. Third, some of the methodological difficulties that may occur when attempts are made to demonstrate validity are considered. The article concludes with a plea for a closer focus on the issue of measurement validity within physical therapy research.

Bias↗

The validity of the reinstatement model of craving and relapse to drug use.

RATIONALE: The reinstatement procedure has been used increasingly as a laboratory model of craving and relapse to drug abuse. With the number of reports involving this procedure growing, its validity as a model of relapse merits discussion. OBJECTIVES: The present commentary addresses the validity of the reinstatement procedure in relation to the following three types of models: 1) formal equivalence models, which are assessed on the basis of how well they resemble some phenomenon outside the laboratory (i.e. face validity); 2) correlational models, which are assessed on the basis of how well they predict outcomes of various interventions (such as drug administration or environmental change) when effected outside the laboratory (i.e. predictive validity); and 3) functional equivalence models, which are assessed on the basis of whether the laboratory phenomenon is mechanistically identical or reasonably similar to the phenomenon outside the laboratory (i.e. content validity). METHODS: In order to evaluate the reinstatement model, we briefly examined its various forms and uses, and compared preclinical outcomes to what is known about relapse from the clinical literature. RESULTS. In its most general form, the reinstatement model has reasonable face validity; that is, there is a general agreement in appearance or form of the behavior in the model and the clinical target, relapse. This face validity is generally absent for the procedure when it is used as a model of craving. The predictive validity of the model has not been established. Evidence from studies of treatments for drug relapse have not supported the validity of the model, however from studies of the effects of the presentation of various types of stimuli (e.g. drug "priming") there is mixed evidence supporting predictive validity. With regard to functional equivalence, there is reasonable evidence supporting functional commonalities between drug self-administration in laboratory animals and human drug abusers, which lends support to the validity of the reinstatement model. However, there are several specific areas of departure between the methods and results using the model and clinical practices and observations about relapse, suggesting a lack of functional equivalence. CONCLUSIONS: There is reasonable evidence to support the face validity of the model, but at this time, neither its predictive validity nor functional equivalence has been fully established, which underscores the need for caution in generalizing results from the model to the clinical condition.

Animals↗

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75&#x2009;161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et&#xa0;al., Nanda et&#xa0;al., Naylor et&#xa0;al., and Van Leeuwen et&#xa0;al., each showing fair discrimination. The Teede et&#xa0;al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et&#xa0;al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et&#xa0;al. and van Leeuwen et&#xa0;al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans↗

Internal and external validation of predictive models: a simulation study of bias and precision in small samples.

We performed a simulation study to investigate the accuracy of bootstrap estimates of optimism (internal validation) and the precision of performance estimates in independent validation samples (external validation). We combined two data sets containing children presenting with fever without source (n=376+179=555; 120 bacterial infections). Random samples were drawn from this combined data set for the development (n=376) and validation (n=179) of logistic regression models. The models included statistically significant predictors for infection selected from a set of 57 candidate predictors. Model development, including the selection of predictors, and validation were repeated in a bootstrapping procedure. The resulting expected optimism estimate in the receiver operating characteristic (ROC) area was compared with the observed optimism according to independent validation samples. The average apparent ROC area was 0.74, which was expected (based on bootstrapping) to decrease by 0.07 to 0.67, whereas the observed decrease in the validation samples was 0.09 to 0.65. Omitting the selection of predictors from the bootstrap procedure led to a severe underestimation of the optimism (decrease 0.006). The standard error of the observed ROC area in the independent validation samples was large (0.05). We recommend bootstrapping for internal validation because it gives reasonably valid estimates of the expected optimism in predictive performance provided that any selection of predictors is taken into account. For external validation, substantial sample sizes should be used for sufficient power to detect clinically important changes in performance as compared with the internally validated estimate.

Bias↗

Rational experimental design for bioanalytical methods validation. Illustration using an assay method for total captopril in plasma.

Generally, bioanalytical chromographic methods are validated according to a predefined programme and distinguish a pre-validation phase, a main validation phase and a follow-up validation phase. In this paper, a rational, total performance evaluation programme for chromatographic methods is presented. The design was developed in particular for the pre-validation and main validation phases. The entire experimental design can be performed within six analytical runs. The first run (pre-validation phase) is used to assess the validity of the expected concentration-response relationship (lack of fit, goodness of fit), to assess specificity of the method and to assess the stability of processed samples in the autosampler for 30 h (benchtop stability). The latter experiment is performed to justify overnight analyses. Following approval of the method after the pre-validation phase, the next five runs (main validation phase) are performed to evaluate method precision and accuracy, recovery, freezing and thawing stability and over-curve control/dilution. The design is nested, i.e., many experimental results are used for the evaluation of several performance characteristics. Analysis of variance (ANOVA) is used for the evaluation of lack of fit and goodness of fit, precision and accuracy, freezing and thawing stability and over-curve control/dilution. Regression analysis is used to evaluate benchtop stability. For over-curve control/dilution, additional to ANOVA, also a paired comparison is applied. As a consequence, the recommended design combines the performance of as few independent validation experiments as possible with modern statistical methods, resulting in optimum use of information. A demonstration of the entire validation programme is given for an HPLC method for the determination of total captopril in human plasma.

Calibration↗

Systematic validation of disease models for pharmacoeconomic evaluations. Swiss HIV Cohort Study.

Pharmacoeconomic evaluations are often based on computer models which simulate the course of disease with and without medical interventions. The purpose of this study is to propose and illustrate a rigorous approach for validating such disease models. For illustrative purposes, we applied this approach to a computer-based model we developed to mimic the history of HIV-infected subjects at the greatest risk for Mycobacterium avium complex (MAC) infection in Switzerland. The drugs included as a prophylactic intervention against MAC infection were azithromycin and clarithromycin. We used a homogenous Markov chain to describe the progression of an HIV-infected patient through six MAC-free states, one MAC state, and death. Probability estimates were extracted from the Swiss HIV Cohort Study database (1993-95) and randomized controlled trials. The model was validated testing for (1) technical validity (2) predictive validity (3) face validity and (4) modelling process validity. Sensitivity analysis and independent model implementation in DATA (PPS) and self-written Fortran 90 code (BAC) assured technical validity. Agreement between modelled and observed MAC incidence confirmed predictive validity. Modelled MAC prophylaxis at different starting conditions affirmed face validity. Published articles by other authors supported modelling process validity. The proposed validation procedure is a useful approach to improve the validity of the model.

AIDS-Related Opportunistic Infections↗

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (&#x2264;&#x2009;12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n&#x2009;=&#x2009;121, 19 events) for training and centers 2-7 (n&#x2009;=&#x2009;207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans↗

Validity and the OSCE.

In preparation for a celebration of '30 years of OSCEs' held during the 2002 meeting of the Association for Medical Education in Europe (AMEE), I was asked to discuss the question, 'Are OSCEs valid to assess competence?". My first instinct was to review work undertaken in ay countries by famous researchers such as Harden, Colliver, Rothman, van der Vleuten, Stillman, Tamblyn and others who have studied and written about the validity of OSCEs. I could then have reviewed the extensive literature produced in Canada by the Medical Council of Canada and in the United States by the Education Commission on Foreign Medical Graduates and National Board of Medical Examiners that has demonstrated the utility of large-scale OSCEs for certification and licensure. I might have tossed in a few papers from my own research on the validity of OSCEs in psychiatry. Indeed, it would have been relatively easy to marshal the medical education literature to answer the question 'Are OSCEs valid to assess competence' strongly in the affirmative. But the more I reflected on the question, the more I confronted concerns that have troubled me for some time. Specifically, I worry that our approaches to validity may themselves not be valid. In this paper, I review what I believe to be three serious problems with our current approaches to showing that 'the OSCE is valid' Let me begin by rethinking the question. What do we mean 'Is the OSCE valid for assessing competence' There are three important problems with this question. First, validity is a property of the application of a test, not of a test itself. Second, we cannot speak of validity without giving consideration to the context in which we use the test. And finally, the concept of validity flounders because the OSCE itself is an important agent in constructing the variables of performance that it is designed to measure. I shall consider each issue in turn.

Canada↗

Content and criterion validity evaluation of National Public Health Performance Standards measurement instruments.

OBJECTIVE: The Centers for Disease Control and Prevention's National Public Health Performance Standards Program (NPHPSP) has developed instruments to measure the performance of local and state public health departments on the 10 "Essential Services of Public Health," which have been tested in several states. This article is a report of the evaluation of the content and criterion validity of the local public health performance assessment instrument, and the content validity of the state public health performance assessment instrument. METHODS: Health department performance is measured using a set of indicators developed for the 10 Essential Services of Public Health and a model standard for each indicator. Content validity of each model standard in the local instrument was addressed by community partners along the following dimensions: the importance of each standard as a measure of the associated Essential Service, its completeness as a measure, and its reasonableness for achievement. All standards for each Essential Service were then judged in terms of their completeness in measuring performance in that service. Content validity of the state instrument was evaluated in a group interview of health department staff members from three states. Criterion validity of the local instrument was assessed for a sample of eight public health departments in Florida and six in New York by examining documentary evidence for selected responses. Criterion validity was also evaluated for a sample of Florida local public health departments and one Hawaii public health department by comparing state health department staffs' judgments of performance against the instrument score. RESULTS: Criterion validity was upheld for a summary performance score on the local instrument, but was not upheld for performance judgments on individual Essential Services. The NPHPSP standards based on the Essential Services have validity for measuring local public health system performance, according to community partners. The model standards are valid measures of state performance, according to state public health departments in three states. CONCLUSIONS: Within the scope of the validity evaluations completed, the NPHPSP state and local performance assessment instruments were found to be valid measures of public health performance.

Attitude of Health Personnel↗