Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

[Validity of an ELISA test for CD4+ T lymphocyte count and validity of total lymphocyte count in the assessment of immunodeficiency status in HIV infection].

A newly available commercial ELISA (TRAx CD4, T Cell Diagnostics USA) for enumerating CD4+ T lymphocytes has been evaluated with blood samples of 105 HIV seropositive and 6 seronegative subjects. Results from the flow cytometric analysis were used as reference. The sensitivity and specificity of the ELISA to identify HIV seropositive subjects having less than 200 CD4+ T lymphocytes/microliters were assessed and studied using the ROC curve. The reproducibility of the ELISA test was analyzed on 40 samples. The results of the ELISA correlated well with these of the flow cytometric analysis (r = 0.79, p < 0.001). However, the ELISA test tends to overestimate the true CD4 count in HIV seropositives. This overestimation could not be explained by the aspecific contribution of monocytic CD4. The threshold for identifying HIV seropositive subjects with less than 200 CD4+ T lymphocytes with a maximum sensitivity and specificity was determined with ROC curve and equalled 400 cell equivalents with the ELISA (sensitivity and specificity were equal to 80%) and 1,450 lymphocytes/microliters with the total absolute lymphocyte count (sensitivity and specificity were equal to 75%). Using this curve, a threshold of 300 cell equivalents for the ELISA test and of 1,100 lymphocytes/microliters for the absolute lymphocyte count was shown to maximize the specificity (> 95%) without a significant loss of sensitivity.

CD4-Positive T-Lymphocytes↗

Lessons learned from validation of in vitro toxicity test: from failure to acceptance into regulatory practice.

As no scientific approach or regulatory guidelines existed for the experimental validation of in vitro toxicity tests, in 1990 a US/European validation workshop agreed in Amden (Switzerland) on a simple definition of the validation process. Several international validation studies failed, although they were conducted according to these recommendations. Taking into account the lessons learned from this experience, a second validation workshop was held by ECVAM in Amden in 1994 to develop a more precisely defined validation concept. Prevalidation and the development of biostatistically defined prediction models were added as essential elements to the validation process. In 1995/1996 the ECVAM validation procedure was officially accepted by EU member countries and at the international level by the US regulatory agencies and the OECD. The improved validation concept was immediately introduced into ongoing validation studies. In 1996 the ECVAM/COLIPA validation study of the in vitro phototoxicity test, which was conducted according to the ECVAM/OECD validation concept, was finished successfully and in 1998 a supporting study on UV-filter chemicals was undertaken. In 1998 the 3T3 NRU PT in vitro phototoxicity test was the first experimentally validated in vitro toxicity test that was recommended for regulatory purposes by ESAC, the ECVAM Scientific Advisory Committee, and by the DG ENV of the EU Commission. Meanwhile, two in vitro skin corrosivity tests have successfully been validated by ECVAM. Finally, in June 2000 the three experimentally validated tests were accepted by EU member states for regulatory purposes as the first in vitro toxicity tests. In addition, ECVAM has funded a successful validation study of three in vitro embryotoxicity tests, which was conducted in 12 European laboratories and finished in July 2000. The three tests validated in this study were the whole embryo culture (WEC) test applied to rat embryos, the micromass (MM) test employing primary cultures of dissociated mouse limb bud cells and the mouse embryonic stem cell test (EST). Examples will be given of successful validation studies during the past decade with particular reference to in vitro toxicity tests that were evaluated for regulatory purposes either by the US validation centre ICCVAM or ECVAM in the fields of sensitisation, phototoxicity and embryotoxicity

Animal Testing Alternatives↗

Development of content-valid technical skill assessment instruments for athletic taping skills.

BACKGROUND AND PURPOSE: The content validity of technical skill assessment instruments (TSAI) for the skills of athletic taping has not been reported. The purpose of this paper is to outline and present the process of content validation for nine TSAIs for athletic taping. Local and national validators were selected from Canadian Athletic Therapists' Association (CATA)-accredited athletic therapy (AT) programs to serve as content validators. METHODS: The process of content validation began with the creation of a detailed task analysis via mail and simple validation by local validators. Subsequently, the detailed task analysis was committee validated by a group of 10 validators from across Canada. Validators judged the importance and difficulty of each item, and a face-to-face committee-validator meeting established consensus on the majority of checklist items. Through a modified Ebel procedure, frequency distribution was used in the formation of the final TSAIs. RESULTS: Initial consensus for pre-taping assessment and technical skill performance items was low. Upon committee discussion and lack of agreement, the decision to remove pretaping assessment items was made. Initial results of importance and difficulty for athletic taping technical skills were low prior to the committee meeting. Results of importance and difficulty improved substantially following the face-to-face committee-validators meeting. Consensus on fail points improved from initial to final committee validation. CONCLUSION: The process of simple and committee validation can be seen as effective methods to establish the content validity of instruments used for the evaluation of athletic taping.

Alberta↗

The reliability, validity, and preliminary responsiveness of the Eye Allergy Patient Impact Questionnaire (EAPIQ).

BACKGROUND: The Eye Allergy Patient Impact Questionnaire (EAPIQ) was developed based on a pilot study conducted in the US and focus groups with eye allergy sufferers in Europe. The purpose of this study was to present the results of the psychometric validation of the EAPIQ. METHODS: One hundred forty six patients from two allergy clinics completed the EAPIQ twice over a two-week period during the fall and winter allergy seasons, along with concurrent measures of health status, work productivity, and utility. Construct validity, reliability (internal consistency and test-retest), concurrent, known-group, and clinical validities, and responsiveness of the EAPIQ were assessed. Known-group validity was assessed by comparing EAPIQ scale scores between patients grouped according to their self-rating of ocular allergy severity (no symptoms, very mild, mild, moderate, severe, very severe). Clinical validity was assessed by assessing differences in EAPIQ scores between groups of patients rated by their clinician as non-symptomatic, mild, moderate, and severe. RESULTS AND DISCUSSION: Results from the validation study suggested the deletion of 14 of 43 items (including embedded questions) that required patients to complete the percentage of time they were troubled by something (daily activity limitations/emotional troubles). These items yielded a significant amount of missing or inconsistent data (50%). The resulting factor analysis suggested four domains: symptoms, daily life impact, psychosocial impact, and treatment satisfaction. When included as separate scales, the symptom-bother and symptom-frequency scales were highly correlated (> 0.9). As a consequence, and due to superior discriminative validity, the symptom bother and frequency items were summed. All items met the tests for item convergent validity (item-scale correlation = 0.4). The success rate for item discriminant validity testing was 97% (item-scale correlation greater with own scale than with any other). The criterion for internal consistency reliability (alpha coefficient > or = 0.70) was met for all EAPIQ scales (range 0.89-0.93), as was the criterion for test-retest reliability (intraclass correlation [ICC] > or = 0.70). Largely moderate correlations between the scales of the EAPIQ and the mini Rhinoconjunctivitis Quality of Life Questionnaire (miniRQLQ) and low correlations with the Health Utilities Index 2/3 (HUI2/3) were indicative of satisfactory concurrent validity. The EAPIQ symptoms, Daily Life Impact, and Psychosocial Impact scales were able to distinguish between patients differing in eye allergy symptom severity, as rated by patients and clinicians, providing evidence of satisfactory known-group and clinical validities, respectively. Preliminary analyses indicated the EAPIQ Symptoms, Daily Life Impact, and Psychosocial Impact scales to be responsive to changes in eye allergies. CONCLUSION: Following item reduction, construct validity, reliability, concurrent validity, known-group validity, and preliminary responsiveness were satisfactory for the EAPIQ in this population of ocular allergy patients.

Adult↗

Measurement validity in physical therapy research.

This article considers the role of measurement validity within physical therapy research. The concept of measurement validity is identified as a component of internal validity, and it is differentiated from the notion of reliability; these concepts are related to systematic and random sources of error, respectively. Using examples from physical therapy and rehabilitation, four main types of validity are reviewed: face validity, criterion-related validity, content validity, and construct validity. The differing implications of these types of validity for quantitative and qualitative research are discussed. Three principal areas of concern are then addressed, based on a critical discussion of selected examples from the literature. First, it is argued that validity is often poorly distinguished from the allied concept of reliability and that purported claims for validity often only demonstrate reliability. Second, it is claimed that validity is too often neglected in favor of reliability, and specific examples relating to gait analysis are put forward to support this argument. Third, some of the methodological difficulties that may occur when attempts are made to demonstrate validity are considered. The article concludes with a plea for a closer focus on the issue of measurement validity within physical therapy research.

Bias↗

The potential for failure in gynecologic regulatory proficiency testing with current slide validation criteria: results from the College of American Pathologists Interlaboratory Comparison in Gynecologic Cytology Program.

CONTEXT: Current regulatory proficiency testing scoring results in an automatic failure for identifying high-grade squamous intraepithelial lesion (HSIL) as negative. OBJECTIVE: The College of American Pathologists Interlaboratory Comparison Program in Cervicovaginal Cytology data from January 2004 to April 2005 were analyzed to estimate the percentage of failure based on negative responses for HSIL and validation criteria. DESIGN: More than 15,000 participants received field-validated and educational slide sets for conventional, ThinPrep, and SurePath modules. Educational sets fulfilled the validation criteria of the Center for Medicare and Medicaid Services, which required the consensus diagnosis of biopsy-proven HSIL (not field-validated) after review by 3 pathologists. The College of American Pathologists field validation required at least 20 responses to the HSIL+ series, with 70% matched to HSIL+ (standard error < or = 0.05). Minimum regulatory proficiency testing failure estimates were based on incorrect negative responses for the reference category of HSIL. RESULTS: For both cytotechnologists and pathologists, there was a statistically significant higher failure rate for slides that were not field-validated versus those that were field-validated. In conventional modules, 5.3% of the slides that were not field-validated were called negative, versus 1.2% of the field-validated slides. In all liquid-based preparations, 4.0% of the non-field-validated versus 2.2% field-validated slides were called negative. Pathologists would have failed more often than cytotechnologists for the slides that were not field-validated, whereas there was no statistical difference in failure performance with field-validated slides. CONCLUSIONS: Failures were significantly greater with the slides that were not field-validated for both conventional and liquid-based preparations (ThinPrep only) and have implications for both regulatory proficiency testing and expert legal review. Poor performance of pathologists relative to that of cytotechnologists may reflect a lack of prescreening of slides or scope of practice issues.

Clinical Competence↗

The validity of the reinstatement model of craving and relapse to drug use.

RATIONALE: The reinstatement procedure has been used increasingly as a laboratory model of craving and relapse to drug abuse. With the number of reports involving this procedure growing, its validity as a model of relapse merits discussion. OBJECTIVES: The present commentary addresses the validity of the reinstatement procedure in relation to the following three types of models: 1) formal equivalence models, which are assessed on the basis of how well they resemble some phenomenon outside the laboratory (i.e. face validity); 2) correlational models, which are assessed on the basis of how well they predict outcomes of various interventions (such as drug administration or environmental change) when effected outside the laboratory (i.e. predictive validity); and 3) functional equivalence models, which are assessed on the basis of whether the laboratory phenomenon is mechanistically identical or reasonably similar to the phenomenon outside the laboratory (i.e. content validity). METHODS: In order to evaluate the reinstatement model, we briefly examined its various forms and uses, and compared preclinical outcomes to what is known about relapse from the clinical literature. RESULTS. In its most general form, the reinstatement model has reasonable face validity; that is, there is a general agreement in appearance or form of the behavior in the model and the clinical target, relapse. This face validity is generally absent for the procedure when it is used as a model of craving. The predictive validity of the model has not been established. Evidence from studies of treatments for drug relapse have not supported the validity of the model, however from studies of the effects of the presentation of various types of stimuli (e.g. drug "priming") there is mixed evidence supporting predictive validity. With regard to functional equivalence, there is reasonable evidence supporting functional commonalities between drug self-administration in laboratory animals and human drug abusers, which lends support to the validity of the reinstatement model. However, there are several specific areas of departure between the methods and results using the model and clinical practices and observations about relapse, suggesting a lack of functional equivalence. CONCLUSIONS: There is reasonable evidence to support the face validity of the model, but at this time, neither its predictive validity nor functional equivalence has been fully established, which underscores the need for caution in generalizing results from the model to the clinical condition.

Animals↗

Externally validated risk prediction models for gestational diabetes mellitus: A systematic review and meta-analysis.

INTRODUCTION: Risk prediction models for gestational diabetes mellitus (GDM) offer potential for early identification and targeted prevention. External validation is crucial to assess model performance across diverse populations. Despite the availability of numerous GDM prediction models, limited evidence exists on their external validation frequency, methodological quality, and clinical applicability. This systematic review evaluated externally validated GDM prediction models, focusing on methodological rigor, reporting standards, and clinical relevance to inform future research and implementation. MATERIAL AND METHODS: Databases including Ovid MEDLINE, Embase, Scopus, Emcare, and CINAHL were searched up to May 1, 2025. Studies reporting external validation of GDM risk prediction models were included. Two reviewers independently screened studies. Data were extracted using the CHARMS framework, and risk of bias and applicability were assessed using PROBAST+AI. The study protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; CRD420251125758). RESULTS: Twenty-six studies validated 33 models, with validation sample sizes ranging from 50 to 75&#x2009;161. Over half used the IADPSG criteria to define GDM. Discrimination metrics were commonly reported, but calibration, overall performance, and clinical utility were often lacking. Meta-analysis was feasible for only four models: Teede et&#xa0;al., Nanda et&#xa0;al., Naylor et&#xa0;al., and Van Leeuwen et&#xa0;al., each showing fair discrimination. The Teede et&#xa0;al. model was the most widely validated, with 11 external validations across six continents and a pooled AUC of 0.72 (95% CI: 0.67-0.76). Despite fewer validations, the Nanda et&#xa0;al. model achieved the highest pooled discrimination (5 validations; pooled AUC 0.77, 95% CI: 0.74-0.80). The Naylor et&#xa0;al. and van Leeuwen et&#xa0;al. models also underwent meta-analysis, as sufficient external validation studies were available to support comparative performance assessment. Notably, 69.23% of studies had a high risk of bias. CONCLUSIONS: While many models showed acceptable predictive performance, most validations were methodologically weak. Future studies should follow best-practice guidelines and promote scalable validation strategies, such as algorithm sharing, to enhance clinical utility.

Humans↗

A modular approach to the ECVAM principles on test validity.

The European Centre for the Validation of Alternative Methods (ECVAM) proposes to make the validation process more flexible, while maintaining its high standards. The various aspects of validation are broken down into independent modules, and the information necessary to complete each module is defined. The data required to assess test validity in an independent peer review, not the process, are thus emphasised. Once the information to satisfy all the modules is complete, the test can enter the peer-review process. In this way, the between-laboratory variability and predictive capacity of a test can be assessed independently. Thinking in terms of validity principles will broaden the applicability of the validation process to a variety of tests and procedures, including the generation of new tests, new technologies (for example, genomics, proteomics), computer-based models (for example, quantitative structure-activity relationship models), and expert systems. This proposal also aims to take into account existing information, defining this as retrospective validation, in contrast to a prospective validation study, which has been the predominant approach to date. This will permit the assessment of test validity by completing the missing information via the relevant validation procedure: prospective validation, retrospective validation, catch-up validation, or a combination of these procedures.

Animal Testing Alternatives↗

Internal and external validation of predictive models: a simulation study of bias and precision in small samples.

We performed a simulation study to investigate the accuracy of bootstrap estimates of optimism (internal validation) and the precision of performance estimates in independent validation samples (external validation). We combined two data sets containing children presenting with fever without source (n=376+179=555; 120 bacterial infections). Random samples were drawn from this combined data set for the development (n=376) and validation (n=179) of logistic regression models. The models included statistically significant predictors for infection selected from a set of 57 candidate predictors. Model development, including the selection of predictors, and validation were repeated in a bootstrapping procedure. The resulting expected optimism estimate in the receiver operating characteristic (ROC) area was compared with the observed optimism according to independent validation samples. The average apparent ROC area was 0.74, which was expected (based on bootstrapping) to decrease by 0.07 to 0.67, whereas the observed decrease in the validation samples was 0.09 to 0.65. Omitting the selection of predictors from the bootstrap procedure led to a severe underestimation of the optimism (decrease 0.006). The standard error of the observed ROC area in the independent validation samples was large (0.05). We recommend bootstrapping for internal validation because it gives reasonably valid estimates of the expected optimism in predictive performance provided that any selection of predictors is taken into account. For external validation, substantial sample sizes should be used for sufficient power to detect clinically important changes in performance as compared with the internally validated estimate.

Bias↗

Preliminary validation of clinical remission criteria using the OMERACT filter for select categories of juvenile idiopathic arthritis.

OBJECTIVE: To begin the validation process of the preliminary criteria for inactive disease (ID), clinical remission on medication (CRM), and clinical remission off medication (CR) in children with select forms of juvenile idiopathic arthritis (JIA). METHODS: We used the OMERACT filter paradigm to estimate the validity of the criteria within each of the filter's 3 components: truth, discrimination, and feasibility, in 5 categories of JIA: systemic arthritis, persistent and extended oligoarthritis, and rheumatoid factor-positive and negative polyarthritis. Data sources for determining validity estimates included a Delphi questionnaire survey sent to 246 pediatric rheumatologists in 34 countries, a consensus conference attended by 20 senior pediatric rheumatologists representing 9 countries, a retrospective chart review of 437 patients with JIA from 3 tertiary care clinics who had been followed between 4 and 22 years, and the literature. RESULTS: Truth component: face and content validity. These aspects of validity were largely established via the Delphi questionnaire exercise and the consensus conference. Using an 80% consensus level, participants felt that a set of non-redundant variables could effectively differentiate the clinical states of ID, CRM, and CR. Criterion validity could not be irrefutably established because no gold standard for inactive disease exists for JIA. As an alternative, published investigations of remission in JIA were used to estimate concurrent and convergent validity, as surrogates for criterion validity and as indicators of overall construct validity. Correlational analyses revealed the new criteria to have good construct validity. Discrimination component: the criteria demonstrated moderate to high levels of classification, prognosis, and responsiveness (sensitivity to change) using data from the chart review. Patients who were able to attain CR remained disease-free for substantially longer periods than did those who attained only ID or CRM. Responsiveness was evidenced by the ability of the criteria to allow movement of most patients between the disease states, consistent with what is known of the course of the disease. Feasibility component: Results of the Delphi and consensus conference produced a set of criteria that are easily, quickly, and inexpensively completed in the physician's office, and present minimal or no risk to the patient. CONCLUSION: The preliminary criteria demonstrated moderate to excellent validity characteristics in some, but not all components of the OMERACT filter. Prospective validation studies are under way.

Arthritis, Juvenile↗

Rational experimental design for bioanalytical methods validation. Illustration using an assay method for total captopril in plasma.

Generally, bioanalytical chromographic methods are validated according to a predefined programme and distinguish a pre-validation phase, a main validation phase and a follow-up validation phase. In this paper, a rational, total performance evaluation programme for chromatographic methods is presented. The design was developed in particular for the pre-validation and main validation phases. The entire experimental design can be performed within six analytical runs. The first run (pre-validation phase) is used to assess the validity of the expected concentration-response relationship (lack of fit, goodness of fit), to assess specificity of the method and to assess the stability of processed samples in the autosampler for 30 h (benchtop stability). The latter experiment is performed to justify overnight analyses. Following approval of the method after the pre-validation phase, the next five runs (main validation phase) are performed to evaluate method precision and accuracy, recovery, freezing and thawing stability and over-curve control/dilution. The design is nested, i.e., many experimental results are used for the evaluation of several performance characteristics. Analysis of variance (ANOVA) is used for the evaluation of lack of fit and goodness of fit, precision and accuracy, freezing and thawing stability and over-curve control/dilution. Regression analysis is used to evaluate benchtop stability. For over-curve control/dilution, additional to ANOVA, also a paired comparison is applied. As a consequence, the recommended design combines the performance of as few independent validation experiments as possible with modern statistical methods, resulting in optimum use of information. A demonstration of the entire validation programme is given for an HPLC method for the determination of total captopril in human plasma.

Calibration↗

Cue validity modulates the neural correlates of covert endogenous orienting of attention in parietal and frontal cortex.

Parietal brain regions have been implicated in reorienting of visuospatial attention in location-cueing paradigms when misleading advance information is provided in form of a spatially invalid cue. The difference in reaction times to invalidly and validly cued targets is termed the "validity effect" and used as a behavioral measure for attentional reorienting. Behavioral studies suggest that the magnitude of the validity effect depends on the ratio of validly to invalidly cued targets (termed cue validity), i.e., on the amount of top-down information provided. Using fMRI, we investigated the effects of a cue validity manipulation upon the neural mechanisms underlying attentional reorienting using valid and invalid spatial cues in the context of 90% and 60% cue validity, respectively. We hypothesized that increased parietal activation would be elicited when subjects need to reorient their attention in a context of high cue validity. Behaviorally, subjects showed significantly higher validity effects in the high as compared to the low cue validity condition, indicating slower reorienting. The neuroimaging data revealed higher activation of right inferior parietal and right frontal cortex in the 90% than in the 60% cue validity condition. We conclude that the amount of top-down information provided by predictive cues influences the neural correlates of reorienting of visuospatial attention by modulating activation of a right fronto-parietal attentional network.

Adult↗

Systematic validation of disease models for pharmacoeconomic evaluations. Swiss HIV Cohort Study.

Pharmacoeconomic evaluations are often based on computer models which simulate the course of disease with and without medical interventions. The purpose of this study is to propose and illustrate a rigorous approach for validating such disease models. For illustrative purposes, we applied this approach to a computer-based model we developed to mimic the history of HIV-infected subjects at the greatest risk for Mycobacterium avium complex (MAC) infection in Switzerland. The drugs included as a prophylactic intervention against MAC infection were azithromycin and clarithromycin. We used a homogenous Markov chain to describe the progression of an HIV-infected patient through six MAC-free states, one MAC state, and death. Probability estimates were extracted from the Swiss HIV Cohort Study database (1993-95) and randomized controlled trials. The model was validated testing for (1) technical validity (2) predictive validity (3) face validity and (4) modelling process validity. Sensitivity analysis and independent model implementation in DATA (PPS) and self-written Fortran 90 code (BAC) assured technical validity. Agreement between modelled and observed MAC incidence confirmed predictive validity. Modelled MAC prophylaxis at different starting conditions affirmed face validity. Published articles by other authors supported modelling process validity. The proposed validation procedure is a useful approach to improve the validity of the model.

AIDS-Related Opportunistic Infections↗

Injury outcome indicators--validation matters.

INTRODUCTION: There is concern that many national non-fatal injury indicators currently in use are misleading. OBJECTIVE: To make the case for the validation of existing unvalidated indicators, as well as the validation of new indicators before they are promulgated. METHOD: The International Collaborative Effort on Injury Statistics (ICE) Criteria were used for investigating the validity of indicators. Examples of indicators that have been found to be valid using these criteria are presented. In contrast, examples of national road safety indicators are also presented, whose validity is questionable. Trends in road safety indicators with and without threats to validity are contrasted. RESULTS: The New Zealand Injury Prevention Strategy (NZIPS) serious injury indicators are presented as indicators with no identifiable threats to validity. National road safety indicators from Canada, New Zealand and the United Kingdom, with identifiable threats to validity, are also presented. When trends for the valid NZIPS motor vehicle traffic crash indicators are compared with the New Zealand national road safety indicators, which have identifiable threats to validity, they show contrasting trends. This raises concerns that the current national indicators are potentially misleading. CONCLUSION: Validation does matter. For any indicator, it is important that it is clearly defined and specified. The specification should make it clear what parameter the indicator aims to reflect. Before use, the indicator should be validated against this target parameter. That parameter, and the indicators aimed to estimate it, should focus attention on important injuries, ie. injuries that are associated with significant mortality, threat-to-life, threat-of-disablement, loss of quality of life, or increased cost.

Accidents, Traffic↗