Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Meeting report: Validation of toxicogenomics-based test systems: ECVAM-ICCVAM/NICEATM considerations for regulatory use.

This is the report of the first workshop "Validation of Toxicogenomics-Based Test Systems" held 11-12 December 2003 in Ispra, Italy. The workshop was hosted by the European Centre for the Validation of Alternative Methods (ECVAM) and organized jointly by ECVAM, the U.S. Interagency Coordinating Committee on the Validation of Alternative Methods (ICCVAM), and the National Toxicology Program (NTP) Interagency Center for the Evaluation of Alternative Toxicological Methods (NICEATM). The primary aim of the workshop was for participants to discuss and define principles applicable to the validation of toxicogenomics platforms as well as validation of specific toxicologic test methods that incorporate toxicogenomics technologies. The workshop was viewed as an opportunity for initiating a dialogue between technologic experts, regulators, and the principal validation bodies and for identifying those factors to which the validation process would be applicable. It was felt that to do so now, as the technology is evolving and associated challenges are identified, would be a basis for the future validation of the technology when it reaches the appropriate stage. Because of the complexity of the issue, different aspects of the validation of toxicogenomics-based test methods were covered. The three focus areas include a) biologic validation of toxicogenomics-based test methods for regulatory decision making, b) technical and bioinformatics aspects related to validation, and c) validation issues as they relate to regulatory acceptance and use of toxicogenomics-based test methods. In this report we summarize the discussions and describe in detail the recommendations for future direction and priorities.

Animal Testing Alternatives↗

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans↗

Construct and face validity and task workload for laparoscopic camera navigation: virtual reality versus videotrainer systems at the SAGES Learning Center.

BACKGROUND: Laparoscopic camera navigation (LCN) training on simulators has demonstrated transferability to actual operations, but no comparative data exist. The objective of this study was to compare the construct and face validity, as well as workload, of two previously validated virtual reality (VR) and videotrainer (VT) systems. METHODS: Attendees (n = 90) of the SAGES 2005 Learning Center performed two repetitions on both VR (EndoTower) and VT (Tulane Trainer) LCN systems using 30 degrees laparoscopes and completed a questionnaire regarding demographics, simulator characteristics, and task workload. Construct validity was determined by comparing the performance scores of subjects with various levels of experience according to five parameters and face validity according to eight. The validated NASA-TLX questionnaire that rates the mental, physical, and temporal demand of a task as well as the performance, effort, and frustration of the subject was used for workload measurement. RESULTS: Construct validity was demonstrated for both simulators according to the number of basic laparoscopic cases (p = 0.005), number of advanced cases (p < 0.001), and frequency of angled scope use (p < 0.001), and only for VT according to training level (p < 0.001) and fellowship training (p = 0.008). Face validity ratings on a 1-20 scale averaged 15.4 +/- 3 for VR vs. 16 +/- 2.6 for VT (p = 0.04). Ninety-six percent of participants rated both simulators as valid educational tools. The NASA-TLX overall workload score was 69.5 +/- 24 for VR vs. 68.8 +/- 20.5 for VT (p = 0.31). CONCLUSIONS: This is the largest study to date that compares two validated LCN simulators. While subtle differences exist, both VR and VT simulators demonstrated excellent construct validity, good face validity, and acceptable workload parameters. These systems thus represent useful training devices and should be widely used to improve surgical performance.

Adult↗

Validity and the OSCE.

In preparation for a celebration of '30 years of OSCEs' held during the 2002 meeting of the Association for Medical Education in Europe (AMEE), I was asked to discuss the question, 'Are OSCEs valid to assess competence?". My first instinct was to review work undertaken in ay countries by famous researchers such as Harden, Colliver, Rothman, van der Vleuten, Stillman, Tamblyn and others who have studied and written about the validity of OSCEs. I could then have reviewed the extensive literature produced in Canada by the Medical Council of Canada and in the United States by the Education Commission on Foreign Medical Graduates and National Board of Medical Examiners that has demonstrated the utility of large-scale OSCEs for certification and licensure. I might have tossed in a few papers from my own research on the validity of OSCEs in psychiatry. Indeed, it would have been relatively easy to marshal the medical education literature to answer the question 'Are OSCEs valid to assess competence' strongly in the affirmative. But the more I reflected on the question, the more I confronted concerns that have troubled me for some time. Specifically, I worry that our approaches to validity may themselves not be valid. In this paper, I review what I believe to be three serious problems with our current approaches to showing that 'the OSCE is valid' Let me begin by rethinking the question. What do we mean 'Is the OSCE valid for assessing competence' There are three important problems with this question. First, validity is a property of the application of a test, not of a test itself. Second, we cannot speak of validity without giving consideration to the context in which we use the test. And finally, the concept of validity flounders because the OSCE itself is an important agent in constructing the variables of performance that it is designed to measure. I shall consider each issue in turn.

Canada↗

Content and criterion validity evaluation of National Public Health Performance Standards measurement instruments.

OBJECTIVE: The Centers for Disease Control and Prevention's National Public Health Performance Standards Program (NPHPSP) has developed instruments to measure the performance of local and state public health departments on the 10 "Essential Services of Public Health," which have been tested in several states. This article is a report of the evaluation of the content and criterion validity of the local public health performance assessment instrument, and the content validity of the state public health performance assessment instrument. METHODS: Health department performance is measured using a set of indicators developed for the 10 Essential Services of Public Health and a model standard for each indicator. Content validity of each model standard in the local instrument was addressed by community partners along the following dimensions: the importance of each standard as a measure of the associated Essential Service, its completeness as a measure, and its reasonableness for achievement. All standards for each Essential Service were then judged in terms of their completeness in measuring performance in that service. Content validity of the state instrument was evaluated in a group interview of health department staff members from three states. Criterion validity of the local instrument was assessed for a sample of eight public health departments in Florida and six in New York by examining documentary evidence for selected responses. Criterion validity was also evaluated for a sample of Florida local public health departments and one Hawaii public health department by comparing state health department staffs' judgments of performance against the instrument score. RESULTS: Criterion validity was upheld for a summary performance score on the local instrument, but was not upheld for performance judgments on individual Essential Services. The NPHPSP standards based on the Essential Services have validity for measuring local public health system performance, according to community partners. The model standards are valid measures of state performance, according to state public health departments in three states. CONCLUSIONS: Within the scope of the validity evaluations completed, the NPHPSP state and local performance assessment instruments were found to be valid measures of public health performance.

Attitude of Health Personnel↗

Validation of numerical ground water models used to guide decision making.

Many sites of ground water contamination rely heavily on complex numerical models of flow and transport to develop closure plans. This complexity has created a need for tools and approaches that can build confidence in model predictions and provide evidence that these predictions are sufficient for decision making. Confidence building is a long-term, iterative process and the author believes that this process should be termed model validation. Model validation is a process, not an end result. That is, the process of model validation cannot ensure acceptable prediction or quality of the model. Rather, it provides an important safeguard against faulty models or inadequately developed and tested models. If model results become the basis for decision making, then the validation process provides evidence that the model is valid for making decisions (not necessarily a true representation of reality). Validation, verification, and confirmation are concepts associated with ground water numerical models that not only do not represent established and generally accepted practices, but there is not even widespread agreement on the meaning of the terms as applied to models. This paper presents a review of model validation studies that pertain to ground water flow and transport modeling. Definitions, literature debates, previously proposed validation strategies, and conferences and symposia that focused on subsurface model validation are reviewed and discussed. The review is general and focuses on site-specific, predictive ground water models used for making decisions regarding remediation activities and site closure. The aim is to provide a reasonable starting point for hydrogeologists facing model validation for ground water systems, thus saving a significant amount of time, effort, and cost. This review is also aimed at reviving the issue of model validation in the hydrogeologic community and stimulating the thinking of researchers and practitioners to develop practical and efficient tools for evaluating and refining ground water predictive models.

Decision Making↗

Variance and confidence limits in validation studies based on comparison between three different types of measurements.

BACKGROUND: The methods used in epidemiological studies to assess exposure are often affected by a conspicuous amount of measurement error. Exposure-measurement error is recognised to cause attenuation in the association between exposure and disease. Among different possible approaches, the validity coefficient of a measurement can be estimated by a comparison of three types of measurements, using either structural equation models or factor analysis (the triads method). These approaches assume that the measurements are linearly related to true intake and have independent random errors. METHODS: In this paper we present an estimator of the variance of the estimated validity coefficient to compute the associated confidence intervals. Standard error for the validity coefficient allows the efficiency of validation studies to be evaluated. Our work was motivated by the fact that existing software does not provide correct standard errors for the estimated validity coefficient. The approach is illustrated using selected examples from dietary validation studies. RESULTS: The accuracy of our formula is evaluated by comparison with the results of a simulation study, which shows that our variance estimator provides good results for sample sizes of at least n = 100 and when the expected value of the validity coefficient is not too close to 1.0, independent of the sample size. Our estimator formula performs better than either a naïve approach, that computes the standard error for a validity coefficient as if it is a straightforward correlation coefficient, or the SAS-CALIS procedure, which uses a maximum likelihood method. CONCLUSIONS: In evaluating the validity of the type of measurement chosen to assess exposure in an epidemiological study, it is important to provide an estimate of the precision of the validity coefficient of the measurement. Our variance estimator may help calculate sample size requirements for validation studies.

Analysis of Variance↗

Influence of the validation method on diagnostic accuracy for caries. A comparison of six digital and two conventional radiographic systems.

OBJECTIVE: To evaluate the influence of the validation method on the diagnostic accuracy and the relative comparison of eight radiographic systems for caries detection. METHODS: Three hundred and thirty-eight approximal and 145 occlusal surfaces were radiographed under standardised conditions using six CCD-based sensor systems: MPDx (Dental/Medical Diagnostic Systems Inc., Woodland Hills, CA, USA), Dixi (Planmeca, Helsinki, Finland), Sidexis (Sirona, Bensheim, Germany), RVG(old) (Trophy, Paris, France, 1994 model), RVG(new) (Trophy, Paris, France, 2000 model) and Visualix (Gendex, Milan, Italy) and two film systems: Ektaspeed Plus and Insight (Eastman Kodak, Rochester, NY, USA). Four observers examined the radiographs for approximal and occlusal caries using a five-point confidence scale. The presence of caries was validated histologically and radiographically. Diagnostic accuracy was evaluated using ROC curve areas (A(z)). RESULTS: For both approximal and occlusal caries the mean A(z) of the eight radiographic systems was significantly higher using radiographic than histological validation (P<0.001). Using histological validation for approximal caries, Dixi (A(z)=0.71) and Ektaspeed Plus (A(z)=0.7) were not significantly different, but Dixi was significantly more accurate than the other digital systems and the Insight film. Using radiographic validation for approximal caries, Ektaspeed Plus (A(z)=0.87) was significantly more accurate than Dixi (A(z)=0.82). Dixi was significantly more accurate than MPDx (A(z)=0.74), RVG(old) (A(z)=0.77), RVG(new) (A(z)=0.77) and Visualix (A(z)=0.76). Corresponding variations were found for occlusal caries depending on the validation method. Using histological validation, MPDx (A(z)=0.76) was significantly less accurate than Dixi (A(z)=0.81), Sidexis (A(z)=0.8), Ektaspeed Plus (A(z)=0.82) and Insight (A(z)=0.81). Using radiographic validation, MPDx (A(z)=0.83) was also significantly less accurate than RVG(old) (A(z)=0.89) and RVG(new) (A(z)=0.9). CONCLUSION: A(z) obtained from radiographic validation was significantly higher than A(z) obtained from histological validation. Comparison of the diagnostic efficacy for caries of the eight radiographic systems was strongly influenced by the validation method. DOI: 10.1038/sj/dmfr/4600645

Analysis of Variance↗

Validity: on meaningful interpretation of assessment data.

CONTEXT: All assessments in medical education require evidence of validity to be interpreted meaningfully. In contemporary usage, all validity is construct validity, which requires multiple sources of evidence; construct validity is the whole of validity, but has multiple facets. Five sources--content, response process, internal structure, relationship to other variables and consequences--are noted by the Standards for Educational and Psychological Testing as fruitful areas to seek validity evidence. PURPOSE: The purpose of this article is to discuss construct validity in the context of medical education and to summarize, through example, some typical sources of validity evidence for a written and a performance examination. SUMMARY: Assessments are not valid or invalid; rather, the scores or outcomes of assessments have more or less evidence to support (or refute) a specific interpretation (such as passing or failing a course). Validity is approached as hypothesis and uses theory, logic and the scientific method to collect and assemble data to support or fail to support the proposed score interpretations, at a given point in time. Data and logic are assembled into arguments--pro and con--for some specific interpretation of assessment data. Examples of types of validity evidence, data and information from each source are discussed in the context of a high-stakes written and performance examination in medical education. CONCLUSION: All assessments require evidence of the reasonableness of the proposed interpretation, as test data in education have little or no intrinsic meaning. The constructs purported to be measured by our assessments are important to students, faculty, administrators, patients and society and require solid scientific evidence of their meaning.

Data Interpretation, Statistical↗

A methodology for global validation of microarray experiments.

BACKGROUND: DNA microarrays are popular tools for measuring gene expression of biological samples. This ever increasing popularity is ensuring that a large number of microarray studies are conducted, many of which with data publicly available for mining by other investigators. Under most circumstances, validation of differential expression of genes is performed on a gene to gene basis. Thus, it is not possible to generalize validation results to the remaining majority of non-validated genes or to evaluate the overall quality of these studies. RESULTS: We present an approach for the global validation of DNA microarray experiments that will allow researchers to evaluate the general quality of their experiment and to extrapolate validation results of a subset of genes to the remaining non-validated genes. We illustrate why the popular strategy of selecting only the most differentially expressed genes for validation generally fails as a global validation strategy and propose random-stratified sampling as a better gene selection method. We also illustrate shortcomings of often-used validation indices such as overlap of significant effects and the correlation coefficient and recommend the concordance correlation coefficient (CCC) as an alternative. CONCLUSION: We provide recommendations that will enhance validity checks of microarray experiments while minimizing the need to run a large number of labour-intensive individual validation assays.

3T3-L1 Cells↗

Content validity of self-report measurement instruments: an illustration from the development of the Brain Tumor Module of the M.D. Anderson Symptom Inventory.

PURPOSE/OBJECTIVES: To illustrate one technique for establishing content validity of measurements using the initial development and testing of the M.D. Anderson Symptom Inventory Brain Tumor Module. DATA SOURCES: Published articles, book chapters, and subjective judgments of experts. DATA SYNTHESIS: Content validity is the essential first step in the development of items to be included in a measurement instrument. Content validity is a criterion-referenced process that is judged by how well each item in a newly developed instrument reflects its respective objective or content domain. The stages in addressing content validity include a developmental stage and a judgment-quantification stage. Steps involved in the developmental stage include domain identification, item generation, and instrument formation. The judgment-quantification stage is when experts review the items and either report validity of the items subjectively or with an empirically referenced method, such as calculation of the content validity index. The content validity of a set of questions designed to measure symptoms in a population of patients with primary brain tumors was ascertained by using the calculation of the content validity index. CONCLUSIONS: The final version of the M.D. Anderson Symptom Inventory Brain Tumor Module consists of the 13 core items and 18 additional items designated as valid by a panel of experts. The instrument will be administered to a group of patients to determine construct validity and reliability of the items. IMPLICATIONS FOR NURSING: Self-report instruments are used to measure various health outcomes in oncology. Oncology nurses are in a key position to develop such instruments to be used in clinical care and research of symptoms associated with cancer. Understanding the process of content validation is an essential first step in developing new instruments.

Brain Neoplasms↗

[Reliability and validity of the Japanese version of the coping inventory for stressful situations (CISS): a contribution to the cross-cultural studies of coping].

OBJECTIVE: There has recently been a dramatic increase in the number of studies on coping behavior as an intervening variable between stress and health. Most of the available measures of coping are, however, psychometrically inadequate. We therefore decided to develop the Japanese version of the Coping Inventory for Stressful Situations (CISS) with special regards to its cross-cultural equivalence, reliability and validity. The CISS is a self-report measure of an individual's typical pattern of coping along three orthogonal dimensions of Task-, Emotion-, and Avoidance-oriented coping; its reliability and validity have been well studied in North America, where it was originally developed. METHOD: We obtained the Japanese version of the CISS (J-CISS) by means of back-translation. In Study 1, we administered the J-CISS and the 12-item General Health Questionnaire (GHQ) to 33 Japanese university students twice with an interval of four weeks. In Study 2,550 Japanese high school students completed the J-CISS and the Maudsley Personality Inventory. RESULTS: The equivalence of the Japanese version with the original was ascertained by means of back-translation involving multiple, independent mental health professionals and by factor congruence between the two versions. A principal component factor analysis (Varimax rotation) of the Study 2 data allowed us to extract three factors, which were virtually identical to the original ones. The high corrected item-remainder correlations, internal consistency reliabilities and test-retest reliabilities all attested to the reliability of the J-CISS. In order to examine its content validity, we compared the J-CISS with two coping questionnaires that have been in use in Japan, and found that the J-CISS covered most of the coping styles in these two questionnaires. However, such coping styles as "giving up," "to lose is to win (a Japanese proverb)," "it is best to do nothing" were not included in the original CISS and hence in the J-CISS. The criterion validity of the J-CISS was examined both in terms of predictive validity and concurrent validity. In Study 1, those who scored below the cut-off of the GHQ at Time 1 but above the cut-off at Time 2 had significantly higher Emotion-oriented coping scores at Time 1 than those who remained below the cut-off of the GHQ at Times 1 and 2 (predictive validity). In Study 2, the J-CISS scales and the MPI scales showed theoretically predicted correlations (concurrent validity). The results of the factor analysis and the corrected item-remainder correlations were suggestive of high construct validity of the J-CISS. Moreover, the mean inter-item correlation was between .20 and .40 for each scale, indicating its homogeneity. Factor analysis of each scale revealed that each scale indeed contained only one factor. Correlations among the three scales of the J-CISS established that the three scales formed multi-dimensional measures of coping. CONCLUSION: The results of our study indicate 1) that the obtained Japanese version of the CISS is to be regarded as final, 2) that coping styles can be measured in a consistent and reliable manner both in Japan and North America, and 3) that this cross-cultural equivalence as well as the other validity studies have further augmented the validity of the CISS itself.

Adaptation, Psychological↗

Validation of the Pediatric Voice-Related Quality-of-Life survey.

OBJECTIVE: To validate the Pediatric Voice-Related Quality-of-Life (PVRQOL) survey, which was designed to assess voice changes over time in the pediatric population. DESIGN: Prospective longitudinal study. SETTING: Outpatient pediatric otolaryngology office practice. PARTICIPANTS: One hundred twenty parents of children aged 2 through 18 years having a variety of otolaryngological diagnoses including disorders that affect the voice. INTERVENTIONS: The previously validated Pediatric Voice Outcomes Survey and the PVRQOL were jointly administered to the parents of the study participants. Test-retest reliability was accomplished by having 70 caregivers repeat the instrument 2 weeks after the initial visit. The Cronbach alpha value was calculated to determine reliability. Instrument validity was determined by examining convergent and discriminant validity. MAIN OUTCOME MEASURE: Correlation of PVRQOL scores with Pediatric Voice Outcomes Survey scores. RESULTS: Reliability of the PVRQOL was established by evaluating the Cronbach alpha value (.96; P<.001) and by test-retest reliability (weighted kappa value, 0.8). Validity of the PVQROL was tested by evaluating its ability to show significant change in voice-related quality-of-life after adenoidectomy (discriminant validity) (P<.001). The PVQROL also proved valid when the overall score was correlated with the previously validated Pediatric Voice Outcomes Survey (r = 0.7; P<.001). CONCLUSION: The PVRQOL is a more comprehensive survey than the previously validated Pediatric Voice Outcomes Survey and is another valid instrument to examine the health-related quality-of-life issues in pediatric voice disorders.

Adolescent↗

Monitoring sedation status over time in ICU patients: reliability and validity of the Richmond Agitation-Sedation Scale (RASS).

CONTEXT: Goal-directed delivery of sedative and analgesic medications is recommended as standard care in intensive care units (ICUs) because of the impact these medications have on ventilator weaning and ICU length of stay, but few of the available sedation scales have been appropriately tested for reliability and validity. OBJECTIVE: To test the reliability and validity of the Richmond Agitation-Sedation Scale (RASS). DESIGN: Prospective cohort study. SETTING: Adult medical and coronary ICUs of a university-based medical center. PARTICIPANTS: Thirty-eight medical ICU patients enrolled for reliability testing (46% receiving mechanical ventilation) from July 21, 1999, to September 7, 1999, and an independent cohort of 275 patients receiving mechanical ventilation were enrolled for validity testing from February 1, 2000, to May 3, 2001. MAIN OUTCOME MEASURES: Interrater reliability of the RASS, Glasgow Coma Scale (GCS), and Ramsay Scale (RS); validity of the RASS correlated with reference standard ratings, assessments of content of consciousness, GCS scores, doses of sedatives and analgesics, and bispectral electroencephalography. RESULTS: In 290-paired observations by nurses, results of both the RASS and RS demonstrated excellent interrater reliability (weighted kappa, 0.91 and 0.94, respectively), which were both superior to the GCS (weighted kappa, 0.64; P<.001 for both comparisons). Criterion validity was tested in 411-paired observations in the first 96 patients of the validation cohort, in whom the RASS showed significant differences between levels of consciousness (P<.001 for all) and correctly identified fluctuations within patients over time (P<.001). In addition, 5 methods were used to test the construct validity of the RASS, including correlation with an attention screening examination (r = 0.78, P<.001), GCS scores (r = 0.91, P<.001), quantity of different psychoactive medication dosages 8 hours prior to assessment (eg, lorazepam: r = - 0.31, P<.001), successful extubation (P =.07), and bispectral electroencephalography (r = 0.63, P<.001). Face validity was demonstrated via a survey of 26 critical care nurses, which the results showed that 92% agreed or strongly agreed with the RASS scoring scheme, and 81% agreed or strongly agreed that the instrument provided a consensus for goal-directed delivery of medications. CONCLUSIONS: The RASS demonstrated excellent interrater reliability and criterion, construct, and face validity. This is the first sedation scale to be validated for its ability to detect changes in sedation status over consecutive days of ICU care, against constructs of level of consciousness and delirium, and correlated with the administered dose of sedative and analgesic medications.

Aged↗

Criterion and content validity of a novel structured haggling contingent valuation question format versus the bidding game and binary with follow-up format.

Contingent valuation question formats that will be used to elicit willingness to pay for goods and services need to be relevant to the area they will be used in order for responses to be valid. A novel contingent valuation question format called the "structured haggling technique" (SH) that resembles the bargaining system in Nigerian markets was designed and its criterion and content validity compared with those of the bidding game (BG) and binary-with-follow-up (BWFU) technique. This was achieved by determining the willingness to pay (WTP) for insecticide-treated nets (ITNs) in Southeast Nigeria. Content validity was determined through observation of actual trading of untreated nets together with interviews with sellers and consumers. Criterion validity was determined by comparing stated and actual WTP. Stated WTP was determined using a questionnaire administered to 810 household heads and actual WTP was determined by offering the nets for sale to all respondents one month later. The phi (correlation) coefficient was used to compare criterion validity across question formats. The phi coefficients were SH (0.60: 95% C.I. 0.50-0.71), BG (0.42: 95% C.I. 0.29-0.54) and the BWFU (0.32: 95% C.I. 0.20-0.44), implying that the BG and SH had similar levels of criterion-validity while the BWFU was the least criterion-valid. However, the SH was the most content-valid. It is necessary to validate the findings in other areas where haggling is common. Future studies should establish the content validity of question formats in the contexts in which they will be used before administering questionnaires.

Attitude to Health↗

Validation of a food-frequency questionnaire assessment of carotenoid and vitamin E intake using weighed food records and plasma biomarkers: the method of triads model.

BACKGROUND: Reliability or validity studies are important for the evaluation of measurement error in dietary assessment methods. An approach to validation known as the method of triads uses triangulation techniques to calculate the validity coefficient of a food-frequency questionnaire (FFQ). OBJECTIVE: To assess the validity of an FFQ estimates of carotenoid and vitamin E intake against serum biomarker measurements and weighed food records (WFRs), by applying the method of triads. DESIGN: The study population was a sub-sample of adult participants in a randomised controlled trial of beta-carotene and sunscreen in the prevention of skin cancer. Dietary intake was assessed by a self-administered FFQ and a WFR. Nonfasting blood samples were collected and plasma analysed for five carotenoids (alpha-carotene, beta-carotene, beta-cryptoxanthin, lutein, lycopene) and vitamin E. Correlation coefficients were calculated between each of the dietary methods and the validity coefficient was calculated using the method of triads. The 95% confidence intervals for the validity coefficients were estimated using bootstrap sampling. RESULTS: The validity coefficients of the FFQ were highest for alpha-carotene (0.85) and lycopene (0.62), followed by beta-carotene (0.55) and total carotenoids (0.55), while the lowest validity coefficient was for lutein (0.19). The method of triads could not be used for beta-cryptoxanthin and vitamin E, as one of the three underlying correlations was negative. CONCLUSIONS: Results were similar to other studies of validity using biomarkers and the method of triads. For many dietary factors, the upper limit of the validity coefficients was less than 0.5 and therefore only strong relationships between dietary exposure and disease will be detected.

Adult↗

Validation of surgical site infection surveillance in the Netherlands.

OBJECTIVES: To describe how continuous validation of data on surgical site infection (SSI) is being performed in the Dutch National Nosocomial Infection Surveillance System (Preventie Ziekenhuisinfecties door Surveillance [PREZIES]), to assess the quality and accuracy of the PREZIES data, and to present the corresponding outcomes of the assessment. DESIGN: Mandatory, 1-day on-site validation visit to participating hospitals every 3 years. The process of surveillance, including the quality of the method of data collection, is validated by means of a structured interview. The use of SSI criteria is validated by review of medical records, with the judgment of the validation team as the criterion standard. SETTING: Hospitals participating in PREZIES. RESULTS: During 1999-2004, the validation team visited 40 hospitals and reviewed 859 medical charts. There was no deviation between reports of SSI by infection control professionals and findings by the PREZIES validation team at 30 hospitals and 1 deviation in each of 10 hospitals; the positive predictive value was 0.97, and the negative predictive value was 0.99. The validation team often gave advice to the hospital, aimed at perfecting the process of surveillance. On 2 occasions, data were removed from the PREZIES database after the validation visit revealed deviations from the SSI surveillance protocol that could have resulted in nonrepresentative SSI rate data. CONCLUSIONS: PREZIES is confident that the assembled Dutch SSI surveillance data are reliable and robust and are sufficiently accurate to be used as a reference for interhospital comparison. PREZIES will continue performing on-site validation visits, to improve the process of surveillance and ensure the reliability of the surveillance data.

Cross Infection↗

Internal and external validation of the NOSEP prediction score for nosocomial sepsis in neonates.

OBJECTIVE: To evaluate the performance of a scoring system (NOSEP) to predict nosocomial sepsis in neonates at the hospital where the score was developed (internal validation) and in an independent data set from other centers (external validation). DESIGN: Multiple center prospective cohort study. SETTING: Six neonatal intensive care units from the Flanders in Belgium. PATIENTS: We analyzed two groups of patients: 62 episodes of presumed nosocomial sepsis in the internal validation cohort and 93 episodes of presumed nosocomial sepsis in a multiple center external validation cohort. INTERVENTIONS: Assessment of the predictive power of the NOSEP score 24 hrs preceding sepsis workup and the patients' basic demographic characteristics and co-morbidity was performed. Diagnosis of nosocomial sepsis and the microbiology results were registered. MAIN RESULTS: The NOSEP score's discriminative capability was very good in the internal validation (area under receiver operating characteristic curve = 0.73 +/- 0.08 [sem]). The NOSEP score performed satisfactory in the external validation (area under receiver operating characteristic curve = 0.66 +/- 0.06). The calibration capability in both validation sets as measured by goodness-of-fit tests (internal validation, p =.56; external validation, p =.48) was good. An improvement of the NOSEP score was obtained for the external centers by redefining the cut-off of the items of the NOSEP score (area under receiver operating characteristic curve for NOSEP-NEW-I = 0.71 +/- 0.05) or adding co-morbidity factors (area under receiver operating characteristic curve for NOSEP-NEW-II = 0.82 +/- 0.04), with good calibration performance (goodness-of-fit test, p >.50). Finally, the fit of the NOSEP score demonstrated no significant variation across subgroups of patients. CONCLUSIONS: The predictive power of the original NOSEP score is very good in neonates at the original neonatal intensive care unit. In other neonatal intensive care units, its discriminatory performance is satisfactory but could be improved after modification of the variables in the model or adding additional variables. To use such a NOSEP score in other neonatal intensive care units, its accuracy has to be validated and adjusted if necessary.

Cross Infection↗