Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Validity and the OSCE.

In preparation for a celebration of '30 years of OSCEs' held during the 2002 meeting of the Association for Medical Education in Europe (AMEE), I was asked to discuss the question, 'Are OSCEs valid to assess competence?". My first instinct was to review work undertaken in ay countries by famous researchers such as Harden, Colliver, Rothman, van der Vleuten, Stillman, Tamblyn and others who have studied and written about the validity of OSCEs. I could then have reviewed the extensive literature produced in Canada by the Medical Council of Canada and in the United States by the Education Commission on Foreign Medical Graduates and National Board of Medical Examiners that has demonstrated the utility of large-scale OSCEs for certification and licensure. I might have tossed in a few papers from my own research on the validity of OSCEs in psychiatry. Indeed, it would have been relatively easy to marshal the medical education literature to answer the question 'Are OSCEs valid to assess competence' strongly in the affirmative. But the more I reflected on the question, the more I confronted concerns that have troubled me for some time. Specifically, I worry that our approaches to validity may themselves not be valid. In this paper, I review what I believe to be three serious problems with our current approaches to showing that 'the OSCE is valid' Let me begin by rethinking the question. What do we mean 'Is the OSCE valid for assessing competence' There are three important problems with this question. First, validity is a property of the application of a test, not of a test itself. Second, we cannot speak of validity without giving consideration to the context in which we use the test. And finally, the concept of validity flounders because the OSCE itself is an important agent in constructing the variables of performance that it is designed to measure. I shall consider each issue in turn.

Canada↗

Content and criterion validity evaluation of National Public Health Performance Standards measurement instruments.

OBJECTIVE: The Centers for Disease Control and Prevention's National Public Health Performance Standards Program (NPHPSP) has developed instruments to measure the performance of local and state public health departments on the 10 "Essential Services of Public Health," which have been tested in several states. This article is a report of the evaluation of the content and criterion validity of the local public health performance assessment instrument, and the content validity of the state public health performance assessment instrument. METHODS: Health department performance is measured using a set of indicators developed for the 10 Essential Services of Public Health and a model standard for each indicator. Content validity of each model standard in the local instrument was addressed by community partners along the following dimensions: the importance of each standard as a measure of the associated Essential Service, its completeness as a measure, and its reasonableness for achievement. All standards for each Essential Service were then judged in terms of their completeness in measuring performance in that service. Content validity of the state instrument was evaluated in a group interview of health department staff members from three states. Criterion validity of the local instrument was assessed for a sample of eight public health departments in Florida and six in New York by examining documentary evidence for selected responses. Criterion validity was also evaluated for a sample of Florida local public health departments and one Hawaii public health department by comparing state health department staffs' judgments of performance against the instrument score. RESULTS: Criterion validity was upheld for a summary performance score on the local instrument, but was not upheld for performance judgments on individual Essential Services. The NPHPSP standards based on the Essential Services have validity for measuring local public health system performance, according to community partners. The model standards are valid measures of state performance, according to state public health departments in three states. CONCLUSIONS: Within the scope of the validity evaluations completed, the NPHPSP state and local performance assessment instruments were found to be valid measures of public health performance.

Attitude of Health Personnel↗

Variance and confidence limits in validation studies based on comparison between three different types of measurements.

BACKGROUND: The methods used in epidemiological studies to assess exposure are often affected by a conspicuous amount of measurement error. Exposure-measurement error is recognised to cause attenuation in the association between exposure and disease. Among different possible approaches, the validity coefficient of a measurement can be estimated by a comparison of three types of measurements, using either structural equation models or factor analysis (the triads method). These approaches assume that the measurements are linearly related to true intake and have independent random errors. METHODS: In this paper we present an estimator of the variance of the estimated validity coefficient to compute the associated confidence intervals. Standard error for the validity coefficient allows the efficiency of validation studies to be evaluated. Our work was motivated by the fact that existing software does not provide correct standard errors for the estimated validity coefficient. The approach is illustrated using selected examples from dietary validation studies. RESULTS: The accuracy of our formula is evaluated by comparison with the results of a simulation study, which shows that our variance estimator provides good results for sample sizes of at least n = 100 and when the expected value of the validity coefficient is not too close to 1.0, independent of the sample size. Our estimator formula performs better than either a naïve approach, that computes the standard error for a validity coefficient as if it is a straightforward correlation coefficient, or the SAS-CALIS procedure, which uses a maximum likelihood method. CONCLUSIONS: In evaluating the validity of the type of measurement chosen to assess exposure in an epidemiological study, it is important to provide an estimate of the precision of the validity coefficient of the measurement. Our variance estimator may help calculate sample size requirements for validation studies.

Analysis of Variance↗

Influence of the validation method on diagnostic accuracy for caries. A comparison of six digital and two conventional radiographic systems.

OBJECTIVE: To evaluate the influence of the validation method on the diagnostic accuracy and the relative comparison of eight radiographic systems for caries detection. METHODS: Three hundred and thirty-eight approximal and 145 occlusal surfaces were radiographed under standardised conditions using six CCD-based sensor systems: MPDx (Dental/Medical Diagnostic Systems Inc., Woodland Hills, CA, USA), Dixi (Planmeca, Helsinki, Finland), Sidexis (Sirona, Bensheim, Germany), RVG(old) (Trophy, Paris, France, 1994 model), RVG(new) (Trophy, Paris, France, 2000 model) and Visualix (Gendex, Milan, Italy) and two film systems: Ektaspeed Plus and Insight (Eastman Kodak, Rochester, NY, USA). Four observers examined the radiographs for approximal and occlusal caries using a five-point confidence scale. The presence of caries was validated histologically and radiographically. Diagnostic accuracy was evaluated using ROC curve areas (A(z)). RESULTS: For both approximal and occlusal caries the mean A(z) of the eight radiographic systems was significantly higher using radiographic than histological validation (P<0.001). Using histological validation for approximal caries, Dixi (A(z)=0.71) and Ektaspeed Plus (A(z)=0.7) were not significantly different, but Dixi was significantly more accurate than the other digital systems and the Insight film. Using radiographic validation for approximal caries, Ektaspeed Plus (A(z)=0.87) was significantly more accurate than Dixi (A(z)=0.82). Dixi was significantly more accurate than MPDx (A(z)=0.74), RVG(old) (A(z)=0.77), RVG(new) (A(z)=0.77) and Visualix (A(z)=0.76). Corresponding variations were found for occlusal caries depending on the validation method. Using histological validation, MPDx (A(z)=0.76) was significantly less accurate than Dixi (A(z)=0.81), Sidexis (A(z)=0.8), Ektaspeed Plus (A(z)=0.82) and Insight (A(z)=0.81). Using radiographic validation, MPDx (A(z)=0.83) was also significantly less accurate than RVG(old) (A(z)=0.89) and RVG(new) (A(z)=0.9). CONCLUSION: A(z) obtained from radiographic validation was significantly higher than A(z) obtained from histological validation. Comparison of the diagnostic efficacy for caries of the eight radiographic systems was strongly influenced by the validation method. DOI: 10.1038/sj/dmfr/4600645

Analysis of Variance↗

Validity: on meaningful interpretation of assessment data.

CONTEXT: All assessments in medical education require evidence of validity to be interpreted meaningfully. In contemporary usage, all validity is construct validity, which requires multiple sources of evidence; construct validity is the whole of validity, but has multiple facets. Five sources--content, response process, internal structure, relationship to other variables and consequences--are noted by the Standards for Educational and Psychological Testing as fruitful areas to seek validity evidence. PURPOSE: The purpose of this article is to discuss construct validity in the context of medical education and to summarize, through example, some typical sources of validity evidence for a written and a performance examination. SUMMARY: Assessments are not valid or invalid; rather, the scores or outcomes of assessments have more or less evidence to support (or refute) a specific interpretation (such as passing or failing a course). Validity is approached as hypothesis and uses theory, logic and the scientific method to collect and assemble data to support or fail to support the proposed score interpretations, at a given point in time. Data and logic are assembled into arguments--pro and con--for some specific interpretation of assessment data. Examples of types of validity evidence, data and information from each source are discussed in the context of a high-stakes written and performance examination in medical education. CONCLUSION: All assessments require evidence of the reasonableness of the proposed interpretation, as test data in education have little or no intrinsic meaning. The constructs purported to be measured by our assessments are important to students, faculty, administrators, patients and society and require solid scientific evidence of their meaning.

Data Interpretation, Statistical↗

[Reliability and validity of the Japanese version of the coping inventory for stressful situations (CISS): a contribution to the cross-cultural studies of coping].

OBJECTIVE: There has recently been a dramatic increase in the number of studies on coping behavior as an intervening variable between stress and health. Most of the available measures of coping are, however, psychometrically inadequate. We therefore decided to develop the Japanese version of the Coping Inventory for Stressful Situations (CISS) with special regards to its cross-cultural equivalence, reliability and validity. The CISS is a self-report measure of an individual's typical pattern of coping along three orthogonal dimensions of Task-, Emotion-, and Avoidance-oriented coping; its reliability and validity have been well studied in North America, where it was originally developed. METHOD: We obtained the Japanese version of the CISS (J-CISS) by means of back-translation. In Study 1, we administered the J-CISS and the 12-item General Health Questionnaire (GHQ) to 33 Japanese university students twice with an interval of four weeks. In Study 2,550 Japanese high school students completed the J-CISS and the Maudsley Personality Inventory. RESULTS: The equivalence of the Japanese version with the original was ascertained by means of back-translation involving multiple, independent mental health professionals and by factor congruence between the two versions. A principal component factor analysis (Varimax rotation) of the Study 2 data allowed us to extract three factors, which were virtually identical to the original ones. The high corrected item-remainder correlations, internal consistency reliabilities and test-retest reliabilities all attested to the reliability of the J-CISS. In order to examine its content validity, we compared the J-CISS with two coping questionnaires that have been in use in Japan, and found that the J-CISS covered most of the coping styles in these two questionnaires. However, such coping styles as "giving up," "to lose is to win (a Japanese proverb)," "it is best to do nothing" were not included in the original CISS and hence in the J-CISS. The criterion validity of the J-CISS was examined both in terms of predictive validity and concurrent validity. In Study 1, those who scored below the cut-off of the GHQ at Time 1 but above the cut-off at Time 2 had significantly higher Emotion-oriented coping scores at Time 1 than those who remained below the cut-off of the GHQ at Times 1 and 2 (predictive validity). In Study 2, the J-CISS scales and the MPI scales showed theoretically predicted correlations (concurrent validity). The results of the factor analysis and the corrected item-remainder correlations were suggestive of high construct validity of the J-CISS. Moreover, the mean inter-item correlation was between .20 and .40 for each scale, indicating its homogeneity. Factor analysis of each scale revealed that each scale indeed contained only one factor. Correlations among the three scales of the J-CISS established that the three scales formed multi-dimensional measures of coping. CONCLUSION: The results of our study indicate 1) that the obtained Japanese version of the CISS is to be regarded as final, 2) that coping styles can be measured in a consistent and reliable manner both in Japan and North America, and 3) that this cross-cultural equivalence as well as the other validity studies have further augmented the validity of the CISS itself.

Adaptation, Psychological↗

Monitoring sedation status over time in ICU patients: reliability and validity of the Richmond Agitation-Sedation Scale (RASS).

CONTEXT: Goal-directed delivery of sedative and analgesic medications is recommended as standard care in intensive care units (ICUs) because of the impact these medications have on ventilator weaning and ICU length of stay, but few of the available sedation scales have been appropriately tested for reliability and validity. OBJECTIVE: To test the reliability and validity of the Richmond Agitation-Sedation Scale (RASS). DESIGN: Prospective cohort study. SETTING: Adult medical and coronary ICUs of a university-based medical center. PARTICIPANTS: Thirty-eight medical ICU patients enrolled for reliability testing (46% receiving mechanical ventilation) from July 21, 1999, to September 7, 1999, and an independent cohort of 275 patients receiving mechanical ventilation were enrolled for validity testing from February 1, 2000, to May 3, 2001. MAIN OUTCOME MEASURES: Interrater reliability of the RASS, Glasgow Coma Scale (GCS), and Ramsay Scale (RS); validity of the RASS correlated with reference standard ratings, assessments of content of consciousness, GCS scores, doses of sedatives and analgesics, and bispectral electroencephalography. RESULTS: In 290-paired observations by nurses, results of both the RASS and RS demonstrated excellent interrater reliability (weighted kappa, 0.91 and 0.94, respectively), which were both superior to the GCS (weighted kappa, 0.64; P<.001 for both comparisons). Criterion validity was tested in 411-paired observations in the first 96 patients of the validation cohort, in whom the RASS showed significant differences between levels of consciousness (P<.001 for all) and correctly identified fluctuations within patients over time (P<.001). In addition, 5 methods were used to test the construct validity of the RASS, including correlation with an attention screening examination (r = 0.78, P<.001), GCS scores (r = 0.91, P<.001), quantity of different psychoactive medication dosages 8 hours prior to assessment (eg, lorazepam: r = - 0.31, P<.001), successful extubation (P =.07), and bispectral electroencephalography (r = 0.63, P<.001). Face validity was demonstrated via a survey of 26 critical care nurses, which the results showed that 92% agreed or strongly agreed with the RASS scoring scheme, and 81% agreed or strongly agreed that the instrument provided a consensus for goal-directed delivery of medications. CONCLUSIONS: The RASS demonstrated excellent interrater reliability and criterion, construct, and face validity. This is the first sedation scale to be validated for its ability to detect changes in sedation status over consecutive days of ICU care, against constructs of level of consciousness and delirium, and correlated with the administered dose of sedative and analgesic medications.

Aged↗

Criterion and content validity of a novel structured haggling contingent valuation question format versus the bidding game and binary with follow-up format.

Contingent valuation question formats that will be used to elicit willingness to pay for goods and services need to be relevant to the area they will be used in order for responses to be valid. A novel contingent valuation question format called the "structured haggling technique" (SH) that resembles the bargaining system in Nigerian markets was designed and its criterion and content validity compared with those of the bidding game (BG) and binary-with-follow-up (BWFU) technique. This was achieved by determining the willingness to pay (WTP) for insecticide-treated nets (ITNs) in Southeast Nigeria. Content validity was determined through observation of actual trading of untreated nets together with interviews with sellers and consumers. Criterion validity was determined by comparing stated and actual WTP. Stated WTP was determined using a questionnaire administered to 810 household heads and actual WTP was determined by offering the nets for sale to all respondents one month later. The phi (correlation) coefficient was used to compare criterion validity across question formats. The phi coefficients were SH (0.60: 95% C.I. 0.50-0.71), BG (0.42: 95% C.I. 0.29-0.54) and the BWFU (0.32: 95% C.I. 0.20-0.44), implying that the BG and SH had similar levels of criterion-validity while the BWFU was the least criterion-valid. However, the SH was the most content-valid. It is necessary to validate the findings in other areas where haggling is common. Future studies should establish the content validity of question formats in the contexts in which they will be used before administering questionnaires.

Attitude to Health↗

Internal and external validation of the NOSEP prediction score for nosocomial sepsis in neonates.

OBJECTIVE: To evaluate the performance of a scoring system (NOSEP) to predict nosocomial sepsis in neonates at the hospital where the score was developed (internal validation) and in an independent data set from other centers (external validation). DESIGN: Multiple center prospective cohort study. SETTING: Six neonatal intensive care units from the Flanders in Belgium. PATIENTS: We analyzed two groups of patients: 62 episodes of presumed nosocomial sepsis in the internal validation cohort and 93 episodes of presumed nosocomial sepsis in a multiple center external validation cohort. INTERVENTIONS: Assessment of the predictive power of the NOSEP score 24 hrs preceding sepsis workup and the patients' basic demographic characteristics and co-morbidity was performed. Diagnosis of nosocomial sepsis and the microbiology results were registered. MAIN RESULTS: The NOSEP score's discriminative capability was very good in the internal validation (area under receiver operating characteristic curve = 0.73 +/- 0.08 [sem]). The NOSEP score performed satisfactory in the external validation (area under receiver operating characteristic curve = 0.66 +/- 0.06). The calibration capability in both validation sets as measured by goodness-of-fit tests (internal validation, p =.56; external validation, p =.48) was good. An improvement of the NOSEP score was obtained for the external centers by redefining the cut-off of the items of the NOSEP score (area under receiver operating characteristic curve for NOSEP-NEW-I = 0.71 +/- 0.05) or adding co-morbidity factors (area under receiver operating characteristic curve for NOSEP-NEW-II = 0.82 +/- 0.04), with good calibration performance (goodness-of-fit test, p >.50). Finally, the fit of the NOSEP score demonstrated no significant variation across subgroups of patients. CONCLUSIONS: The predictive power of the original NOSEP score is very good in neonates at the original neonatal intensive care unit. In other neonatal intensive care units, its discriminatory performance is satisfactory but could be improved after modification of the variables in the model or adding additional variables. To use such a NOSEP score in other neonatal intensive care units, its accuracy has to be validated and adjusted if necessary.

Cross Infection↗

Evolution of a companywide validation program.

A successful computer validation program requires a solid foundation of policies, guidelines, and procedures. An unstructured approach to validation may yield adequate results for a specific computerized system, but it will not sustain ongoing, consistent computer validation activities over time. The evolution of a strong computer validation program starts with an awareness of the need for such a program and builds on that established base. One approach is a step-by-step method in which each new step builds on the results of one or more of the previous steps. However, the order of events is not as important as the ultimate completion of all the building blocks. A strong computer validation program requires a companywide policy and an administrator who provides oversight and chairs the computer validation committee. The committee should be established to provide departmental leadership and to develop company guidelines for computer validation. Committee members will also create system inventories and classify and prioritize those systems in terms of the need for validation. Departmental standard operating procedures must be developed to provide standardized methods for routine validation activities. Finally the quality assurance unit should have procedures in place that provide for its involvement in the computer validation process.

Computer Systems↗

Validation of an outcomes instrument for tonsil and adenoid disease.

OBJECTIVE: To design and validate a disease-specific health status instrument-the Tonsil and Adenoid Health Status Instrument-for use in children with tonsil and adenoid disease. DESIGN: Prospective psychometric and clinimetric instrument validation in 3 stages. SETTINGS: A tertiary academic pediatric specialty hospital and a tertiary academic hospital, in 2 different cities. PATIENTS/OTHER PARTICIPANTS: Children with tonsil and adenoid disease presenting for evaluation and treatment (n = 224). INTERVENTION/METHOD: Prospective instrument validation. Stage 1 consisted of initial item testing, reduction, and subscale construction; stage 2, reliability and validity testing, factor analysis, and final item reduction; and stage 3, responsiveness analysis. MAIN OUTCOME MEASURES: Test-retest and internal consistency reliability; content, construct, and criterion validity; orthogonal principal components factor analysis; and response sensitivity analysis. RESULTS: Factor analysis and item analysis confirmed 6 distinct subscales measuring different constructs (aspects) of disease-specific health status that are affected by tonsil and adenoid disease: eating and swallowing, airway and breathing, infections, health care utilization, cost of care, and behavior. For each subscale, the Tonsil and Adenoid Health Status Instrument demonstrated excellent test-retest reliability (r = 0.72-0.88) and internal consistency reliability (Cronbach alpha = .73-.87). Content validity was ensured during the design process. Construct validity was demonstrated by means of convergent and divergent validity with a global quality-of-life instrument (the Child Health Questionnaire, version PF28). Criterion validity was also satisfactory. Finally, the instrument was appropriately sensitive, with high standardized response means and effect sizes. CONCLUSIONS: The Tonsil and Adenoid Health Status Instrument is a valid, reliable, and sensitive instrument with 6 distinct subscales. This instrument has significant utility for outcomes research in children with tonsil and adenoid disease.

Adenoids↗

The importance of validating the diagnosis of coronary heart disease when measuring secondary prevention: a cross-sectional study in general practice.

PURPOSE: To compare levels of recorded risk factors and drug treatment between patients with validated and non-validated diagnoses of coronary heart disease (CHD) in Northern Ireland. METHODS: Patients with a nitrate prescription in the previous year or a CHD Read code were identified from computer records of 25 practices, stratified by partnership size and area board. Computer and paper records of a random sample of 10% of these were searched for specified criteria to validate the diagnosis of CHD. The diagnosis was considered valid if the patient was found to have one or more positive investigations for CHD. Records of blood pressure, cholesterol, blood sugar, body mass index and drugs prescribed were taken into account. RESULTS: The combined practice population was 151,071; 7338 (4.86%) were identified by the computer search as meeting the defined entry criteria for CHD. Among the 10% random sample the diagnosis of CHD could not be validated for 36.5% (265/727). Significantly more patients with a validated than non-validated diagnosis had recorded cholesterol levels below 5.0 mmol/l (55.8 vs. 34.5%, p < 0.001) and were prescribed aspirin (75.3 vs. 40.8%, p < 0.001), beta-blockers (51.5 vs. 28.3%, p < 0.001), angiotensin-converting-enzyme inhibitors (29.2 vs. 15.5%, p < 0.001) and lipid-lowering drugs (50.9 vs. 23.0%, p < 0.001). A recent nitrate prescription had a higher predictive value for validated CHD than a Read code for CHD alone (71.2 vs. 53.1%, p < 0.001). No other significant differences were found between the two groups regarding the extent or levels of recorded risk factors. CONCLUSIONS: Patients with a validated diagnosis of CHD appear to be better managed than those whose diagnosis has not been confirmed. Validation of diagnosis has important implications for assessing the provision of secondary prevention and for clinical governance.

Adolescent↗

Constructing and Validating Motive Bridging Inferences

Understanding Jane left early for the birthday party, She spent an hour shopping at the mall requires detecting that the first statement motivates the second. The validation model states that before accepting this bridging inference, the reader validates it with reference to relevant knowledge. In particular, a mediating idea is first derived from the text outcome and its candidate motive. If the mediating idea is supported by general knowledge, then the inference has been validated. In tests of this anaylsis, experimental subjects read motive or control sequences and then answered questions probing the knowledge hypothesized to validate the motive inferences, such as Do birthday parties involve presents? Five experiments confirmed that understanding motive sequences facilitates validating knowledge. A control procedure also refuted a priming counterexplanation of these effects (Experiment 1). Validation processing obtained for motive-outcome statements separated by two to four sentences in coherent sequences (Experiments 2 to 4). Inferred and explicit validating knowledge had a similar representational status (Experiment 3). Whereas proofreading abolished the validation effect, a reading strategy promoting causal processing did not enhance it (Experiment 4). A delayed priming procedure indicated that validating knowledge is integrated with the text representation (Experiment 5). The implications of these findings for the constructionist and minimal inference analyses were explored. The validation effects were simulated using construction-integration model.

Journal Article↗

Approaches for assessing the validity of a functional observational battery.

As neurobehavioral assessments during the preliminary stages of chemical testing are more widely undertaken, it is critical that the screening procedures utilized be valid indicators of neurobehavioral function and that they be sensitive, specific, and reliable. Efforts in this laboratory have been directed towards assessing these features in the use of a functional observational battery (FOB). For the purpose of assessing validity, we have examined FOB data which addresses the issues of criterion, predictive, concurrent, and construct validities. The FOB appears to be valid for detecting chemical-induced neurological dysfunction in rats, i.e., shows a good degree of criterion validity. Furthermore, in many instances the effects observed with the FOB may be predictive of symptomatology in humans. When comparisons can be made between effects detected with the FOB and other methods of measuring neurotoxicity (e.g., neuropathology), concurrent validity can also be established. To assess construct validity, effects of neurotoxicants can be classified into functional domains which are described by various measures in the FOB. Approaches for assessing the validity of the test method thus include answering specific research questions directed at assessing criterion, predictive, concurrent, and construct validity. Available data indicate that, in these aspects, the FOB is a valid screening method for the detection of neurotoxicity.

Animals↗

The validation of three human reliability quantification techniques--THERP, HEART and JHEDI: Part III--Practical aspects of the usage of the techniques.

This is the third paper in a series of three dealing with the detailed investigation of the empirical validity of three human reliability assessment (HRA) techniques. The first paper introduced the need for validation and specified the three techniques most requiring validation. The second paper detailed the results of an extensive independent validation experiment. This experimental validation involved 30 UK assessors using the techniques THERP, HEART and JHEDI (10 assessors per technique) to estimate the human error probabilities (HEPs) for 30 nuclear power and reprocessing (NP&R) tasks. The results for all three techniques were positive in terms of significant correlations, and general precision levels of 72% of all HEP estimates within a factor of 10 of the true value (unknown to the assessors). These results lend support to the empirical validity of these techniques in particular, and to HRA in general. However, the results were not all positive. In particular the consistency of usage of the techniques was variable. Additionally, subjects were generally not good at knowing their own uncertainty, i.e. they were not able to accurately predict when they were accurate nor when they were inaccurate. This desirable parameter is known as calibration, and the results from the validation suggested that subjects were not well-calibrated. This paper aims to determine how consistency of usage can be improved and to discern whether certain task types are, in practice, not well-assessed by the techniques, and hence are effectively currently beyond these techniques' abilities. Such information is aimed at aiding the HRA practitioner, or the ergonomist, interested in using these techniques. Recommendations for improving calibration are also discussed in this paper. A subsidiary but important focus of this paper is of a more fundamental nature, and of more general interest to the ergonomist. It concerns the validity of the techniques from an error reduction perspective. Currently these techniques may be used to identify how to reduce error probability, which is generally (in the qualitative sense) within the domain of ergonomics. One major mechanism for HRA-based error reduction is the utilisation of Performance Shaping Factor (PSF) information. This paper considers the validity of these PSF as ergonomics constructs. Drawing results from the validation exercise, it is seen how different PSF can be applied to the same scenario and can result in the same error probability, but will result in different error reduction guidance. It is therefore recommended that error reduction guidance must be based on a composite analysis of the results of the task, error identification and quantification analyses, with most weighting given to the qualitative analyses.

Evaluation Studies as Topic↗

Teasing apart quality and validity in systematic reviews: an example from acupuncture trials in chronic neck and back pain.

The objectives of the study were (1) to carry out a systematic review to assess the analgesic efficacy and the adverse effects of acupuncture compared with placebo for back and neck pain and (2) to develop a new tool, the Oxford Pain Validity Scale (OPVS), to measure validity of findings from randomized controlled trials (RCTs), and to enable ranking of trial findings according to validity within qualitative reviews. Published RCTs (of acupuncture at both traditional and non-traditional points) were identified from systematic searching of bibliographic databases (e.g. MEDLINE) and reference lists of retrieved reports. Pain outcome data were extracted with preference given to standardized outcomes such as pain intensity. Information on adverse effects was also extracted. All included trials were scored using a five-item 0-16 point validity scale (OPVS). The individual RCTs were ranked according to their OPVS score to enable more weight to be placed on the trials of greater validity when drawing an overall conclusion about the efficacy of acupuncture for relieving neck and back pain. Statistical analyses were carried out on the OPVS scores to assess the relationship between trial finding (positive or negative) and validity. Thirteen RCTs met the inclusion criteria. Five trials concluded that acupuncture was effective, and eight concluded that it was not effective for relieving back or neck pain. There was no obvious difference between the findings of trials using traditional and non-traditional points. Using the new OPVS scale, the validity scores of the included trials ranged from 4 to 14. There was no significant relationship between OPVS score and trial finding (positive versus negative). Authors' conclusions did not always agree with their data. We drew our own conclusions (positive/negative) based on the data presented in the reports. Re-analysis using our conclusions showed a significant relationship between OPVS score and trial finding, with higher validity scores associated with negative findings. OPVS is a useful tool for assessing the validity of trials in qualitative reviews. With acupuncture for chronic back and neck pain, we found that the most valid trials tended to be negative. There is no convincing evidence for the analgesic efficacy of acupuncture for back or neck pain.

Acupuncture Therapy↗

Plateletapheresis: instrumentation validation.

Plateletapheresis instrumentation validation is required to document that a new or modified instrument or technique is capable of consistently producing acceptable products at the production center using their equipment, personnel, and counting techniques even though the instrument or technique may already have FDA or equivalent approval for use. To pursue the process of validation, several questions need to be addressed: when is it required, what products are validated, what parameters are monitored, and how many products are required. Validation is required when a new instrument or technique (process) is used that could affect the quality of the product. According to the FDA, each apheresis system (e.g., Spectra LRS, Amicus) and each type of product (e.g., single, double, triple) need to be validated separately. Parameters to be validated vary, but usually platelet (plt) yield, white blood cell (WBC) content (if products are labeled "leukoreduced"), and 5-day storage pH are monitored. The number of procedures monitored is also quite variable, but we use 20 samples for highly variable parameters such as platelet yield and WBC content and five samples for less variable parameters such as 5-day storage pH. As an example, we validated the Fenwal Amicus (Baxter Biotech) for single apheresis platelet products. With 20 samples, we found that: 85% of the products contained > or = 3 x 10(11) plt (requirement was at least 75% contain > or = 3 x 10(11) plt); platelet concentration of all products was < or = 1.515 x 10(6) plt/microL (requirement was < or = 2.435 x 10(6) plt/microL), and WBC content was < 1 x 10(6) WBC in all products (requirement was all products contain < 5 x 10(6) WBC). In addition, in five samples, the 5-day storage pH was 6.89-7.25 (requirement was all products should be > or = 6.2 pH). Once validation is complete and acceptable, the process should be monitored on a regular basis using some form of process control. Statistical process control programs are available that can assist in documenting validation and ongoing process control. With the use of process validation and ongoing process control, the plateletapheresis center can assure that acceptable products are consistently being produced.

Forms and Records Control↗

Incremental validity of new clinical assessment measures.

The authors address conceptual and methodological foundations of incremental validity in the evaluation of newly developed clinical assessment measures. Incremental validity is defined as the degree to which a measure explains or predicts a phenomenon of interest, relative to other measures. Incremental validity can be evaluated on several dimensions, such as sensitivity to change, diagnostic efficacy, content validity, treatment design and outcome, and convergent validity. Indices of incremental validity can vary depending on the criterion measures, comparison measures, and individual differences in samples. The authors review the rationale for, principles, and methods of incremental validation, including the selection of comparison and criterion measures, and address data analytic strategies and the conditional nature of incremental validity evaluations in the selection of measures. Incremental validity contributes to, but is different from, cost-benefits, which reflect the cost of acquiring the data and the benefits from the data. The impact of an incremental validity index on whether a measure is selected will be moderated by the cost of acquiring the new data, the importance of the measured phenomenon, and the clinical utility of the new data.

Humans↗