Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Constructing and Validating Motive Bridging Inferences

Understanding Jane left early for the birthday party, She spent an hour shopping at the mall requires detecting that the first statement motivates the second. The validation model states that before accepting this bridging inference, the reader validates it with reference to relevant knowledge. In particular, a mediating idea is first derived from the text outcome and its candidate motive. If the mediating idea is supported by general knowledge, then the inference has been validated. In tests of this anaylsis, experimental subjects read motive or control sequences and then answered questions probing the knowledge hypothesized to validate the motive inferences, such as Do birthday parties involve presents? Five experiments confirmed that understanding motive sequences facilitates validating knowledge. A control procedure also refuted a priming counterexplanation of these effects (Experiment 1). Validation processing obtained for motive-outcome statements separated by two to four sentences in coherent sequences (Experiments 2 to 4). Inferred and explicit validating knowledge had a similar representational status (Experiment 3). Whereas proofreading abolished the validation effect, a reading strategy promoting causal processing did not enhance it (Experiment 4). A delayed priming procedure indicated that validating knowledge is integrated with the text representation (Experiment 5). The implications of these findings for the constructionist and minimal inference analyses were explored. The validation effects were simulated using construction-integration model.

Journal Article↗

Approaches for assessing the validity of a functional observational battery.

As neurobehavioral assessments during the preliminary stages of chemical testing are more widely undertaken, it is critical that the screening procedures utilized be valid indicators of neurobehavioral function and that they be sensitive, specific, and reliable. Efforts in this laboratory have been directed towards assessing these features in the use of a functional observational battery (FOB). For the purpose of assessing validity, we have examined FOB data which addresses the issues of criterion, predictive, concurrent, and construct validities. The FOB appears to be valid for detecting chemical-induced neurological dysfunction in rats, i.e., shows a good degree of criterion validity. Furthermore, in many instances the effects observed with the FOB may be predictive of symptomatology in humans. When comparisons can be made between effects detected with the FOB and other methods of measuring neurotoxicity (e.g., neuropathology), concurrent validity can also be established. To assess construct validity, effects of neurotoxicants can be classified into functional domains which are described by various measures in the FOB. Approaches for assessing the validity of the test method thus include answering specific research questions directed at assessing criterion, predictive, concurrent, and construct validity. Available data indicate that, in these aspects, the FOB is a valid screening method for the detection of neurotoxicity.

Animals↗

The validation of three human reliability quantification techniques--THERP, HEART and JHEDI: Part III--Practical aspects of the usage of the techniques.

This is the third paper in a series of three dealing with the detailed investigation of the empirical validity of three human reliability assessment (HRA) techniques. The first paper introduced the need for validation and specified the three techniques most requiring validation. The second paper detailed the results of an extensive independent validation experiment. This experimental validation involved 30 UK assessors using the techniques THERP, HEART and JHEDI (10 assessors per technique) to estimate the human error probabilities (HEPs) for 30 nuclear power and reprocessing (NP&R) tasks. The results for all three techniques were positive in terms of significant correlations, and general precision levels of 72% of all HEP estimates within a factor of 10 of the true value (unknown to the assessors). These results lend support to the empirical validity of these techniques in particular, and to HRA in general. However, the results were not all positive. In particular the consistency of usage of the techniques was variable. Additionally, subjects were generally not good at knowing their own uncertainty, i.e. they were not able to accurately predict when they were accurate nor when they were inaccurate. This desirable parameter is known as calibration, and the results from the validation suggested that subjects were not well-calibrated. This paper aims to determine how consistency of usage can be improved and to discern whether certain task types are, in practice, not well-assessed by the techniques, and hence are effectively currently beyond these techniques' abilities. Such information is aimed at aiding the HRA practitioner, or the ergonomist, interested in using these techniques. Recommendations for improving calibration are also discussed in this paper. A subsidiary but important focus of this paper is of a more fundamental nature, and of more general interest to the ergonomist. It concerns the validity of the techniques from an error reduction perspective. Currently these techniques may be used to identify how to reduce error probability, which is generally (in the qualitative sense) within the domain of ergonomics. One major mechanism for HRA-based error reduction is the utilisation of Performance Shaping Factor (PSF) information. This paper considers the validity of these PSF as ergonomics constructs. Drawing results from the validation exercise, it is seen how different PSF can be applied to the same scenario and can result in the same error probability, but will result in different error reduction guidance. It is therefore recommended that error reduction guidance must be based on a composite analysis of the results of the task, error identification and quantification analyses, with most weighting given to the qualitative analyses.

Evaluation Studies as Topic↗

Teasing apart quality and validity in systematic reviews: an example from acupuncture trials in chronic neck and back pain.

The objectives of the study were (1) to carry out a systematic review to assess the analgesic efficacy and the adverse effects of acupuncture compared with placebo for back and neck pain and (2) to develop a new tool, the Oxford Pain Validity Scale (OPVS), to measure validity of findings from randomized controlled trials (RCTs), and to enable ranking of trial findings according to validity within qualitative reviews. Published RCTs (of acupuncture at both traditional and non-traditional points) were identified from systematic searching of bibliographic databases (e.g. MEDLINE) and reference lists of retrieved reports. Pain outcome data were extracted with preference given to standardized outcomes such as pain intensity. Information on adverse effects was also extracted. All included trials were scored using a five-item 0-16 point validity scale (OPVS). The individual RCTs were ranked according to their OPVS score to enable more weight to be placed on the trials of greater validity when drawing an overall conclusion about the efficacy of acupuncture for relieving neck and back pain. Statistical analyses were carried out on the OPVS scores to assess the relationship between trial finding (positive or negative) and validity. Thirteen RCTs met the inclusion criteria. Five trials concluded that acupuncture was effective, and eight concluded that it was not effective for relieving back or neck pain. There was no obvious difference between the findings of trials using traditional and non-traditional points. Using the new OPVS scale, the validity scores of the included trials ranged from 4 to 14. There was no significant relationship between OPVS score and trial finding (positive versus negative). Authors' conclusions did not always agree with their data. We drew our own conclusions (positive/negative) based on the data presented in the reports. Re-analysis using our conclusions showed a significant relationship between OPVS score and trial finding, with higher validity scores associated with negative findings. OPVS is a useful tool for assessing the validity of trials in qualitative reviews. With acupuncture for chronic back and neck pain, we found that the most valid trials tended to be negative. There is no convincing evidence for the analgesic efficacy of acupuncture for back or neck pain.

Acupuncture Therapy↗

Plateletapheresis: instrumentation validation.

Plateletapheresis instrumentation validation is required to document that a new or modified instrument or technique is capable of consistently producing acceptable products at the production center using their equipment, personnel, and counting techniques even though the instrument or technique may already have FDA or equivalent approval for use. To pursue the process of validation, several questions need to be addressed: when is it required, what products are validated, what parameters are monitored, and how many products are required. Validation is required when a new instrument or technique (process) is used that could affect the quality of the product. According to the FDA, each apheresis system (e.g., Spectra LRS, Amicus) and each type of product (e.g., single, double, triple) need to be validated separately. Parameters to be validated vary, but usually platelet (plt) yield, white blood cell (WBC) content (if products are labeled "leukoreduced"), and 5-day storage pH are monitored. The number of procedures monitored is also quite variable, but we use 20 samples for highly variable parameters such as platelet yield and WBC content and five samples for less variable parameters such as 5-day storage pH. As an example, we validated the Fenwal Amicus (Baxter Biotech) for single apheresis platelet products. With 20 samples, we found that: 85% of the products contained > or = 3 x 10(11) plt (requirement was at least 75% contain > or = 3 x 10(11) plt); platelet concentration of all products was < or = 1.515 x 10(6) plt/microL (requirement was < or = 2.435 x 10(6) plt/microL), and WBC content was < 1 x 10(6) WBC in all products (requirement was all products contain < 5 x 10(6) WBC). In addition, in five samples, the 5-day storage pH was 6.89-7.25 (requirement was all products should be > or = 6.2 pH). Once validation is complete and acceptable, the process should be monitored on a regular basis using some form of process control. Statistical process control programs are available that can assist in documenting validation and ongoing process control. With the use of process validation and ongoing process control, the plateletapheresis center can assure that acceptable products are consistently being produced.

Forms and Records Control↗

Incremental validity of new clinical assessment measures.

The authors address conceptual and methodological foundations of incremental validity in the evaluation of newly developed clinical assessment measures. Incremental validity is defined as the degree to which a measure explains or predicts a phenomenon of interest, relative to other measures. Incremental validity can be evaluated on several dimensions, such as sensitivity to change, diagnostic efficacy, content validity, treatment design and outcome, and convergent validity. Indices of incremental validity can vary depending on the criterion measures, comparison measures, and individual differences in samples. The authors review the rationale for, principles, and methods of incremental validation, including the selection of comparison and criterion measures, and address data analytic strategies and the conditional nature of incremental validity evaluations in the selection of measures. Incremental validity contributes to, but is different from, cost-benefits, which reflect the cost of acquiring the data and the benefits from the data. The impact of an incremental validity index on whether a measure is selected will be moderated by the cost of acquiring the new data, the importance of the measured phenomenon, and the clinical utility of the new data.

Humans↗

Experience with two validation methods in a prevalence survey on nosocomial infections.

OBJECTIVE: To determine whether an investigator effect remained on the first German study on the prevalence of nosocomial infections Nosokomiale Infektionen in Deutschland Erfassung und Prävention (NIDEP), despite extensive validation efforts. DESIGN: Two validation methods were applied: bedside validation and validation by case studies. In both cases, the results of the four investigators were compared with the diagnosis of gold standard observers. SETTING: Validation measures were applied before, intermittently, during, and at the end of the surveillance period in 72 acute-care hospitals with 14,966 patients. RESULTS: The overall sensitivity in the bedside-validation periods was 89.0%; the overall specificity was 99.5%. For validation by case studies, overall sensitivity was 95.6%, and overall specificity was 92.8%. At the end of the surveillance, a remarkable investigator effect was found. CONCLUSION: Despite validation results that were assessed as satisfactory, based on available literature, an investigator effect was observed. This underlines the need for data validation and the formulation of recommendations for data validation. Clarification of the Centers for Disease Control and Prevention criteria for pneumonia and primary bloodstream infection and the inclusion of some diagnostic test results may reduce or prevent an investigator effect in future studies.

Bias↗

How valid and reliable are patient satisfaction data? An analysis of 195 studies.

OBJECTIVE: To assess the properties of validity and reliability of instruments used to assess satisfaction in a broad sample of health service user satisfaction studies, and to assess the level of awareness of these issues among study authors. DESIGN: Examination and analysis of 195 papers published in 1994 in 139 journals. The following databases were searched: British Nursing Index, CINAHL, EMBASE, MedLine, Popline, and PsycLIT. MAIN MEASURES: Number and types of strategies used for content, criterion, and construct validity, and for stability and internal consistency. Associations between validity/reliability and other study characteristics. RESULTS: Eighty-nine (46%) of the 195 studies reported some validity or reliability data; 76 reported some element of content validity; 14 reported criterion validity, with patient's intent to return the most commonly used criterion; four reported construct validity. Thirty-four studies reported internal consistency reliability, 31 of which used Cronbach's coefficient alpha; eight studies reported test-retest reliability. Only 11 studies (6% of the 181 quantitative studies) reported content validity and criterion or construct validity and reliability. 'New' instruments designed specifically for the reported study demonstrated significantly less evidence for reliability/validity than did 'old' instruments. CONCLUSION: With few exceptions, the study instruments in this sample demonstrated little evidence of reliability or validity. Moreover, study authors exhibited a poor understanding of the importance of these properties in the assessment of satisfaction. Researchers must be aware that this is poor research practice, and that lack of a reliable and valid assessment instrument casts doubt on the credibility of satisfaction findings.

Data Collection↗

Testing for predictive validity in health care education research: a critical review.

BACKGROUND: The assessment of predictive validity is the essential core from which a sound model of prediction is built. METHOD: Three methods for assessing predictive validity in health care education research were reviewed: longitudinal profile development, cross-validation, and inspection of the adjusted R2. A total of 47 articles published between 1973 and 1993 in nine health care disciplines were critically reviewed to determine whether the studies tested for predictive validity by using these methods. RESULTS: Very few of the 47 studies used at least one of the three methods for assessing predictive validity. Furthermore, the proportion of variance explained that is reported in the articles is typically small even before assessment of predictive validity. It is sobering to note that these small values may be inflated, since shrinkage is likely to occur when assessing predictive validity on a second, or cross-validation, sample. CONCLUSION: The scarcity of testing of predictive validity in the studies reviewed highlights the necessity of future research to establish the degree of predictive validity, if improvements in predicting success in health care education research are to be realized.

Education↗

The reliability and validity of three dimensional ultrasound volumetric measurements using an in vitro balloon and in vivo uterine model.

OBJECTIVE: To evaluate the reliability and validity of two and three dimensional ultrasound volumetric measurements using balloon and uterine models. DESIGN: Prospetive observational study. SETTING: Obstetric ultrasound department at a university teaching hospital. METHOD: Two and three dimensional ultrasound volumetric measurements (with 5, 10 and 15 ultrasonic slices) were performed on 30 different sets of ultrasound images obtained from 15 water filled balloons with volumes ranging from 19 to 697mL. The measurements were performed independently by two observers who were blinded to the true volumes of the balloons. For the uterine model, only three dimensional ultrasonic volume measurements were performed independently on 16 uteri by two observers who were again unaware of the definitive uterine volumes. OUTCOME MEASURE: For the assessment of intra-and inter-rater reliability, the intraclass correlation coefficient was used. The index of concordance between the ultrasonic volumes and those obtained by the reference standard (validity) was assessed with the conventional Pearson's correlation coefficient, limits of agreement method and the intra-class correlation coefficient. RESULTS: High levels of reliability and validity were obtained for both two and three dimensional ultrasound balloon volume measurements. For two dimensional ultrasonic volume measurements, the intra-class correlation coefficient ranged from 0.992 to 0.998 for reliability and validity whereas the Pearson's correlation coefficient for validity was 0.996. With three dimensional ultrasonic volume measurements, the intra-class correlation coefficient ranged from 0.991 to 0.999 for reliability and validity whereas the Pearson's correlation coefficient for validity was 0.999. Both two and three dimensional ultrasonic measurements tended to underestimate the true balloon volume with the largest observed mean difference obtained with three dimensional ultrasound measurements using five ultrasonic slices and the smallest value obtained with three dimensional ultrasound measurements employing 15 ultrasonic slices. The mean difference in volume measurement for two dimensional ultrasound was intermediate between these two values. However, two dimensional ultrasound volume measurement generated the largest range between the limits of agreement whereas the smallest range was obtained with three dimensional ultrasound using 10 ultrasonic slices. The intra-class correlation coefficient for reliability and validity with three dimensional ultrasonic uterine volume estimation ranged from 0.956 to 0.996 whereas the Pearson's correlation coefficient for validity ranged from 0.993 to 0.999). The use of three dimensional ultrasound also consistently under-estimated the actual uterine volumes. The larger the number of ultrasonic slices employed for three dimensional ultrasound, the smaller was the mean difference between the ultrasonic and true uterine volume measurements and the smaller the limits of agreement. CONCLUSIONS: The reliability and validity of balloon and uterine volume measurement by three dimensional ultrasound is high. This allows further research on three dimensional ultrasound for measuring pelvic organ volumes in the prediction of pelvic pathology.

Female↗

External validity, generalizability, and knowledge utilization.

PURPOSE: To examine the concepts of external validity and generalizability, and explore strategies to strengthen generalizability of research findings, because of increasing demands for knowledge utilization in an evidence-based practice environment. FRAMEWORK: The concepts of external validity and generalizability are examined, considering theoretical aspects of external validity and conflicting demands for internal validity in research designs. Methodological approaches for controlling threats to external validity and strategies to enhance external validity and generalizability of findings are discussed. CONCLUSIONS: Generalizability of findings is not assured even if internal validity of a research study is addressed effectively through design. Strict controls to ensure internal validity can compromise generalizability. Researchers can and should use a variety of strategies to address issues of external validity and enhance generalizability of findings. Enhanced external validity and assessment of generalizability of findings can facilitate more appropriate use of research findings.

Data Collection↗

Validation of the Rockall risk scoring system in upper gastrointestinal bleeding.

BACKGROUND: Several scoring systems have been developed to predict the risk of rebleeding or death in patients with upper gastrointestinal bleeding (UGIB). These risk scoring systems have not been validated in a new patient population outside the clinical context of the original study. AIMS: To assess internal and external validity of a simple risk scoring system recently developed by Rockall and coworkers. METHODS: Calibration and discrimination were assessed as measures of validity of the scoring system. Internal validity was assessed using an independent, but similar patient sample studied by Rockall and coworkers, after developing the scoring system (Rockall's validation sample). External validity was assessed using patients admitted to several hospitals in Amsterdam (Vreeburg's validation sample). Calibration was evaluated by a chi2 goodness of fit test, and discrimination was evaluated by calculating the area under the receiver operating characteristic (ROC) curve. RESULTS: Calibration indicated a poor fit in both validation samples for the prediction of rebleeding (p<0.0001, Vreeburg; p=0.007, Rockall), but a better fit for the prediction of mortality in both validation samples (p=0.2, Vreeburg; p=0.3, Rockall). The areas under the ROC curves were rather low in both validation samples for the prediction of rebleeding (0.61, Vreeburg; 0.70, Rockall), but higher for the prediction of mortality (0.73, Vreeburg; 0.81, Rockall). CONCLUSIONS: The risk scoring system developed by Rockall and coworkers is a clinically useful scoring system for stratifying patients with acute UGIB into high and low risk categories for mortality. For the prediction of rebleeding, however, the performance of this scoring system was unsatisfactory.

Adolescent↗

Pragmatic controlled clinical trials in primary care: the struggle between external and internal validity.

BACKGROUND: Controlled clinical trials of health care interventions are either explanatory or pragmatic. Explanatory trials test whether an intervention is efficacious; that is, whether it can have a beneficial effect in an ideal situation. Pragmatic trials measure effectiveness; they measure the degree of beneficial effect in real clinical practice. In pragmatic trials, a balance between external validity (generalizability of the results) and internal validity (reliability or accuracy of the results) needs to be achieved. The explanatory trial seeks to maximize the internal validity by assuring rigorous control of all variables other than the intervention. The pragmatic trial seeks to maximize external validity to ensure that the results can be generalized. However the danger of pragmatic trials is that internal validity may be overly compromised in the effort to ensure generalizability. We are conducting two pragmatic randomized controlled trials on interventions in the management of hypertension in primary care. We describe the design of the trials and the steps taken to deal with the competing demands of external and internal validity. DISCUSSION: External validity is maximized by having few exclusion criteria and by allowing flexibility in the interpretation of the intervention and in management decisions. Internal validity is maximized by decreasing contamination bias through cluster randomization, and decreasing observer and assessment bias, in these non-blinded trials, through baseline data collection prior to randomization, automating the outcomes assessment with 24 hour ambulatory blood pressure monitors, and blinding the data analysis. SUMMARY: Clinical trials conducted in community practices present investigators with difficult methodological choices related to maintaining a balance between internal validity (reliability of the results) and external validity (generalizability). The attempt to achieve methodological purity can result in clinically meaningless results, while attempting to achieve full generalizability can result in invalid and unreliable results. Achieving a creative tension between the two is crucial.

Blood Pressure↗

[Clinical assessment of schizophrenic syndrome (CASS): validity evaluation of the new diagnostic tool].

AIM: The aim of the study was an evaluation of validity measures of the CASS (Clinical Assessment of Schizophrenic Syndromes)--a new multi-purpose and multi-level clinical diagnostic instrument consisting of a diagnostic questionnaire (CASS-D) allowing to analyze a diagnosis of schizophrenia according to DSM-IV and ICD-10 criteria as well as of three rating scales designed for description and intensity evaluation of schizophrenic syndromes on the global (CASS-G), dimensional (CASS-P, a profile of 13 basic dimensions) or symptomatological (CASS-S, a set of 31 symptoms) level. SUBJECTS: 194 inpatients consecutively admitted to the Department within approximately 6 months were assessed twice (at the start and end of their hospitalization) by 12 trained diagnosticians. METHOD: Several measures of validity were analyzed. Results obtained by means of CASS were compared with results of the SANS/SAPS, BPRS, and PANSS as reference rating scales (diagnostic validity). Characteristics of frequency, intensity, dynamics and specificity of scale items were used to analyze content validity. Factorial structure of CASS scales was applied as a measure of construct validity. FINDINGS: Diagnostic validity of the new instrument seems to be confirmed by its very high correlation coefficients with rating scales recognized as international standards: BPRS, PANSS, and SANS/SAPS. Reasonable characteristics of frequency, intensity, dynamics and specificity of individual items (dimensions, symptoms) and sum scores of CASS scales and relationships between their values strongly suggest their content validity. Both sum (CASS-P, CASS-S) and global (CASS-G) scores of scales under study revealed some specificity--they had significantly higher values in patients with schizophrenia than in patients with other diagnoses. It allows to distinguish in a schizophrenic syndrome described by CASS components which are specific and not specific for this disorder. The latter have been left also in the final version for their practical and clinical importance. Construct validity of the CASS was studied separately for different scales (CASS-P, CASS-S) and different groups (all or only schizophrenic patients) by means of several factor analyses, and performed along identical statistical procedure (principal component method of extraction with criterion eigenvalue > 1, followed by Equamax rotation). Resulting solutions could be interpreted reasonably and consistently with contemporary attempts to find adequate factorial models of intrinsic structure of schizophrenic syndrome. Thus they support confidence for constructive aspect of the CASS validity. CONCLUSIONS: Ultimately, results obtained in the study suggest that the CASS may be considered as an instrument with some promising indices of diagnostic, content and construct validity, which may be potentially useful for clinical and research purposes.

Humans↗

[OECD ist accepting test guidelines for validated in vitro toxicity tests in 1996]

Since 1990 in Europe a scientific concept for the validation of in vitro toxicity tests has been developed to facilitate regulatory acceptance of the new methods at the international level. ERGATT and ECVAM have promoted the concept of validation on two workshops in 1990 and 1994. A pre-validation stage is an essential part of this validation concept to achieve a better standardisation of in vitro tests before entering formal validation. Within the NTP in 1995/96 Federal Agencies of the USA represented by the validation center ICCVAM have accepted a validation concept, which basically agrees with the essentials of ECVAM"s European validation concept. Subsequently, in January of 1996 the major industrial nations have at the OECD level agreed to the European/US validation concept. This will at the international level allow mutual acceptance of data obtained with in vitro toxicity tests rather than with animal tests. It seems very likely that in 1996 in vitro testing for skin penetration with human skin will be the first in vitro toxicity test accepted by the OECD. Several in vitro tests for local irritancy testing will follow in 1997/98, since they are currently undergoing validation according to OECD criteria.

Journal Article↗

Establishing the internal and external validity of experimental studies.

The information needed to determine the internal and external validity of an experimental study is discussed. Internal validity is the degree to which a study establishes the cause-and-effect relationship between the treatment and the observed outcome. Establishing the internal validity of a study is based on a logical process. For a research report, the logical framework is provided by the report's structure. The methods section describes what procedures were followed to minimize threats to internal validity, the results section reports the relevant data, and the discussion section assesses the influence of bias. Eight threats to internal validity have been defined: history, maturation, testing, instrumentation, regression, selection, experimental mortality, and an interaction of threats. A cognitive map may be used to guide investigators when addressing validity in a research report. The map is based on the premise that information in the report evolves from one section to the next to provide a complete logical description of each internal-validity problem. The map addresses experimental mortality, randomization, blinding, placebo effects, and adherence to the study protocol. Threats to internal validity may be a source of extraneous variance when the findings are not significant. External validity is addressed by delineating inclusion and exclusion criteria, describing subjects in terms of relevant variables, and assessing generalizability. By using a cognitive map, investigators reporting an experimental study can systematically address internal and external validity so that the effects of the treatment are accurately portrayed and generalization of the findings is appropriate.

Double-Blind Method↗

Validity of work-related assessments.

Insufficient evidence of the validity of work-related assessments is frequently reported as a major concern in occupational rehabilitation. Despite this concern, and the continuing development of new and old assessments, no comprehensive evaluation of the evidence has been conducted. OBJECTIVES: The purpose of this study was to first determine the extent and quality of available evidence for the validity of work-related assessments, and then where sufficient evidence was available, determine the level of validity. STUDY DESIGN: This study examined available literature and sources in order to review the extent to which validity has been established for 28 work-related assessments. RESULTS: The levels of evidence and validity are presented for each assessment. Most work-related assessments have limited evidence of validity. Of those that had adequate evidence, validity ranged from poor to good. There was no instrument that demonstrated moderate to good validity in all areas. Very few work-related assessments were able to demonstrate adequate validity in more than one area, or with more than one study, even when contributory evidence was included. CONCLUSION: With this study clinicians will be able to examine their options with regard to the validity of the assessments they choose to use.

Journal Article↗

Development and validation of the Observation List for early signs of Dementia (OLD).

OBJECTIVE: Development and validation of a short Observation List of possible early signs of Dementia (OLD) for use in general practice. DESIGN: Stepwise development using reviews of publications and expert consensus. Field study for evaluation of reliability. Validation study (interviews, family forms) using existing valid and reliable measures. Use of data reduction techniques to construct a short version. Setting of field study Twenty-two GPs in 19 Dutch practices. PARTICIPANTS: The first two patients seen on 15 working days (n = 470) were observed. Inclusion: age > 75, without a known diagnosis of dementia. Exclusion: psychiatric treatment, severe depression, acute illness with confusion. Division of patients into three groups with no, intermediate, and the most signs (total of interviewed patients, n = 60; family forms, n = 39). Outcome measures Reliability (Cronbach's alpha and factor-analysis). Convergent validity using the Cognitive Screening Test (CST), the Word Learning Test (WLT, total and retention), the Informant Questionnaire on Cognitive Decline in the Elderly (IQCODE), the Groningen Activities Restriction Scale (GARS), and an IADL scale. Discriminant validity using the geriatric depression scale (GDS). Construct validity using a Principal Component Analysis (PRINCALS). Incremental validity using the intuitive opinion of the GP (McNemar test). RESULTS: Reliability in the total group 0.88, first factor explained variance 42.5%. Convergent validity (two-way ANOVA) results: CST (p = 0.00), WLT-total (p = 0.001), WLT retention (p = 0.00), IQCODE (p = 0.09). No statistically significant differences for GARS and IADL. GDS (p = 0.30) not different. PRINCALS first factor explained 48% of variance. The OLD added to the GP opinion (McNemar p = 0.00). Reliability short version 0.89 (interviewed group), 0.86 (total group). CONCLUSIONS: The OLD is a valid and reliable method to detect early signs of dementia in general practice that can indicate when it may be useful to employ existing screening instruments.

Aged↗