Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

[The adaptation and validation of the food security scale in a community of Caracas, Venezuela].

This paper describes the process of modifying and validating a hunger index developed in the United States by Wehler et al (1992). It is part of a research whose main objective is to develop and validate a simple method that measures both quantitative (food sufficiency) and qualitative (female self-perception) dimensions of household food insecurity. In a pilot study, the original instrument was modified from a 2 point 8 item to a 4 point 12 item scale. Precision measured with Alpha Chronbach's coefficient was high (0.871) suggesting consistency in the scale's items. The instrument was applied to a sample of 238 poor and very poor households in a peri-urban barrio of Caracas. To determine overall internal validity of the scale the relationship between possible economic, social and behavioral determinants and food security level measured with the scale, was analyzed. Construct validity of the scale was established with factor and principal components analysis. Finally, with multiple regression analysis evidence is presented for overall validity of the scale. Four determinants: predictors of food sufficiency score, monthly income per capita, social class, and number of children in the household predict in the expected direction self-perceived food security level (R2 = 0.343). Results suggest that this instrument, together with an abbreviated measure of food sufficiency, based on strategic foods may be a valid, precise, and simple method for identifying and monitoring households that suffer from some degree of food insecurity in poor urban communities.

Adult↗

The validity of general health questionnaires, GHQ-12 and GHQ-28, in mental health studies of working people.

The purpose of the study was to determine such cut-off points in the scores of General Health Questionnaires (GHQ-12 and GHQ-28) that allow for optimal identification of people with mental health disorders in the Polish working population attending primary health care settings. The groups under the study covered 419 and 392 patients for GHQ-12, and GHQ-28, respectively. In the GHQ-12 group, 90 and in the GHQ-28 group, 80 subjects filled in the questionnaires and agreed to participate in the second stage of the study--a psychiatric interview. The criterion validity of the GHQs was a mental health diagnosis, based on the Munich version of Composite International Diagnostic Interview. The complete computerized version of interview, covering all diagnostic sections, has been adopted. In the mental health diagnosis only disorders, which currently troubled patients were taken into consideration and disorders which created problems in the distant past were excluded. In the group covered by GHQ-12 examination, 55.6% of persons had at least one type of mental disorder diagnosed, based on the criteria of both Diagnostic and Statistical Manual for Mental Disorders (DSM-IV) and the International Classification of Diseases (ICD-10). In the GHQ-28 group, the percentage of persons with mental disorders was 47.5%. After excluding patients with nicotine dependence disorder only, the frequency of mental health problems decreased to 45.5% and 33.8%, respectively. The proposed cut-off points, 2/3 points for GHQ-12 and 5/6 points for GHQ-28, were established at the level of the highest possible sensitivity and specificity not lower than 75%. These principals have been accepted for a practical reason, as the acceptance of the lower level of specificity forces medical practitioners to devote too much time to practically healthy people. At the above mentioned cut-off points for GHQ-12 sensitivity is 64% and specificity--79%, while for GHQ-28 the values are 59% and 75%, respectively. These validity coefficients were calculated from distributions of groups, from which persons with nicotine dependence as the only disorder were excluded. Incorporation of these people in the whole sample reduced the questionnaires' validity. Modification of responses scoring from the standard one--GHQ to CGHQ has not improved the validity of questionnaires. Lower validity coefficients of GHQ-28, in comparison to GHQ-12 validity are the effect of greater influence of somatic disease on the results acquired in this scale version of the questionnaire.

Adult↗

[Validation of a clinical protocol for the detection of dementia in the population].

AIMS: To analyse the validity of a set of neuropsychological and functional tests, and to study their value in detecting and diagnosing dementia through a pilot study. PATIENTS AND METHODS: A total of 131 subjects (101 controls and 30 with dementia) were evaluated using a comprehensive neuropsychological and functional battery. Validity analyses were conducted using ROC curves in accordance with the definitions of diagnostic test validation. Finally, a discriminant analysis was performed with the tests that showed greater diagnostic validity in the study of the ROC curves. RESULTS: The case and control groups were not significantly different as regards age, sex and level of schooling. The ROC curves analyses showed the following to be the tests with the highest diagnostic validity: the MMSE, delayed recall of a short story, delayed recall of six pictures, the Spanish version of the S IQCODE (shortened) and Pfeffer s FAQ. The discriminant analysis evidenced the fact that the joint utilisation of all the foregoing tests, except delayed recall of six pictures, classified 96.55% of our sample correctly. CONCLUSIONS: By combining direct cognitive evaluation of the subject and functional performance evaluated by a trustworthy informer, the vast majority of participants in the pilot study were correctly classified into patients with and without dementia. The high diagnostic validity of four relatively short tests lends support to their use in broader clinical or population studies.

Aged↗

A practical approach to quality improvement: the experience of the RNZCGP practice standards validation field trial.

AIM: This paper describes the development, implementation and validation of general practice standards, supported by a continuous quality improvement (CQI) process that teaches practice teams how to work together to identify and enhance the quality of care they provide. METHODS: Practice standards were developed through consensus by key stakeholders in general practice, pre-tested in four practices, and refined and piloted in 20 practices throughout New Zealand during 1999. A further field trial was undertaken to validate the standards and test the process of practice assessment. During 2000-2001, 74 practices volunteered to be assessed against the standards. Sixty one general practitioners, practice nurses and practice managers, nominated from independent practitioner associations (IPAs) or primary care organisations (PCOs), were trained to undertake the assessments. RESULTS: On five of 13 variables, no statistically significant differences at the 0.05 level were identified between the practices in the field trial and a random sample of practices studied by Kljakovic. The Royal New Zealand College of General Practitioners (RNZCGP) standards were found to have excellent face validity and content validity, and good construct validity. Internal consistency was fair. Lessons from the evaluation have informed an improved version of the practice assessment tool. CONCLUSIONS: The validation field trial provided the RNZCGP with a framework and tool for an accreditation process based on the principles of CQI. The tool offers patients and other stakeholders a credible measure of quality and safety at the practice level through a process bridging quality control and quality improvement.

Accreditation↗

[Translation and validation of a French version of the Young Mania Rating Scale (YMRS)].

Both the Young Mania Rating Scale (YMRS) and the Mania Assessment Scale (MAS) have been widely used during the last decade for the evaluation of severity of mania in clinical trials. For both scales good inter-rater reliability, validity and sensitivity to change have been reported. The French version of the MAS has been validated. To our know-ledge, the YMRS has not yet been translated into French and validated. The main objective of the present study was to validate a French version of the YMRS and to test its use in manic patients entering a study on the effectiveness of valproic acid and olanzapine combination. After translating the items in French, we tested this version of the YMRS on two samples of psychiatric patients recruited in a ward of adult inpatients (18 to 65 Years old) at the Department of Psychiatry, Geneva University Hospital. The first sample included 18 (hypo) manic inpatients (10 males, 8 females). Mean age was 37.0 (standard deviation 10.1). Interviews were video taped and assessed by three different judges on both scales (YMRS and MAS). The second sample included 20 inpatients (5 males, 15 females) who provided written informed consent to enter a study on the association of valproic acid and olanzapine in the treatment of mania. Mean age was 40.0 (standard deviation 11.3). Patients were followed over four weeks and assessed on both scales (YMRS and MAS) every seven days (day 0, 7, 14, 21 and 28). On day 7, patients were assessed during a joint interview by two of three judges who independently administered both scales in permuted order. On days 0, 14, 21 and 28, patients were evaluated by one of the same three raters. Inter-rater reliability was assessed by comparing item scores and total scores assigned by different judges with intra-class correlation coefficient ICC (2,1). Three judges were considered for patients in sample 1. Two judges were considered for patients in sample 2 (day 7 assessment). Concurrent validity with the MAS was analysed in sample 2 on days 0, 7, 14, 21 and 28 using Spearman rank-order correlation coefficient. Sensitivity to change was assessed in sample 2 by comparing total score at inclusion and at last observation using Wilcoxon signed ranks test. For both the MAS and YMRS, intraindividual change was calculated as the difference between total scores at inclusion and discharge (last observation carried forward approach). The relationship between changes on the two scales was analysed through Spearman correlation coefficient. Significance level was set to 0.05 for each test. Ranges of YMRS total scores were 2 to 32 in sample 1 and 1 to 28 in sample 2, indicating symptom severity from euthymic to moderately manic. Inter-rater reliability was very good for the total scores in both samples, both for the MAS and the YMRS (ICC>0.89). When considering YMRS individual items, correlation coefficient varied from 0.61 to 0.96 in the first sample. In the second sample, 9 of 11 items displayed values above 0.63. The remaining two items, increased motor activity and energy and Language-thought disorder, presented modest inter-rater reliability (ICC=0.54 and 0.50 respectively). This was largely attributable to a single patient, who was perceived very differently by the two judges (scores 0-2 for increased motor activity and energy; 1-4 for Language-thought disorder). When this patient was excluded, intra-class correlation coefficients were above 0.69 for both items. Overall, inter-rater reliability of the YMRS items was in the same range as for the MAS items (0.61-0.96 vs 0.61-0.93 in sample 1; 0.50-0.93 vs 0.54-0.83 in sample 2). Correlation between the two instruments was very high and statistically significant at each weekly assessment (rs>0.91, p<0.001) except for day 21 which displayed a somewhat lower correlation (rs=0.75, p<0.01). This latter result was attributed to a reduced spread of values and number of patients on day 21. YMRS and MAS total scores as a function of time in patients receiving combined treatment with olanzapine and valproic acid (sample 2) show that for both at for both scales, total scores significantly decreased from day 0 to last observation (Wilcoxon signed ranks test, p<0.001), with median decrease of 18 points both on the YMRS (range 9-32) and MAS (range 10-33). Median relative decrease was 67% for the YMRS and 69% for the MAS. When analysing the relationship between intraindividual changes on the YMRS and MAS, highly significant correlation was observed (Spearman rs=0.93, p<0.001), showing that the two scales were virtually interchangeable in assessing treatment efficacy. In conclusion, the YMRS is a simple and easy-to-use instrument for measuring severity of manic symptoms The newly translated French version was satisfactory in terms of inter-rater reliability, concurrent validity with the MAS, and sensitivity to change in patients receiving treatment for manic symptoms. This should allow its future use for international comparison studies.

Adolescent↗

[Development and validation of a modified quality of life questionnaire for patients treated by antireflux surgery].

INTRODUCTION: Past decade witnessed a growing interest in scientific medical publications on health related quality of life (HRQoL), which has yielded an increasing number of generic and disease-specific instruments. To date, several studies have evaluated the impact of GERD on HRQoL. AIMS: To develop a QoL questionnaire for patients with GERD underwent laparoscopic fundoplication (LF). This questionnaire was developed to be more comprehensive than existing measures. MATERIALS AND METHODS: We undertook a retrospective analysis of 116 patients underwent laparoscopic fundoplication for GERD between 1994 and 2002 in the 1st Department of Surgery, Semmelweis University. These patients--included 55 men and 61 women, with mean age of 46 years (14-77)--were used in the psychometric evaluation. Our questionnaire was developed using internationally accepted, valid QoL instruments and scales. RESULTS: Internal-consistency reliability was high (alpha value overall 0.95; dimensions 0.74-0.96). Using convergent and divergent validity, construct validity was evaluated by examining Pearson correlation coefficients between items and scales. Construct validity was demonstrated based on observed correlations. Known-groups validity of questionnaire is proved to. CONCLUSIONS: Our questionnaire is a short and user-friendly instrument. It has excellent reliability, construct and known-groups validity.

Adolescent↗

Reliability, validity and responsiveness of instruments to assess disabilities in personal care in patients with rheumatic disorders. A systematic review.

OBJECTIVES: The first aim was to make an inventory of available instruments and questionnaires for the assessment of disabilities in personal care in patients with rheumatic disorders. The second aim was to investigate which of these instruments have acceptable, methodological quality with regard to reliability, validity and responsiveness. The third aim was to investigate the assumption that convergent validity results in stronger correlations when validated against a more similar construct. METHODS: A computer-aided literature search (1982-2001) in several databases was performed to identify studies focusing on the clinimetric properties of instruments to assess impairments in function in patients with rheumatic disorders. Data were extracted in a standardised way and compared to a priori defined criteria. RESULTS: In total, 19 measurement instruments were included. Five out of these 19 were found to have acceptable reliability, while 12 had acceptable validity. Only three questionnaires met both criteria. Results concerning the responsiveness of these three questionnaires were conflicting. No difference was found in the strength of correlation between validation against the most similar construct versus validation against the least similar construct. CONCLUSION: It is concluded that the Arthritis Impact Measurement Scale (AIMS) is the most suitable instrument for the assessment of disabilities in personal care.

Activities of Daily Living↗

[The appraisal of reliability and validity of subjective workload assessment technique and NASA-task load index].

OBJECTIVE: To test the reliability and validity of two mental workload assessment scales, i.e. subjective workload assessment technique (SWAT) and NASA task load index (NASA-TLX). METHODS: One thousand two hundred and sixty-eight mental workers were sampled from various kinds of occupations, such as scientific research, education, administration and medicine, etc, with randomized cluster sampling. The re-test reliability, split-half reliability, Cronbach's alpha coefficient and correlation coefficients between item score and total score were adopted to test the reliability. The test of validity included structure validity. RESULTS: The re-test reliability coefficients of these two scales and their items were ranged from 0.516 to 0.753 (P < 0.01), indicating the two scales had good re-test reliability; the split-half reliability of SWAT was 0.645, and its Cronbach's alpha coefficient was more than 0.80, all the correlation coefficients between its items score and total score were more than 0.70; as for NASA-TLX, both the split-half reliability and Cronbach's alpha coefficient were more than 0.80, the correlation coefficients between its items score and total score were all more than 0.60 (P < 0.01) except the item of performance. Both scales had good inner consistency. The Pearson correlation coefficient between the two scales was 0.492 (P < 0.01), implying the results of the two scales had good consistency. Factor analysis showed that the two scales had good structure validity. CONCLUSION: Both SWAT and NASA-TLX have good reliability and validity and may be used as a valid tool to assess mental workload in China after being revised properly.

Adolescent↗

Measuring functional disability in early rheumatoid arthritis: the validity, reliability and responsiveness of the Recent-Onset Arthritis Disability (ROAD) index.

OBJECTIVE: Disability has been identified as a core outcome measure in rheumatoid arthritis (RA). The aim of this study was to test the Recent-Onset Arthritis Disability (ROAD) questionnaire for validity, reliability and responsiveness in Italian patients with early RA. METHODS: The psychometric properties of ROAD were tested in 159 patients with early RA, mean age 54.7 (+/- 8.8), 74.3% women, mean disease duration 14.5 months (+/- 1.9 months). All completed the ROAD, the Medical Outcomes Study SF-36 Health Survey (SF-36), the Health Assessment Questionnaire (HAQ) and the patient global assessment (PGA) of functional disability twice, in order to test for validity and responsiveness. Of the 159 patients who completed the health status instruments on two occasions, 121 were included in the responsiveness analyses. The test-retest reliability of the ROAD questionnaire was calculated using intraclass correlation coefficients (ICCs) and the Bland and Altman method on 77 patients who completed the questionnaire twice over an interval of one week. Construct validity was assessed using Spearman's correlations, while responsiveness was evaluated by 3 different methods: (1) effect size (the mean difference between the baseline scores and thefollow-up scores divided by the standard deviation of the baseline scores); (2) standardized response mean (the mean change in scores divided by the standard deviation of the change in scores); (3) receiver operating characteristics (ROC) curve analysis. RESULTS: ROAD fulfilled the established criteria for validity, reliability and responsiveness. In comparison with the SF-36, the expected correlations were found when comparing items measuring similar constructs, thus supporting the convergent construct validity. Significant correlations were seen between ROAD scores and HAQ scores (rho = 0.372), SF-36 physical component summary (PCS) (rho = -0.413), PGA functional disability (rho = 0.417), pain (rho = 0.639), Ritchie index (rho = 0.357), number of swollen joints (rho = 0.387), patient and physician assessment of disease activity (rho = 0.467 and 0.323, respectively), and Disease Activity Score (rho = 0.476). Test-retest reliability was satisfactory, with ICCs of 0.927 (upper extremity function), 0.892 (lower extremity function), and 0.851 (activity of daily living/work). Bland-Altman plots confirmed this finding. The results of responsiveness analysis indicate that the ROAD subscales were slightly more sensitive to perceived change in functional disability than those of HAQ, SF-36 PCS, and PGA offunctional disability. CONCLUSION: Our data suggest that the ROAD index is a reliable, valid and responsive tool for measuring physical functioning in patients with early RA, and is suitable for use in clinical trials and daily clinical practice. Its generalizability and utility for assessing aggressive treatment and functional outcomes must now be evaluated in broader settings.

Arthritis, Rheumatoid↗

The validity of A new practical quality of life measure in patients on renal replacement therapy.

OBJECTIVE: A new quality of life measure, apart of the National Health and Welfare 2003 survey, is a promising tool for outcome evaluation of clinical practice due to its brevity, validity, reliability, and providing easy interpretation against general population norm-based scores. The measure consisting of 9-items, and so called 9-item Thai Health status Assessment Instrument (9-THAI) was used to assess its validity and reliability in patients on renal replacement therapy (RRT). MATERIAL AND METHOD: Three hundred and two patients on RRT who visited Srinagarind Hospital from March to May 2005 were studied Convergent and divergent validity were assessed using SF-36 as the concurrent measure. Concurrent validity was also assessed using hematocrit level and hospitalization history in the last year as concurrent clinical measures. Test-retest reliability was studied by repeated measure within one 1 month. Responsiveness of 9-THAI was studied in patients who reported health improvement. RESULTS: Results of correlations between 9-THAI and SF-36 domains were as hypothesized 9-THAI scores were significantly correlated with hematocrit level and hospitalization history. The results confirmed the validity of 9-THAI for use as a quality of life measure. Intraclass correlation coefficients of 9-THAI scores in stable patients were satisfactory. Among patients on RRT who reported overall health improvement, 9-THAI scores significantly increased, thus adding further evidence of the responsiveness of 9-THAI. CONCLUSION: The 9-THAI is a valid and reliable generic health status measure that can be used as an ideal core in a battery of quality of life measures in clinical practice for patients on RRT.

Female↗

Validation procedures for the Hamilton Thorne Integrated Visual Optical System sperm and cell analyzer.

The Hamilton Thorne Integrated Visual Optical System (IVOS) analyzer is a combined internal optical and computer system widely applied to laboratory animals, such as rat, rabbit, and canine, sperm in reproductive toxicology. It measures sperm motility, progressive motility, velocity, motion parameters, and concentration. The system is specially designed to facilitate on-site validation. Digital encoding allows image storage with absolute replay fidelity for later reexamination, enabling on-site validation complying with Good Laboratory Practices. A NIST-certified scale provides the length standard on which validation is based. Sperm typically move in wavy tracks, which must be characterized and validated. Playback of exact sperm tracks allows visual validation of sperm head position, given as Cartesian coordinates, and visual determination of motility from the playback screen. A cursor checks coordinate values, and manual confirmation of sperm motion parameters may be performed directly from the validated coordinates. Agreement to within 0.2% is obtained between manual and IVOS computation. Concentration is determined using a specific DNA stain that enables discrimination between sperm and somatic cell nuclei. The IVOS concentration of sperm nuclei in homogenized rat testis and cauda epididymis has been determined to be within 5% of manual hemacytometer counts of the same homogenate.

Humans↗

Development and validation of the SDDS-PC screen for multiple mental disorders in primary care.

OBJECTIVE: To develop, validate, and cross-validate a patient-completed screen for multiple mental disorders in primary care. DESIGN: Comparison of a patient self-report screen with an independent diagnostic assessment by mental health professionals using the Structured Clinical Interview for DSM-III-R diagnoses as criterion standard. SETTING: Three Rhode Island family practices and a South Carolina family medicine residency. SUBJECTS: In the initial validation study, 937 patients in Rhode Island were screened; 388 were interviewed. In the cross-validation study, 775 patients were screened in Rhode Island and South Carolina, and 257 were interviewed. SCREEN ITEMS: Sixty-two questions pertaining to nine mental disorders and suicidal ideation. RESULTS: A 16-item screen remained after analysis of item and scale performance. Sensitivity, specificity, and positive predictive value, respectively, were calculated for the following scales: alcohol abuse or dependence (62%, 98%, and 54%), generalized anxiety disorder (90%, 54%, and 5%), major depression (90%, 77%, and 40%), obsessive-compulsive disorder (65%, 73%, and 5%), panic disorder (78%, 80%, and 21%), and suicidal ideation (43%, 91%, and 51%). Replication in a new sample showed attenuated but acceptable operating characteristics for cross-validation. CONCLUSIONS: The Symptom-Driven Diagnostic System for Primary Care screen assesses multiple mental disorders that are common to primary care. It serves as a sensitive, valid, and patient-friendly first step in a new approach to recognizing and managing mental disorders in primary care. Finally, it aids the primary care clinician in selecting an appropriate diagnostic interview module for the disease for which the patient screened positive.

Adult↗

Screening for cognitive impairment in older individuals. Validation study of a computer-based test.

OBJECTIVE: This study examined the validity of a computer-based cognitive test that was recently designed to screen the elderly for cognitive impairment. DESIGN: Criterion-related validity was examined by comparing test scores of impaired patients and normal control subjects. Construct-related validity was computed through correlations between computer-based subtests and related conventional neuropsychological subtests. SETTING: University center for memory disorders. PARTICIPANTS: Fifty-two patients with mild cognitive impairment by strict clinical criteria and 50 unimpaired, age- and education-matched control subjects. Control subjects were rigorously screened by neurological, neuropsychological, imaging, and electrophysiological criteria to identify and exclude individuals with occult abnormalities. RESULTS: Using a cut-off total score of 126, this computer-based instrument had a sensitivity of 0.83 and a specificity of 0.96. Using a prevalence estimate of 10%, predictive values, positive and negative, were 0.70 and 0.96, respectively. Computer-based subtests correlated significantly with conventional neuropsychological tests measuring similar cognitive domains. Thirteen (17.8%) of 73 volunteers with normal medical histories were excluded from the control group, with unsuspected abnormalities on standard neuropsychological tests, electroencephalograms, or magnetic resonance imaging scans. CONCLUSIONS: Computer-based testing is a valid screening methodology for the detection of mild cognitive impairment in the elderly, although this particular test has important limitations. Broader applications of computer-based testing will require extensive population-based validation. Future studies should recognize that normal control subjects without a history of disease who are typically used in validation studies may have a high incidence of unsuspected abnormalities on neurodiagnostic studies.

Aged↗

Reliability and validity in binary ratings: areas of common misunderstanding in diagnosis and symptom ratings.

Confusion may exist between the reliability of a binary rating (for example, schizophrenia versus not-schizophrenia) and its implications for validity. High reliability does not guarantee validity, but paradoxically, low reliability does not imply poor validity in all contexts. Changes in the base rate or in experimental design may indicate high validity even when the reliability was thought to be low. Attempts to improve the psychiatric nomenclature by increasing only reliability run the risk of the "attenuation paradox" where further increases in reliability will make the ratings less valid. Finally, the assumption of random error in making diagnoses does not always hold, so that statistical analyses must be adjusted accordingly. New statistical methods are needed to index only false-positive or false-negative rates in order to quantify the error that will reduce some validity coefficients.

Bipolar Disorder↗

A simulation study of cross-validation for selecting an optimal cutpoint in univariate survival analysis.

Continuous measurements are often dichotomized for classification of subjects. This paper evaluates two procedures for determining a best cutpoint for a continuous prognostic factor with right censored outcome data. One procedure selects the cutpoint that minimizes the significance level of a logrank test with comparison of the two groups defined by the cutpoint. This procedure adjusts the significance level for maximal selection. The other procedure uses a cross-validation approach. The latter easily extends to accommodate multiple other prognostic factors. We compare the methods in terms of statistical power and bias in estimation of the true relative risk associated with the prognostic factor. Both procedures produce approximately the correct type I error rate. Use of a maximally selected cutpoint without adjustment of the significance level, however, results in a substantially elevated type I error rate. The cross-validation procedure unbiasedly estimated the relative risk under the null hypothesis while the procedure based on the maximally selected test resulted in an upward bias. When the relative risk for the two groups defined by the covariate and true changepoint was small, the cross-validation procedure provided greater power than the maximally selected test. The cross-validation based estimate of relative risk was unbiased while the procedure based on the maximally selected test produced a biased estimate. As the true relative risk increased, the power of the maximally selected test was about 10 per cent greater than the power obtained using cross-validation. The maximally selected test overestimated the relative risk by about 10 per cent. The cross-validation procedure produced at most 5 per cent underestimation of the true relative risk. Finally, we report the effect of dichotomizing a continuous non-linear relationship between covariate and risk. We compare using a linear proportional hazard model to using models based on optimally selected cutpoints. Our simulation study indicates that we can have a substantial loss of statistical power when we use cutpoint models in cases where there is a continuous relationship between covariate and risk.

Humans↗

What do we mean by validating a prognostic model?

Prognostic models are used in medicine for investigating patient outcome in relation to patient and disease characteristics. Such models do not always work well in practice, so it is widely recommended that they need to be validated. The idea of validating a prognostic model is generally taken to mean establishing that it works satisfactorily for patients other than those from whose data it was derived. In this paper we examine what is meant by validation and review why it is necessary. We consider how to validate a model and suggest that it is desirable to consider two rather different aspects - statistical and clinical validity - and examine some general approaches to validation. We illustrate the issues using several case studies.

Accidental Falls↗

Reliability and validity of an observer-rated disfigurement scale for head and neck cancer patients.

BACKGROUND: Facial disfigurement is considered to be one of the most distressing aspects of head and neck cancer and its treatment, but it has been the focus of little systematic study. Existing studies have yielded conflicting results about the psychosocial impact of disfigurement. No studies to date have examined disfigurement using a valid and reliable observer-rated measure. The purpose of the current study was to examine the validity (convergent and discriminant) and the inter-rater reliability of a novel nine-point observer-rated disfigurement scale. METHODS: The sample consisted of 74 ambulatory head and neck cancer patients more than 6 months post treatment. Ratings of disfigurement were assigned independently by surgical and nonsurgical raters. Validity was assessed by comparing the association between disfigurement ratings and sociodemographic and illness treatment variables. Reliability was assessed by examining the concordance between the surgical and nonsurgical ratings. RESULTS: Disfigurement ratings were not associated with several sociodemographic variables, supporting the discriminant validity of the scale. Disfigurement was significantly related to a diagnosis of oral cancer, a history of adjunctive radiation, the type of surgical procedure performed, the degree of physical dysfunction, and the presence of postoperative complications. Observer ratings of disfigurement were significantly related to patient ratings of disfigurement. These findings support the convergent validity of the disfigurement scale. Inter-rater reliability of the scale was high (intraclass correlation coefficient =.91). CONCLUSION: The study provides preliminary evidence for the validity and inter-rater reliability of a novel nine point observer-rated disfigurement scale that may be useful in evaluating the impact of disfigurement on quality of life in head and neck cancer.

Adaptation, Psychological↗

Relationships among clinical and validity scales of the Basic Personality Inventory.

In interpreting the results of a self-report inventory it is important to evaluate the extent to which stylistic distortion may have been operative. This task is complicated because validity measures frequently are confounded with content measures. In order to evaluate the potential utility of the validity measures for the Basic Personality Inventory (BPI) the relationships between validity measures and content scales were evaluated in a sample of 71 inmates. While some validity indices had significant correlations with content scales, other validity indices were relatively independent of the content scales. Recommendations are provided for using the BPI validity scales.

Adolescent↗