Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Derivation and validation of a pulmonary tuberculosis prediction model.

OBJECTIVE: To describe the derivation and validation of a pulmonary tuberculosis (TB) prediction model that would enable early discontinuation of unnecessary respiratory isolation. DESIGN: Patients placed in isolation for suspected pulmonary TB were studied retrospectively (derivation cohort) and prospectively (validation cohort). Independent predictors of pulmonary TB in the derivation cohort (January 1992-March 1994) were identified by retrospective analysis. Predictors in the model were assigned weights on the basis of the results of the multivariate analysis in order to quantitate the risk of TB in an individual patient. The prospective validation consisted of application of the model to patients placed in isolation during the period April 1994 to June 1995. The predictability of the model in the derivation and validation cohorts was evaluated using receiver operating characteristics (ROC), curve analysis, and calculation of the area under the ROC curve (AUC). SETTING: A university-affiliated, urban, public hospital with a large population of prison inmates and patients with human immunodeficiency virus infection. INTERVENTIONS: Prospective application of the prediction model to patients placed in isolation during the validation period. RESULTS: Four factors were found to be independent predictors of pulmonary TB among 296 isolation episodes in the derivation cohort; positive acid-fast sputum smear (odds ratio [OR], 5.8; 95% confidence interval [CI95], 3.0-11.0; weight = 3 points), localized chest radiograph findings (OR, 2.5; CI95, 1.3-4.9; weight = 2 points), residence in a correctional facility (OR, 2.3; CI95, 1.2-4.4; weight = 2 points), and history of weight loss (OR, 1.8; CI95, 1.0-3.2; weight = 1 points). Infection control practitioners applied the model prospectively to 220 isolation episodes. The mean (+/-SE) AUCs of the ROC curve for the derivation and validation cohorts were not significantly different (.86 +/- .04 vs .86 +/- .07; P = .90). There was a significant decline in the mean duration of isolation from the onset of an automatic TB isolation policy in August 1992 to the end of the study (P = .045 by analysis of variance). CONCLUSIONS: A pulmonary TB prediction model was derived and validated prospectively in a hospital with a moderately high prevalence of TB. The model quantitated the risk of TB in an individual patient and aided infection control practitioners and primary-care physicians in their decisions to discontinue isolation during the validation period. Utilization of the model was responsible, in part, for a decrease in the mean duration of isolation during the study period. Although the model may not have general applicability due to the uniqueness of the patient population studied, this study illustrates how prediction models can be developed and used effectively to deal with a clinical problem.

Adult↗

Validity of cephalometric landmarks. An experimental study on human skulls.

Cephalometric landmark validity (the difference between the estimated landmark and the true landmark) has surprisingly not previously been comprehensively evaluated, and no previous study has examined the validity of cephalometric angles and distances. The aim of this study was to investigate the validity of 15 commonly used skeletal and dental cephalometric landmarks, and the subsequent effects on 17 angles and distances. Small steel balls were glued on to 30 Chinese dry skulls to represent the true anatomical landmarks. The skulls were mounted in a purpose-designed skull holder and two cephalograms recorded of each skull, one with and one without the steel balls on the landmarks. Validity was expressed as the difference in the measurements between the assessments made with and without the steel ball markers. Measurements were made relative to X and Y co-ordinates which were constructed from reference points (steel balls) glued intracranially to the skulls. Seven out of the 10 skeletal landmarks and all five dental landmarks, were found to be non-valid along the X or the Y axes (P < 0.05). The standard deviations of the validity errors were large, being 1.0-2.5 mm, along at least one axis, for eight of the skeletal landmarks and three of the dental landmarks. Four of the cephalometric angles (SNA, SN/MnP, MxP/MnP, and LI/MnP) and three of the distances (N-Me, MxP-Me, and lower incisor edge to APg) were also found to be invalid (P < 0.05). The validity errors were greater for angles involving dental landmarks and for angles dependent on four landmarks compared to those dependent on three. The standard deviations of the validity errors for the skeletal angles ranged from 0.9 to 1.8 degrees, except for ANB (0.4 degrees), and for the dental angles from 3.2 to 5.8 degrees.

Adult↗

Validation of measures of food insecurity and hunger.

The most recent survey effort to determine the extent of food insecurity and hunger in the United States, the Food Security Supplement, included a series of questions to assess this complex phenomenon. The primary measure developed from this Food Security Supplement was based on measurement concepts, methods and items from two previously developed measures. This paper presents the evidence available that questionnaire-based measures, in particular the national food security measure, provide valid measurement of food insecurity and hunger for population and individual uses. The paper discusses basic ideas about measurement and criteria for establishing validity of measures and then uses these criteria to structure an examination of the research results available to establish the validity of food security measures. The results show that the construction of the national food security measure is well grounded in our understanding of food insecurity and hunger, its performance is consistent with that understanding, it is precise within usual performance standards, dependable, accurate at both group and individual levels within reasonable performance standards, and its accuracy is attributable to the well-grounded understanding. These results provide strong evidence that the Food Security Supplement provides valid measurement of food insecurity and hunger for population and individual uses. Further validation research is required for subgroups of the population, not yet studied for validation purposes, to establish validity for monitoring population changes in prevalence and to develop and validate robust and contextually sensitive measures in a variety of countries that reflect how people experience and think about food insecurity and hunger.

Female↗

The SF-36 Arthritis-Specific Health Index (ASHI): II. Tests of validity in four clinical trials.

OBJECTIVE: The SF-36 Arthritis-Specific Health Index (ASHI) was constructed to improve the responsiveness of the SF-36 Health Survey to changes in the severity of arthritis through the use of arthritis-specific scoring algorithms. This study compared the responsiveness of the ASHI and other generic scales and summary measures scored from the SF-36 in clinical trials of health outcomes for patients with arthritis. METHODS: Longitudinal data for patients (n = 835) participating in four placebo-controlled trials were analyzed. Study participants had at least a 6-month history of moderate to severe osteoarthritis or rheumatoid arthritis of the knee or hip. All had undergone a washout period of 3 to 14 days before baseline assessment to bring about a flare state in osteoarthritis or rheumatoid arthritis symptoms. Their average age was 60 years, and 72% were female. Responders and nonresponders were classified on the basis of physician assessments of changes in arthritis severity, with blinding as to treatment group; treated and untreated (placebo) groups were also compared. For the SF-36 ASHI, generic physical (PCS) and mental (MCS) component summary measures and each of eight subscales scored from the SF-36 (acute version) change scores were computed by subtracting scores before treatment from scores at 2-week follow-up. To evaluate empirical validity, analyses of variance were performed. For each measure, an F-ratio was computed for the comparison between clinically defined groups of responders and nonresponders and between groups of patients assigned to placebo versus drug therapy. Relative validity (RV) coefficients were computed for the ASHI in comparison with PCS, MCS, and the best SF-36 scale to determine which was more responsive. RESULTS: In analyses of each of the four trials and all trials combined, RV coefficients for the ASHI were higher than those for both of the generic SF-36 summary measures and for the most valid SF-36 scale (Bodily Pain), with only one exception. Across 40 tests of validity in distinguishing treated from untreated patients, the ASHI was 5% to 19% more valid than the best SF-36 scale (RV = 1.05-1.19; RV = 1.10 in all trials combined). The generic summary measures (PCS and MCS) were much less valid in these tests (RV = 0.67 and 0.27, respectively). In analyses of responders and nonresponders, RV coefficients for the ASHI ranged from 0.70 to 1.22 (RV = 1.04 in all trials combined), in comparison with the best SF-36 subscale, which was always Bodily Pain. RV coefficients were lower for PCS (RV = 0.75) and much lower than the MCS (RV = 0.18) in comparisons of treatment outcomes based on all trials combined. CONCLUSION: The ASHI appears to be more valid than the eight SF-36 scales and PCS and MCS summary measures for purposes of distinguishing between treated and untreated patients and between clinical responders and nonresponders. This study demonstrates the feasibility of improving the validity of the SF-36 through the use of arthritis-specific scoring while retaining the option of generic scoring, which makes it possible to also compare results across diseases and treatments.

Arthritis, Rheumatoid↗

Maintaining study validity in a changing clinical environment.

BACKGROUND: Nurse scientists who conduct intervention research in a variety of clinical settings find themselves facing numerous challenges posed by today's changing and sometimes complex health care environment. Maintaining study validity thus becomes a major focus of interventional research, but existing literature does not fully address challenges to study validity nor offer potential solutions. OBJECTIVES: The purposes of this paper are to 1) discuss methodologic challenges to maintaining study validity of intervention research that is conducted in a changing clinical environment, and 2) share strategies for maximizing study validity. METHODS: A recently completed intervention study is used as an example to discuss two specific areas that affected study validity, provide examples of selected threats to validity, and outline strategies used to minimize these threats. RESULTS: Careful definition of goals, thoughtful decision making, and implementation of specific strategies to maintain study validity helped increased the rigor of the research. CONCLUSIONS: Investigators conducting intervention research in changing clinical settings can reduce threats to study validity and increase design rigor by considering clinical realities (e.g., clinician-researcher role conflict) when making methodologic decisions, becoming familiar with the setting, and involving clinicians in the research.

Decision Making↗

The Pediatric Multiple Organ Dysfunction Score (P-MODS): development and validation of an objective scale to measure the severity of multiple organ dysfunction in critically ill children.

OBJECTIVE: To develop and then prospectively validate an objective scale to grade multiple organ system dysfunction in a large population of critically ill children. DESIGN: Prospective, observational cohort study. SETTING: A pediatric intensive care unit at a tertiary care pediatric teaching hospital. PATIENTS: A total of 6,456 pediatric consecutive admissions (mean age 4.62 yrs) admitted to the pediatric intensive care unit. INTERVENTIONS: a) Identification of variables that could define organ dysfunction in children; b) development of a Pediatric Multiple Organ Dysfunction Score (P-MODS); c) correlation of the score with outcome at pediatric intensive care unit discharge; d) subsequent prospective validation. MEASUREMENTS AND MAIN RESULTS: A computer system randomly separated patients into two groups: a development set to create the scoring system and a validation set to evaluate score performance and reproducibility. Survivors and nonsurvivors were compared to define variables that were significantly more abnormal in nonsurvivors. Those variables were correlated with pediatric intensive care unit mortality rate. Optimal intervals for each variable were defined on the development set, and their performance was evaluated in the validation set. Descriptors for organ dysfunction were identified in five organ systems: cardiovascular (lactic acid), respiratory (Pa(O(2))/Fi(O(2)) ratio), hepatic (bilirubin), hematologic (fibrinogen), and renal (blood urea nitrogen). A grading scale for each variable was set from 0 to 4, corresponding to mortality rates of <5% and >50%, respectively. P-MODS was calculated by summing the worst score for all variables. Overall performance of the score was evaluated by generating receiver operating characteristic curves for both study sets. The score correlated strongly and in a graded fashion with pediatric intensive care unit mortality rate. In both sets (development and validation), mortality rate was <5% when the score was 0 and >70% at the highest score. Overall mortality rate was 5.9% (development set) and 5.3% (validation set). The score showed excellent discrimination reflected in areas under the curve: 0.81 (development set) and 0.78 (validation set). CONCLUSIONS: P-MODS correlated strongly with pediatric intensive care unit mortality in both study sets and can provide an objective measure for assessing organ dysfunction in the pediatric intensive care unit. With further study and validation across many centers, it is likely that P-MODS could function as a quantitative, clinically relevant surrogate outcome measure for future therapeutic trials.

Child, Preschool↗

Self-reports by alcohol and drug abuse inpatients: factors affecting reliability and validity.

The reliability and validity of self-report data regarding substance abuse has often been questioned. To determine how best to enhance the veracity of self-report, three factors which might affect self-report veracity were examined: alcohol status at time of interview; level of cognitive functioning; and method of self-report data collection. Subjects were 234 admissions to an inpatient substance abuse treatment unit. Self-report data were collected via both personal interview on the day of admission and and questionnaire within the first week of stay. Self-reports concerned use of alcohol, cocaine, and marijuana in the days preceding admission. Test-retest reliability for the questionnaire data produced reliability coefficients of 0.88, 0.91, and 0.88, for alcohol, cocaine, and marijuana, respectively. Variation in inter-test interval had virtually no effect upon reliability coefficients. Interview data were compared to toxicologic analyses of blood and urine samples collected on admission. Overall, this comparison showed self-reports to be valid, with a 97% agreement between verbal report and laboratory data for alcohol, 93% for cocaine, and 84% for marijuana. The comparison of interview data with questionnaire responses also showed self-reports to be valid: 90% agreement for alcohol, 93% for cocaine, and 81% for marijuana. Level of cognitive function did not influence the validity of self-reports for any of the three substances. Recent consumption of alcohol also had no statistically significant effect on the validity of self-reported marijuana use, regardless of the operational form of validity tested. However, BAC-negative subjects produced a significantly greater validity coefficient for self-reported cocaine use (kappa = 0.87) than did BAC-positive patients (kappa = 0.43), when interview data were compared with toxicologic measures. A similar finding was not uncovered when interview and questionnaire data were compared. An interaction between admission alcohol status and cognitive function was uncovered for cocaine self-reports when interview data was compared with toxicologic measures. The rate of agreement for alcohol-negative subjects is quite high for both cognitively impaired and unimpaired subjects (M = 93% and M = 94%, respectively) as well as for alcohol-positive, cognitively unimpaired subjects (M = 94%), but not for alcohol-positive, cognitively impaired subjects (M = 67%). Results are discussed in terms of threats to the validity of self-report and strategies for the optimization of response accuracy.

Adult↗

The validity of Qualpacs.

The validity of the quality assessment instrument Qualpacs is discussed. The literature reviewed addresses the sensitivity, scoring and scope of the instrument, operational decisions and cost and policy implications. Specific attention is given to the reliability and validity of Qualpacs and methodological challenges of convergent validity testing. The study, which was funded by the Department of Health, England, employed a multiple triangulation research design to assess the validity of the instrument. Qualpacs was compared to two other instruments (Monitor and Senior Monitor) and to other quality of nursing care data derived from: observation of patients' activities and nurse-patient interaction, and interviews with patients and nurses on their perceptions of quality of care. The results support Qualpacs as a relatively valid instrument in that it displayed greater convergent validity than Monitor or Senior Monitor. Convergent validity was stronger when compared to Senior Monitor and the Monitor DG3 schedule than the other Monitor schedules. Comparison with the observation of nurse and patient activities and interactions supported Qualpacs for medical and surgical wards but not for elderly care wards. Congruence between Qualpacs' items and the views of patients and nurses provides evidence in support of construct validity of the instrument.

Data Collection↗

French adaptation and preliminary validation of a questionnaire to evaluate understanding of informed consent documents in phase I biomedical research.

The content of informed consent documents (ICD) is a crucial element in the process of providing information to participants in biomedical research. Clear comprehension of the information, i.e. the ability to understand its meaning and its consequences, is of utmost importance. The objective of this study was to describe the different steps in the French adaptation and preliminary validation of the Qualité de Compréhension des Formulaires d'information et de consentement (QCFic) questionnaire (http://www.lyon.inserm.fr/cic-grenoble) based on the American Quality of Informed Consent (QuIC) questionnaire. Adaptation and preliminary validation of the QuIC for use in France was composed of five principal steps: translation, scientific validation, lexical validation, edition of gold-standard answers and a pilot study. Each stage was conducted by independent groups of experts, under the coordination of the study board. Thirteen questions were added and one was suppressed. Two steps were required for the scientific validation and for lexical validation, 21 modifications were proposed. Relative to gold-standard answers, the three experts gave the same answer for 24 questions and for nine other questions, two of the three gave identical answers, which were validated by the study board. Results of a pilot study showed a global QCFic score of 88.99 (84.13-90.92) and no specific commentary was made about the content of the questions, so no more modification needed to be made. A preliminary validated French questionnaire, the QCFic, is now available to evaluate the quality of an informed consent document in phase I clinical trials. It is quick and easy to use.

Adult↗

Integrating validity theory with use of measurement instruments in clinical settings.

OBJECTIVE: To present validity concepts in a conceptual framework useful for research in clinical settings. PRINCIPAL FINDINGS: We present a three-level decision rubric for validating measurement instruments, to guide health services researchers step-by-step in gathering and evaluating validity evidence within their specific situation. We address construct precision, the capacity of an instrument to measure constructs it purports to measure and differentiate from other, unrelated constructs; quantification precision, the reliability of the instrument; and translation precision, the ability to generalize scores from an instrument across subjects from the same or similar populations. We illustrate with specific examples, such as an approach to validating a measurement instrument for veterans when prior evidence of instrument validity for this population does not exist. CONCLUSIONS: Validity should be viewed as a property of the interpretations and uses of scores from an instrument, not of the instrument itself: how scores are used and the consequences of this use are integral to validity. Our advice is to liken validation to building a court case, including discovering evidence, weighing the evidence, and recognizing when the evidence is weak and more evidence is needed.

Data Collection↗

Validation of the smoking self-efficacy survey for Taiwanese children.

PURPOSE: To determine the psychometric properties of the Chinese version of the smoking self-efficacy (SSE) survey. DESIGN AND METHODS: The SSE survey was translated into Chinese then was back-translated into English, reviewed for content validity, pilot tested, and administered to 401 children between December 1998 and August 1999. A random cluster sampling method was used in this study. FINDINGS: Reliability was indicated by Cronbach's alpha coefficient, .98. The validity of the SSE scale was determined by face validity, item-total correlation coefficient, content validity index, and concurrent validity. Principal component analysis was done to determine the construct validity of the SSE scale. The revised SSE scale had three components accounting for 74.3% of the total variance with alpha of .96. The correlation coefficient between the SSE and revised SSE scale was .99. CONCLUSIONS: Findings show that the revised SSE scale is not only parsimonious but also it is as reliable and valid as is the original SSE scale. This translated instrument is appropriate for use in studies of smoking behavior in Taiwanese children aged 11 to 14 years. Further research will be needed to validate the SSE scale with different populations and settings in Taiwan.

Child↗

The reliability and validity of birth certificates.

OBJECTIVES: To summarize the reliability and validity of birth certificate variables and encourage nurses to spearhead data improvement. DATA SOURCES: A Medline key word search of reliability and validity of birth certificate, and a reference review of more than 60 articles were done. STUDY SELECTION: Twenty-four primary research studies of U.S. birth certificates that involved validity or reliability assessment. DATA EXTRACTION: Studies were reviewed, critiqued, and organized as either a reliability or a validity study and then grouped by birth certificate variable. DATA SYNTHESIS: The reliability and validity of birth certificate data vary considerably by item. Insurance, birthweight, Apgar score, and delivery method are more reliable than prenatal visits, care, and maternal complications. Tobacco and alcohol use, obstetric procedures, and delivery events are unreliable. Birth certificates are not valid sources of information on tobacco and alcohol use, prenatal care, maternal risk, pregnancy complications, labor, and delivery. CONCLUSIONS: Birth certificates are a key data source for identifying causes of increasing U.S. infant mortality but have serious reliability and validity problems. Nurses are with mothers and infants at birth, so they are in a unique position to improve data quality and spread the word about the importance of reliable and valid data. Recommendations to improve data are presented.

Alcohol Drinking↗

Validity of information concerning the use of dental services obtained in interviews.

Two hundred and fifty-two persons out of a population of 358 were interviewed concerning their use of dental services. The validity of the information was tested by comparing the answers from each respondent with the contents of his/her dental treatment record. Replies to a question about the time interval since the last dental visit showed a high degree of validity. The validity of information concerning the type of treatment received at the last course of dental visits showed high validity for a single treatment and low validity when the treatment services were mixed. Responses about the regularity of treatment attendance demonstrated decreasing degree of validity with increasing number of dental visits during the last 5 years. The demographic and socioeconomic characteristics of the respondents showed little relation to the validity of their answers. However, the degree of validity decreased with increasing number of teeth.

Adult↗

Taxonomic validation: an overview.

A valid taxonomy legitimizes the elements that make up the taxonomy and increases trust in its generalizability and predictability. There is a concern that the NANDA Taxonomy is not a valid taxonomic structure. Despite on-going work to validate individual nursing diagnoses, there is little research that focuses on validation of groups of diagnoses (taxons) within the NANDA taxonomy. This last article in a series of four will familiarize the readers with why, what, and how a taxonomy of nursing diagnoses can be validated. This article highlights assessment of the validity and reliability of a taxonomy, compares the process of taxonomic validation to the research process, and explores examples of validation design.

Cluster Analysis↗

Validation of Monte Carlo generated phase-space descriptions of medical linear accelerators.

The accuracy of Monte Carlo codes in dose calculation systems relies on the correctness of the input data. Monte Carlo calculations are performed to generate phase-space descriptions of the Varian 2100C accelerator at 6 and 18 MeV. Before these data can be reliably used as the input for dose calculations in patients, they must be properly validated. This validation consists of three different stages: validation of the coding of the geometry, validation of the user code for the Monte Carlo code, and validation of calculated results. Geometric validation is performed by isolating and testing treatment head components independently. The user code is checked by testing for energy conservation and the variance reduction schemes incorporated into the user code are checked by comparison of results calculated with and without their employment. Validation of the phase-space description is performed by calculation of depth dose curves and lateral profiles for dose deposition in phantom, with difference plots used to illustrate any discrepancies. Calculated and experimental in-phantom output is also determined. After complete validation, the calculated data can then be reliably used as the input for dose calculations.

Biophysical Phenomena↗

Validity, reliability, and applicability of seven definitions of hip osteoarthritis used in epidemiological studies: a systematic appraisal.

OBJECTIVE: To summarise and review articles addressing the quality (validity, reliability, applicability) of seven commonly used definitions of hip osteoarthritis (OA) for epidemiological studies in order to use it primarily as a classification criterion. METHODS: Medline and Embase were searched and articles studying the validity, reliability, or applicability of the definitions of hip OA were selected. Two reviewers independently extracted data on the quality of the seven definitions. RESULTS: Review of the literature showed the validity of the various definitions of hip OA, in particular, has barely been investigated. Minimal joint space (MJS) demonstrated the highest (intra- and interrater) reliability, and showed the highest association with hip pain and restricted internal rotation compared with the other definitions of hip OA. The reliability of the Kellgren and Lawrence grade and the index according to Lane is comparable with that of the MJS, but the construct validity should be investigated more thoroughly. The reliability and validity according to the Croft grade were inferior to the MJS, the Kellgren and Lawrence grade, and the index according to Lane. Despite precise and extensive development, the ACR criteria showed poor reliability and poor cross-validity (agreement between three ACR criteria sets) in a primary care setting. CONCLUSIONS: The reliabilities of MJS, Kellgren and Lawrence, and the index according to Lane were comparable, but the MJS had the highest relationship with hip pain in a male population. Considering how often definitions of hip OA are used, it is surprising that the validity has been so poorly investigated, and the validity needs to be studied more thoroughly.

Adult↗

Validation successes: chemicals.

The ECVAM validation concept, which was defined at two validation workshops held in Amden (Switzerland) in 1990 and 1994, and which takes into account the essential elements of prevalidation and biostatistically defined prediction models, has been officially accepted by European Union (EU) Member States and by the Federal regulatory agencies of the USA and the OECD. The ECVAM validation concept was introduced into the ongoing ECVAM/COLIPA validation study of in vitro phototoxicity tests, which ended successfully in 1998. The 3T3 neutral red uptake in vitro phototoxicity test was the first experimentally validated in vitro toxicity test recommended for regulatory purposes by the ECVAM Scientific Advisory Committee (ESAC). It was accepted by the EU into the legislation for chemicals in the year 2000. From 1996 to 1998, two in vitro skin corrosivity tests were successfully validated by ECVAM, and they were also officially accepted into the EU regulations for chemicals in the year 2000. Meanwhile, in 2002, the OECD Test Guidelines Programme is considering the worldwide acceptance of the validated in vitro phototoxicity and corrosivity tests. Finally, from 1997 to 2000, an ECVAM validation study on three in vitro embryotoxicity tests was successfully completed. Therefore, the three in vitro embryotoxicity tests, the whole embryo culture (WEC) test on rat embryos, the micromass (MM) test on limb bud cells of mouse embryos, and the embryonic stem cell test (EST) including a permanent embryonic mouse stem cell line, are considered for routine use in laboratories of the European pharmaceutical and chemicals industries.

3T3 Cells↗

Further validation of the Elderly Mobility Scale for measurement of mobility of hospitalized elderly people.

OBJECTIVE: To further assess the validity and inter-rater reliability of the Elderly Mobility Scale (EMS). Also whether the scale reflects elderly people's perceptions regarding their mobility, and whether it can predict discharge destination, or likelihood of falling. DESIGN: Questionnaire-based study completed on admission and weekly after this on all patients referred to physiotherapy for mobility problems over the course of one month. SETTING: Care of the elderly wards in the Bristol General Hospital. SUBJECTS: Sixty-six patients (ages 66-69 years, 66% female) were included in the validity study. Nineteen patients (ages 71-95 years, 47% female) were included in the inter-rater reliability study. INTERVENTIONS: EMS, Barthel and patients' perceptions of mobility were tested with routine physiotherapy treatment carried out as necessary. MAIN OUTCOME MEASURES: Concurrent validity was assessed by correlating EMS scores with Barthel scores using Spearman's test. Inter-rater reliability was also tested using a Spearman's correlation. EMS scores of patients were also evaluated in conjunction with whether or not they fell and their destination on discharge. RESULTS: A significant correlation between EMS and Barthel scores indicated concurrent validity. Inter-rater reliability was demonstrated on 19 patients with a significant correlation between scores. No predictive validity could be ascribed to EMS in terms of discharge destination or likelihood of falling. Results do indicate a possible predictive validity of the functional reach component of the EMS regarding the risk of future falls. CONCLUSIONS: The EMS was found to be a valid scale with good inter-rater reliability that could be readily applied during daily clinical work. However, it was found to have no predictive validity in terms of falling or discharge destination.

Activities of Daily Living↗