Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

Using probability vs. nonprobability sampling to identify hard-to-access participants for health-related research: costs and contrasts.

This article compares the recruitment costs and participant characteristics associated with the use of probability and nonprobability sampling strategies in a longitudinal study of older hemodialysis patients and their spouses. Contrasts were made of people who accrued to the study based on probability and nonprobability sampling strategies. Probability-based sampling was more time-efficient and cost-effective than nonprobability sampling. There were no significant differences between the respondents identified through probability and nonprobability sampling on age, gender, years married, education, work status, and professional job status. Respondents from the probability sample were more likely to be Protestant and less likely to be Catholic than those from the nonprobability sample. Respondents from the probability sample were more likely to be Black, whereas those from the nonprobability sample were more likely to be White. There are strengths and shortcomings associated with both nonprobability and probability sampling. Researchers need to consider representativeness and external validity issues when designing sampling and related recruitment plans for health-related research.

Aged↗

The diagnostic efficiency of the Rorschach Depression Index and the Schizophrenia Index: a review.

This review focuses on the diagnostic efficiency of the new versions of the Rorschach Comprehensive System Depression Index (DEPI) and the Schizophrenia Index (SCZI). Clinical diagnosis according to the Diagnostic and Statistical Manual of Mental Disorders was chosen as the external validation criteria. The sensitivity, specificity, and overall classification rates for the indices were presented from the studies or computed from the data when possible. The positive and negative predictive validity was estimated at three different base rates. As regards the DEPI the results showed a large variation in diagnostic performance as the index seemed to have relatively more success in identifying nonpsychotic and unipolar depression than psychotic and bipolar depression. The DEPI did not successfully identify depression among adolescent patients. As regards the SCZI the results more consistently indicated that the index effectively discriminates between psychotic and nonpsychotic patients and the predictive validity of both a positive and negative SCZI was found to be high.

Adult↗

An examination and replication of the psychometric properties of the Massachusetts Youth Screening Instrument--second edition (MAYSI-2) among adolescents in detention settings.

There is a high prevalence of psychological disorders among adolescents in detention facilities. The need for a simple, effective screening tool led to the development of the Massachusetts Youth Screening Instrument (MAYSI) and its successor, the MAYSI-2. This study evaluated the MAYSI-2 psychometric properties based on the records of 704 youths evaluated at intake to detention facilities. In addition to factor structure, the study evaluated test-retest reliability and concurrent external validity. Results were generally encouraging in terms of the use of MAYSI-2 in detention facilities, and directions for future research are discussed.

Adolescent↗

External correlates of the MMPI-2 content component scales in mental health inpatients.

External correlates of the Minnesota Multiphasic Personality Inventory-2 (MMPI-2) Content Component Scales were identified using an inpatient sample of 544 adults. The Brief Psychiatric Rating Scale (BPRS) and Symptom Checklist 90-Revised (SCL-90-R) produced correlates of the Content Component Scales, demonstrating external validity with clinician-rated and self-report scales. Relationships between MMPI-2 Content Component Scales and patient hospital chart variables were also examined. Results demonstrated cross-criterion validity in that most MMPI-2 Content Component Scales correlated with appropriate BPRS dimensions, SCL-90-R items, and other patient variables. Discriminant validity of the Content Component Scales was also demonstrated. In addition, the finding that similar patterns of correlates are produced when the component scales are correlated with a self-report measure, as well as clinician ratings and medical chart variables, provides converging lines of evidence supporting the construct validity of the Content Component Scales.

Adult↗

Access to health care and hospitalization for ambulatory care sensitive conditions.

Hospitalization for Ambulatory Care Sensitive Conditions (ACSH) is an accepted indicator of access to health care and avoidable morbidity. Accessible care of reasonable quality should reduce ACSH. Little research has examined the indicator's external validity. We calculated standardized ACSH rates for 32 areas of Victoria, Australia (population 4.4 million). A representative survey measured access, disease prevalence, propensity to seek care, disease burden, social determinants of health services use, and behavioral risk factors. Regression analyses compared self-rated access with ACSH rates. Independent of prevalence, propensity to seek care, disease burden, and physician supply, better access was associated with lower ACSH rates. Results provide support for the ACSH indicator. When rural residence was considered, the covariate measuring access was not significant. However, rural residence also may contribute importantly to access. Results suggest both the complexity of the meaning of access and the desirability of further research to validate the ACSH indicator.

Adolescent↗

Process evaluations of the 5-a-day projects.

Process evaluation is an important, but infrequently conducted, component of evaluating the impact of health promotion interventions. The process evaluation results from the nine 5-a-Day projects were overviewed. Process evaluation helped explain some of the weaker aspects of program performance, process indicators occasionally declined over time and varied by demographic characteristics, and some process measures were related to mediating variables and program outcomes. Future development of process evaluation must include further development of concepts, more consistent and thorough conduct of process evaluation, appropriate methodological work, and assessment of the relations among the process evaluation components and the program mediators and outcomes. Further development in this area promises refinement in understanding how and why interventions achieve their effects, how best to conduct intervention programs to maximize effects, and enhancement of the internal and external validity of the studies.

Adolescent↗

Work site health promotion research: to what extent can we generalize the results and what is needed to translate research to practice?

Information on external validity of work site health promotion research is essential to translate research findings to practice. The authors provide a literature review of work site health behavior interventions. Using the RE-AIM framework, they summarize characteristics and results of these studies to document reporting of intervention reach, adoption, implementation, and maintenance. The authors reviewed a total of 24 publications from 11 leading health behavior journals. They found that participation rates among eligible employees were reported in 87.5% of studies; only 25% of studies reported on intervention adoption. Data on characteristics of participants versus nonparticipants were reported in fewer than 10% of studies. Implementation data were reported in 12.5% of the studies. Only 8% of studies reported any type of maintenance data. Stronger emphasis is needed on representativeness of employees, work site settings studied, and longer term results. Examples of how this can be done are provided.

Health Behavior↗

Patient preferences in randomised trials: threat or opportunity?

OBJECTIVES: To assess whether it is feasible to elicit patients' preferences for treatments and then to proceed with randomisation which may allocate those with preferences to their less preferred treatment; and to describe which prognostic variables were associated with such preferences within the context of a randomised trial of an exercise programme for back pain. METHODS: The first 97 patients enrolled in a randomised controlled trial (RCT) for the treatment of back pain were asked about their preferences, health characteristics and other prognostic variables. RESULTS: Fifty-eight (60%) patients preferred to be allocated to the exercise programme whilst 38 (39%) were indifferent; one patient preferred conventional general practitioner (GP) management. No patient refused randomisation. Comparing patients preferring the exercise programme with indifferent patients showed that the former had a higher belief in the effectiveness of the new treatment (P < 0.01), tended to have worse back pain (P = 0.09), had back pain for a shorter duration (P = 0.04), and tended to have had more GP home visits (P = 0.06). CONCLUSIONS: For many randomised trials preference may be an important prognostic variable. In such circumstances, preference should be taken into account in the final analysis. This study demonstrates it is sometimes feasible to randomise patients to their less preferred treatment, thus allowing more robust statistical comparisons between randomised groups. This modification may make RCTs more rigorous and improve their external validity.

Adult↗

Threats to applicability of randomised trials: exclusions and selective participation.

BACKGROUND: Although the randomised controlled trial (RCT) is regarded as the 'gold standard' in terms of evaluating the effectiveness of interventions, it is susceptible to challenges to its external validity if those participating are unrepresentative of the reference population for whom the intervention in question is intended. In the past, reporting on numbers and types of potential subjects that have been excluded by design, and centres, clinicians or patients that have elected not to participate, has generally been poor, and the threat to inference posed by possible selection bias is unclear. METHODS: A systematic review was undertaken, based largely on MEDLINE and EMBASE with follow-up of cited references, to assess the extent, nature and importance of excluding potential subjects or the unwillingness of particular centres, clinicians or patients to participate. RESULTS: RCTs vary widely in the extent to which potential future recipients of treatment are included. The reasons cited for excluding certain categories of patient may be medical or scientific. Medical reasons include a high risk of adverse effects and the belief that benefit will be relatively small or absent (or has already been established) in the groups in question. Scientific reasons include more precise estimates of treatment effect because of a relatively homogeneous sample and the reduction of potential bias by excluding those individuals most likely to be lost to follow-up. Many RCTs have blanket exclusions, such as the elderly, women and ethnic minorities, but reasons for these exclusions are seldom given. Evaluative research is undertaken predominantly in university or teaching centres. Non-randomised studies are more likely than RCTs to include non-teaching centres. The effect of patient non-participation appears to depend on whether the RCT is concerned with treatment of an existing condition or with disease prevention. Participants in treatment trials tend to be more severely ill than those who do not participate. In contrast, those who participate in prevention trials are more likely to have adopted a healthy lifestyle than those who decline. Most evaluative studies fail to document adequately the characteristics of those who, while eligible, do not participate. However, subjects included in RCTs (i.e. eligible and participating) tend to have a different prognosis than patients identified from clinical databases. CONCLUSIONS: Narrow inclusion criteria may offer benefits such as increased precision and reduced loss to follow-up, but there are important disadvantages, such as uncertainty about extrapolation of results, which may result in denial of effective treatment to groups who might benefit, and delay in obtaining definitive results because of reduced recruitment rate. Selective participation by teaching centres and sicker patients in treatment RCTs may exaggerate the measured treatment effect. Prevention trials, on the other hand, may underestimate effects as participants have less capacity to benefit.

Australia↗

Retention of survivors of acute lymphoblastic leukemia in a longitudinal study of bone mineral density.

Attrition in longitudinal studies of survivors of childhood cancer reduces these studies' statistical power, introduces bias and threatens internal and external validity. This study investigated the variables associated with dropout of survivors of acute lymphoblastic leukemia in a trial investigating the effect of vitamin D and calcium supplementation and nutritional counseling on bone mineral density (BMD). Twenty-five participants withdrew from the study. Common reasons given for withdrawing were intolerance of the study drug, family hardship and schedule conflicts. Few statistically and clinically significant differences identified participants who completed the study. Nurses need to be aware of the reasons that participants withdraw from clinical trials, as they are in a strategic position to encourage patients to participate in health promotion studies.

Adaptation, Psychological↗

Presumed consent and other predictors of cadaveric organ donation in Europe.

CONTEXT: Few studies on presumed consent and environmental predictors of cadaveric organ donation in Europe have been published. OBJECTIVE: To determine if a presumed consent policy and other variables can be used to predict the cadaveric organ donation rate per million population. DESIGN: Secondary analysis of published data. SETTING: Europe. PARTICIPANTS: The unit of analysis for this study is the individual country. MAIN OUTCOME MEASURE: Cadaveric organ donation rate per million population. RESULTS: Original and transformed data were subjected to ordinary least-squares regression. All 4 independent variables were significant predictors of cadaveric donation rate, including (1) having a presumed consent (opting-out) policy in practice, (2) number of transplant centers per million population, (3) percentage of the population enrolled in third-tier education, and (4) percentage of population that is Roman Catholic. CONCLUSION: Findings may be useful to academics and professionals responsible for organ procurement. Additional research is necessary for practical application of findings. Generalizing these findings beyond Europe may be problematic because of external validity constraints.

Cadaver↗

Predictive value of ischemic lesion volume assessed with magnetic resonance imaging for neurological deficits and functional outcome poststroke: A critical review of the literature.

OBJECTIVE: Ischemic lesion volume is assumed to be an important predictor of poststroke neurological deficits and functional outcome. This critical review examines the methodological quality of MRI studies and the predictive value of hemispheric infarct volume for neurological deficits (at body function level) and functional outcome (at activities level). METHODS: Using Medline, PiCarta, and Embase to identify studies, 13 of the 747 identified studies met the authors' inclusion criteria. Subsequently, studies were tested for adherence to the key methodological criteria for internal, statistical, and external validity. Each criterion was weighted binary, and studies with 6 points or more were judged to be valid for assessing the predictive value of MRI for outcome. RESULTS: The 13 included studies had several methodological weaknesses with respect to internal validity, and none of them took lesion location into account. Only a few used outcome measures according to the International Classification of Functioning, Disability and Health and followed patients beyond 6 months. Correlation coefficients between MRI lesion volume and outcomes were higher for outcomes defined at body function level (National Institutes of Health Stroke Scale; median 0.67; range: 0.57-0.91) than for those defined at the level of activities (Barthel Index; median -0.49; range: -0.33 to -0.74). CONCLUSIONS: Methodological shortcomings of most studies confound the prognostic value of MRI in predicting stroke outcome, and few studies have focused on functional outcome. Future studies should investigate the added value of MRI volume over clinical neurological variables in predicting functional outcome beyond 6 months poststroke.

Brain↗

Generalizing results of randomized trials to clinical practice: reliability and cautions.

BACKGROUND: Well designed randomized controlled trials provide reliable evidence of treatment effects, but there is no consensus on how best to apply these results to clinical practice. The main concerns are that populations enrolled in trials are more selected than those treated in a clinical setting, and whether the treatment effects observed in trials will also be observed in clinical practice. METHODS: An informal literature review was undertaken to find studies analysing the issue of generalizing trial results (external validity) to clinical practice. RESULTS: Most of the studies focused on differences in patients characteristics (age, gender, severity of disease, concomitant treatments and so on) between the clinical trial population and a 'real world' clinical population. None provided good evidence of a reduction in the treatment effect in the trial compared to what might happen in clinical practice for simple pharmacological treatments. However complex treatments like surgery or percutaneous interventional procedures, had a greater potential for variation. Extrapolating treatments to different health care settings from the trial can result in important variations in treatment effects. CONCLUSIONS: Complex therapies need careful consideration before they can be applied routinely from trials into practice, and applying results from one health care environment to a different one should be carried out with caution. Generalizing results from well conducted trials to clinical practice can mostly be carried out with confidence, especially for simple therapies with good evidence of benefit.

Humans↗

EpiJen: a server for multistep T cell epitope prediction.

BACKGROUND: The main processing pathway for MHC class I ligands involves degradation of proteins by the proteasome, followed by transport of products by the transporter associated with antigen processing (TAP) to the endoplasmic reticulum (ER), where peptides are bound by MHC class I molecules, and then presented on the cell surface by MHCs. The whole process is modeled here using an integrated approach, which we call EpiJen. EpiJen is based on quantitative matrices, derived by the additive method, and applied successively to select epitopes. EpiJen is available free online. RESULTS: To identify epitopes, a source protein is passed through four steps: proteasome cleavage, TAP transport, MHC binding and epitope selection. At each stage, different proportions of non-epitopes are eliminated. The final set of peptides represents no more than 5% of the whole protein sequence and will contain 85% of the true epitopes, as indicated by external validation. Compared to other integrated methods (NetCTL, WAPP and SMM), EpiJen performs best, predicting 61 of the 99 HIV epitopes used in this study. CONCLUSION: EpiJen is a reliable multi-step algorithm for T cell epitope prediction, which belongs to the next generation of in silico T cell epitope identification methods. These methods aim to reduce subsequent experimental work by improving the success rate of epitope prediction.

Computer Simulation↗

Diagnostic accuracy of the Eurotest for dementia: a naturalistic, multicenter phase II study.

BACKGROUND: Available screening tests for dementia are of limited usefulness because they are influenced by the patient's culture and educational level. The Eurotest, an instrument based on the knowledge and handling of money, was designed to overcome these limitations. The objective of this study was to evaluate the diagnostic accuracy of the Eurotest in identifying dementia in customary clinical practice. METHODS: A cross-sectional, multi-center, naturalistic phase II study was conducted. The Eurotest was administered to consecutive patients, older than 60 years, in general neurology clinics. The patients' condition was classified as dementia or no dementia according to DSM-IV diagnostic criteria. We calculated sensitivity (Sn), specificity (Sp) and area under the ROC curves (aROC) with 95% confidence intervals. The influence of social and educational factors on scores was evaluated with multiple linear regression analysis, and the influence of these factors on diagnostic accuracy was evaluated with logistic regression. RESULTS: Sixteen neurologists recruited a total of 516 participants: 101 with dementia, 380 without dementia, and 35 who were excluded. Of the 481 participants who took the Eurotest, 38.7% were totally or functionally illiterate and 45.5% had received no formal education. Mean time needed to administer the test was 8.2+/-2.0 minutes. The best cut-off point was 20/21, with Sn = 0.91 (0.84-0.96), Sp = 0.82 (0.77-0.85), and aROC = 0.93 (0.91-0.95). Neither the scores on the Eurotest nor its diagnostic accuracy were influenced by social or educational factors. CONCLUSION: This naturalistic and pragmatic study shows that the Eurotest is a rapid, simple and useful screening instrument, which is free from educational influences, and has appropriate internal and external validity.

Aged↗

Approaches to the evaluation of outbreak detection methods.

BACKGROUND: An increasing number of methods are being developed for the early detection of infectious disease outbreaks which could be naturally occurring or as a result of bioterrorism; however, no standardised framework for examining the usefulness of various outbreak detection methods exists. To promote comparability between studies, it is essential that standardised methods are developed for the evaluation of outbreak detection methods. METHODS: This analysis aims to review approaches used to evaluate outbreak detection methods and provide a conceptual framework upon which recommendations for standardised evaluation methods can be based. We reviewed the recently published literature for reports which evaluated methods for the detection of infectious disease outbreaks in public health surveillance data. Evaluation methods identified in the recent literature were categorised according to the presence of common features to provide a conceptual basis within which to understand current approaches to evaluation. RESULTS: There was considerable variation in the approaches used for the evaluation of methods for the detection of outbreaks in public health surveillance data, and appeared to be no single approach of choice. Four main approaches were used to evaluate performance, and these were labelled the Descriptive, Derived, Epidemiological and Simulation approaches. Based on the approaches identified, we propose a basic framework for evaluation and recommend the use of multiple approaches to evaluation to enable a comprehensive and contextualised description of outbreak detection performance. CONCLUSION: The varied nature of performance evaluation demonstrated in this review supports the need for further development of evaluation methods to improve comparability between studies. Our findings indicate that no single approach can fulfil all evaluation requirements. We propose that the cornerstone approaches to evaluation identified provide key contributions to support internal and external validity and comparability of study findings, and suggest these be incorporated into future recommendations for performance assessment.

Benchmarking↗

The Knee Clinical Assessment Study-CAS(K). A prospective study of knee pain and knee osteoarthritis in the general population: baseline recruitment and retention at 18 months.

BACKGROUND: Selective non-participation at baseline (due to non-response and non-consent) and loss to follow-up are important concerns for longitudinal observational research. We investigated these matters in the context of baseline recruitment and retention at 18 months of participants for a prospective observational cohort study of knee pain and knee osteoarthritis in the general population. METHODS: Participants were recruited to the Knee Clinical Assessment Study-CAS(K)--by a multi-stage process involving response to two postal questionnaires, consent to further contact and medical record review (optional), and attendance at a research clinic. Follow-up at 18-months was by postal questionnaire. The characteristics of responders/consenters were described for each stage in the recruitment process to identify patterns of selective non-participation and loss to follow-up. The external validity of findings from the clinic attenders was tested by comparing the distribution of WOMAC scores and the association between physical function and obesity with the same parameters measured directly in the target population as whole. RESULTS: 3106 adults aged 50 years and over reporting knee pain in the previous 12 months were identified from the first baseline questionnaire. Of these, 819 consented to further contact, responded to the second questionnaire, and attended the research clinics. 776 were successfully followed up at 18 months. There was evidence of selective non-participation during recruitment (aged 80 years and over, lower socioeconomic group, currently in employment, experiencing anxiety or depression, brief episode of knee pain within the previous year). This did not cause significant bias in either the distribution of WOMAC scores or the association between physical function and obesity. CONCLUSION: Despite recruiting a minority of the target population to the research clinics and some evidence of selective non-participation, this appears not to have resulted in significant bias of cross-sectional estimates. The main effect of non-participation in the current cohort is likely to be a loss of precision in stratum-specific estimates e.g. in those aged 80 years and over. The subgroup of individuals who attended the research clinics and who make up the CAS(K) cohort can be used to accurately estimate parameters in the reference population as a whole. The potential for selection bias, however, remains an important consideration in each subsequent analysis.

Aged↗

Comparison of artificial neural network and logistic regression models for prediction of mortality in head trauma based on initial clinical data.

BACKGROUND: In recent years, outcome prediction models using artificial neural network and multivariable logistic regression analysis have been developed in many areas of health care research. Both these methods have advantages and disadvantages. In this study we have compared the performance of artificial neural network and multivariable logistic regression models, in prediction of outcomes in head trauma and studied the reproducibility of the findings. METHODS: 1000 Logistic regression and ANN models based on initial clinical data related to the GCS, tracheal intubation status, age, systolic blood pressure, respiratory rate, pulse rate, injury severity score and the outcome of 1271 mainly head injured patients were compared in this study. For each of one thousand pairs of ANN and logistic models, the area under the receiver operating characteristic (ROC) curves, Hosmer-Lemeshow (HL) statistics and accuracy rate were calculated and compared using paired T-tests. RESULTS: ANN significantly outperformed logistic models in both fields of discrimination and calibration but under performed in accuracy. In 77.8% of cases the area under the ROC curves and in 56.4% of cases the HL statistics for the neural network model were superior to that for the logistic model. In 68% of cases the accuracy of the logistic model was superior to the neural network model. CONCLUSIONS: ANN significantly outperformed the logistic models in both fields of discrimination and calibration but lagged behind in accuracy. This study clearly showed that any single comparison between these two models might not reliably represent the true end results. External validation of the designed models, using larger databases with different rates of outcomes is necessary to get an accurate measure of performance outside the development population.

Adolescent↗