Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Telepsychiatry treatment outcome research methodology: efficacy versus effectiveness.

The use of videoconferencing technology to provide mental health services (telepsychiatry) offers hope for addressing longstanding problems regarding work force shortages and access to care, especially in remote or rural areas. However, data on treatment outcomes (i.e., based on randomized clinical trials) from telepsychiatry applications are virtually nonexistent, representing an important gap in the literature. An important methodological decision point in developing treatment outcome research is whether to take an efficacy or effectiveness approach. Efficacy approaches offer enhanced internal validity; however, they may have limited generalizability to real-world settings. Effectiveness approaches offer enhanced external validity. But, they are typically less controlled than efficacy studies, thereby limiting the assumptions that can be made about causality. The current state of telepsychiatry research necessitates efficacy studies, the outcomes from which can be used to inform future effectiveness studies.

Health Services Research↗

The Sickness Impact Profile for nursing homes (SIP-NH)

BACKGROUND: Valid, feasible measures of functional status are needed to evaluate the expanding nursing home population. This study attempts to increase relevance and reduce respondent burden of the Sickness Impact Profile (SIP) for nursing home residents while maintaining internal consistency and validity. METHODS: 231 residents from one academic and four community nursing homes, aged > or = 60 with a Mini-Mental State Exam score > or = 11, were study participants. Nominal group process was used to identify items and/or categories for removal. Candidate items were those that: represented restrictions of the nursing home environment, had weak item-total score correlations, and/or made minimal contribution to category internal consistency. Reduction was constrained by: minimum correlation of r = .90 between SIP and Sickness Impact Profile for Nursing Homes (SIP-NH) scores, coefficients alpha that fell within 95% confidence regions about predicted alpha. Convergent and discriminant validity were evaluated with the Katz Activities of Daily Living, Physical Disability Index, Geriatric Depression Scale, and Folstein Mini-Mental State Exam. RESULTS: The SIP-NH contains 66 items, a 51.5% reduction. Correlations between the SIP-NH and SIP were: total score r = .98, Physical dimension r = .97, and Psychosocial dimension r = .97. Alpha coefficients all fell within the 95% confidence regions. The SIP and the SIP-NH did not differ in correlations with validating instruments. CONCLUSIONS: The SIP-NH reduces respondent burden and has acceptable internal consistency and external validity. Potentially useful for discriminatory and predictive purposes, responsiveness to change will require longitudinal evaluation.

Activities of Daily Living↗

Schizophrenia and related disorders: experience with current diagnostic systems.

Schizophrenia and related disorders include a variety of psychotic disorders in the major classification systems, ICD-10 and DSM-IV, with only partial concordance between the two systems. They both rely on demonstrated reliability, but which disorders are the most valid still has to be determined. Particularly for the ICD-10 disorders, only few studies examining external validity have appeared. Disorders of uncertain validity include 'schizo-affective disorders' which in ICD-10 contain the DSM-IV psychotic mood disorders with first-rank symptoms or bizarre delusions; ICD-10 'schizotypal disorder' which in DSM-IV is a personality disorder; the ICD-10 'acute and transient psychotic disorders' and the DSM-IV 'brief psychotic disorder'. Concerning diagnostic criteria, the reliability and validity of Schneiderian first-rank symptoms, 'bizarre' delusions and the Bleulerian 'negative' symptoms have been questioned. Validity studies in these areas are needed before it will be possible to provide major reconstructions for future diagnostic systems. One may hope that, eventually, one common worldwide psychiatric classification will be available.

Diagnosis, Differential↗

Validation of probabilistic predictions.

Current advances in high-speed computing and increased availability of statistical software have led to widespread use of statistical methods for the development of computerized protocols predictive of binary health outcomes. If these predictive algorithms are to be used in settings other than those for which they were developed, e.g., applied in a different geographic setting or extrapolated for use in a slightly different population, then they should be carefully validated to ensure appropriate application. Miller et al. (Stat Med. 1991) provided a comprehensive methodology for external validation of logistic prediction models, and applied these methods in a temporal validation setting. In this article, the authors emphasize how these methods can be applied to general forms of probabilistic predictions and provide several SAS macros for computation of the desired statistics.

Adolescent↗

MELPREDICT: a logistic regression model to estimate CDKN2A carrier probability.

BACKGROUND: Heritable alterations in CDKN2A account for a subset of familial melanoma cases although no robust method exists to identify those at risk of being a mutation carrier. METHODS: We set out to construct a model for estimating CDKN2A mutation carrier probability using a cohort of 116 consecutive familial cutaneous melanoma patients evaluated at Massachusetts General Hospital Pigmented Lesion Center between April 2001 and September 2004. Germline CDKN2A and CDK4 status on the familial melanoma cases and clinical features associated with mutational status were then used to build a multiple logistic regression model to predict carrier probability and performance of model on external validation. RESULTS: From the 116 kindreds prone to melanoma in the Boston area, 13 CDKN2A mutation carriers were identified and 12 were subsequently used in the modeling. Proband age at diagnosis, number of proband primaries, and number of additional family primaries were most closely associated with germline mutations. The estimated probability of the proband being a mutation carrier based on the logistic regression model (MELPREDICT) is given by e(L)/(1 + e(L) where L = 1.99+[0.92x(no. of proband primaries)]+[0.74x(no. of additional family primaries)]-[2.11xln(age)]. The mean estimated probabilities for subjects in the Boston dataset were 55.4% and 5.1% for the mutation carriers and non-carriers respectively. In a receiver operator characteristic analysis, the area under the curve was 0.881 (95% confidence interval 0.739 to 1.000) for the Boston model set (n = 116) and 0.803 (0.729 to 0.877) for an external Toronto hereditary melanoma cohort (n = 143). CONCLUSIONS: These results represent the first-iteration logistic regression model to approximate CDKN2A carrier probability. Validation of this model with an external dataset revealed relatively robust performance.

Adolescent↗

Illness severity and self-efficacy as course predictors of DSM-IV alcohol dependence in a multisite clinical sample.

Illness severity and self-efficacy are two constructs of growing interest as predictors of clinical response in alcoholism. Using alternative measures of illness severity (DSM-IV symptom count, Alcohol Dependence Scale, and Addiction Severity Index) and self-efficacy (brief version of the Situational Confidence Questionnaire) rigorously controlled for theoretically important background variables, we studied their unique contribution to multiple indices of relapse, relapse latency, and use of alternative coping behaviors in a large, heterogeneous clinical sample. The Alcohol Dependence Scale contributed to the prediction of 4 of 5 relapse indicators. The SCQ failed to predict relapse behavior or its precursor, coping response. The findings emphasize the predictive validity of severity of dependence as a course specifier and underline the need for more sensitive and externally valid measures of cognitive processes such as self-efficacy for application in future studies of posttreatment behavior.

Adaptation, Psychological↗

Toward a validation of a new definition of agitated depression as a bipolar mixed state (mixed depression).

PURPOSE: As psychotic agitated depression is now a well-described form of mixed state during the course of bipolar I disorder, we sought to investigate the diagnostic validity of a new definition for agitated (mixed) depression in bipolar II (BP-II) and major depressive disorder (MDD). MATERIALS AND METHODS: Three hundred and thirty six consecutive outpatients presenting with major depressive episodes (MDE) but without history of mania were evaluated with the Structured Clinical Interview for DSM-IV when presenting for the treatment of MDE. On the basis of history of hypomania they were assigned to BP-II (n = 206) vs. MDD (n = 130). All patients were also examined for hypomania during the current MDE. Mixed depression was operationally defined by the coexistence of a MDE and at least two of the following excitatory signs and symptoms as described by Koukopoulos and Koukopoulos (Koukopoulos A, Koukopoulos A. Agitated depression as a mixed state and the problem of melancholia. In: Akiskal HS, editor. Bipolarity: beyond classic mania. Psychiatr Clin North Am 1999;22:547-64): inner psychic tension (irritability), psychomotor agitation, and racing/crowded thoughts. The validity of mixed depression was investigated by documenting its association with BP-II disorder and with external variables distinguishing it from unipolar MDD (i.e., younger age at onset, greater recurrence, and family history of bipolar disorders). We analyzed the data with multivariate regression (STATA 7). RESULTS: MDE plus psychic tension (irritability) and agitation accounted for 15.4%, and MDE plus agitation and crowded thoughts for 15.1%. The highest rate of mixed depression (38.6%) was achieved with a definition combining MDE with psychic tension (irritability) and crowded thoughts: 23.0% of these belonged to MDD and 76.9% to BP-II. Moreover, any of these permutations of signs and symptoms defining mixed depression was significantly and strongly associated with external validators for bipolarity. The mixed irritable-agitated syndrome depression with racing-crowded thoughts was further characterized by distractibility (74-82%) and increased talkativeness (25-42%); of expansive behaviors from the criteria B list for hypomania, only risk taking occurred with some frequency (15-17%). CONCLUSIONS: These findings support the inclusion of outpatient-agitated depressions within the bipolar spectrum. Agitated depression is validated herein as a dysphorically excited form of melancholia, which should tip clinicians to think of such a patient belonging to or arising from a bipolar substrate. Our data support the Kraepelinian position on this matter, but regrettably this is contrary to current ICD-10 and DSM-IV conventions. Cross-sectional symptomatologic hints to bipolarity in this mixed/agitated depressive syndrome are virtually absent in that such patients do not appear to display the typical euphoric/expansive characteristics of hypomania-even though history of such behavior may be elicited by skillful interviewing for BP-II. We submit that the application of this diagnostic entity in outpatient practice would be of considerable clinical value, given the frequency with which these patients are encountered in such practice and the extent to which their misdiagnosis as unipolar MDD could lead to antidepressant monotherapy, thereby aggravating it in the absence of more appropriate treatment with mood stabilizers and/or atypical antipsychotics.

Adult↗

Functional outcome measures to assess interventions for spasticity.

PURPOSE: Clinicians use functional loss as a criterion to treat spasticity, but the connection between function and severity of spasticity is not well established for monitoring spasticity treatment effect. Studies were reviewed which have implemented outcome measures to assess functional changes relative to changes in spasticity. Criteria for review included the reliability and internal validity of the functional measures used and the strengths/weaknesses of the study designs that likely affected the external validity of the measures for this application. Guidelines are provided for the use and development of functional outcome measures in futures studies of spasticity treatment based on this review. DATA IDENTIFICATION: An English-language literature search using MEDLINE and bibliographies of published articles and textbooks was conducted. RESULTS: Very few functional measures demonstrated changes concurrent to a reduction in spasticity. There were multiple potential confounding factors in study protocols, reporting of results, and data analysis that might account for the limited number of measures shown to be valid for this application. Selected standardized ordinal functional outcome scales (the PECS and PEDI) and specific functional tasks were identified as measures that show promise for assessing changes concurrent with altered spasticity level. CONCLUSION: Based on a review of previous studies, functional measures involving posture, positioning, balance, and certain mobility skills have potential, with further test development, to provide needed information regarding the impact of spasticity on functional outcome.

Disability Evaluation↗

On the reliability and validity of physician ratings for vulvodynia and the discriminant validity of its subtypes.

OBJECTIVE: This study aimed to test the reliability and validity of physician ratings in a broadly defined sample of women with vulvodynia and to examine the external validity of the vulvodynia subtypes. DESIGN: Participants were 50 women who were independently diagnosed with vulvodynia by two study gynecologists. Physician ratings corresponding to Friedrich's three criteria for vulvar vestibulitis were taken at the two examinations. Each participant's diagnosis was subtyped as vulvar vestibulitis (VV) or dysesthetic vulvodynia (DV) based upon the physician ratings. Participants completed standardized measures of pain, sexual function, psychological function, and quality of life to examine the discriminant validity of the subtypes. RESULTS: Test-retest reliability for the physician ratings of Friedrich's three criteria was stable for two of the three criteria (i.e., pain on attempted vaginal entry and tenderness to pressure localized within the vulvar vestibule). When these criteria were used to categorize participants as having VV or DV, the subtypes were not statistically different for measures used to examine the discriminant validity of the subtypes. While the distribution of patients changed when premenopausal state was added to the inclusion criteria for VV, the subtypes differed little on the outcome measures. CONCLUSIONS: Findings from the present study suggest that physician ratings for Friedrich's criteria can be operationalized and found to be reliable and valid in a wide range of women with vulvodynia. The absence of differences between subtypes on measures of pain, sexual function, psychological function, and quality of life challenge the clinical significance of these subtypes and support the theory that vulvodynia represents a continuum of chronic vulvar pain rather than two distinct entities.

Adult↗

An evaluation of the Extended Barthel Index with acute ischemic stroke patients.

OBJECTIVE: To evaluate the Extended Barthel Index with acute ischemic stroke patients. METHODS: This prospective 1- to 6-week poststroke follow-up study was carried out using 33 newly diagnosed acute ischemic stroke patients who were admitted to the University Medical Centre Ljubljana, Department of Neurology. Measures used were Barthel Index (BI), Extended Barthel Index (EBI), Fugl-Meyer Motor Impairment Scale, 1-5 Self-Assessment scale, Rivermead Behavioural Memory Test. RESULTS: The EBI is a reliable scale in terms of internal consistency. The cognitive part is less reliable than the physical part of the EBI. It is a 3-dimensional scale as calculated by factor analysis (factor 1 with eigen value 8.2, factor 2 with eigen value 2.7 and factor 3 with eigen value 0.9). Criterion validity to the BI and the Fugl-Meyer Motor Impairment Scale was supported (P=0.1-0.001). External validity to the Self-Assessment scale was also supported (P<0.001). It is more sensitive to the changes in functional status that occur in the 1st 6 weeks poststroke than the original BI, although the ceiling effect was not really explained in this follow-up period. CONCLUSION: The EBI is a valid, reliable, 2- to 3-dimensional outcome measure of disability/activity for stroke patients. To some extent, it also reveals the level of patients' perception of their functional status.

Activities of Daily Living↗

A unidimensional measure of Hong's psychological reactance scale.

Research using Hong's Psychological Reactance Scale has been fraught with methodological concerns. Researchers have been unable to find a stable, a nd replicable factor structure. Here, results suggested t hat Hong's Psychological Reactance Scale is a unidimensional one with an average alpha of .74 (SD=.46). This value was attained by first analyzing correlation matrices reproduced from three reports on Hong's Psychological Reactance Scale and then verifying this new factor structure with original data. Tests for internal consistency supported a 1-factor solution. Tests for external consistency supported prior findings in relation to Psychological Reactance and offer evidence that the 1-factor solution is externally valid. While the authors contend that a 1-factor solution is appropriate, further testing is needed for external consistency and refinement of the measure.

Factor Analysis, Statistical↗

Negative v positive schizophrenia. Definition and validation.

We developed criteria for dividing the schizophrenic syndrome into three subtypes: positive, negative, and mixed schizophrenia. Positive schizophrenia is characterized by prominent delusions, hallucinations, positive formal thought disorder, and persistently bizarre behavior; negative schizophrenia, by affective flattening, alogia, avolition, anhedonia, and attentional impairment. In mixed schizophrenia either both negative and positive symptoms are prominent, or neither is prominent. We explored the validity of these criteria in a variety of ways. Significant differences between the three types were noted using external validators such as premorbid adjustment, indices of cognitive dysfunction, ventricular brain ratio, and course in hospital. The correlational structure of the symptom complexes also provided further support for our approach to subtyping.

Adult↗

Classifying psychotic disorders: issues regarding categorial vs. dimensional approaches and time frame to assess symptoms.

The study's aims were to empirically derive classes of disorders and dimensional syndromes within psychotic disorders on the basis of the three time frames of symptom assessment and to comparatively examine their external validity. The level of concordance among classes and among dimensions across the time frames was generally low. The external correlates of psychopathological syndromes differed as a function of both type of assessment and the dimensional or categorical approach used. The dimensional approach was more effective than the categorical approach in predicting a set of clinical variables, irrespective of the time frame used to assess the symptoms. It is concluded that classification of psychotic disorders is highly dependent upon the time frame considered to assess symptoms and that dimensional classifications do have higher predictive power than categorical ones.

Adult↗

Systematic reviews of diagnostic research. Considerations about assessment and incorporation of methodological quality.

The aim of this paper is to provide a theoretical background for performing and reading systematic reviews of diagnostic studies. We first discuss items for assessment of methodological quality in diagnostic studies and then present methods on how to incorporate these quality measures in systematic reviews. The items of internal validity determine whether the presented results of the individual studies are unbiased and can be trusted. Items of external validity determine to what extent the results are applicable outside the population in which the study was performed. The issues concern the adequacy of the study population, the performance and interpretation of the diagnostic tests and the presentation of the results. Several methods exist for incorporation of issues of methodological quality into systematic reviews, such as subgroup analyses, meta-regression analysis, and methodological scores. Publications of diagnostic studies should provide sufficient information to enable assessment of the methodological quality. Furthermore, publication of results of subgroup analyses should be promoted. Methodological criteria lists might help to improve the quality of systematic reviews of diagnostic research. With the items of methodological quality in mind the general practitioner might be better equipped to critically read and interpret diagnostic reviews.

Epidemiologic Studies↗

A depressive symptom scale for the California Psychological Inventory: construct validation of the CPI-D.

To facilitate life span research on depressive symptomatology, a depressive symptom scale for the California Psychological Inventory (CPI) is needed. The authors constructed such a scale (the CPI-D) and compared its psychometric properties with 2 widely used self-report depression scales: the Beck Depression Inventory and the Center for Epidemiological Studies Depression Scale. Construct validity of the CPI-D was examined in 3 studies. Study 1 established content validity, classifying CPI-D items into Diagnostic and Statistical Manual of Mental Disorders-Fourth Edition depressive symptoms. Study 2 used 3 large samples to gather evidence for reliability and validity: correlational analyses demonstrated alpha reliability and convergent and discriminant validity; factor analysis provided evidence for discriminant validity with anxiety; and regression analyses demonstrated comparative validity with existing standard PI scales. Study 3 used clinician ratings of depression and anxiety as criteria for external validity.

Adult↗

What is required to evaluate the impact of pharmaceutical reference pricing?

OBJECTIVE: To describe empirical studies evaluating the impact of reference pricing (RP) interventions in pharmaceutical markets in order to discuss the requirements for these evaluations. METHODS: Ten studies were included in this review. For each study, the nature of the intervention, the nature of the data available, the nature of the question to be answered and the requirements of the evaluation method were examined through a questionnaire. The most frequently used evaluation method was the conventional before-after estimator, and only three studies used the difference-in-differences method. RESULTS: Nine studies evaluated a therapeutic RP system and one evaluated a generic RP system. All of the papers reported how the reference price level was established, but only one study directly reported the updating frequency and criteria of the RP system. In four studies, details of simultaneous interventions were not reported. There is no paper providing evidence on overall social welfare impact. Four papers estimated the impact of intervention on the consumer price of drugs covered by the RP system. Only one provided information about the impact on the price of related drugs not covered by the system. Three studies included an outcome variable for the use of health services. The impact of RP intervention on the level of competition in the market for those medicines covered by the system was reported in only three of ten papers. DISCUSSION/CONCLUSION: Despite the rigorous effort made to evaluate the impact of RP policies in some countries, several limitations may affect both their internal validity (the nature of the data available, statistical problems common to non-experimental data, etc.) and their external validity (heterogeneity in the nature of the intervention).

Cost Control↗

Clinical practice guidelines for the care and treatment of breast cancer: 13. Sentinel lymph node biopsy.

OBJECTIVE: To provide information and recommendations to women with breast cancer and their physicians regarding what is now known about sentinel lymph node (SLN) biopsy. OPTIONS: Axillary dissection; SLN biopsy followed by backup axillary dissection; SLN biopsy. OUTCOMES: Accurate determination of cancer stage, resulting in better-informed therapeutic decisions. EVIDENCE: Systematic review of English-language literature published from January 1991 to December 2000 retrieved primarily from MEDLINE and CANCERLIT. RECOMMENDATIONS: Axillary dissection is the standard of care for the surgical staging of operable breast cancer. If a patient requests or is offered SLN biopsy, the benefits and risks as well as what is and is not known about the procedure should be outlined. Patients should be informed of the number of SLN biopsies performed by the surgeon and the surgeon's success rate with the procedure, as determined by the identification of the SLN and the false-negative rate (the presence of tumour cells in the axillary nodes when the SLN biopsy result is negative). Before surgeons replace axillary dissection by SLN biopsy as the staging procedure at their institution, they should (a) familiarize themselves with the literature on the topic and the techniques needed to perform the procedure, (b) follow a defined protocol for all 3 aspects of the procedure (nuclear medicine, surgery, pathology) and (c) perform backup axillary dissection until an acceptable success rate (as determined by the identification of the SLN and the false-negative rate) is achieved. A surgeon who performs breast cancer surgery infrequently should not perform SLN biopsy. A positive SLN biopsy result or failure to identify an SLN should prompt full axillary dissection. SLN biopsy is contraindicated in women who have clinically palpable nodes, locally advanced breast cancer, multifocal tumours, previous breast surgery or previous irradiation of the breast. Staining of tissue sections with hematoxylin and eosin, and not immunohistochemical analysis for cytokeratin, should determine adjuvant therapy. Participation in randomized clinical trials is encouraged. [A patient version of these guidelines appears in Appendix 1.] VALIDATION: Internal validation within the Steering Committee on Clinical Practice Guidelines for the Care and Treatment of Breast Cancer; no external validation.

Axilla↗

Clinical practice guidelines for the care and treatment of breast cancer: 14. The role of hormone replacement therapy in women with a previous diagnosis of breast cancer.

OBJECTIVE: To provide information and recommendations to women with a previous diagnosis of breast cancer and their physicians regarding hormone replacement therapy (HRT). OUTCOMES: Control of menopausal symptoms, quality of life, prevention of osteoporosis, prevention of cardiovascular disease, risk of recurrence of breast cancer, risk of death from breast cancer. EVIDENCE: Systematic review of English-language literature published from January 1990 to July 2001 retrieved from MEDLINE and CANCERLIT. RECOMMENDATIONS: * Routine use of HRT (either estrogen alone or estrogen plus progesterone) is not recommended for women who have had breast cancer. Randomized controlled trials are required to guide recommendations for this group of women. Women who have had breast cancer are at risk of recurrence and contralateral breast cancer. The potential effect of HRT on these outcomes in women with breast cancer has not been determined in methodologically sound studies. However, in animal and in vitro studies, the development and growth of breast cancer is known to be estrogen dependent. Given the demonstrated increased risk of breast cancer associated with HRT in women without a diagnosis of breast cancer, it is possible that the risk of recurrence and contralateral breast cancer associated with HRT in women with breast cancer could be of a similar magnitude. * Postmenopausal women with a previous diagnosis of breast cancer who request HRT should be encouraged to consider alternatives to HRT. If menopausal symptoms are particularly troublesome and do not respond to alternative approaches, a well-informed woman may choose to use HRT to control these symptoms after discussing the risks with her physician. In these circumstances, both the dose and the duration of treatment should be minimized. VALIDATION: Internal validation within the Steering Committee on Clinical Practice Guidelines for the Care and Treatment of Breast Cancer; no external validation. SPONSOR: The steering committee was convened by Health Canada. COMPLETION DATE: October 2001.

Animals↗