Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Evaluation of proscriptive health care policy implementation in screening mammography.

PURPOSE: To evaluate the potential effect of proscriptive health care policies directed toward improving screening mammogram interpretation in the United States. MATERIALS AND METHODS: Percentiles of accuracy based on a random sample of 110 U.S. radiologists were used to examine the number of radiologists who would need to be restricted from providing mammographic interpretation to increase median accuracy from 66% to 67%, 71%, and 76%. In addition, reading volume data recorded for the sampled readers were used to project the percentage reduction in service volume (mammograms per year) that would result from restriction. Characteristics of participating radiologists were compared with those of nonparticipating radiologists by using chi2 testing and analysis of variance to assess the external validity of the results. RESULTS: To increase median accuracy by 1% (from 66% to 67%) would require prohibiting about 2,200 U.S. radiologists (ie, the 11% in the lowest quantile for accuracy) from performing mammographic interpretation and would result in a reduction of yearly service volume of approximately 10%. An increase in median accuracy of 5% (to 71%) would require prohibiting about 6,000 U.S. radiologists (ie, 30%) from performing this service, with an accompanying volume reduction of 25%. An increase in median accuracy of 10% (to 77%) would require prohibiting about 11,400 practicing U.S. radiologists (ie, 57%) from performing this service and would diminish the national service capacity by 50%. CONCLUSION: These data show that implementation of proscriptive health care policies based on accuracy would diminish the service capacity of screening mammography in the United States.

Adult↗

[Low reporting of clinical characteristics of patients with diabetes mellitus included in the main clinical trials on hypertension].

BACKGROUND AND OBJECTIVE: Subjects with diabetes mellitus (DM) are at high risk of cardiovascular events. However, their cardiovascular risk varies according to several clinical characteristics like age, sex, ethnicity, type and duration of the DM, quality of glycemic control, type of hypoglycemic treatment, or the presence of nephropathy or previous cardiovascular events. It is not known if these characteristics have been sufficiently reported in clinical trials on hypertension including subjects with DM. MATERIAL AND METHOD: We analyzed randomized controlled clinical trials about treatment of hypertension published before May 2003 with the following characteristics: a) inclusion of subjects with DM and b) a primary end-point including at least one of the following events: myocardial infarction, cerebrovascular disease or total mortality. In these trials we evaluated the above-mentioned clinical characteristics concerning cardiovascular risk. RESULTS: Sixteen trials were eventually analyzed. These trials were classified into: a) trials designed only for subjects with DM (RENAAL, UKPDS 38, UKPDS 39, IDNT); b) trials with a subgroup analysis for subjects with DM (SHEP, Syst-Eur, MICRO-HOPE, HOT, STOP-2, LIFE), and c) trials without a subgroup analysis for subjects with DM (CAPP, NORDIL, INSIGHT, ALLHAT, SANBPSG, CONVINCE). All these trials included a total of 33984 subjects with DM, most of them > or = 55 years old. The percentage of women evaluated ranged from 33.5 to 71.9% although in 2 trials no data were found regarding this percentage (12.5%). There was no information on ethnicity in 10 trials (62.5%), on the type of DM in 7 (43.8%), on the duration of the disease in 13 (81.3%), on the degree of glycemic control in 10 (62.5%), on the type of hypoglycemic treatment in 11 (68.8%), on the presence or absence of nephropathy in 11 (68.8%) and on the presence or absence of previous cardiovascular events in 3 (18.8%). Trials without subgroup analysis for subjects with DM had a lower reporting of the clinical characteristics of the subjects evaluated than the rest of the trials. CONCLUSIONS: The reporting of relevant clinical characteristics of cardiovascular risk in subjects with DM in the main clinical trials on hypertension is low, diminishing their external validity. To optimize the therapeutic recommendations in the treatment of hypertension in DM these deficiencies should be overcome.

Adult↗

Clinical dimensions of delusional beliefs: a factor-analytic study.

We investigated the factorial composition as well as the demographic, anamnestic and clinical correlates of 13 features of delusional beliefs in a sample of 127 psychiatric inpatients with active delusions at the time of their assessment: Six factors were extracted, jointly accounting for 67.3% of the variance, which were interpreted as the dimensions of cognitive disintegration, volitional dyscontrol, doxastic strength, distress, incongruence and expansiveness, respectively. Moreover, most factors exhibited distinctive profiles of demographic, anamnestic or clinical correlates. Overall, our findings provide supportive evidence for both the clinical multidimensionality of delusional beliefs at the factor-analytic level and the - at least partial - external validity of the factorial solution obtained.

Adult↗

Effectiveness of a therapeutic community treatment in Spain: a long-term follow-up study.

In this paper, the effectiveness of the treatment program developed by Proyecto Hombre ('Project Mankind') in Asturias, Spain, is evaluated. In a long-term follow-up (range from 73 days to 8 years) with a sample of 249 subjects, the results obtained by subjects completing the treatment (194) were compared with pre-treatment results and with those of the group that dropped out (55). The measurements used were relapses in illegal drugs, alcohol, changes in family situation, educational level, employment, criminal involvement and state of health. External validation of self-report measures given in the questionnaire was carried out. Findings support the effectiveness of the treatment in all measures and the validity of self-report items. Relapse rate in 'treatment-completed' group was 10.3%, whilst in the non-completers group it reached 63.6% (significant difference, p < 0.001). Relapses of non-completers were more severe, occurred sooner after leaving the program (they stayed abstinent for shorter periods) and lasted longer than those of subjects completing the treatment.

Adolescent↗

Long-term follow-up of patients in four prospective studies of the German Breast Cancer Study Group (GBSG): A summary of key results.

Evaluation of treatment modalities and prognostic factors in early breast cancer requires a long-term follow-up of patients in prospective studies. The German Breast Cancer Study Group (GBSG) started four nationwide studies in 1983 in which a total of 2,746 patients have been enrolled and followed for about 10 years. Questions that have been addressed in these studies are still relevant today and comprise the role of breast-conserving therapy, the duration of adjuvant chemotherapy, and whether adjuvant radiotherapy is needed. The key results of these studies are highlighted including some important findings on prognosis, e.g., the role of isolated locoregional recurrence and the prognosis of patients with 10 or more positive lymph nodes. The data of all randomized patients were regularly included into the overviews of the Early Breast Cancer Trialists' Collaborative Group; the data of the nonrandomized patients have been used to examine the external validity of treatment comparisons. Overall, it can be concluded that the GBSG studies have made valuable and internationally recognized contributions to the prognosis and treatment of patients with early breast cancer.

Adult↗

Empirie und Dogma in den < > medizinischer Wissenschaft.

Empiric and Dogmatic 'Seasons' in the Progress of Medical ScienceThe current 'postmodern' times are not simply permissive towards any theory whatever. In medicine, the 'external validity' of theories and the exact place of various healing schools has become an important subject. 'Critical appraisal' of the empiric evidence for treatment benefit in all areas of medicine, conventional or unconventional, will separate the wheat from the chaff. Pluralistic evidence-based medicine constitutes the flail for achieving this task. This happens in late summer. Later in winter, the chaff of useless medicine will decay; a new therapeutic spring will only be possible when the current task of identifying effective and beneficial treatments in all areas and of abandoning useless therapies will be fulfilled. The current primacy of the empiric study of medicine, rather than the dogmatic, should not be misunderstood as a pure empirism. Rather, we are in the empiric 'season'. Questions about mechanisms of action will remain, but will retreat to the back of the mind for some time.

Journal Article↗

An example on the value of non-randomisation in clinical trials in complementary medicine.

BACKGROUND: Randomised clinical trials may in principle show a small external validity. Non-randomised clinical trials therefore are sometimes regarded as an appropriate alternative when complementary and conventional treatments are compared. OBJECTIVES: To assess the value of advanced statistical methods in the process of estimating differences between a complementary and a conventional treatment of acute sinusitis in a non-randomised clinical trial. METHODS: Multicentre, non-randomised, controlled clinical trial comparing 2 complementary and 3 conventional ENT centres. Patients were free to choose the physician (and hence the therapy). Treatment differences were estimated by controlling for confounders in analyses of covariance or by propensity score techniques. RESULTS: Most potential confounders (sex, age, life-style parameters) did not have significant effects on the choice of therapy. Disease severity and previous ENT surgery were the main confounding factors. At study onset they almost cause a defined separation of both treatment groups. As a result estimated treatment differences vary substantially depending on the chosen statistical model. CONCLUSIONS: When comparing complementary and conventional treatments, non-randomised clinical trials may be misleading. Results may be strongly biased even when advanced statistical methods are used. Trials of complex statistical designs are needed to give valid results.

Acute Disease↗

Epidemiological perspectives for the development of future diagnostic systems.

Three uses of epidemiological data for diagnostic revisions are suggested. (1) Psychometric analysis of community cases over the full spectrum of symptom severity can help determine whether true discrete disease entities underlie particular symptom profiles. (2) Analysis of external validators in community samples can help select optimal cut-points for defining diagnostic criteria. (3) Clinical epidemiological studies can provide quick and comparatively inexpensive preliminary checks, prior to carrying out expensive experimental treatment trials, on the plausibility of proposed subtyping distinctions based on nonexperimental evidence regarding differential treatment response. Correction of a number of basic conceptual and methodological flaws that have hampered previous studies would allow epidemiological research of these three sorts to play a much more important part in future diagnostic revisions than they have in past revisions.

Epidemiologic Research Design↗

Annual rate and predictors of conversion to dementia in subjects presenting mild cognitive impairment criteria defined according to a population-based study.

Elderly subjects diagnosed with mild cognitive impairment (MCI) are becoming the target of intervention trials. The criteria used for MCI are principally issued from prospective clinical studies, although longitudinal population-based studies having identified several cognitive predictors of dementia can be of great contribution in the definition of these criteria. This study was conducted to explore the external validity of MCI criteria issued from a longitudinal population-based study, and subsequently to identify the best predictors of the short-term conversion to Alzheimer's disease 2 years after the MCI diagnosis. Ninety elderly volunteers with memory complaint diagnosed with MCI on the basis of their functional and neuropsychological performances were followed up within 2 years. The potential predictors of the conversion to dementia collected at baseline included age, gender, educational level, size of temporal lobe, apolipoprotein E genotype and a series of neuropsychological measures (Mac Nair Scale, Mini-Mental State Examination, Benton Visual Retention Test, Isaacs Set Test, Digit Symbol Substitution Task, Letter Cancellation Task, digit span tasks and finger-tapping test). Within the 2 years, 29 subjects (32.2%) presented a conversion to dementia. The risk of conversion to dementia was associated with age and size of temporal lobe but not with gender, education, or apolipoprotein E4 genotype. Several neuropsychological measures were associated with the risk of conversion to dementia, but in a logistic regression performed with the significant variables found in the univariate analysis, only the Letter Cancellation Test was shown to be an independent predictor. In conclusion, the quite elevated conversion rates obtained show the usefulness, when defining MCI criteria, of considering not only memory impairment but also impairment in other cognitive areas, as well as mild impairment on higher-order activities of daily living. Among the variables considered, the Letter Cancellation Test proved to be a major predictor of short-term conversion to dementia.

Age Factors↗

Traditional Chinese medicine (phytotherapy): Health Technology Assessment report - selected aspects.

OBJECTIVE: A summary of main aspects from a Health Technology Assessment report on Traditional Chinese Medicine (TCM) in Switzerland concerning effectiveness and safety is given. MATERIALS AND METHODS: Literature search was performed through 13 databases, by scanning reference lists of articles and by contacting experts. Assessed were quality of documentation, internal and external validity. RESULTS: Effectiveness: 43 articles concerning 'gastrointestinal tract and liver' were assessed. The studies covering 7,436 patients were undertaken in China (35), Japan (3), USA (2) and Australia (3); 33/43 being controlled studies. 34/40 show significantly better results in the TCM-treated group. A comparison of studies on results of treatment based on a diagnosis according to TCM criteria and studies on results of treatment according to Western diagnosis shows that treatment based on TCM diagnosis improves the result. The comparison of treatment by individual medication and standard medication showed a trend in favor of individual medication. SAFETY: TCM training and practice for physicians in Switzerland are officially regulated. Side effects occur, but no severe effects have been registered up to now in Switzerland. TCM medicinals are imported; admission regulations are being installed. Problems due to production abroad, Internet trade, self-medication or admixtures are possible. CONCLUSION: The evaluation of the literature search provides evidence for a basic clinical effectiveness of TCM therapy. Severe side effects were not observed in Switzerland. Regulations for trading and use of medicinals prevent treatment risks. Further clinical studies in a Western context are required.

Drugs, Chinese Herbal↗

A categorical approach to depression by a three-dimensional system.

Depressed moods do not arrive already hallmarked 'endogenous'. Depressions do, however, arrive as perceivable phenomena and the domain of clinically important and conceptually connected items includes about 20 items. This domain can be subdivided into items relevant for the severity of the depression and items relevant for the diagnosis of depression. It has been argued that Rasch models rather than multivariate analysis are the adequate models when studying the internal consistency of rating scales. Applying the Rasch models, it was found that 10 items have an additive relationship for the severity of depression and that 5 items were additively related for endogenous depression, whereas five other items were additively related for reactive depression. Moreover, it was found that the dimension of severity should be transformed to three categories: no depression, minor depression and major depression. The two dimensions: endogenous and reactive depression, respectively, should be transformed to the categories: definite endogenous depression, combined endogenous and reactive depression, definite reactive depression, probable endogenous depression and probable reactive depression. These categories for severity and for the diagnosis of depression have a high internal consistency. However, we still need studies to verify the external validity of these concepts.

Depression↗

Delusional disorder and mood disorder: can they coexist?

DSM III-R acknowledges that delusional disorder and mood disturbance can coexist. The aim of the study is to analyze mood disturbances occurring within a delusional disorder. An external validator, as the increased familial risk of psychiatric disorder, is also added to clarify the relationship between the two diagnostic areas. We found a high frequency of mood disturbances in our patients (50.7%). We were able to identify a proportion of delusional patients affected by a recurrent form of mood disturbance (35.2%); in about 42% of these patients the onset of the mood disturbance preceded the onset of the delusional disorder by a considerable interval of time. Our hypothesis is that, in these cases, the observed mood disturbance could represent a codiagnosis of true mood disorder. This hypothesis is partly supported by familial data.

Adult↗

Positive/negative symptomatology and experimental measures of attention in schizophrenic patients.

In a search for an external validation of the negative syndrome construct and the attentional impairment item on Andreasen's scale, 49 unmedicated schizophrenic patients were administered the Span of Apprehension Test and a Continuous Performance Test with two levels of difficulty. This schizophrenic sample performed significantly more poorly on the attentional tests than a comparable group of 27 healthy control subjects. Depending on the difficulty of the test we found a number of significant correlations between the SANS composite score and the pertaining attentional impairment item on the one hand and experimental indices of attentional functioning on the other hand, which might corroborate the psychopathological assumptions. The implications of these results for further attempts to validate clinical concepts of the positive/negative dichotomy by experimental means will be discussed.

Adolescent↗

Methods for assessing positive and negative symptoms.

Once adequate reliability has been achieved, however, it is important to document internal consistency and external validity. This chapter has summarized some of our work with the former topic. It suggests that the SANS has high internal consistency, while the SAPS is somewhat less internally consistent. What are the clinical and research implications of these results concerning internal consistency? On a superficial and statistical level, these results might be considered to suggest that the SAPS is less valid than the SANS, and perhaps that therefore it is not useful. On the other hand, it might be argued that it is not necessary to measure all the items on the SANS, since they are highly intercorrelated with one another and any one might 'stand in place' for the others. While in a sense statistically correct, these conclusions are probably misleading in both clinical and research settings. The SANS and SAPS were designed primarily as descriptive instruments that are useful for encoding symptoms commonly observed in psychiatric patients. Essentially, these results document clinical common sense. Patients with one negative symptom tend to have several others, while patients with one positive symptom do not necessarily have others. In other words, affective flattening seems to be related to alogia and avolition, but delusions and hallucinations do not necessarily occur together. In a comprehensive description of an individual patient, it is important to document all the types of symptoms that are present. This is particularly crucial in pharmacologic studies, where one may wish to document that some symptoms or groups of symptoms are more responsive to treatment than are other symptoms. For example, although anhedonia and alogia are statistically correlated with one another in a population of schizophrenics, in an individual patient anhedonia might be more responsive to treatment with a specific medication than is alogia. In spite of the high intercorrelations, it remains important in clinical settings to evaluate all relevant symptoms. The results of the factor analyses in this second study do not suggest as clean and strong a separation between positive and negative symptoms as was indicated in our original study. Factor analysis is notoriously sample-dependent, but there is no reason to suspect that the sample in the second study was different in any way from that of the first. Both involved consecutive admissions of DSM-III schizophrenics to the Iowa Psychiatric Hospital. The individuals doing the clinical evaluation did change, however.(ABSTRACT TRUNCATED AT 400 WORDS)

Affective Symptoms↗

TIMI risk score for ST-elevation myocardial infarction: A convenient, bedside, clinical score for risk assessment at presentation: An intravenous nPA for treatment of infarcting myocardium early II trial substudy.

BACKGROUND: Considerable variability in mortality risk exists among patients with ST-elevation myocardial infarction (STEMI). Complex multivariable models identify independent predictors and quantify their relative contribution to mortality risk but are too cumbersome to be readily applied in clinical practice. METHODS AND RESULTS: We developed and evaluated a convenient bedside clinical risk score for predicting 30-day mortality at presentation of fibrinolytic-eligible patients with STEMI. The Thrombolysis in Myocardial Infarction (TIMI) risk score for STEMI was created as the simple arithmetic sum of independent predictors of mortality weighted according to the adjusted odds ratios from logistic regression analysis in the Intravenous nPA for Treatment of Infarcting Myocardium Early II trial (n=14 114). Mean 30-day mortality was 6.7%. Ten baseline variables, accounting for 97% of the predictive capacity of the multivariate model, constituted the TIMI risk score. The risk score showed a >40-fold graded increase in mortality, with scores ranging from 0 to >8 (P:<0.0001); mortality was <1% among patients with a score of 0. The prognostic discriminatory capacity of the TIMI risk score was comparable to the full multivariable model (c statistic 0. 779 versus 0.784). The prognostic performance of the risk score was stable over multiple time points (1 to 365 days). External validation in the TIMI 9 trial showed similar prognostic capacity (c statistic 0.746). CONCLUSIONS: The TIMI risk score for STEMI captures the majority of prognostic information offered by a full logistic regression model but is more readily used at the bedside. This risk assessment tool is likely to be clinically useful in the triage and management of fibrinolytic-eligible patients with STEMI.

Aged↗

Time period required for transcranial Doppler monitoring of embolic signals to predict recurrent risk of embolic transient ischemic attack and stroke from arterial stenosis.

BACKGROUND AND PURPOSE: We aimed to investigate whether the time period of transcranial Doppler monitoring for embolic signals can be reduced without loss of clinical yield compared with routinely performed 1-hour monitoring. METHODS: Investigations on the basis of a post hoc analysis of a previously published cohort of 86 patients (55 men, 31 women; mean age 60.6 years) with a nondisabling arterioembolic ischemic event in the anterior circulation within the last 30 days (mean 7.3) and an ipsilateral medium-grade or high-grade stenosis of the carotid or middle cerebral artery. Patients underwent 1-hour monitoring for embolic signals and were followed up prospectively for 6 weeks to evaluate the relationship between embolic signals and risk of an early ischemic recurrence. Risk was also calculated after fictitious reduction of the monitoring period from 60 minutes to 50, 40, 30, 20, and 10 minutes, respectively, and compared with the results obtained from the 1-hour period. RESULTS: The number of patients positive for embolic signals decreased with the decreasing monitoring period. By this, the odds ratio of embolic signals for an early ischemic recurrence "decreased" from 40 (derived from the 1-hour monitoring) to 10 when the monitoring lasted < or =30 minutes. The relationship between the rate of embolic signals per hour and risk of a recurrent stroke is described by an S-shaped curve. As a consequence, risk estimated from reduced monitoring periods can differ considerably from that derived from the 1-hour monitoring if the signal frequency lies within a medium range (eg, between 3 and 15 signals in 30 minutes). CONCLUSIONS: The time period of monitoring for embolic signals may be reduced without loss of clinical relevant information when signal frequency is low or already high during the reduced monitoring period, but it should be prolonged to maximally an hour at signal numbers within a medium range. However, our results need to be externally validated on an independent cohort of patients or confirmed by a prospective study before this modification can be recommended in general.

Aged↗

Comparison of six depression rating scales in geriatric stroke patients.

We compared three self-rating scales (the Geriatric Depression Scale, the Zung Scale, and the Center for Epidemiologic Studies Depression Scale) with three examiner-rating scales (the Hamilton Rating Scale, the Comprehensive Psychopathological Rating Scale-Depression, and the Cornell Scale), to see which was best for 40 elderly (mean age 80 years) stroke patients, 17 of whom were depressed according to clinical examination. External validity and concurrent validity were good for all except the Cornell Scale. Reliability (internal consistency) showed that some items were not significantly correlated, which might be explained by our selection of the patients. The Geriatric Depression Scale, the Zung Scale, and the Comprehensive Psychopathological Rating Scale-Depression had the highest sensitivity, and the Zung Scale had the highest positive predictive value (93%). With regard to internal consistency, sensitivity, and predictive value, the best self-rating scales were the Geriatric Depression and the Zung scales and the best examiner-rating scale was the Comprehensive Psychopathological Rating Scale-Depression.

Aged↗

Effects of intensity of rehabilitation after stroke. A research synthesis.

BACKGROUND AND PURPOSE: A research synthesis was performed to (1) critically review controlled studies evaluating effects of different intensities of stroke rehabilitation in terms of disabilities and impairments and (2) quantify patterns by calculating summary effect sizes. The influences of organizational setting of rehabilitation management, blind recording, and amount of rehabilitation on the summary effect sizes were calculated. METHODS: A Medline literature search was performed for a critical review of the literature. The internal and external validity of the studies was evaluated. In addition, a meta-analysis was performed by applying the fixed (Hedges's g) effects model. RESULTS: The effects of different intensities of rehabilitation were studied in nine controlled studies involving 1051 patients. Analysis of the methodological quality revealed scores varying from 14% to 47% of the maximum feasible score. Meta-analysis demonstrated a statistically significant summary effect size for activities of daily living (0.28 +/- 0.12). Lower summary effect sizes (0.19 +/- 0.17) were found for studies in which experimental and control groups were treated in the same setting compared with studies in which the two groups of patients were treated in different settings (0.40 +/- 0.19). Variables defined on a neuromuscular level (0.37 +/- 0.24) showed larger summary effect sizes than variables defined on a functional level (0.10 +/- 0.21). Weighting individual effect sizes for the difference in amount of rehabilitation between experimental and control groups resulted in larger summary effect sizes for activities of daily living and functional outcome parameters for studies that were not confounded by organizational setting. CONCLUSIONS: A small but statistically significant intensity-effect relationship in the rehabilitation of stroke patients was found. Insufficient contrast in the amount of rehabilitation between experimental and control conditions, organizational setting of rehabilitation management, lack of blinding procedures, and heterogeneity of patient characteristics were major confounding factors.

Activities of Daily Living↗