Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Using prognostic models in clinical infertility.

The chance that a couple who have tried to conceive for 12 months will succeed without assisted conception treatment is still higher than the chance that the same couple will benefit from treatment. In this context, it is important to assess the chance that a treatment-independent or 'spontaneous' pregnancy will occur in a couple whose wish for a child is unfulfilled. Prognostic models can be useful in this assessment. In recent years, prognostic models have been published both for the occurrence of 'spontaneous' pregnancy and for pregnancy after in vitro fertilization. This article discusses the theoretical aspects of prognostic modelling and assesses whether the current prognostic models are good enough to justify their use in clinical practice. The performance of existing models for the prediction of spontaneous conception was found to be acceptable on internal as well as on external validation. However, the performance of the existing models predicting IVF outcome was found to be disappointing on the few occasions on which such external validation has been performed.

Journal Article↗

Predictive validity of the strain index in manufacturing facilities.

The Strain Index is a job analysis method for determining if workers are exposed to increased risk of developing distal upper extremity disorders. Its predictive and external validity was initially demonstrated in a pork processing plant. The purpose of this study was to evaluate its predictive validity in two manufacturing plants. While blinded to health outcomes, investigators analyzed the right and left sides of 28 single-task jobs using the Strain Index and classified them as "hazardous" or "safe" based on the Strain Index score. Subsequently, OSHA 200 logs were used to ascertain the occurrence of distal upper extremity disorders retrospectively. If at least one such disorder occurred on the right or left side during the prior three years, that side was classified as "positive." If no such disorder was reported during the prior three years, that side was classified as "negative." When comparing sides, symmetry between morbidity and hazard classification was required. When comparing jobs, such symmetry was not required. Evidence of association between the hazard classifications and the morbidity classifications for the 56 sides and the 28 jobs was evaluated using 2 x 2 contingency tables. For the sides, the association between hazard classification and morbidity classification was statistically significant with an empirical odds ratio of 73.2. The sensitivity, specificity, positive predictive value, and negative predictive value were 1.00, 0.84, 0.47, and 1.00. Similar results were noted for the jobs--the empirical odds ratio was 106.6, and the sensitivity, specificity, positive predictive value, and negative predictive value were 1.00, 0.91, 0.75, and 1.00. While these results provide additional evidence of the Strain Index's external validity and predictive validity, it should be noted that these jobs involved the performance of single tasks.

Arm Injuries↗

Predictive validity of the Strain Index in turkey processing.

The Strain Index is a job analysis method for determining if workers are exposed to increased risk of developing distal upper extremity disorders. Its predictive and external validity was initially demonstrated in a pork processing plant. The purpose of this study was to evaluate the predictive validity of the Strain Index in one turkey processing plant. While blinded to health outcomes, investigators analyzed the right and left sides of workers in 28 jobs using the Strain Index and classified them as "hazardous" or "safe" based on the Strain Index score. Subsequently, OSHA 200 logs were used to ascertain the occurrence of distal upper extremity disorders retrospectively. If at least one such disorder had occurred on the right or left side during the previous 3 years, that side was classified as "positive." If no such disorder was reported during the previous 3 years, that side was classified as "negative." When comparing sides, symmetry between morbidity and hazard classification was required. When comparing jobs, such symmetry was not required. Evidence of association between the hazard classifications and the morbidity classifications for the 56 sides and the 28 jobs was evaluated using 2 x 2 contingency tables. For the sides, the association between hazard classification and morbidity classification was statistically significant, with an odds ratio of 22.0. The sensitivity, specificity, positive predictive value, and negative predictive value were 0.86, 0.79, 0.92, and 0.65, respectively. Similar results were noted for the jobs--the odds ratio was 50.0, and the sensitivity, specificity, positive predictive value, and negative predictive value were 0.91, 0.83, 0.95, and 0.71. These results provide additional evidence of the external validity and predictive validity of the Strain Index.

Animals↗

Development and validation of a risk score for somatic erectile dysfunction: combined results from three cross-sectional surveys.

OBJECTIVES: Some men with erectile dysfunction (ED) have difficulties discussing their condition with their physicians. Existing screening and diagnostic tools for ED often require the administration of personal questions regarding the condition. We present a simple risk score to estimate the individual likelihood of somatic ED, based on age and existing health conditions. METHODS: Data from the Cologne Male Survey (n = 4396) were used to develop a multivariable logistic regression model for the individual ED likelihood. The regression equation was both internally and externally validated using data from a national study (Berlin study) and a multinational cross-sectional study (MALES study). RESULTS: A final regression equation including age, pelvic surgery, diabetes mellitus, arterial circulatory disorder, heart disease, smoking, and hypertension reached an area under the receiver operating characteristic curve of 0.84 (0.5 means random and 1.0 perfect discrimination). Internal validation did not indicate any relevant overfit and the external validation results (national data: AUC = 0.75; multinational data: AUC = 0.67) are similar to those of other popular risk scores. CONCLUSIONS: The validated ED risk score developed from the regression equation can be used as a screening tool to identify patients who are at a high risk of somatic ED. This tool can facilitate entering into discussions between physicians and patients regarding erectile function.

Adult↗

Australian data and psychometric properties of the Strengths and Difficulties Questionnaire.

OBJECTIVE: We examine the Australian psychometric properties of the Strengths and Difficulties Questionnaire (SQD), a brief screening measure of behavioural and emotional problems in children and adolescents. METHOD: Using a large community sample (n = 1359) of young Australian children (4-9 years), we assessed the internal consistency, stability, and external validity of the parent-report SDQ. Normative data and cut-offs were also produced. RESULTS: Moderate to strong internal reliability was exhibited across all SDQ subscales, and support was found for the original five-factor structure of the measure. Adequate validity was evidenced in the relationship of these scales to one another, while correlations between the SDQ subscales, teacher ratings, and diagnostic interviews demonstrated sound external validity. SDQ total difficulties scores were associated with concurrent treatment status and scores over a 12-month period were stable. CONCLUSIONS: The current study of the SDQ with Australian children presents evidence of sound psychometric properties. Being the first study to empirically support the use of the SDQ in Australia, it is recommended that the youth and teacher-report forms of the measure receive similar attention in the future.

Adolescent↗

Development and validation of the Headache Needs Assessment (HANA) survey.

OBJECTIVE: To develop and validate a brief survey of migraine-related quality-of-life issues. The Headache Needs Assessment (HANA) questionnaire was designed to assess two dimensions of the chronic impact of migraine (frequency and bothersomeness). METHODS: Seven issues related to living with migraine were posed as ratings of frequency and bothersomeness. Validation studies were performed in a Web-based survey, a clinical trial responsiveness population, and a retest reliability population. Headache characteristics (eg, frequency, severity, and treatment), demographic information, and the Headache Disability Inventory were used for external validation. RESULTS: The HANA was completed in full by 994 adults in the Web survey, with a mean total score of 77.98 +/- 40.49 (range, 7 to 175). There were no floor or ceiling effects. The HANA met the standards for validity with internal consistency reliability (Cronbach alpha =.92, eigenvalue for the single factor = 4.8, and test-retest reliability = 0.77). External validity showed a high correlation between HANA and Headache Disability Inventory total scores (0.73, P<.0001), and high correlations with disease and treatment characteristics. CONCLUSIONS: These data demonstrate the psychometric properties of the HANA. The brief questionnaire may be a useful screening tool to evaluate the impact of migraine on individuals. The two-dimensional approach to patient-reported quality of life allows individuals to weight the impact of both frequency and bothersomeness of chronic migraines on multiple aspects of daily life.

Activities of Daily Living↗

Validity of the sleep subscale of the Diagnostic Assessment for the Severely Handicapped-II (DASH-II).

Currently there are no available sleep disorder measures for individuals with severe and profound intellectual disability. We, therefore, attempted to establish the external validity of the Diagnostic Assessment for the Severely Handicapped-II (DASH-II) sleep subscale by comparing daily observational sleep data with the responses of direct care staff to the sleep subscale of the DASH-II. Participants included 25 individuals with severe intellectual disability and 25 individuals with profound intellectual disability who reside in a large developmental center in central Louisiana. Four of the five items of the DASH-II were shown to have external validity. Implications of these findings for future research and practice are discussed.

Child↗

Early rheumatoid arthritis: toward tailor-made therapy.

Therapeutic possibilities for the treatment of early rheumatoid arthritis (RA) have expanded largely. New treatment modalities appear very effective with respect to relevant outcomes, such as radiographic progression. At the same time, the costs of disease-modifying antirheumatic drugs (DMARDs) have exponentially increased so that--given the rather high prevalence of RA--cost may become a limiting factor in the treatment of patients with RA. Therefore, there is a need to define the profile of those patients that should be treated with the most effective, and, unfortunately, the most costly, DMARDs. The authors describe herewith the heterogeneity of RA with respect to its most important outcomes, as well as the inability to predict those outcomes appropriately at the individual patient level. This heterogeneity of RA is not acknowledged in the modern landmark clinical trials that the authors base therapeutic decisions on, and the external validity of those trials is at stake. In this article, the authors discuss the consequences of the heterogeneity of RA in light of the perceived lack of external validity of evidence-generating landmark trials. The authors propose the following solutions to overcome this discrepancy: 1) earlier recognition of RA, and 2) appropriate prediction of treatment efficacy, because the most challenging scientific efforts may be taken in the near future in order to arrive at a tailor-made therapy for every individual presenting with RA.

Antirheumatic Agents↗

A European's perspective of COX-2 drug safety.

Regulators require less evidence to take action on drug safety concerns than would be required to grant drug efficacy claims. Such an asymmetrical view of risk points to the need for large studies to judge the safety of novel medicines compared with standard therapy as smaller sample sizes might produce unreliable and misleading toxicity signals. Large studies conducted in the classical manner are expensive and difficult to conduct. In addition, tight entry criteria and strictly protocolized care results in such trials having poor external validity. If the safety of medicines is to be judged efficiently, novel, easier to conduct, and less-expensive solutions are required. Such trials could have improved external validity were they incorporated into normal care. This paper discusses the safety of cyclooxygenase-2 inhibitors and describes the sort of studies that could be carried out to gather better safety information on their use versus standard nonsteroidal anti-inflammatory drugs.

Anti-Inflammatory Agents, Non-Steroidal↗

Clinical prediction model to characterize pulmonary nodules: validation and added value of 18F-fluorodeoxyglucose positron emission tomography.

BACKGROUND: The added value of 18F-fluorodeoxyglucose (FDG) positron emission tomography (PET) scanning as a function of pretest risk assessment in indeterminate pulmonary nodules is still unclear. OBJECTIVE: To obtain an external validation of the prediction model according to Swensen and colleagues, and to quantify the potential added value of FDG-PET scanning as a function of its operating characteristics in relation to this prediction model, in a population of patients with radiologically indeterminate pulmonary nodules. DESIGN, SETTING, AND PATIENTS: Between August 1997 and March 2001, all patients with an indeterminate solitary pulmonary nodule who had been referred for FDG-PET scanning were retrospectively identified from the database of the PET center at the VU University Medical Center. RESULTS: One hundred six patients were eligible for the study, and 61 patients (57%) proved to have malignant nodules. The goodness-of-fit statistic for the model (according to Swensen) indicated that the observed proportion of malignancies did not differ from the predicted proportion (p = 0.46). PET scan results, which were classified using the 4-point intensity scale reading, yielded an area under the evaluated receiver operating characteristic curve of 0.88 (95% confidence interval [CI], 0.77 to 0.91). The estimated difference of 0.095 (95% CI, -0.003 to 0.193) between the PET scan results classified using the 4-point intensity scale reading and the area under the curve (AUC) from the Swensen prediction was not significant (p = 0.058). The PET scan results, when added to the predicted probability calculated by the Swensen model, improves the AUC by 13.6% (95% CI, 6 to 21; p = 0.0003). CONCLUSION: The clinical prediction model of Swensen et al was proven to have external validity. However, especially in the lower range of its estimates, the model may underestimate the actual probability of malignancy. The combination of visually read FDG-PET scans and pretest factors appears to yield the best accuracy.

Aged↗

[Validity in a community (with outside verification) of primary prevention studies on hypercholesterolemia].

OBJECTIVES: The main objective of this study was to determine the degree of similarity between large primary prevention trials of hypercholesterolemia and our population of patients with dyslipidemia, in order to evaluate the external validity of these studies and their applicability to the general population. DESIGN: Descriptive retrospective study. SETTING: Tafalla Health Center in Navarra (Northern Spain), serving a population of 11 500 inhabitants.Participants. All patients older than 18 years assigned to our health center who had dyslipidemia with no antecedents of ischemic heart disease. RESULTS: The percentage of patients in our sample who satisfied the inclusion criteria used in large clinical trials ranged from 2.4% to 46%, depending on the study: AFCAPS/TexCAPS 1998, 46.2%; HPS 2002, 46.1%; WOSCOPS 1995, 10.9%; HHS 1987, 10.6%; LRC-CPPT 1984, 2.4%. CONCLUSIONS: Many of our patients (54%-97%) with dyslipidemia would not have been eligible for inclusion in earlier studies of hyperlipidemia and primary prevention. The external validity (applicability to the general population) of these studies is questionable. Decision-making in clinical practice for the primary prevention of hypercholesterolemia should be based on the risk/benefit ratio of pharmacological treatment.

Adult↗

Validation of a prediction model for the follicle-stimulating hormone response dose in women with polycystic ovary syndrome.

OBJECTIVE: To validate a published model for the prediction of the individual FSH response dose for gonadotropin induction of ovulation in polycystic ovary syndrome (PCOS). DESIGN: Structured, complete, and carefully monitored patient-based data collection to test the external validity of the prediction model. SETTING: Twenty-nine hospitals in The Netherlands. PATIENT(S): Eighty-five clomiphene citrate (CC)-resistant women with PCOS. INTERVENTION(S): Ovulation induction in a chronic low-dose step-up FSH regimen. MAIN OUTCOME MEASURE(S): Predicted individual FSH response dose, defined as follicle growth >10 mm in diameter on ultrasound. RESULT(S): The model, using the women's body mass index, CC response, initial serum FSH level, and initial serum insulin-to-glucose ratio was studied in the validation sample. Overall, the FSH response dose predicted by the model was higher than the observed response dose. The predictive performance of the model was poor, with an R(2) of 0.11, and the average prediction error was 35 IU. CONCLUSION(S): The external validity of the model predicting the individual FSH response dose was inadequate in women with CC-resistant PCOS undergoing ovulation induction with recombinant FSH in a low-dose step-up regimen.

Adult↗

Subject attrition in prevention research.

Subject attrition threatens the internal validity of substance abuse prevention studies because differences in the rate of attrition and the substance use behavior of remaining subjects in the different conditions could account for any differences found in substance use rates. Attrition threatens the external validity of prevention studies because, to the extent that study dropouts are different from remaining subjects, the results of the study may not be generalizable to study dropouts. Analysis of these threats to the validity of prevention studies should be routinely conducted. However, studies of alcohol and drug abuse prevention have generally failed to report or analyze subject attrition. Smoking prevention studies have more frequently reported attrition, and they have recently begun to analyze the degree to which attrition may affect the internal and external validity of the study. Evidence thus far suggests that differences in attrition across conditions do occur occasionally. The evidence is substantial that study dropouts are systematically more likely to smoke, to use other substances, and to score highly on other risk-taking measures.

Alcoholism↗

The 'number needed to sample' in primary care research. Comparison of two primary care sampling frames for chronic back pain.

BACKGROUND: Sampling for primary care research must strike a balance between efficiency and external validity. For most conditions, even a large population sample will yield a small number of cases, yet other sampling techniques risk problems with extrapolation of findings. OBJECTIVE: To compare the efficiency and external validity of two sampling methods for both an intervention study and epidemiological research in primary care--a convenience sample and a general population sample--comparing the response and follow-up rates, the demographic and clinical characteristics of each sample, and calculating the 'number needed to sample' (NNS) for a hypothetical randomized controlled trial. METHODS: In 1996, we selected two random samples of adults from 29 general practices in Grampian, for an epidemiological study of chronic pain. One sample of 4175 was identified by an electronic questionnaire that listed patients receiving regular analgesic prescriptions--the 'repeat prescription sample'. The other sample of 5036 was identified from all patients on practice lists--the 'general population sample'. Questionnaires, including demographic, pain and general health measures, were sent to all. A similar follow-up questionnaire was sent in 2000 to all those agreeing to participate in further research. We identified a potential group of subjects for a hypothetical trial in primary care based on a recently published trial (those aged 25-64, with severe chronic back pain, willing to participate in further research). RESULTS: The repeat prescription sample produced better response rates than the general sample overall (86% compared with 82%, P < 0.001), from both genders and from the oldest and youngest age groups. The NNS using convenience sampling was 10 for each member of the final potential trial sample, compared with 55 using general population sampling. There were important differences between the samples in age, marital and employment status, social class and educational level. However, among the potential trial sample, there were no demographic differences. Those from the repeat prescription sample had poorer indices than the general population sample in all pain and health measures. CONCLUSIONS: The repeat prescription sampling method was approximately five times more efficient than the general population method. However demographic and clinical differences in the repeat prescription sample might hamper extrapolation of findings to the general population, particularly in an epidemiological study, and demonstrate that simple comparison with age and gender of the target population is insufficient.

Adult↗

Validation of the inflammatory bowel disease questionnaire in Swedish patients with ulcerative colitis.

BACKGROUND: The Inflammatory Bowel Disease Questionnaire (IBDQ) is a disease-specific health-related quality of life (HRQOL) questionnaire including four dimensions and a sum score. The aim of this study was to assess the internal and external validity, reliability, and sensitivity of a Swedish version of the IBDQ. METHODS: Three hundred consecutive patients with ulcerative colitis completed the IBDQ and three other health-related quality of life questionnaires (the Rating Form of IBD Patient Concerns (RFIPC), the Short Form-36 (SF-36) and the Psychological General Well-Being (PGWB) index). Disease activity was evaluated using a 1-week symptom diary, blood tests and rigid sigmoidoscopy. One hundred and fourteen patients filled in the questionnaire a second time, of whom 75 had been in stable remission for over 6 months and 39 had a significant clinical change in disease activity. RESULTS: Factor analysis of the 32 IBDQ items did not support the four dimensional scores. The dimensional scores had sufficient convergent validity, but low discriminative validity and homogeneity. The homogeneity was also low for the sum score. The inter-dimensional correlations were high. The concurrent validity was supported by correlations between the dimensional scores and other measures of disease activity and HRQOL. Patients in relapse scored significantly less on the sum score and the four dimensions compared to patients in remission. The test-retest correlations for the dimensional scores were 0.40-0.76. Patients with a change in disease activity during the 6-month follow-up period had a significant change in IBDQ scores not found in those who remained in remission. CONCLUSIONS: The Swedish version of the IBDQ had external validity and was shown to be a reliable and sensitive measure of HRQOL in ulcerative colitis, though there are some concerns regarding the internal validity. The use of a sum score was not supported and the questionnaire may benefit from a redivision of items into dimensions with better homogeneity and discriminative validity.

Colitis, Ulcerative↗

Research progress and application prospects of multi-omics integration strategies in precision risk stratification of type 1 diabetes mellitus.

Type 1 diabetes (T1D) is a chronic metabolic disease mediated by autoimmunity. Its pathogenesis involves complex interactions between genetic susceptibility and environmental factors. Conventional T1D risk stratification primarily relies on genetic markers, islet autoantibodies, and glycemic indicators. Although these biomarkers remain indispensable in current clinical practice, they are often insufficient when used alone to accurately identify ultra-early high-risk individuals, predict disease progression rates, or support individualized preventive strategies. Consequently, more comprehensive molecular approaches are needed to improve precision risk stratification. In recent years, the rapid development of multi-omics technologies has provided new strategies for precise risk stratification of T1D. This narrative review critically evaluates how multi-omics integration strategies can improve precision risk stratification throughout the T1D disease continuum by integrating complementary molecular information from genomics, transcriptomics, proteomics, metabolomics, epigenomics, and the microbiome. Particular emphasis is placed on stage-specific biomarker discovery, multi-omics data integration frameworks, artificial intelligence-assisted prediction models, biomarker validation, and the opportunities and challenges associated with clinical translation. Current evidence suggests that integrated multi-omics approaches have the potential to improve risk prediction accuracy, distinguish heterogeneous disease trajectories, identify individuals at imminent risk of progression, and provide biologically informed targets for precision intervention. However, important challenges remain, including data harmonization, external validation, model interpretability, cost-effectiveness, and integration into routine clinical screening programs. Future research should prioritize prospective multicenter cohorts, standardized analytical pipelines, externally validated prediction models, and clinically interpretable multi-omics frameworks to facilitate the translation of precision risk stratification into routine T1D prevention and management.

Humans↗

Cyclophosphamide versus methylprednisolone for the treatment of neuropsychiatric involvement in systemic lupus erythematosus.

BACKGROUND: Neuropsychiatric involvement in systemic lupus erythematosus is complex and several clinical presentations are related to this disease such as: convulsions, chronic headache, transverse myelitis, vascular brain disease, psychosis and neural cognitive dysfunction. OBJECTIVES: To assess the efficacy and safety of cyclophosphamide and methylprednisolone in the treatment of neuropsychiatric manifestations of systemic lupus erythematosus on mortality and side effects. SEARCH STRATEGY: We searched EMBASE, LILACS, Cochrane Controlled Trials Register and MEDLINE up to and including December 1999, additional articles were sought through handsearching in relevant journals, using the search strategy described in the Cochrane Handbook [Dickersin 1994]. There were no language restrictions. SELECTION CRITERIA: All randomized controlled trials which compared cyclophosphamide to methylprednisolone were to be included. Patients of any age and gender were included if they fulfilled the criterion of the American Rheumatology Association for the diagnosis of systemic lupus erythematosus and presented with any one of the following neuropsychiatric events; convulsions, organic brain syndrome; cranial neuropathy. Outcome measures included the following: a) Overall mortality (primary event); b) Motor and psychiatric deficit (primary event); c) Clinical improvement (secondary event). DATA COLLECTION AND ANALYSIS: The analysis planned was to do the following: Data would be independently extracted by the two reviewers and cross-checked. The methodological quality of each trial would be assessed by the same two reviewers. Details of the randomisation (generation and concealment), blinding, and the number of patients lost on follow-up would be recorded. The results of each RCT would be summarised on an intention-to-treat basis in 2 x 2 tables for each outcome. External validity would be defined by characteristics of the participants, the interventions and the outcomes. If appropriate, RCTs would be stratified based on control group and category of disease in accordance to the clinical homogeneity (external validity). The results obtained from these different methods are very similar, and therefore, only the results from the Risk Difference method, with the corresponding 95% confidence interval would be presented in this review. The fixed effects model would be used if there was no significant statistical heterogeneity. MAIN RESULTS: We found no randomised controlled trials comparing cyclophosphamide versus methylprednisolone for the treatment of neuropsychiatric involvement in the systemic lupus erythematosus. REVIEWER'S CONCLUSIONS: Cyclophosphamide regimen treatment is a form of care in neuropsychiatric involvement in systemic lupus erythematosus with no evidence to prove better effectiveness and safety when compared with methylprednisolone. This systematic review found no randomised controlled trials and its findings must be interpreted as 'no evidence of effect' and not as 'evidence of no effect'.

Antirheumatic Agents↗

Ensemble DNA methylation clock demonstrates Immune-metabolic aging signatures associated with mortality.

Aging is a multifactorial process that is best described in terms of the progressive acquisition of multiple layers of phenotypic changes, such as epigenetic modifications, inflammation, and metabolic dysregulation. DNA methylation clocks have been extensively used to construct epigenetic clocks based on the DNAm profiles that can be used to estimate biological age and predict age-associated outcomes. Nevertheless, the vast majority of clocks constructed so far have been based on linear models, which are unlikely to fully account for the heterogeneity and non-linearity of survival-related DNAm signatures. In this work, we constructed a heterogeneous stacked ensemble survival model based on DNAm data obtained from the Framingham Heart Study. We first identified 190 CpG loci using elastic net Cox regression and subsequently constructed a survival prediction model based on the fusion of five complementary survival models by means of a neural network meta-learner. The prediction power of the survival model was evaluated in an external validation cohort, where we observed strong performance for predicting all-cause mortality that significantly exceeded PhenoAge and was statistically comparable to GrimAge. These performance estimates were derived in cohorts of European ancestry and externally validated in postmenopausal women aged 50-79 years, and should therefore be interpreted as applicable only to demographically similar populations.

Humans↗