Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,225 records · Page 68Linked to original sources

Childhood inattention-overactivity, aggression, and stimulant medication history as predictors of young adult outcomes.

This study examined the contributions of childhood symptom dimensions and aspects of methylphenidate (MPH) treatment to the prediction of young adult outcomes in boys who were referred to a child psychiatry outpatient clinic. They were diagnosed with hyperkinetic reaction of childhood/minimal brain dysfunction, and given MPH for an average of 30 months. Including significant effects and statistical trends, childhood Inattention-Over-activity was uniquely associated with fewer than 10% of adult outcomes such as schizotypic features, impairment on the Global Assessment Scale (GAS), and unemployment. Childhood aggression was uniquely associated with 38% of adult outcomes such as lifetime diagnoses of major depression, drug abuse disorder, and antisocial personality disorder; MMPI PD, PA, and SC scores; and six additional measures of adult impairment and life circumstances-extending external validation of the two-factor model to young adulthood. For 20 young adult outcomes (63%), aspects of childhood treatment with MPH had no lasting effects. For one adult outcome (3%), a lasting negative effect of childhood drug treatment was found; better initial response to medication was associated with not graduating from high school. For 11 young adult outcomes (34%), however, aspects of childhood MPH treatment had positive effects that lasted long after treatment was discontinued. Higher dosage was associated with fewer diagnoses of alcoholism or suicide attempts. Better response to medication was associated with lower MMPI D scores and better social functioning. Longer medication duration was associated with fewer schizotypic features, lower MMPI MA scores, higher WAIS Performance and Full Scale IQs, and better WRAT Reading and Arithmetic performance.

Adult↗

Absolute fat mass, percent body fat, and body-fat distribution: which is the real determinant of blood pressure and serum glucose?

Associations of body mass index (BMI), absolute fat mass, percent body fat, and regional fat distribution with concentrations of fasting blood glucose and blood pressure were examined cross-sectionally in 1551 men and women aged 15-79 y from two study centers. Measurements included height, weight, multiple skinfold thicknesses, body density by underwater weighing, and waist and hip girths. Three principal findings emerged: 1) Absolute overall body mass and fat mass were stronger predictors of blood pressure and blood glucose than were relative fat mass, after age, height, and current cigarette-smoking status were adjusted for; 2) when diastolic blood pressure and serum glucose were used as the external validity criteria, densitometry was not a "gold standard" for body composition associated with risk for increased blood pressure and serum glucose; and 3) BMI was as good a predictor of blood pressure and glucose as was any other measure of body fat in nearly all analyses.

Adipose Tissue↗

Prediction of body cell mass, fat-free mass, and total body water with bioelectrical impedance analysis: effects of race, sex, and disease.

The inability to precisely estimate body composition with simple, inexpensive, and easily applied techniques is an impediment to clinical investigations in nutrition. In this study, predictive equations for body cell mass (BCM), fat-free mass (FFM), and total body water (TBW) were derived from direct measurements through use of single-frequency bioelectrical impedance analysis (BIA) in 332 subjects, including white, black, and Hispanic men and women, who were both healthy control subjects and patients infected with the human immunodeficiency virus (HIV). Preliminary studies showed more accurate predictions of BCM when parallel-transformed values of reactance were used rather than the values reported by the bioelectrical impedance analyzer. Modeling equations derived after logarithmic transformation of height, reactance, and impedance were more accurate predictors than equations using height2/resistance, and the use of sex-specific equations further improved accuracy. The effect of adding weight to the modeling equation was less important than the BIA measurements. The resulting equations were validated internally, and race and disease (HIV infection) were shown not to affect the predictions. The equation for FFM was validated externally against results derived from hydrodensitometry in 440 healthy individuals; the SEE was < 5%. These results indicate that body composition can be estimated with simple and easily applied techniques, and that the estimates are sufficiently precise for use in clinical investigation and practice.

Adult↗

Quality of reporting of observational longitudinal research.

Observational longitudinal research is particularly useful for assessing etiology and prognosis and for providing evidence for clinical decision making. However, there are no structured reporting requirements for studies of this design to assist authors, editors, and readers. The authors developed and tested a checklist of criteria related to threats to the internal and external validity of observational longitudinal studies. The checklist criteria concerned recruitment, data collection, biases, and data analysis and descriptive issues relevant to study rationale, study population, and generalizability. Two raters independently assessed 49 randomly selected articles describing stroke research published from 1999 to 2003 in six journals: American Journal of Epidemiology, Journal of Epidemiology and Community Health, Stroke, Annals of Neurology, Archives of Physical Medicine and Rehabilitation, and American Journal of Physical Medicine and Rehabilitation. On average, 17 of the 33 checklist criteria were reported. Criteria describing the study design were better reported than those related to internal validity. No relation was found between study type (etiologic or prognostic) or word count and quality of reporting. A flow diagram for summarizing participant flow through a study was developed. Editors and authors should consider using a checklist and flow diagram when reporting on observational longitudinal research.

Epidemiology↗

Motivational intervention: an individual counselling vs a group treatment approach for alcohol-dependent in-patients.

AIMS: The present study aimed to evaluate whether individual counselling for alcohol-dependent patients in three sessions is as effective as a 2-week group treatment programme as part of an in-patient stay in a psychiatric hospital which was to foster motivation to seek further help and to strengthen the motivation to stay sober. Of particular importance was the external validity of the results, i.e. a 'normal' intake load of in-patients in detoxification and a wide variety of motivation to stop drinking were to be investigated. METHODS: Subjects eligible for the study were all patients with alcohol problems admitted to a psychiatric hospital, but without psychosis, as the main diagnosis, and with a maximum of 10 detoxification treatments in the past. A randomized-controlled trial was conducted with 161 alcohol-dependent in-patients who received three individual counselling sessions on their ward in addition to detoxification treatment and 161 in-patients who received 2 weeks of in-patient treatment and four out-patient group sessions in addition to detoxification. Both interventions followed the principles and strategies of motivational interviewing. RESULTS: Six months after intervention, group-treatment patients showed a higher rate of participation in self-help groups; however, this difference had disappeared 12 months after treatment. The abstinence rate among the former patients did not differ between the two intervention groups. CONCLUSION: Group treatment may lead to a higher rate of participation in self-help groups, but does not increase the abstinence rate 6 months after treatment.

Alcoholism↗

Evaluation of the HSE COSHH Essentials exposure predictive model on the basis of BAuA field studies and existing substances exposure data.

This paper presents an in-house BAuA study on the evaluation of the COSHH Essentials exposure predictive model. External validation is based on measurement data obtained in BAuA field studies performed in various industries, e.g. printing industry and textile industry. In addition, measurement data and information on industrial hygiene provided by the chemical industry within the framework of the Existing Substances Risk Assessment programme are used. Although the evaluated exposure data cover a wide variety of activities and workplace scenarios, there is still a considerable lack of appropriate exposure data, especially for the more stringent control strategies. It was found that the level of agreement between the measurements for solid substances (powders, dusts) and the predicted ranges is reasonably good. The situation is in part different for liquids. In workplaces where organic solvents are used in litre quantities, exposure levels are within the predicted ranges or are often lower. For small-scale uses of liquids (millilitre scale), e.g. in carpenters' workshops, there were indications that the exposure levels can exceed the predicted ranges. However, it must be noted that the database is rather small.

Hazardous Substances↗

Determinants of dermal exposure relevant for exposure modelling in regulatory risk assessment.

Risk assessment of chemicals requires assessment of the exposure levels of workers. In the absence of adequate specific measured data, models are often used to estimate exposure levels. For dermal exposure only a few models exist, which are not validated externally. In the scope of a large European research programme, an analysis of potential dermal exposure determinants was made based on the available studies and models and on the expert judgement of the authors of this publication. Only a few potential determinants appear to have been studied in depth. Several studies have included clusters of determinants into vaguely defined parameters, such as 'task' or 'cleaning and maintenance of clothing'. Other studies include several highly correlated parameters, such as 'amount of product handled', 'duration of task' and 'area treated', and separation of these parameters to study their individual influence is not possible. However, based on the available information, a number of determinants could clearly be defined as proven or highly plausible determinants of dermal exposure in one or more exposure situation. This information was combined with expert judgement on the scientific plausibility of the influence of parameters that have not been extensively studied and on the possibilities to gather relevant information during a risk assessment process. The result of this effort is a list of determinants relevant for dermal exposure models in the scope of regulatory risk assessment. The determinants have been divided into the major categories 'substance and product characteristics', 'task done by the worker', 'process technique and equipment', 'exposure control measures', 'worker characteristics and habits' and 'area and situation'. To account for the complex nature of the dermal exposure processes, a further subdivision was made into the three major processes 'direct contact', 'surface contact' and 'deposition'.

Humans↗

A leakage-aware genomic prediction pipeline for meropenem resistance in Klebsiella pneumoniae using transformer-based resistome representation learning.

MOTIVATION: Antimicrobial resistance (AMR) in Klebsiella pneumoniae, particularly to carbapenems such as meropenem, is a major global health problem. Machine learning is increasingly used to predict resistance from genomic markers; however, many models fail to capture high-level gene-gene interactions and may exhibit inflated performance due to lineage-biased prediction. Existing genomic prediction models largely rely on flat feature representations that fail to capture epistatic gene interactions, and commonly suffer from inflated performance estimates due to phylogenetic data leakage. To address these limitations simultaneously, a leakage-aware hybrid TabTransformer-CatBoost pipeline was developed, combining self-attention-based resistome representation learning with gradient boosting classification under clade-aware data partitioning. A self-attention encoder converts sparse gene presence-absence profiles into contextualized latent embeddings, which are subsequently classified using gradient boosting to capture lineage-aware AMR patterns. RESULTS: The proposed architecture outperformed classical baselines including Logistic Regression, Random Forest, XGBoost, and optimized CatBoost models. Internal accuracy reached 92.59% for the Chained Hybrid configuration (area under the receiver operating characteristic curve, AUROC = 0.8670, F1&#x2009;=&#x2009;0.8537). Performance gains primarily originated from the embedding stage, as confirmed by ablation analysis. External validation across independent multinational cohorts (n&#x2009;=&#x2009;305) demonstrated generalizability (AUROC = 0.8105; F1&#x2009;=&#x2009;0.7552). Permutation testing produced near-zero Matthews Correlation Coefficient (MCC)&#x2009;=&#x2009;0.0091, indicating predictions reflect genuine biological signal rather than noise. These results establish attention-based genomic embedding with gradient boosting as a scalable, interpretable, and leakage-aware framework for clinical AMR prediction. AVAILABILITY AND IMPLEMENTATION: The source code for the TabTransformer-CatBoost framework, including preprocessing pipelines and pre-trained embeddings, is available at https://github.com/SibelKervanci/kp-meropenem-tabtransformer.

Journal Article↗

Kernel hierarchical gene clustering from microarray expression data.

MOTIVATION: Unsupervised analysis of microarray gene expression data attempts to find biologically significant patterns within a given collection of expression measurements. For example, hierarchical clustering can be applied to expression profiles of genes across multiple experiments, identifying groups of genes that share similar expression profiles. Previous work using the support vector machine supervised learning algorithm with microarray data suggests that higher-order features, such as pairwise and tertiary correlations across multiple experiments, may provide significant benefit in learning to recognize classes of co-expressed genes. RESULTS: We describe a generalization of the hierarchical clustering algorithm that efficiently incorporates these higher-order features by using a kernel function to map the data into a high-dimensional feature space. We then evaluate the utility of the kernel hierarchical clustering algorithm using both internal and external validation. The experiments demonstrate that the kernel representation itself is insufficient to provide improved clustering performance. We conclude that mapping gene expression data into a high-dimensional feature space is only a good idea when combined with a learning algorithm, such as the support vector machine that does not suffer from the curse of dimensionality. AVAILABILITY: Supplementary data at www.cs.columbia.edu/compbio/hiclust. Software source code available by request.

Algorithms↗

Assessment of depression in patients with chronic fatigue syndrome.

Assessment of the relationship of depression to chronic fatigue syndrome (CFS) is a complicated but important topic. This relationship may range from the misdiagnostic (i.e., depression misidentified as CFS) to the etiologic (i.e., CFS causes an organic affective syndrome). Assessment should focus on the symptoms and syndromes of depressive disorder, utilization of a single rating scale to assess presumed depression is discouraged, and alternate approaches to classification that allow for symptomatic overlap of a major depressive disorder and CFS are suggested. Careful attention needs to be given to the use of external validating criteria in empiric studies, such as natural history, clinical course (including treatment response), and family history.

Depression↗

The usefulness of area-based socioeconomic measures to monitor social inequalities in health in Southern Europe.

BACKGROUND: The study objective was to investigate the association between health outcomes and several small-area-based socioeconomic measures and also with individual socioeconomic measures as a check on external validity. METHODS: Cross-sectional design based on the analysis of the Barcelona Health Interview Survey of 1992. A representative stratified sample of the non-institutionalised population resident in Barcelona city (Spain) was obtained. The present study refers to the 4171 respondents aged over 14. We studied perceived health status, presence of chronic conditions and smoking as health outcomes. Area socioeconomic measures (1991 census) were generated at census tract level and individual socioeconomic measures were educational level and social class obtained through the survey. RESULTS: With individual socioeconomic measures we observed that the lower the educational level or social class, the higher the probability of reporting a perceived health status of fair, poor or very poor and of presenting some chronic condition. With regard to smoking, among men this trend was similar [odds ratio (OR) = 1.5; 95% confidence interval (CI) = 1.2-1.9 in social classes IV-V with respect to social classes I-II], while among women it was reversed (OR = 0.7; 95% CI = 0.5-0.9). With the different area-based socioeconomic indicators differences were also observed in this sense, with the exception of smoking in women for which these indicators do not show any differences by socioeconomic level. CONCLUSIONS: With several census area-based socioeconomic measures similar effects on inequalities in health have been observed. In general, these inequalities were in the same sense as those obtained with individual-based measures. Small-area-based socioeconomic measures from the Spanish census could greatly enhance analysis of social inequalities in health, overcoming the absence of socioeconomic data in public health registries and in medical records.

Adolescent↗

Using video-recorded consultations for research in primary care: advantages and limitations.

BACKGROUND: Video-recording primary care consultations is an established technique for primary care research. Despite the widespread use of video-recording to help answer a variety of research questions, little is known about how this recording technique influences the findings of studies in which it is employed. OBJECTIVE: This article investigates how video-recorded consultations have been used in research and discusses how this technique may influence both the internal and external validity of studies. CONCLUSION: Using video-recorded consultations for research purposes may cause bias in the characteristics of doctors and patients who agree to participate in research. There is little evidence, however, that video-recording influences the behaviour of either GPs or patients. Recommendations are made for researchers who are considering using video-recorded consultations in their research.

Bias↗

Recruiting family physicians and patients for a clinical trial: lessons learned.

BACKGROUND: The randomized controlled trial (RCT) is the most definitive tool for evaluating an intervention. However, methodological deficiencies may limit the internal or external validity of the RCT. OBJECTIVE: Our aim was to describe the tactics used and the resources required randomly to select and recruit family physicians (FPs) and their patients aged 65 and older (seniors) for a community-based cluster RCT in primary care. METHODS: We randomly selected 48 FPs in 24 urban and rural sites in Southern Ontario, and 889 of their community-dwelling seniors (approximately 20 per FP) taking five or more medications daily. To accomplish this, the principal investigator (an FP) contacted the eligible FPs. The participating FPs' office staff then generated and contacted the roster of eligible seniors, with support provided by the research staff. RESULTS: Of the 163 randomly selected FPs telephoned, 94 were ineligible and 48 (69.6%) of the remaining 69 participated. The rosters were generated with the assistance of the research staff (taking 1.5-8.0 hours) in each of the 48 practices, using electronic appointment records (n = 26), electronic billing records (n = 17), electronic medical records (n = 2) or written charts or file cards (n = 3). Of the 2078 seniors approached, 799 were ineligible and 889 (69.5%) of the remaining 1279 participated. Seniors' refusal rates among practices ranged from 4.8 to 62.3%. CONCLUSIONS: Recruitment of a representative sample and generalizability of results are possible in RCTs in primary care. Involvement of an FP in physician recruitment and clinical research nurses who provided assistance to office staff were keys to success.

Aged↗

Judging outcomes in psychosocial interventions for dementia caregivers: the problem of treatment implementation.

PURPOSE: In published dementia caregiver intervention research, there is widespread failure to measure the level at which treatment was implemented as intended, thereby introducing threats to internal and external validity. The purpose of this article is to discuss the importance of inducing and assessing treatment implementation (TI) strategies in caregiving trials and to propose Lichstein's TI model as a potential guide. DESIGN AND METHODS: The efforts of a large cooperative research study of caregiving interventions, Resources for Enhancing Alzheimer's Caregiver Health (REACH), illustrates induction and assessment of the three components of TI: delivery, receipt, and enactment. RESULTS: The approaches taken in REACH vary with the intervention protocols and include using treatment manuals, training and certification of interventionists, and continuous monitoring of actual implementation. IMPLICATIONS: Investigation and description of treatment process variables allows researchers to understand which aspects of the intervention are responsible for therapeutic change, potentially resulting in development of more efficacious and efficient interventions.

Adaptation, Psychological↗

Context matters: interpreting impact findings in child survival evaluations.

Appropriate consideration of contextual factors is essential for ensuring internal and external validity of randomized and non-randomized evaluations. Contextual factors may confound the association between delivery of the intervention and its potential health impact. They may also modify the effect of the intervention or programme, thus affecting the generalizability of results. This is particularly true for large-scale health programmes, for which impact may vary substantially from one context to another. Understanding the nature and role of contextual factors may improve the validity of study results, as well as help predict programme impact across sites. This paper describes the experience acquired in measuring and accounting for contextual factors in the Multi-Country Evaluation of the IMCI (Integrated Management of Childhood Illness) strategy in five countries: Bangladesh, Brazil, Peru, Uganda and Tanzania. Two main types of contextual factors were identified. Implementation-related factors include the characteristics of the health systems where IMCI was implemented, such as utilization rates, basic skills of health workers, and availability of drugs, supervision and referral. Impact-related factors include baseline levels and patterns of child mortality and nutritional status, which affect the scope for programme impact. We describe the strategies used in the IMCI evaluation in order to obtain data on relevant contextual factors and to incorporate them in the analyses. Two case studies - from Tanzania and Peru - show how appropriate consideration of contextual factors may help explain apparently conflicting evaluation results.

Child Health Services↗

Ten methodological lessons from the multi-country evaluation of integrated Management of Childhood Illness.

OBJECTIVE: To describe key methodological aspects of the Multi-Country Evaluation of the Integrated Management of Childhood Illness strategy (MCE-IMCI) and analyze their implications for other public health impact evaluations. DESIGN: The MCE-IMCI evaluation designs are based on an impact model that defined expectations in the late 1990s about how IMCI would be implemented at country level and below, and the outcomes and impact it would have on child health and survival. MCE-IMCI studies include: feasibility assessments documenting IMCI implementation in 12 countries; in-depth studies using compatible designs in five countries; and cross-site analyses addressing the effectiveness of specific subsets of IMCI activities. The MCE-IMCI was designed to evaluate the impact of IMCI, and also to see that the findings from the evaluation were taken up through formal feedback sessions at national, sub-national and local levels. RESULTS: Issues that arose early in the MCE-IMCI included: (1) defining the scope of the evaluation; (2) selecting study sites and developing research designs; (3) protecting objectivity; and (4) developing an impact model. Issues that arose mid-course included: (5) anticipating and addressing problems with external validity; (6) ensuring an appropriate time frame for the full evaluation cycle; (7) providing feedback on results to policymakers and programme implementers; and (8) modifying site-specific designs in response to early findings about the patterns and pace of programme implementation. Two critical issues could best be addressed only near the close of the evaluation: (9) factors affecting the uptake of evaluation results by policymakers and programme decision makers; and (10) the costs of the evaluation. CONCLUSIONS: Large-scale effectiveness evaluations present challenges that have not been addressed fully in the methodological literature. Although some of these challenges are context-specific, there are important lessons from the MCE that can inform future designs. Most of the issues described here are not addressed explicitly in research reports or evaluation textbooks. Describing and analyzing these experiences is one way to promote improved impact evaluations of new global health strategies.

Child↗

Evaluating the impact of health promotion programs: using the RE-AIM framework to form summary measures for decision making involving complex issues.

Current public health and medical evidence rely heavily on efficacy information to make decisions regarding intervention impact. This evidence base could be enhanced by research studies that evaluate and report multiple indicators of internal and external validity such as Reach, Effectiveness, Adoption, Implementation and Maintenance (RE-AIM) as well as their combined impact. However, indices that summarize the combined impact of, and complex interactions among, intervention outcome dimensions are not currently available. We propose and discuss a series of composite metrics that combine two or more RE-AIM dimensions, and can be used to estimate overall intervention impact. Although speculative and, at this point, there have been limited empirical data on these metrics, they extend current methods and are offered to yield more integrated composite outcomes relevant to public health. Such approaches offer potential to help identify interventions most likely to meaningfully impact population health.

Decision Making↗

A scale without anthropometric measurements can be used to identify low weight-for-age in children less than five years old.

Malnutrition and morbidity have a synergistic association that often leads to death. However, malnutrition in children who die is largely underreported, because anthropometry of the deceased child is rarely known. This study had two purposes: i) to develop a scale that would help determine if a child had low weight-for-age (w/a), in the absence of anthropometric measures; and ii) to select an appropriate cut-off that would give the best sensitivity (Se) and specificity (Sp) of the proposed scale when contrasted with actual w/a measurement. The study was designed as a diagnostic test, and carried out in a rural area in central Mexico. We included 132 children under 5 y old with w/a under -2 Z score and 284 children with marginal or no w/a deficit as a control group. The proposed scale included potential predictive variables from clinical, socioeconomic and family factors. The best logistic regression model to predict low w/a included: birth weight less than 2,800 g, introduction of weaning foods after the sixth month of life, introduction of animal protein after the sixth month of life, low socioeconomic status, low w/a in siblings and more than three morbidity episodes in the previous 6 mon. Selecting a cut-off of 4 for this model to identify children with low w/a showed a Se and Sp of 85 and 95%, respectively. We tested the external validity of the scale in a different locale, and included 877 children under 5 y old from 10 rural communities. In this population, the scale showed Se of 84% and Sp of 81% to identify low w/a. Based on these results, we propose that the scale be included as a means of identifying low w/a in children who have died. We believe that this should be done in verbal autopsies, which, based on our previous research, the Ministry of Health adopted as part of the regular activities to monitor problems in the disease to health-seeking to death process.

Birth Weight↗