Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Estimating sample size in clinical studies: basic methodological principles.

In order to be valid, clinical studies must be methodologically rigorous. The internal validity of a study is of crucial importance: a study is valid if its results are an unbiased estimation of the true result. In this case, the validity is internal because it refers to the group of patients under study and not necessarily different ones (external validity or applicability). Internal validity in clinical research is achieved through rigorous design, data collection and appropriate analysis, and is threatened by bias (systematic errors) or chance (random variation of the phenomena under study). Regardless of the type of study (analytic, descriptive, etc.), the characteristics of its sample are fundamental for the validity of the results. The sampling methods are crucial if the study patients are to be representative of the population to which one desires to extrapolate the results. One of the most fundamental characteristics of a sample is its size. Even the best executed study may fail to answer the research question if the sample size is too small. On the other hand, a study with too large a sample is harder to conduct and more costly. The goal of planning the sample size is to estimate the appropriate number of research subjects for the study. In this paper we will present and discuss the methodological principles underlying calculation of sample size: outcomes, type I and II error, alpha and beta, study power and variability.

Clinical Trials as Topic↗

[Depression scales in schizophrenia: a critical review].

Depressive syndromes frequently occur during the evolution of schizophrenia. The evaluation of depression in schizophrenic patients is difficult because of an overlap between depressive, negative and extrapyramidal symptoms. The scales usually employed to evaluate depression have not been validated in schizophrenic populations, therefore, some authors developed specific depression scales for schizophrenics. In this work, we present the available data about the metrologic and psychometric properties of the Hamilton (HDRS), Montgomery Asberg (MADRS) and Widlöcher (ERD) depression scales in schizophrenic populations. We further present the validation works of the Psychotic Depression Scale (PDS) and the Calgary Depression Scale (CDSS). Non specific depression scales are unsatisfactory, since negative and extrapyramidal symptoms overlap with depressive symptoms. The ERD allows a distinction of the three symptom groups, when motor, ideic and subjective subscores are used. The two specific scales actually distinguish depression with a minimal level of contamination. Nevertheless, the factorial structure of the PDS comprises several non depressive factors that are of questionable interest. The CDSS is a well documented and validated tool, with an unidimensional structure, a good internal consistency, a high inter-rater reliability, and good external validity and specificity.

Depressive Disorder↗

[Validation in Havana City, Cuba of ENZYMEBA, an immunoassay for detecting Entamoeba histolytica in feces].

Parasitologists from different medical institutions in Havana City carried out the external validation of ENZYMEBA, a diagnostic procedure of intestinal amebiasis developed at the "Pedro Kourí" Institute of Tropical Medicine. To this end, serial faeces specimens from 212 individuals were collected and observed on the microscope (reference test). The ENZYMEBA immunoassay (validation test) was also made. On comparing ENZYMEBA with the microscopic examination, satisfactory indexes of sensitivity and specificity were found. No cross-reactions were detected in faeces specimens, where found other parasites were present, too. Taking into account that only one faeces specimen per patient is enough for the diagnosis of intestinal amebiasis with ENZYMEBA, this procedure may be useful in studies of therapeutical efficacy and prevalence.

Animals↗

Validation of population-based ADHD subtypes and identification of three clinically impaired subtypes.

Statistically based classification methods have successfully refined ADHD into homogenous and heritable subtypes. External validity and impairment of these subtypes was examined using the Child Behavior Checklist (CBCL). We compared mean CBCL syndrome and competency t-scores across ADHD subtypes defined by latent class analysis in a sample of 1,346 individual twins from Missouri. The potential for comorbidity with conduct disorder (CD), oppositional defiant disorder (ODD), or major depression (MD) to increase impairment in specific ADHD subtypes was also examined. CBCL profiles confirm differences in severity, with more severe classes having increased syndrome scale and decreased competency scale CBCL scores. Clinically significant impairment was found for severe inattentive and combined subtypes and the mild combined subtype. Overall, the presence of comorbid CD, ODD, or MD did not result in increased ADHD subtype impairment. CBCL scores distinguish impairment in ADHD subtypes created through LCA. Comorbidity with CD, ODD, or MD does not significantly increase impairment among ADHD subtypes. The mild combined ADHD subtype represents a clinically significant but under-studied form of ADHD.

Adolescent↗

Neuropsychological studies of late onset and subthreshold diagnoses of adult attention-deficit/hyperactivity disorder.

BACKGROUND: Diagnosing attention-deficit/hyperactivity disorder (ADHD) in adults is difficult when the diagnostician cannot establish an onset prior to the DSM-IV criterion of age 7 or if the number of symptoms recalled does not achieve the DSM-IV threshold for diagnosis. Because neuropsychological deficits are associated with ADHD, we addressed the validity of the DSM-IV age at onset and symptom threshold criteria by using neuropsychological test scores as external validators. METHODS: We compared four groups of adults: 1) full ADHD subjects met all DSM-IV criteria for childhood-onset ADHD; 2) late-onset ADHD subjects met all criteria except the age at onset criterion; 3) subthreshold ADHD subjects did not meet full symptom criteria; and 4) non-ADHD subjects did not meet any of the above criteria. RESULTS: Late-onset and full ADHD subjects had similar patterns of neuropsychological dysfunction. By comparison, subthreshold ADHD subjects showed few neuropsychological differences with non-ADHD subjects. CONCLUSIONS: Our results showing similar neuropsychological underpinning in subjects with late-onset ADHD suggest that the DSM-IV age at onset criterion may be too stringent. Our data also suggest that ADHD subjects who failed to ever meet the DSM-IV threshold for diagnosis have a milder form of the disorder.

Adolescent↗

The measurement of effort-reward imbalance at work: European comparisons.

Using comparative data from five countries, this study investigates the psychometric properties of the effort-reward imbalance (ERI) at work model. In this model, chronic work-related stress is identified as non-reciprocity or imbalance between high efforts spent and low rewards received. Health-adverse effects of this imbalance were documented in several prospective and cross-sectional investigations. The internal consistency, discriminant validity and factorial structure of 'effort', 'reward', and 'overcommitment' scales are evaluated, using confirmatory factor analysis. Moreover, content (or external) validity is explored with respect to a measure of self-reported health. Data for the analysis is derived from epidemiologic studies conducted in five European countries: the Somstress Study (Belgium; n = 3796), the GAZEL-Cohort Study (France; n = 10,174), the WOLF-Norrland Study (Sweden; n = 960), the Whitehall II Study (UK; n = 3697) and the Public Transport Employees Study (Germany; n = 316). Internal consistency of the scales was satisfactory in all samples, and the factorial structure of the scales was consistently confirmed (all goodness of fit measures were > 0.92). Moreover, in 12 of 14 analyses, significantly elevated odds ratios of poor health were observed in employees scoring high on the ERI scales. In conclusion, a psychometrically well-justified measure of work-related stress (ERI) grounded in sociological theory is available for comparative socioepidemiologic investigations. In the light of the importance of work for adult health such investigations are crucial in advanced societies within and beyond Europe.

Adolescent↗

Standardized physical examination protocol for low back disorders: feasibility of use and validity of symptoms and signs.

A standardized examination protocol was developed for the assessment of low back disorders in primary care. The protocol was found feasible in the occupational health service setting. The interexaminer repeatability between an occupational physician and an occupational physiotherapist was good for most items. The predictive validity of different symptoms and signs was investigated with regard to future sick leaves due to low back disorders. Relief of pain when lying, severe trouble at work caused by pain, continuous pain, and pain in the leg or numbness or diminished sensitivity in the foot predicted sick leaves. Of physical signs, pain in the low back or buttock during lateral flexion and a side difference > or = 20 degrees in the straight-leg-raising angle were predictors for sick leaves. The predictive validity of the protocol items should be tested in another patient population before conclusions can be drawn concerning the external validity of our results.

Adult↗

A self-report typology of behavioral adjustment for young children.

A large, national U.S. sample of children rated their own behavior and emotions using the Self-Report of Personality-Child version (SRP-C) of the Behavior Assessment System for Children (C. R. Reynolds & R. W. Kamphaus, 1992). Cluster analysis was used to group 4,981 self-reports (SRP-C) of children between the ages of 8 and 11 years. Theoretical and empirical considerations were used to identify a 10-cluster solution. Internal validation procedures revealed that the 10-cluster solution was well replicated by independently classifying 2 large subsamples of participants. External validation evidence revealed that only 2 of the 10 clusters could be differentiated by parent and teacher ratings of behavior problems. Peer ratings of social status and behavior, however, proved far better than adult ratings at differentiating the clusters. These findings suggest that the realm of intraindividual adjustment is not well understood by parents and teachers of these same children.

Adaptation, Psychological↗

A systematic review of the quality of homeopathic clinical trials.

BACKGROUND: While a number of reviews of homeopathic clinical trials have been done, all have used methods dependent on allopathic diagnostic classifications foreign to homeopathic practice. In addition, no review has used established and validated quality criteria allowing direct comparison of the allopathic and homeopathic literature. METHODS: In a systematic review, we compared the quality of clinical-trial research in homeopathy to a sample of research on conventional therapies using a validated and system-neutral approach. All clinical trials on homeopathic treatments with parallel treatment groups published between 1945-1995 in English were selected. All were evaluated with an established set of 33 validity criteria previously validated on a broad range of health interventions across differing medical systems. Criteria covered statistical conclusion, internal, construct and external validity. Reliability of criteria application is greater than 0.95. RESULTS: 59 studies met the inclusion criteria. Of these, 79% were from peer-reviewed journals, 29% used a placebo control, 51% used random assignment, and 86% failed to consider potentially confounding variables. The main validity problems were in measurement where 96% did not report the proportion of subjects screened, and 64% did not report attrition rate. 17% of subjects dropped out in studies where this was reported. There was practically no replication of or overlap in the conditions studied and most studies were relatively small and done at a single-site. Compared to research on conventional therapies the overall quality of studies in homeopathy was worse and only slightly improved in more recent years. CONCLUSIONS: Clinical homeopathic research is clearly in its infancy with most studies using poor sampling and measurement techniques, few subjects, single sites and no replication. Many of these problems are correctable even within a "holistic" paradigm given sufficient research expertise, support and methods.

Clinical Trials as Topic↗

Questions from practice: a basis for research.

Occupational health nurses in clinical practice are in an excellent position to identify unanswered questions that affect the health and well being of employees. Once these questions have been asked, the occupational health nurses may proceed with structured research to find answers. The research begins with a thorough review of existing literature to learn the background of the issue and clearly define a research question. This question is then framed conceptually to guide the study. The theoretical framework may be supported or refuted by the research project. The study method is dictated by the research question and the theoretical framework. Two basic research methods are intervention and descriptive studies. Intervention studies, with treatment and control groups, attempt to show differences between groups when one receives a certain treatment and the other group does not. Descriptive studies generally survey attitudes or activities, then test for associations. Data analysis is determined by the type of study and type of data collected. Descriptive statistics generate frequencies and correlations; inferential statistics yield stronger information about associations of the variables. Interpretation of research findings must include consideration of threats to validity. Internal validity allows the investigator to state with confidence that the intervention was responsible for the difference between the groups. External validity allows for generalizability of the findings to other populations. The purpose of nursing research is to advance the discipline of nursing. Research refines nursing theories and guides practice. It is through research that occupational health nurses gain confidence to alter procedures and provide interventions that have been shown to be effective.

Clinical Nursing Research↗

Speed of onset.

Speed of onset can be an important attribute of drug product performance. In studies employing appropriate designs and outcome measures, it is theoretically possible to make internally valid comparisons of the distribution of onset times associated with the use of two or more drug products. The external validity of such estimates is still potentially arguable. Steps that may be taken to enhance the likelihood that the results of such comparative studies will be accepted as valid bases for comparative claims are discussed.

Humans↗

Results of diagnostic accuracy studies are not always validated.

BACKGROUND AND OBJECTIVE: Internal validation of a diagnostic test estimates the degree of random error, using the original data of a diagnostic accuracy study. External validation requires a new study in an independent but similar population. Here we describe whether diagnostic research is validated, which technique is used, and to what extent the validation study results differ from the original. STUDY DESIGN AND SETTING: All original diagnostic accuracy studies published in 1993 in a predefined set of journals were selected. Validation of these studies was assessed in the original article and in articles published within a period of 10 years, through a literature search and contacting the authors. RESULTS: None of the original studies reported any form of validation. Validation studies published later could be identified for 7 of the 11 original studies. Test characteristics were difficult to compare. Despite what was generally believed, not every validation study showed results inferior to the original. We found more studies that evaluated the test in a different population than in a similar one. CONCLUSION: Not every diagnostic accuracy study is validated. Diagnostic tests are more often repeated in different populations.

Databases, Bibliographic↗

Automated CEAP Classification of Venous Duplex Reports Using Multimodal Artificial Intelligence.

OBJECTIVE: To develop and internally validate a prototype multimodal artificial intelligence system for automated CEAP (Clinical, Etiological, Anatomical and Pathophysiological) classification of venous duplex ultrasound (VDUS) reports, integrating natural language processing of free-text components with computer vision analysis of hand-drawn anatomical diagrams. METHODS: Single centre retrospective observational study using routinely collected clinical data. One thousand consecutive venous duplex ultrasound reports from Cambridge University Hospitals NHS Foundation Trust, UK (July 2024 - May 2025) were labelled according to the CEAP classification, excluding the Etiological component, which could not be reliably determined from duplex reports alone. Transfer learning was applied using ClinicalBERT for text and MobileNetV3 for diagrammatic data. Clinical classes were predicted from request line text. Text- and image-based pathophysiological models were developed for four anatomical territories (Great Saphenous Vein, Small Saphenous Vein, Deep system, Perforators), combined using late fusion with probability averaging. RESULTS: The clinical CEAP model achieved accuracy of 0.91, macro-F1 of 0.82, and macro-AUC of 0.98. Pathophysiological prediction varied, with text models broadly outperforming image models. Fusion yielded heterogeneous benefits, improving SSV performance but reducing Deep system accuracy. The performance of the final pathophysiological CEAP fusion models varied across anatomical territories: accuracy ranged from 0.70-0.92 and macro-AUC from 0.80-0.92. CONCLUSION: This study demonstrates the feasibility of automated CEAP classification from VDUS reports. Despite class imbalance affecting minority class predictions, the strong discriminatory performance validates this multimodal ML model for extracting clinically meaningful information from real-world data. This approach offers potential, pending external validation, to streamline vascular services through automated triage and guideline-compliant decision making.

Artificial intelligence↗

Education bias in the mini-mental state examination.

Education is correlated with cognitive status assessment. Concern for test bias has led to questions of equivalent construct validity across education groups. Following the work of previous researchers, we submitted Mini-Mental State Examination (MMSE) responses to external validation analyses. Subjects were older participants in the Epidemiologic Catchment Area study (age 50-98). Little evidence for test bias against those with low education was found. The correlation of MMSE scores and age was equivalent across high- and low-education groups (-.29 vs. -.27, p = .48), as was the correlation of MMSE scores and activities of daily living (ADL) functioning (-.23 vs. -.27, p = .42). The MMSE displayed significantly higher internal consistency reliability in the low-education group (.75 vs. .72, p = .04). The MMSE did not predict functional decline over 1 year or mortality over 13 years differently by level of educational attainment. Evidence for sex bias was found. The MMSE was more highly correlated with age among women than among men (-.28 vs. -.21, p < .001). The MMSE was more highly correlated with ADL impairment among women than among men (-.30 vs. -.17, p = .01). The MMSE predicted mortality differently according to participant sex (p = .053). The lack of evidence for bias provides little support to proposals to adjust MMSE scores according to level of education.

Activities of Daily Living↗

The development of a comprehensive headache diary--verbal descriptor scales.

A headache diary was developed to allow the ongoing assessment of subjective and behavioural components of headache pain. A modified version of Philips & Hunter's (1981) pain behaviour checklist which assesses three types of pain behaviour was incorporated for the latter. A scale based on pain descriptors was developed for the subjective component. This scale was factor analysed to reveal five aspects of headache pain, two reactive and three sensory. When the descriptors were scaled it was found that the amount of pain represented by some words was significantly greater in headache than non-headache subjects. Within each descriptor factor, the words do not appear to vary in intensity, with the exception of one reactive factor which appears to reflect overall evaluation of pain. Internal validity of the diary was investigated showing distinctive patterns of association between the subjective factors, pain behaviour, and pain intensity. External validity was assessed by the use of the diary in an intervention study and in delineating differences between migraine and tension headache groups.

Adolescent↗

Validity of messages from quadriplegic persons with cerebral palsy.

Interpreting gestures from retarded, nonvocal subjects is scientifically risky. Investigators must construe the subject's meaning without reference to an external validity measure. A procedure was devised in which message content was provided to nonvocal, severely palsied quadriplegic subjects in advance. Subjects' responses were limited to yes/no gestures. Another investigator elicited the messages without prior knowledge of their content. Results, which indicated reasonably high correspondence between stimulus messages and messages elicited, suggest that such subjects can present the content of their own phenomenal field accurately and that investigators' interpretations need not be considered imaginary.

Adult↗

Use of near-infrared reflectance spectroscopy in predicting nitrogen, phosphorus and calcium contents in heterogeneous woody plant species.

Near-infrared reflectance spectroscopy was applied to determine nitrogen (N), phosphorus (P) and calcium (Ca) content in leaf samples of 18 woody species. A total of 183 samples from mountain, riparian and dry areas from the Central-Western Iberian Peninsula were collected for this purpose. The wide intervals of variation observed in nutrient concentrations (6.6-45.0 g kg(-1) for N, 0.24-2.97 g kg(-1) for P, and 1.00-20.06 g kg(-1) for Ca) were due to the great heterogeneity of the samples. To develop calibration equations, multiple linear regression, and partial least-squares regression (PLSR) were used. In both cases, three mathematical transformations of the data were applied: log1/R and first and second derivatives. The best calibration statistics were obtained using PLSR and derivative transformations (second derivative for N and first derivative for P and Ca). The following coefficients of multiple determination (R2) and standard errors of cross validation were obtained: 0.99 and 0.93 for N, 0.94 and 0.15 for P, and 0.95 and 0.88 for Ca. In the external validation the standard errors of prediction obtained were 0.76 (N), 0.11 (P) and 0.60 (Ca).

Calcium↗

Validation by simulation of a clinical trial model using the standardized mean and variance criteria.

OBJECTIVE: To develop and validate a model of a clinical trial that evaluates the changes in cholesterol level as a surrogate marker for lipodystrophy in HIV subjects under alternative antiretroviral regimes, i.e., treatment with Protease Inhibitors vs. a combination of nevirapine and other antiretroviral drugs. METHODS: Five simulation models were developed based on different assumptions, on treatment variability and pattern of cholesterol reduction over time. The last recorded cholesterol level, the difference from the baseline, the average difference from the baseline and level evolution, are the considered endpoints. Specific validation criteria based on a 10% minus or plus standardized distance in means and variances were used to compare the real and the simulated data. RESULTS: The validity criterion was met by all models for considered endpoints. However, only two models met the validity criterion when all endpoints were considered. The model based on the assumption that within-subjects variability of cholesterol levels changes over time is the one that minimizes the validity criterion, standardized distance equal to or less than 1% minus or plus. CONCLUSION: Simulation is a useful technique for calibration, estimation, and evaluation of models, which allows us to relax the often overly restrictive assumptions regarding parameters required by analytical approaches. The validity criterion can also be used to select the preferred model for design optimization, until additional data are obtained allowing an external validation of the model.

Anti-Retroviral Agents↗