Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Clinical trials in the elderly. Pivotal points.

A clinical trial, that is, the scientific assessment of drug action in humans, must be undertaken only if there is reference to an expectation of benefit from the trial. Before the trial starts, a series of questions should be answered, including the need for the trial in elderly patients, particularly in view of the possible vulnerability of the elderly study subjects. Patient selection, randomization, follow-up, analysis, and interpretation must be scientifically valid. Of major importance is the external validity of the trial, that is, the generalizability. Any trial, particularly those involving the elderly, should be designed to develop or contribute to generalizable knowledge; that is, it should be possible to extend the conclusions from the trial beyond the study population to the population at large. When elderly are involved, as they should be when a drug is proposed for use mainly in the elderly population, it is especially important that the study be scientifically valid, medically important, and ethically sound. Studies involving the elderly should have sufficient numbers of females and minorities, the former because most elderly patients are females, the latter because minorities now reach the age of 65 years and beyond more so than in the past. Both risk and benefit should be addressed in terms of potential magnitude and duration. If the study drug has a narrow therapeutic window, it should be intensively studied. When the study is completed, there ought to be clear guidelines for the clinician to design initial, individualized, optimal dosage regimens or for subsequent adjustment of the regimen. The final report, which should be easily evaluable by clinicians, should fully discuss reasons for dropouts, inappropriate patient inclusion, number and types of adverse reactions, and defects in design and conduct of the study.

Aged↗

Identifying candidate subunit vaccines using an alignment-independent method based on principal amino acid properties.

Subunit vaccine discovery is an accepted clinical priority. The empirical approach is time- and labor-consuming and can often end in failure. Rational information-driven approaches can overcome these limitations in a fast and efficient manner. However, informatics solutions require reliable algorithms for antigen identification. All known algorithms use sequence similarity to identify antigens. However, antigenicity may be encoded subtly in a sequence and may not be directly identifiable by sequence alignment. We propose a new alignment-independent method for antigen recognition based on the principal chemical properties of protein amino acid sequences. The method is tested by cross-validation on a training set of bacterial antigens and external validation on a test set of known antigens. The prediction accuracy is 83% for the cross-validation and 80% for the external test set. Our approach is accurate and robust, and provides a potent tool for the in silico discovery of medically relevant subunit vaccines.

Algorithms↗

Temporal lobe signs: electroencephalographic validity and enhanced scores in special populations.

Internal and external validity tests were completed for an inventory that has been used to infer signs of temporal lobe lability. Strong, positive correlations were reported for a normal (reference) population between the numbers of responses that referred to paranormal experiences (including feelings of a "presence") and separately to religious beliefs and the numbers of spikes per minute within electroencephalographic recordings from the temporal lobe. Numbers of spikes were also correlated with the subjects' scores on the hysteria, schizophrenia, and psychasthenia scales from the MMPI. These clusters of items were not correlated with electrical activity from the occipital lobe (the comparison region). Numbers of responses to control clusters of mundane experiences were not correlated with the temporal lobe measures. A group of student poets scored higher on different subclusters of temporal lobe signs and on the schizophrenia and mania scales of the MMPI than the reference group. For both groups, there were positive correlations between the amount of alpha activity in the temporal lobe only and answers to items such as "hearing inner voices" and "feeling as if things were not real." These results demonstrate that quantitative measures of electrical changes in the temporal lobe are correlated with (or with the report of) specific experiences that are prevalent during surgical or epileptic stimulation of this brain region.

Adult↗

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (≥54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans↗

Multi-national, multi-lingual, multi-professional CATs: (Curriculum Analysis Tools).

A consortium of dental schools and allied dental programs was established in 1991 with the expressed purpose of creating a curriculum database program that was end-user modifiable [1]. In April of 1994, a beta version (Beta 2.5 written in FoxPro(TM) 2.5) of the software CATs, an acronym for Curriculum Analysis Tools, was released for use by over 30 of the consortium's 60 member institutions, while the remainder either waited for the Macintosh (TM) or Windows (TM) versions of the program or were simply not ready to begin an institutional curriculum analysis project. Shortly after this release, the design specifications were rewritten based on a thorough critique of the Beta 2.5 design and coding structures and user feedback. The result was Beta 3.0 which has been designed to accommodate any health professions curriculum, in any country that uses English or French as one of its languages. Given the program's extensive use of screen generation tools, it was quite easy to offer screen displays in a second language. As more languages become available as part of the Unified Medical Language System, used to document curriculum content, the program's design will allow their incorporation. When the software arrives at a new institution, the choice of language and health profession will have been preselected, leaving the Curriculum Database Manager to identify the country where the member institution is located. With these 'macro' end-user decisions completed, the database manager can turn to a more specific set of end-user questions including: 1) will the curriculum view selected for analysis be created by the course directors (provider entry of structured course outlines) or by the students (consumer entry of class session summaries)?; 2) which elements within the provided course outline or class session modules will be used?; 3) which, if any, internal curriculum validation measures will be included?; and 4) which, if any, external validation measures will be included. External measures can include accreditation standards, entry-level practitioner competencies, an index of learning behaviors, an index of discipline integration, or others defined by the institution. When data entry, which is secure to the course level, is complete users may choose to browse a variety of graphic representations of their curriculum, or either preview or print a variety of reports that offer more detail about the content and adequacy of their curriculum. The progress of all data entry can be monitored by the database manager over the course of an academic year, and all reports contain extensive missing data reports to ensure that the user knows whether they are studying complete or partial data. Institutions using the beta version of the program have reported considerable satisfaction with its functionality and have also offered a variety of design and interface enhancements. The anticipated release date for Curriculum Analysis Tools (CATs) is the first quarter of 1995.

Curriculum↗

A new set of amino acid descriptors and its application in peptide QSARs.

In this work, a new set of amino acid descriptors, i.e., VHSE (principal components score Vectors of Hydrophobic, Steric, and Electronic properties), is derived from principal components analysis (PCA) on independent families of 18 hydrophobic properties, 17 steric properties, and 15 electronic properties, respectively, which are included in total 50 physicochemical variables of 20 coded amino acids. Using the stepwise multiple regression (SMR) method combined with partial least squares (PLS), the VHSE scales are then applied to QSAR studies of bitter-tasting dipeptides (BTD), angiotensin-converting enzyme (ACE) inhibitors, and bradykinin-potentiating pentapeptides (BPP). To validate the predictive power of resulting models, external validation are also performed. A comparison of the results to those obtained with z scores and other two-dimensional (2D) or three-dimensional(3D) descriptors shows that the VHSE scales are comparable for parameterizing the structural variability of the peptide series.

Amino Acid Sequence↗

Ecological validity and cultural sensitivity for outcome research: issues for the cultural adaptation and development of psychosocial treatments with Hispanics.

This article has two objectives. The first is to provide a culturally sensitive perspective to treatment outcome research as a resource to augment the ecological validity of treatment research. The relationships between external validity, ecological validity, and culturally sensitive research are reviewed. The second objective is to present a preliminary framework for culturally sensitive interventions that strengthen ecological validity for treatment outcome research. The framework, consisting of eight dimensions of treatment interventions (language, persons, metaphors, content, concepts, goals, methods, and context) can serve as a guide for developing culturally sensitive treatments and adapting existing psychosocial treatments to specific ethnic minority groups. Examples of culturally sensitive elements for each dimension of the intervention are offered. Although the focus of the article is on Hispanic populations, the framework may be valuable to other ethnic and minority groups.

Cross-Cultural Comparison↗

Determining the basic psychometric properties of the Greek KDQOL-SF.

The aim of this study was to determine the basic psychometric properties, i.e. reliability and validity, of the Greek version of the Kidney Disease Quality of Life Short Form (KDQOL-SF). The instrument was self-administered to a homogenous group of 665 end stage renal disease patients in 20 dialysis units throughout Greece and the overall response rate was 72.6%. Reliability was demonstrated by Cronbach's alpha exceeding the recommended minimum value of 0.70 in all, except one, scales. Tests of item-internal consistency, after correction for overlap, resulted in correlations between items and their hypothesized scales, which exceeded the 0.40 standard in 94.5% of the cases. Item discriminant validity tests indicated 100% scaling success for six out of eight generic and disease-targeted scales. Validity was supported by the confirmation of expected correlations between scales and the overall health-rating item included in the instrument and with sociodemographic and self-reported health variables. Multiple stepwise linear regression analysis demonstrated that all disease-targeted scales were important predictors of SF-36 general health scales and the variance explained ranged from 37% to 57%. Overall, the psychometric properties of the KDQOL-SF, resulting from this first-time administration of the instrument to a Greek dialysis population, were good and the disease targeted scales were informative and of high internal consistency reliability. Cross-sectional construct validity is demonstrated, despite the lack of external validity criteria based on clinical ratings of severity. The results support administering the Greek KDQOL-SF in studies evaluating dialysis therapy and contribute to transnational comparison of findings.

Aged↗

Comparison of three generic questionnaires measuring quality of life in adolescents and adults with cystic fibrosis: the 36-item short form health survey, the quality of life profile for chronic diseases, and the questions on life satisfaction.

OBJECTIVE: To compare different generic instruments in measuring quality of life and to demonstrate dimensions of quality of life (QL) in patients with cystic fibrosis (CF). METHODS: The short-form-36 health survey (SF-36), the quality of life profile for chronic diseases (PLC), and the questions on life satisfaction (FLZ(M)) were simultaneously employed in a cross-sectional study with 70 adolescents and adults with CF. The different concepts of the measures were compared. Internal consistency (Cronbach's alpha), convergent and construct validity (correlation patterns, common factor analysis), and external validity (correlations with symptom and pulmonary function scores, with intensity of therapy; comparisons with healthy peers) of the three instruments were investigated. RESULTS: Similar reliability, but different validity of the questionnaires are demonstrated. Seventy-three percent of the total variance across the three measures could be explained with a seven-factor-solution: (1) physical functioning (19.3% of total variance), (2) mental health (19.3%), (3) social integration (7.5%), (4) role function/pain (7.5%), (5) economic/material living conditions (7.5%), (6) partnership/family (6.7%) and (7) anxiety (5.2%). DISCUSSION: The different validity of the instruments has to be considered in chosing a questionnaire appropriate to the purpose of measuring. Shortcomings of each instrument can be overcome by multimethod designs and by developing disease-specific scales.

Adolescent↗

DSM-IV: empirical guidelines from psychometrics.

This commentary addresses the use of psychometric theory and methodology in the development of the 4th edition of the Diagnostic and Statistical Manual of Mental Disorders (DSM-IV). Reliability issues include interdiagnostician reliability, temporally consistent diagnoses, and the relations of diagnostic criteria within categories. Validity issues include content validity of the diagnostic criteria, criterion-related validity (the relation between different criterion sets or their algorithms and alternative diagnostic criteria), and construct validity (the relation between diagnostic categories and external validators). Specific questions and methodology to investigate its utility vary with the different uses proposed for the diagnostic system. Specific psychometric methodologies that may be useful in developing the DSM-IV are noted, as are the limitations of psychometrics and their applicability to DSM-IV.

Humans↗

A formal cognitive model of the go/no-go discrimination task: evaluation and implications.

This article proposes and tests a formal cognitive model for the go/no-go discrimination task. In this task, the performer chooses whether to respond to stimuli and receives rewards for responding to certain stimuli and punishments for responding to others. Three cognitive models were evaluated on the basis of data from a longitudinal study involving 400 adolescents. The results show that a cue-dependent model presupposing that participants can differentiate between cues was the most accurate and parsimonious. This model has 3 parameters denoting the relative impact of rewards and punishments on evaluations, the rate that contingent payoffs are learned, and the consistency between learning and responding. Commission errors were associated with increased attention to rewards; omission errors were associated with increased attention to punishments. Both error types were associated with low choice consistency. The parameters were also shown to have external validity: Attention to rewards was associated with externalizing behavior problems on the Achenbach scale, and choice consistency was associated with low Welsh anxiety. The present model can thus potentially improve the sensitivity of the task to differences between clinical populations.

Adolescent↗

The hospital social work self-efficacy scale: a partial replication and extension.

The Hospital Social Work Self-Efficacy Scale (HSWSE) was developed to assess social workers' confidence about their ability to perform hospital social work. The scale was originally tested with master's-level social work students. This partial replication and extension included both students and professional social workers. In addition, we used a new scale to assess construct validity and added another hospital to assess external validity. In general, findings continue to support the use of the HSWSE. In addition, anecdotal evidence suggests that this is an outcome measure that practitioners see as easy to use, specific, and highly relevant.

Educational Measurement↗

Evaluation of a Dutch version of the AIMS2 for patients with rheumatoid arthritis.

DUTCH-AIMS2, a Dutch version of AIMS2 and successor to DUTCH-AIMS, is an instrument to assess health status among patients with rheumatic diseases. It provides measurements of 12 areas of health status on scales for health status proper, satisfaction, attribution and arthritis impact. We assessed the reliability of its scales in terms of internal consistency and their validity according to both internal standards and external standards. Correctly completed questionnaires were returned by 231 RA patients and 131 controls. Internal consistency coefficients for the health status scales ranged from 0.66 and 0.89, but most exceeded 0.80. Within-scale factor analyses produced single factors in all composite health status scales for both patients and controls, with only two exceptions. Factor analysis also identified a physical, social and psychological dimension among 11 areas of health. External validity was established by strong correlations between DUTCH-AIMS2 health status scales and functional class, laboratory parameters, and self-assessments of fatigue, loneliness, pain, functional disability and social support. DUTCH-AIMS2 is acceptably reliable and valid for use in a variety of settings.

Activities of Daily Living↗

Upper extremity function in hemiplegia. A cross-validation study of two assessment methods.

The methods devised by DeSouza et al. and by Fugl-Meyer et al. for description of upper extremity function after stroke were compared by parallel assessments in a consecutive series of 50 patients with hemimotor deficit. Very close positive associations between both methods indicated a high degree of cross-validity. As both methods appear to be externally valid, have good inter-rater reliability and as the time needed for assessing the arm function of a hemiplegic or hemiparetic patient rarely exceeds 10 min, it appears that the two methods possess about equal descriptive power.

Arm↗

General purpose of research designs.

The purposes and criteria for formulating a design of research, conditions for judging causality, and use of research design as a control of variance are discussed. The purpose of a research design is to provide a plan of study that permits accurate assessment of cause and effect relationships between independent and dependent variables. The classic controlled experiment is an ideal example of good research design. Factors that jeopardize the evaluation of the effect of experimental treatment (internal validity) and the generalizations derived from it (external validity) are identified. Sources of variance can be controlled by eliminating a variable, randomization, matching, or including a variable as part of the design. A research project should be so designed that (1) it answers the questions being investigated, (2) extraneous factors are controlled, and (3) the degree of generalization that can be made is valid.

Analysis of Variance↗

Measuring quality of life of persons with spinal cord injury: external and structural validity.

STUDY DESIGN: Measurement evaluation of the external and structural components of validity. OBJECTIVES: To examine the relationships between quality of life (QOL) as measured by the spinal cord injury (SCI) version of the Ferrans and Powers Quality of Life Index (QLI) and other constructs represented within the model of disablement; and to examine the domains and scoring model of the QLI by exploring item and overall score/section score relationships. SETTING: Community, Alberta, Canada. METHODS: A convenience sample of 98 individuals with SCI living in the community completed the QLI and measures representing the model of disablement including the ASIA motor index, Functional Independence Measure, Reintegration to Normal Living index, Rosenberg's Self-Esteem Scale and Rotter's Internal-External Locus of Control scale. RESULTS: Four of the five a priori hypotheses were supported. Locus of control was not significantly related to QOL as expected. Factor analysis resulted in a five-factor structure that differed from the four-domain model of the original QLI. Scoring relationships indicated that both the satisfaction and importance ratings contribute to the overall score, although not equally. CONCLUSION: There is support for the external component of validity although further examination regarding locus of control for persons with SCI is warranted. The structural component of validity requires further investigation to elucidate the domains of the SCI version of the QLI and the contribution of the importance scores.

Activities of Daily Living↗

Eligibility of real-world patients for aspirin primary prevention trials in cardiovascular disease.

BACKGROUND: Evidence for the net benefit of aspirin for primary prevention of cardiovascular disease (CVD) is finely balanced, leading to variation in guideline recommendations internationally. External validity of randomised clinical trial (RCT) evidence may therefore be of particular importance. The aim of this study is to characterise real-world patients according to their eligibility for guideline-cited aspirin RCTs for primary CVD prevention. METHODS: Eligibility criteria from 14 RCTs were applied to a linked primary care/hospital discharge dataset of people&#x2009;&#x2265;&#x2009;40 years without CVD. Proportions eligible for each trial were calculated, and characteristics of eligible and ineligible patients compared for each trial, including Cox regression analysis of event rates for major adverse cardiovascular events (MACE), major bleeding events, and non-cardiovascular mortality. RESULTS: Of 570,211 included patients (300,500 [52.7%] women, 336,877 [59%]&#x2009;<&#x2009;60 years), the median proportion ineligible for 14 RCTs was 90.7% (range 42.5-99.4%) and 24.0% of patients were ineligible for all RCTs. On average, trial-ineligible populations were younger (median age trial-ineligible 57.8 vs trial-eligible 62.6 years, p&#x2009;=&#x2009;0.008) and a lower proportion had hypertension (23.9% vs 50.9%, p&#x2009;=&#x2009;0.004), diabetes (6.4% vs 11.5%, p&#x2009;=&#x2009;0.015), or a regular statin prescription (11.8% vs 26.7%, p&#x2009;=&#x2009;0.001). Trial-ineligible populations had a higher hazard of MACE compared to trial-eligible in four RCTs and lower in ten (hazard ratio [HR] range across all RCTs 0.45 [95%CI 0.40-0.51] to 2.78 [95%CI 2.61-2.96]). Hazards of bleeding events in the trial-ineligible were lower than the trial-eligible in eight RCTs and higher in four (HR range across all RCTs 0.63 [95%CI, 0.59-0.66] to 1.69 [95%CI, 1.53-1.86]), and time-varying hazards of non-CVD death were consistently lower in four RCTs and higher in five (HR range across all RCTs and time points 0.29 [95%CI 0.24-0.36] to 11.42 [95%CI 9.91-13.17]). CONCLUSIONS: Compared with trial-ineligible populations within the same age and sex strata, RCTs recruited people of varying CVD risk but often excluded people at high risk of bleeding or non-CVD death, highlighting that many trials may overestimate the net benefit of aspirin for primary prevention.

Humans↗

Diagnostic model for sensitization in workers exposed to occupational high molecular weight allergens.

BACKGROUND: Occupational allergy has great impact on workers exposed to high molecular weight (HMW) allergens. The present study is aimed to develop and validate a generic diagnostic model for sensitization to HMW allergens, defined as positive IgE. METHODS: The model was developed in pooled data from Dutch laboratory animal (LA) workers and bakers using logistic regression analysis. Validity was assessed internally by bootstrapping procedure, and externally in British LA workers. RESULTS: The model included working hours/week, work-related symptoms, total IgE, and IgE to common allergen. Significant interactions between the type of work and the predictors resulted in different scores for LA workers and bakers. Internal and external validation showed that the model was satisfactorily calibrated and discriminated between workers at high and low risk of being sensitized. CONCLUSIONS: It is possible to develop a generic model for sensitization to occupational HMW allergens. However, the weighing of predictors differs across specific work environments.

Adult↗