Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Evidence for developmentally based diagnoses of oppositional defiant disorder and conduct disorder.

This paper compares the validity of DSM-III-R diagnoses of oppositional defiant disorder (ODD) and conduct disorder (CD) and an alternative option which is subdivided into three levels according to developmental sequence and severity: modified oppositional disorder (MODD), intermediate CD (ICD), and advanced CD (ACD). Using a sample of 177 boys followed over 3 years, both the DSM-III-R and the alternative diagnostic constructs are evaluated on three criteria: symptom discriminative validity, and diagnostic external and predictive validity. Most DSM-III-R ODD and CD symptoms discriminated between ODD and CD, but exceptions are noted. Additional analyses demonstrated considerable overlap among DSM-III-R oppositional symptoms. The majority of the symptoms proposed for the alternative option could be assigned to a specific level based on acceptable symptom discrimination. External validity lent support to the distinctions between DSM-III-R ODD and CD, and between MODD, ICD, and ACD. MODD was a better predictor than ODD of which MODD, ICD, and ACD. MODD was a better predictor than ODD of which boys received a later diagnosis of CD. Suggestions are made for the inclusion and exclusion of symptoms for developmentally based diagnoses of oppositional and conduct disorders.

Antisocial Personality Disorder↗

External cross-validation for unbiased evaluation of protein family detectors: application to allergens.

Key issues in protein science and computational biology are design and evaluation of algorithms aimed at detection of proteins that belong to a specific family, as defined by structural, evolutionary, or functional criteria. In this context, several validation techniques are often used to compare different parameter settings of the detector, and to subsequently select the setting that yields the smallest error rate estimate. A frequently overlooked problem associated with this approach is that this smallest error rate estimate may have a large optimistic bias. Based on computer simulations, we show that a detector's error rate estimate can be overly optimistic and propose a method to obtain unbiased performance estimates of a detector design procedure. The method is founded on an external 10-fold cross-validation (CV) loop that embeds an internal validation procedure used for parameter selection in detector design. The designed detector generated in each of the 10 iterations are evaluated on held-out examples exclusively available in the external CV iterations. Notably, the average of these 10 performance estimates is not associated with a final detector, but rather with the average performance of the design procedure used. We apply the external CV loop to the particular problem of detecting potentially allergenic proteins, using a previously reported design procedure. Unbiased performance estimates of the allergen detector design procedure are presented together with information about which algorithms and parameter settings that are most frequently selected.

Allergens↗

Psychometric properties of the Mini-Mental State Examination in patients with acquired brain injury in Turkey.

OBJECTIVE: To evaluate the psychometric properties of Mini-Mental State Examination (MMSE) in patients with acquired brain injury in Turkey. METHODS: A total of 207 patients with acquired brain injury were assessed. Reliability was tested by internal consistency and the person separation index; internal construct validity by Rasch analysis; external construct validity by correlation with cognitive disability; and cross-cultural validity by differential item functioning analysis compared with Italian MMSE data. RESULTS: Reliability was adequate with a Cronbach's alpha of 0.75 and person separation index of 0.76. After collapsing some categories, and adjustment for differential item functioning, internal construct validity was supported by fit of the data to Rasch model. Differential item functioning for culture was found in 2 items and after adjustment, data could be pooled between Turkey and Italy. External construct validity was supported by expected associations. CONCLUSION: The Turkish version of the Mini-Mental State Examination can be used as a cognitive screening tool in acquired brain injury. Cross-cultural validity between Italy and Turkey is supported, given appropriate adjustment for differential item functioning. However, shortfalls in reliability at the individual level, as well as the presence of differential item functioning suggest that a better instrument should be developed to screen for cognitive deficits following acquired brain injury.

Adult↗

Effects of response sets on NEO-PI-R scores and their relations to external criteria.

Validity scales indicate the extent to which the results of a self-report inventory are a valid indicator of the test taker's psychological functioning. Validity scales generally are designed to detect the common response sets of positive impression management (underreporting, or faking good), negative impression management (overreporting, or faking bad), and random responding. The revised NEO Personality Inventory (NEO-PI-R; Costa & McCrae, 1992b) is a popular personality assessment tool based on the 5-factor model of personality and is used in a variety of settings. The NEO-PI-R does not include objective validity scales to screen for positive or negative impression management. The purpose of this study was to examine the utility of recently proposed validity scales for detecting these response sets on the NEO-PI-R (Schinka, Kinder, & Kremer, 1997) and to examine the effects of positive and negative impression management on correlations between the NEO-PI-R and external criteria (the Interpersonal Adjective Scale-Revised-B5 [Wiggins & Trapnell, 1997] and the NEO-PI-R Form R). The validity scales discriminated with reasonable accuracy between standard responding and the 2 response sets. Additionally, most correlations between the NEO-PI-R and external criteria were significantly lower when participants were dissimulating than when responding to standard instructions. It appears that response sets of positive and negative impression management may pose a significant threat to the external validity of the NEO-PI-R and that validity scales for their detection might be a useful addition to the inventory.

Adolescent↗

Adaptation and validation of the Turkish version of the Rheumatoid Arthritis Quality of Life Scale.

OBJECTIVE: The aim of this study was to adapt the Rheumatoid Arthritis Quality of Life (RAQoL) questionnaire for use in Turkey and to test its reliability and validity. METHODS: The translation process included the recent guidelines for cross-cultural adaptation. Reliability of the Turkish RAQoL was assessed by internal consistency and test-retest reliability, internal construct validity by Rasch analysis, and external construct validity by associations with impairments, disability, and general health status. Cross-cultural validity was tested through analysis of differential item functioning (DIF) by comparison with data from the UK version of the RAQoL. RESULTS: Reliability of the adapted version was good, with high internal consistency (Cronbach's alpha 0.95 and 0.96 at times 1 and 2, respectively) and test-retest reliability (Spearman's rho 0.874). Internal construct validity was confirmed by excellent fit to the Rasch model (mean item fit 0.236, SD 1.113) and external construct validity by expected associations. The DIF for culture was found in four items. CONCLUSIONS: Adaptation of the RAQoL for use in Turkey was successful. The instrument can be used in both national and international studies for cross-cultural comparison with the UK, as long as adjustments are made for the few items displaying DIF for culture.

Aged↗

Blood pressure changes validate phase related external suction, a controlled method for stimulation of human baroreceptors.

Phase related external suction (PRES), a new controlled method for manipulating activity in human baroreceptors, applies precisely timed bursts of suction and pressure within the cardiac cycle through an external neck cuff. Seven healthy adult men participated in 32 pseudo-random trials of baroreceptor stimulation and inhibition. Blood pressure was assessed both intra-arterially and with a noninvasive device. In the present study, PRES baroreceptor stimulation elicited invasively measured blood pressure decreases of about 2.5 mmHg (0.33 kPa) and heart rate decreases of about 5 beats,min-1, while baroreceptor inhibition increased invasively measured blood pressure by about 1.5 mmHg (0.20 kPa) and heart rate about 2.5 beats.min-1. It was concluded that PRES is an effective method for baroreceptor manipulation with weaker size effect but better control of nonspecific factors in human subjects than other baroreceptor manipulation techniques. The noninvasive blood pressure measurement device was less sensitive to experimental variation than was the invasive device.

Adult↗

Accurate prediction of VO2max in cycle ergometry.

Numerous equations exist for predicting VO2max from the duration (an analog of maximal work rate, Wmax) of a treadmill graded exercise test (GXT). Since a similar equation for cycle ergometry (CE) was not available, we saw the need to develop such an equation, hypothesizing that CE VO2max could be accurately predicted due to its more direct relationship with W. Thus, healthy, sedentary males (N = 115) and females (N = 116), aged 20-70 yr, were given a 15 W.min-1 CE GXT. The following multiple linear regression equations which predict VO2max (ml.min-1) from the independent variables of Wmax (W), body weight (kg), and age (yr) were derived from our subjects: Males: Y = 10.51 (W) + 6.35 (kg) - 10.49 (yr) + 519.3 ml.min-1; R = 0.939, SEE = 212 ml.min-1. Females: Y = 9.39 (W) + 7.7 (kg) - 5.88 (yr) + 136.7 ml.min-1; R = 0.932, SEE = 147 ml.min-1 Using the 95% confidence limits as examples of worst case errors, our equations predict VO2max to within 10% of its true value. Internal (double cross-validation) and external cross-validation analyses yielded r values ranging between 0.920 and 0.950 for the male and female regression equations. These results indicate that use of the equations generated in this study for a 15 W.min-1 CE GXT provides accurate estimates of VO2max.

Adult↗

Validation of the Tunisian version of the Roland-Morris questionnaire.

Our aim was to validate a culturally adapted, Tunisian-language version of the Roland-Morris Disability Questionnaire (RMDQ), which is a reliable evaluation instrument for low-back-pain disability. A total of 62 patients with low back pain were assessed by the questionnaire. Reliability for the 1-week test/re-test was assessed by a construction of a Bland Altman plot. Internal construct validity was assessed by Cronbach's alphatest. External construct validity was assessed by association with pain, the Schober test and the General Function Score. Sensitivity to change was determined using a t-test for paired data to compare RMDQ scores at inclusion and at completion of the therapeutic sequence of local corticosteroid injections. We also compared the questionnaire score with the General Function Score, both taken after completion of the therapeutic sequence. The constructed Bland Altman plot showed good reliability. Internal consistency of the RMDQ was found to be very good and the Cronbach's alpha test was 0.94, indicating a good internal construct validity. The questionnaire is correlated with the pain visual analogue scale (r=33; p=0.0001), with the Schober test (r=0.27; p=0.0001) and the General Function Score (r=56; p=0.0001) indicating an adequate external construct validity. The RMDQ administered after the therapeutic sequence is sensitive to change (r=0.83; p=0.000). Comparison of the questionnaire score to the General Function Score, after completion of the therapeutic sequence, was satisfactory (r=0.75; p=0.000). We conclude that the Tunisian version of the Roland-Morris questionnaire has good reliability and internal consistency. Furthermore, it has a good internal- and external construct validity and high sensitivity to change. It is an adequate and useful tool for assessing low-back-pain disability.

Disability Evaluation↗

RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware probability analysis.

Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.

Atrial Fibrillation↗

Validation of the Turkish version of the Roland-Morris Disability Questionnaire for use in low back pain.

STUDY DESIGN: A reliability and validity study of a previously translated version of the Roland-Morris Disability Questionnaire (RMDQ). OBJECTIVES: To validate the Turkish version of the RMDQ for use in low back pain. SUMMARY OF BACKGROUND DATA: Clinical and epidemiologic research related to low back pain in the Turkish population would be facilitated by the availability of well-established outcome measures. METHODS: A total of 81 outpatients with low back pain, 64 of whom were followed up on a second occasion, were assessed by the RMDQ. Reliability was assessed using internal consistency and the intraclass correlation coefficient. Internal construct validity was assessed by Rasch analysis; external construct validity was assessed by association with pain and spinal movement. Responsiveness was tested by both the nonparametric and parametric effect sizes. RESULTS: Internal consistency of the RMDQ is found to be adequate (>0.85) at both times, with high intraclass correlation coefficient also at both time points. Internal construct validity of the scale is good, indicating a single underlying construct. Expected associations with pain confirm external construct validity. There is little evidence of differential item functioning. The scale is at the ordinal level. Responsiveness of the RMDQ is good and greater than observed change in spinal movement. CONCLUSIONS: The RMDQ is a robust unidimensional ordinal measure, largely free of differential item functioning, which works well in the Turkish population. Nonparametric effect sizes of ordinal scales are found to overestimate or underestimate the true effect size depending on the nature of the scale and the distribution of patients at baseline.

Adult↗

Effect of patient withdrawal on a study evaluating pharmacist management of hypertension.

STUDY OBJECTIVES: To examine potential threats to internal and external study validity caused by differential patient withdrawal from a randomized controlled trial evaluating pharmacist management of hypertension, to compare the characteristics of patients who withdrew with those of patients who completed the study, and to identify characteristics that predispose patients to withdraw from hypertension management. DESIGN: Prospective, randomized, comparative study. SETTING: Network of primary care clinics. PATIENTS: Four hundred sixty-three patients with a diagnosis of hypertension and a last documented systolic blood pressure of 160 mm Hg or greater and/or diastolic blood pressure of 100 mm Hg or greater. INTERVENTION: Patients were randomly allocated to the pharmacist intervention or usual-care (control) group. Those in the pharmacist intervention group were collaboratively managed by a primary care clinical pharmacy specialist and their primary care provider. Patients in the control group received usual care from only their primary care provider. MEASUREMENTS AND MAIN RESULTS: Of the 463 patients, 191 (41%) withdrew from the study after randomization and 272 (59%) completed the study. Patients who withdrew from the pharmacist intervention group were similar to patients who withdrew from the usual-care group with respect to age, sex, insurance status, and chronic conditions. Patients who smoked or had commercial insurance were more likely to withdraw from the study than the other participants. However, multivariate analysis of all variables, when adjusted for the effect of the intervention, revealed that insurance status was the only variable associated with a heightened probability of withdrawal (p=0.002). CONCLUSION: Although this study had a high withdrawal rate, between-group patient characteristics remained balanced. Therefore, internal validity was preserved, and outcomes from the study groups could be reliably compared. A lack of significant differences between patients who withdrew versus those who completed, with the exception of insurance status, suggests that external validity was not jeopardized.

Aged↗

Construct validity and reliability of the Rivermead Post-Concussion Symptoms Questionnaire.

OBJECTIVES: To provide further evidence of reliability and internal and external construct validity of the Rivermead Post-Concussion Symptoms Questionnaire (RPQ), which measures severity of postconcussion symptoms following head injury. DESIGN AND SETTING: A cross-sectional study of consecutive patients presenting with a head injury in two urban teaching hospitals and a community trust. PATIENTS: Three hundred and sixty-nine patients returned a questionnaire from 1689 consecutive adult patients (18 years and above) referred to radiology for a skull X-ray following a head injury, and those who were currently under the care of a community-based multidisciplinary head injury team. METHOD: Internal construct validity tested by fit to the Rasch Measurement model; external construct validity tested by correlations with Rivermead Head Injury Follow-up Questionnaire (RHFUQ); test-retest reliability tested by correlations at two-week intervals. OUTCOME MEASURES: Rivermead Post-Concussion Symptoms Questionnaire and Rivermead Head Injury Follow-up Questionnaire. MAIN RESULTS: RPQ scores ranged from 0 to 64 (17.3% floor, 0.3% ceiling). Overall fit to the Rasch model was poor (item fit mean -0.416, SD = 1.989, chi-squared= 172.486, p<0.01) suggesting a lack of unidimensionality. The items headaches, dizziness and forgetful displayed misfitting residuals and the first two items also displayed significant item trait fit statistics (p < 0.0006). After removing the items headaches, dizziness and subsequently nausea the RPQ demonstrated good fit at overall and individual item levels, both for the remaining 13 items (RPQ-13) and the three items (RPQ-3) which now formed a subsidiary scale. All items functioned consistently across age and gender. The RPQ-13 and RPQ-3 scales showed test-retest reliability coefficients of 0.89 and 0.72 (both p-values < 0.01) and positive correlations with RHFUQ scores (0.83 for RPQ-13, 0.62 for RPQ-3, both p-values < 0.01). CONCLUSIONS: As currently used, the RPQ does not meet modern psychometric standards. Its 16 items do not tap into the same underlying construct and should not be summated in a single score. When the RPQ is split into two separate scales, the RPQ-13 and the RPQ-3, each set of items forms a unidimensional construct for people with head injury at three months post injury. These scales show good test-retest reliability and adequate external construct validity.

Adolescent↗

External otitis caused by infection with Pseudomonas aeruginosa or Candida albicans cured by use of a topical group III steroid, without any antibiotics.

CONCLUSIONS: Irrespective of the microbial agent, group III steroid solution cured external otitis efficiently in a rat model. The addition of antibiotic components to steroid solutions for the treatment of external otitis is of questionable validity. OBJECTIVE: External otitis, caused by infection with either Pseudomonas aeruginosa or Candida albicans, was established in a rat model and the treatment efficacy of a group III steroid solution was studied. MATERIAL AND METHODS: Three treatments were studied: (i) a group III steroid solution; (ii) a group I steroid combined with two antibiotic components; and (iii) a saline solution. A scoring scale was used to evaluate the characteristics of the ear canal skin. Bacteriological and fungal samples were collected for culturing and ear canal skin biopsies were taken for structural analyses. RESULTS: It was possible to cause P. aeruginosa and C. albicans infections in an animal model. In the P. aeruginosa-infected animals, only the group III steroid treatment cured all the animals. In the C. albicans-infected animals, group III steroid treatment resolved external otitis faster than the other treatment modalities.

Administration, Topical↗

Issues in cross-cultural validity: example from the adaptation, reliability, and validity testing of a Turkish version of the Stanford Health Assessment Questionnaire.

OBJECTIVE: Guidelines have been established for cross-cultural adaptation of outcome measures. However, invariance across cultures must also be demonstrated through analysis of Differential Item Functioning (DIF). This is tested in the context of a Turkish adaptation of the Health Assessment Questionnaire (HAQ). METHODS: Internal construct validity of the adapted HAQ is assessed by Rasch analysis; reliability, by internal consistency and the intraclass correlation coefficient; external construct validity, by association with impairments and American College of Rheumatology functional stages. Cross-cultural validity is tested through DIF by comparison with data from the UK version of the HAQ. RESULTS: The adapted version of the HAQ demonstrated good internal construct validity through fit of the data to the Rasch model (mean item fit 0.205; SD 0.998). Reliability was excellent (alpha = 0.97) and external construct validity was confirmed by expected associations. DIF for culture was found in only 1 item. CONCLUSIONS: Cross-cultural validity was found to be sufficient for use in international studies between the UK and Turkey. Future adaptation of instruments should include analysis of DIF at the field testing stage in the adaptation process.

Activities of Daily Living↗

Single-cell expression quantitative trait locus Mendelian randomization reveals immune cell-specific causal regulatory networks and actionable targets in polycystic ovary syndrome.

ObjectiveTo systematically investigate whether the pathogenesis of polycystic ovary syndrome (PCOS) is causally related to dysregulated gene expression in specific immune cell subsets, and to evaluate the potential of these causal genes as actionable drug targets.MethodsThis study employed a two-sample Mendelian randomization (MR) framework using publicly available genome-wide association study (GWAS) summary statistics. The participant data included 797 PCOS cases and 140,558 controls (no direct patient recruitment was involved). Instrumental variables were derived from high-resolution immune cell-specific single-cell expression quantitative trait locus (sc-eQTL) data (OneK1K project) across 14 immune cell types. Primary analyses utilized the inverse-variance weighted (IVW) method. Shared causal variants were validated using Bayesian colocalization. Phenome-wide association analysis (PheWAS), external transcriptomic dataset validation (GSE8157), and DrugBank database screening were conducted for pleiotropy assessment and drug repositioning.ResultsMR analysis revealed genome-wide significant causal associations for GLIPR1 in non-classical monocytes (Mono NC) and XBP1 in CD4+ effector memory T cells (CD4 ET) with PCOS risk. Higher GLIPR1 expression was associated with a decreased PCOS risk (OR = 0.669, P = 4.34&#xd7;10-6), whereas higher XBP1 expression was associated with an increased risk (OR = 1.406, P = 9.53&#xd7;10-8). Colocalization analysis confirmed that GLIPR1 shares a causal variant with PCOS (PP.H4 = 96.73%). PheWAS and external validation confirmed the safety profile and significant upregulation (P = 0.03) of GLIPR1. Drug repositioning identified SOT-107, a Phase III protein therapy drug, as a potential interacting agent for GLIPR1.ConclusionsThis sc-eQTL MR study reveals immune cell-specific causal regulatory networks in PCOS. GLIPR1 in non-classical monocytes represents a high-confidence protective target, while XBP1 provides suggestive evidence for immune-mediated pathogenesis. The candidate drug SOT-107 highlights theoretical repositioning opportunities, though rigorous preclinical validation remains required.

Female↗

Extending McNemar's test: estimation and inference when paired binary outcome data are misclassified.

McNemar's test is popular for assessing the difference between proportions when two observations are taken on each experimental unit. It is useful under a variety of epidemiological study designs that produce correlated binary outcomes. In studies involving outcome ascertainment, cost or feasibility concerns often lead researchers to employ error-prone surrogate diagnostic tests. Assuming an available gold standard diagnostic method, we address point and confidence interval estimation of the true difference in proportions and the paired-data odds ratio by incorporating external or internal validation data. We distinguish two special cases, depending on whether it is reasonable to assume that the diagnostic test properties remain the same for both assessments (e.g., at baseline and at follow-up). Likelihood-based analysis yields closed-form estimates when validation data are external and requires numeric optimization when they are internal. The latter approach offers important advantages in terms of robustness and efficient odds ratio estimation. We consider internal validation study designs geared toward optimizing efficiency given a fixed cost allocated for measurements. Two motivating examples are presented, using gold standard and surrogate bivariate binary diagnoses of bacterial vaginosis (BV) on women participating in the HIV Epidemiology Research Study (HERS).

Biometry↗

Evaluation of a short scale to assess female sexual functioning.

This article provides information on the external and concurrent validity, test-retest reliability, and sensitivity to change of a short form of the Personal Experiences Questionnaire, which was adapted from the McCoy Female Sexuality Questionnaire. We drew participants from convenience samples of women attending three different clinic settings: family planning clinics, psychiatrists, and sex therapists. We chose the psychiatry and sex therapy clinics as samples likely to show poor sexual functioning in order to assist with external validity assessment and to establish a cut-off score indicating sexual dysfunction. Satisfactory external criterion validity, concurrent validity, reliability on re-test, and validation of the composite score were demonstrated. A cut-off score of 7 or below distinguishes with 79% specificity and sensitivity those with sexual dysfunction.

Adult↗