Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

The incorporation of potential confounding variables in Markov models.

OBJECTIVE: To improve the quality of the methods used in Markov modelling studies by increasing the external validity by means of the incorporation of confounding variables. STUDY DESIGN: The concepts were illustrated using a hypothetical Markov model for Parkinson's disease. METHODS: The methodology consisted of incorporation of an extra explanatory variable in the Markov health states by means of health state-specific relationships between this explanatory variable and costs as well as time-dependent values of the extra explanatory variable. In addition, we determined the relevance of the incorporation of an extra explanatory variable by means of various sensitivity analyses. RESULTS: The results showed that the outcomes of a health economic model may be severely biased, when a confounding effect of an extra explanatory variable is not taken into account. Hence the external validity of Markov models may be limited, and consequently the results of the model are not an accurate reflection of reality. CONCLUSION: This study proves the need for the incorporation of all relevant explanatory variables in a health economic model.

Antiparkinson Agents↗

Conceptual framework and systematic review of the effects of participants' and professionals' preferences in randomised controlled trials.

OBJECTIVES: To develop a conceptual framework of preferences for interventions in the context of randomised controlled trials (RCTs), as well as to examine the extent to which preferences affect recruitment to RCTs and modify the measured outcome in RCTs through a systematic review of RCTs that incorporated participants' and professionals' preferences. Also to make recommendations on the role of participants' and professionals' preferences in the evaluation of health technologies. DATA SOURCES: Electronic databases. REVIEW METHODS: The conceptual review was carried out on published papers in the psychology and economics literature concerning concepts of relevance to patient decision-making and preferences, and their measurement. For the systematic review, studies across all medical specialities meeting strict criteria were selected. Data were then extracted, synthesised and analysed. RESULTS: Key elements for a conceptual framework were found to be that preferences are evaluations of an intervention in terms of its desirability and these preferences relate to expectancies and perceived value of the process and outcome of interventions. RCTs differed in the information provided to patients, the complexity of techniques used to provide that information and the degree to which preference elicitation may simply produce pre-existing preferences or actively construct them. Most current RCTs used written information alone. Preference can be measured in many different ways and most RCTs did not provide quantitative measures of preferences, and those that did tended to use very simple measures. The second part of the study, the systematic review included 34 RCTs. The findings gave support to the hypothesis that preferences affect trial recruitment. However, there was less evidence that external validity was seriously compromised. There was some evidence that preferences influenced outcome in a proportion of trials. However, evidence for preference effects was weaker in large trials and after accounting for baseline differences. Preference effects were also inconsistent in direction. There was no evidence that preferences influenced attrition. Therefore, the available evidence does not support the operation of a consistent and important 'preference effect'. Interventions cannot be categorised consistently on degree of participation. Examining differential preference effects based on unreliable categories ran the risk of drawing incorrect conclusions, so this was not carried out. CONCLUSIONS: Although patients and physicians often have intervention preferences, our review gives less support to the hypothesis that preferences significantly compromise the internal and external validity of trials. This review adds to the growing evidence that when preferences based on informed expectations or strong ethical objections to an RCT exist, observational methods are a valuable alternative. All RCTs in which participants and/or professionals cannot be masked to treatment arms should attempt to estimate participants' preferences. In this way, the amount of evidence available to answer questions about the effect of treatment preferences within and outwith RCTs could be increased. Furthermore, RCTs should routinely attempt to report the proportion of eligible patients who refused to take part because of their preferences for treatment. The findings also indicate a number of approaches to the design, conduct and analysis of RCTs that take account of participants' and/or professionals' preferences. This is referred to as a methodological tool kit for undertaking RCTs that incorporate some consideration of patients' or professionals' preferences. Future research into the amount and source of information available to patients about interventions in RCTs could be considered, with special emphasis on the relationship between sources inside and outside the RCT context. Qualitative research undertaken as part of ongoing RCTs might be especially useful. The processes by which this information leads to preferences in order to develop or extend the proposed expectancy--value framework could also be examined. Other areas for consideration include: how information about interventions changes participants' preferences; a comparison of the feasibility and effectiveness of different informed consent procedures; how strength of preference varies for different interventions within the same RCT and how these differences can be taken account of in the analysis; the differential effects of patients' and professionals' preferences on evidence arising from RCTs; and whether the standardised measurement of preferences within all RCTs (and analysis of the effect on outcome) would allow the rapid development of a significant evidence base concerning patient preferences, albeit in relation to a single preference design.

Attitude of Health Personnel↗

Which children could benefit from additional diagnostic tools in case of suspected appendicitis?

BACKGROUND: New diagnostic tools such as ultrasound scan, computed tomography (CT) scan, and diagnostic laparoscopy, have become available for children with suspected appendicitis but should be reserved for equivocal cases. The aim of this study was to develop a scoring system to identify this subgroup of children. METHODS: Patients from 2 different periods (period 1, 99 consecutive children [group 1] and period 2, 62 consecutive children [group 2] with suspected appendicitis) were prospectively evaluated. Variables predicting appendicitis were obtained from group 1. By means of a regression analysis, a scoring system was created and applied to the patients of group 2. Missed appendicitis and negative appendectomy rates obtained by clinical practice were compared with the results that would have been accomplished based on the scoring system. Thereafter, the scoring system was externally validated in a group of children presented at another hospital (group 3, n = 114). RESULTS: The variables, leukocyte count > or = 10.10(9)/L (2 points); rebound tenderness (2 points); and temperature > or = 38 degrees C (1 point) correlated significantly with appendicitis. The scoring system was used to categorize patients into 3 groups: appendicitis unlikely, doubtful appendicitis, and suspected appendicitis. The specificity and sensitivity of the scoring system were, respectively, 85% and 89%. Applying the scoring system would lead to comparable negative appendectomy rates of 8% versus 6% using clinical judgement and a comparable number of performed laparoscopies (26% v 31%). However, it could lead to a lower missed appendicitis rate (1% v 6%) and a lower perforation rate (0% v 11%). External validation showed comparable performed laparoscopies (32%) and missed appendicitis (2%) rates but a higher negative appendectomy rate (19%), probably owing to a lower percentage of appendicitis in hospital (2, 47%) compared with hospital (1, 71%). CONCLUSIONS: Children can be observed if leukocyte count is less than 10.10(9)/L and rebound tenderness is absent; a diagnostic laparoscopy should be performed if one of these is present, and if both are present one could perform an appendectomy.

Abdominal Pain↗

Integrated Genomic and Proteomic Analysis Reveals T-B Lymphocyte Signatures in the MYCN Driven "Immune Desert" of Specific Neuroblastoma Subtypes.

AIMS: This study aims to systematically dissect how MYCN amplification shapes the immunosuppressive tumor microenvironment (TME) in high-risk neuroblastoma, elucidating key mechanisms underlying immune evasion. METHODS: We performed an integrated multi-omics analysis of bulk RNA-seq (n = 721), single-cell RNA-seq (n = 9), proteomic data (n = 49) and spatial transcriptomics (Visium, with external validation in melanoma). Analyses included unsupervised clustering, cell-cell communication inference, transcriptional regulatory network reconstruction, and spatial proximity assessment to map the immune landscape. RESULTS: A distinct molecular subtype (Class C), defined by MYCN amplification and poor prognosis, exhibited a comprehensive "immune desert" phenotype characterized by low immune scores and minimal leukocyte infiltration. Single-cell analysis confirmed significant depletion of T and B lymphocytes within the Class C TME. Dysregulated transcriptional networks were identified, including upregulation of REL and EOMES in T cells-with EOMES potentially driving exhaustion via regulation of Transient Receptor Potential (TRP) genes, and REL inhibition enhancing cytotoxic function in vitro. A unique immunosuppressive B-cell subset (B7) engaged in enhanced crosstalk with exhausted T cells and harbored a MYC-centered network linked to cell cycle dysregulation and poor survival. Spatial transcriptomics revealed significant proximity between B7-active regions and Treg/exhaustion-enriched areas, externally validated in melanoma. Proteomic data validated elevated REL expression in MYCN-amplified tumors. CONCLUSION: This work delineates the immunosuppressive architecture of MYCN-driven neuroblastoma, revealing novel regulatory nodes within specific lymphocyte compartments. Integrating single-cell, spatial, and proteomic evidence, we propose REL inhibition as a therapeutic candidate, the EOMES/TRP axis as a bioinformatically supported hypothesis, and the B7/MYC hub as a hypothesis supported by transcriptomic and spatial evidence.

Humans↗

Integrating genetic predictors into subsequent breast cancer risk prediction in survivors of childhood cancer.

PURPOSE: Female survivors of childhood cancer are at high risk for developing breast cancer. The contributions of most general population primary breast cancer genetic predictors to this risk have not been explored. METHODS: Analyses included females who survived &#x2265;5 years after their childhood cancer diagnosis with available array (N&#x2009;=&#x2009;2096, subsequent breast cancer [SBC]=218) or whole-genome sequencing (WGS; N&#x2009;=&#x2009;3292, SBC=101) data from the Childhood Cancer Survivor Study and St. Jude Lifetime Cohort. We computed 99 externally-validated primary breast cancer polygenic risk scores (PRS). Using deep-coverage WGS, ClinVar-annotated pathogenic/likely pathogenic (P/LP) variants in breast cancer susceptibility genes were identified. Cox proportional hazards models assessed associations with SBC risk, adjusting for treatments and genetic ancestry. RESULTS: Among 5388 female survivors (genetic ancestry, European: N&#x2009;=&#x2009;4,752; African: N&#x2009;=&#x2009;444; East Asian: N&#x2009;=&#x2009;192), 319 developed SBC. Most (90.9%) PRSs were nominally associated with SBC risk (P&#x2009;<&#x2009;0.05), but effect sizes varied substantially. PRSs with superior discriminatory ability had greater genome-wide coverage (e.g., 6.4 million-variant PRS, HR per SD&#x2009;=&#x2009;1.71, 95% CI&#x2009;=&#x2009;1.43 to 2.05; P&#x2009;=&#x2009;4.2x10-9) and 7.7-fold higher odds (P&#x2009;=&#x2009;7.0x10-4) of including variants in multiple DNA damage repair pathways compared with PRSs with weaker risk associations. Among survivors with WGS, 1.6% carried P/LP variants in clinical testing panel genes, which was associated with a 7.4-fold greater risk (95% CI&#x2009;=&#x2009;3.16 to 17.19). Including genetic factors improved SBC risk prediction by age 40 (P&#x2009;<&#x2009;0.001) compared to treatment exposures alone. CONCLUSIONS: Externally-validated primary breast cancer genetic susceptibility predictors are relevant for SBC risk prediction and should be prioritized for risk stratification in survivors.

Journal Article↗

Comparison of the Radiation Therapy Oncology Group recursive partitioning classification and Union Internationale Contre le Cancer TNM classification for patients with head and neck carcinoma.

BACKGROUND: Prognostic models need to be tested in external validation studies to assess generalizability. Recursive partitioning analysis (RPA), a prognostic system based on the creation of a classification tree, has been proposed as a classification method in patients with head and neck carcinoma. The aim of this study was to compare the RPA and Union Internationale Contre le Cancer (UICC) TNM classification systems in patients with head and neck carcinoma treated consecutively in a single center. METHODS: A total of 2166 patients with carcinomas of the oral cavity, oropharynx, hypopharynx, and larynx was classified according to both the RPA and the TNM classification systems, and the results were compared. The endpoints considered were observed survival and survival free of locoregional tumor. The two methods of classification were evaluated objectively by use of measures of intrastage homogeneity (hazard consistency), interstage heterogeneity (hazard discrimination), predictive power (outcome prediction), and patient distribution between stages (balance). RESULTS: When the endpoint considered was observed survival, there were no clinically relevant differences between the two classifications. However, when the endpoint was locoregional control, the RPA system was sensitive to the type of treatment used, and it was not generalizable. CONCLUSIONS: To evaluate generalizability, new classification proposals need external validation studies that objectively measure the quality of the model. The performance of the RPA system was not reproducible in our cohort of patients when the endpoint evaluated was locoregional control.

Head and Neck Neoplasms↗

A postoperative prognostic nomogram predicting recurrence for patients with conventional clear cell renal cell carcinoma.

PURPOSE: Few published studies have simultaneously analyzed multiple prognostic factors to predict recurrence after surgery for conventional clear cell renal cortical carcinomas. We developed and performed external validation of a postoperative nomogram for this purpose. We used a prospectively updated database of more than 1,400 patients treated at a single institution. MATERIALS AND METHODS: From January 1989 to August 2002, 833 nephrectomies (partial and radical) for renal cell carcinoma of conventional clear cell histology performed at Memorial Sloan-Kettering Cancer Center were reviewed from the center's kidney database. Patients with von Hippel-Lindau disease or familial syndromes, as well as patients presenting with synchronous bilateral renal masses, or distant metastases or metastatic regional lymph nodes before or at surgery were excluded from study. We modeled clinicopathological data and disease followup for 701 patients with conventional clear cell renal cell carcinoma. Prognostic variables for the nomogram included pathological stage, Fuhrman grade, tumor size, necrosis, vascular invasion and clinical presentation (ie incidental asymptomatic, locally symptomatic or systemically symptomatic). RESULTS: Disease recurrence was noted in 72 of 701 patients. Those patients without evidence of disease had a median and maximum followup of 32 and 120 months, respectively. The 5-year probability of freedom from recurrence for the patient cohort was 80.9% (95% confidence interval 75.7% to 85.1%). A nomogram was designed based on a Cox proportional hazards regression model. Following external validation predictions by the nomogram appeared accurate and discriminating, and the concordance index was 0.82. CONCLUSIONS: A nomogram has been developed that can be used to predict the 5-year probability of freedom from recurrence for patients with conventional clear cell renal cell carcinoma. This nomogram may be useful for patient counseling, clinical trial design and effective patient followup strategies.

Carcinoma, Renal Cell↗

Recruitment issues, health habits, and the decision to participate in a health promotion program.

To understand the external validity of experimental studies, it is important to estimate the extent to which the participants are representative of the general population. This paper describes recruitment methods and considers the representativeness of participants in the San Diego Family Health Project. The study was designed to experimentally evaluate the effectiveness of a family-based behavior change intervention in Anglo and Mexican-American families. Initial contact with the families was made through a household health survey that was sent home with all fifth- and sixth-grade children in 12 participating elementary schools. The survey asked about a variety of demographic characteristics, dietary habits, and physical activity habits. Parents were also asked if they were interested in participating in the project. Respondents were classified by level of participation into one of three groups: not interested, expressed initial interest but did not attend the recruitment meeting, and volunteered to participate. Level of participation was the independent variable in the analyses. In separate analyses for Anglo and Mexican-American responders, our data suggested many similarities and a few differences among participant groups. The differences that were observed suggest that participants may already have healthier diets than nonparticipants, although only one of four dietary variables differed by participation status in each ethnic group. The external validity of these data and general recruitment issues are discussed.

Adolescent↗

An external construct validity study of Rorschach personality variables.

This study examined (a) hypothesized relationships between Rorschach variables and self-report test measures relating to nominally similar aspects of personality functioning and (b) interrelationships among Rorschach variables. Sixty-two undergraduates were administered the Rorschach, Barron Ego Strength Scale, Kaplan Self-Derogation Scale, Eagly Self-Esteem Scale, Multiple Affective Adjective Checklist (MAACL), Marlowe-Crowne Social Desirability Scale, and the Rotter Locus of Control Scale. Only a few of the predictions received confirmation: inanimate movement (m) correlated, as expected, with MAACL anxiety and hostility, the egocentricity index (3r + 2)/R (R = total responses) correlated significantly with self-esteem, and human movement with minus form level (M-) correlated (inversely) with ego strength. Among the unpredicted findings were some that appear inconsistent with standard Rorschach interpretation. Rorschach variables human movement (M), and experience actual (EA), generally interpreted as reflecting coping resources, related significantly with self-report measures of poor coping and of dysphoric affect. In general, the Rorschach appears better at identifying weaknesses in the ego rather than strengths.

Adolescent↗

Non-randomised patients in a cholecystectomy trial: characteristics, procedures, and outcomes.

BACKGROUND: Laparoscopic cholecystectomy is now considered the first option for gallbladder surgery. However, 20% to 30% of cholecystectomies are completed as open operations often on elderly and fragile patients. The external validity of randomised trials comparing mini-laparotomy cholecystectomy and laparoscopic cholecystectomy has not been studied. The aim of this study is to analyse characteristics, procedures, and outcomes for all patients who underwent cholecystectomy without being included in such a trial. METHODS: Characteristics (age, sex, co-morbidity, and ASA-score), operation time, hospital stay, and mortality were compared for patients who underwent cholecystectomy outside and within a randomised controlled trial comparing mini-laparotomy and laparoscopic cholecystectomy. RESULTS: During the inclusion period 1719 patients underwent cholecystectomy. 726 patients were randomised and 724 of them completed the trial; 993 patients underwent cholecystectomy outside the trial. The non-randomised patients were older--and had more complications from gallstone disease, higher co-morbidity, and higher ASA--score when compared with trial patients. They were also more likely to undergo acute surgery and they had a longer postoperative hospital stay, with a median 3 versus 2 days (p < 0.001 for all comparisons). Standardised mortality ratio within 90 days of operation was 3.42 (mean) (95% CI 2.17 to 5.13) for non-randomised patients and 1.61 (mean) (95%CI 0.02 to 3.46) for trial patients. For non-randomised patients, operation time did not differ significantly between mini-laparotomy and open cholecystectomy in multivariate analysis. However, the operation for laparoscopic cholecystectomy lasted 20 minutes longer than open cholecystectomy. Hospital stay was significantly shorter for both mini-laparotomy and laparoscopic cholecystectomy compared to open cholecystectomy. CONCLUSION: Non-randomised patients were older and more sick than trial patients. The assignment of healthier patients to trials comparing mini-laparotomy cholecystectomy and laparoscopic cholecystectomy limits the external validity of conclusions reached in such trials.

Adult↗

Irritable bowel syndrome: attrition rates of patients identified at primary care centers during a 50-week period versus those identified in hospitals in a phase II clinical trial.

Although most irritable bowel syndrome (IBS) patients are managed in primary care centers, most trials are performed in the hospital setting. A successful recruitment strategy is important for trial completeness and its external validity. To this end, it was important to assess the attrition rates of patients identified at primary care centers and hospitals in this phase II trial conducted in the outpatient clinics of 16 hospitals. IBS patients were identified through review of the centers' files (prescreening). After a 2-week single-blind placebo screening phase, the patients were randomized to receive an investigational drug or a matched placebo for 24 weeks (dose-ranging study). Thereafter, the patients were invited to participate in a 24-week, double-blind extension study. The attrition rates among patients identified at hospitals (group A) and at primary care centers (group B) were compared in each study phase and during the 50-week period by bivariate and multivariate regression analyses. Group A and B patients were identified in 13 hospitals and 51 primary care centers, respectively. Of 1,001 prescreened patients, 302 started the screening phase with attrition rates of 35% (of 132 patients) and 25% (of 170) for groups A and B, respectively (p = 0.054). The attrition rate during the double-blind phase was 14%. Of the 184 patients who completed the dose-ranging study, 39 (group A: 32%; group B: 14%; p = 0.005) did not wish to participate in the extension study. The attrition rate in the extension study was 15%. The overall attrition rates during the 50-week period were 67% and 53% (p = 0.016) for groups A and B, respectively. The multivariate regression analysis showed that the screening phases (p = 0.000) and unwillingness to participate in the extension study (p = 0.007) had the highest impact on attrition rates. We conclude that referring patients identified in primary care centers to hospitals seems appropriate to ensure potentially eligible patients for IBS trials. Patients identified in primary care centers are more likely than those identified in hospitals to participate in a 24-week extension study, which may be due to their positive feelings about being treated in a hospital rather than being referred back to their original primary care center. This strategy may be considered in future trials since, with a reasonably low attrition rate, it would enhance the external validity of the results obtained.

Adult↗

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans↗

Rational selection of training and test sets for the development of validated QSAR models.

Quantitative Structure-Activity Relationship (QSAR) models are used increasingly to screen chemical databases and/or virtual chemical libraries for potentially bioactive molecules. These developments emphasize the importance of rigorous model validation to ensure that the models have acceptable predictive power. Using k nearest neighbors (kNN) variable selection QSAR method for the analysis of several datasets, we have demonstrated recently that the widely accepted leave-one-out (LOO) cross-validated R2 (q2) is an inadequate characteristic to assess the predictive ability of the models [Golbraikh, A., Tropsha, A. Beware of q2! J. Mol. Graphics Mod. 20, 269-276, (2002)]. Herein, we provide additional evidence that there exists no correlation between the values of q2 for the training set and accuracy of prediction (R2) for the test set and argue that this observation is a general property of any QSAR model developed with LOO cross-validation. We suggest that external validation using rationally selected training and test sets provides a means to establish a reliable QSAR model. We propose several approaches to the division of experimental datasets into training and test sets and apply them in QSAR studies of 48 functionalized amino acid anticonvulsants and a series of 157 epipodophyllotoxin derivatives with antitumor activity. We formulate a set of general criteria for the evaluation of predictive power of QSAR models.

Algorithms↗

Determination of total sulfur in diesel fuel employing NIR spectroscopy and multivariate calibration.

A method for sulfur determination in diesel fuel employing near infrared spectroscopy, variable selection and multivariate calibration is described. The performances of principal component regression (PCR) and partial least square (PLS) chemometric methods were compared with those shown by multiple linear regression (MLR), performed after variable selection based on the genetic algorithm (GA) or the successive projection algorithm (SPA). Ninety seven diesel samples were divided into three sets (41 for calibration, 30 for internal validation and 26 for external validation), each of them covering the full range of sulfur concentrations (from 0.07 to 0.33% w/w). Transflectance measurements were performed from 850 to 1800 nm. Although principal component analysis identified the presence of three groups, PLS, PCR and MLR provided models whose predicting capabilities were independent of the diesel type. Calibration with PLS and PCR employing all the 454 wavelengths provided root mean square errors of prediction (RMSEP) of 0.036% and 0.043% for the validation set, respectively. The use of GA and SPA for variable selection provided calibration models based on 19 and 9 wavelengths, with a RMSEP of 0.031% (PLS-GA), 0.022% (MLR-SPA) and 0.034% (MLR-GA). As the ASTM 4294 method allows a reproducibility of 0.05%, it can be concluded that a method based on NIR spectroscopy and multivariate calibration can be employed for the determination of sulfur in diesel fuels. Furthermore, the selection of variables can provide more robust calibration models and SPA provided more parsimonious models than GA.

Journal Article↗

Threats to validity in the longitudinal study of psychological effects: the case of short stature.

In all studies of health-related problems and their effects on well-being, research design issues threaten to compromise the validity of findings. This is particularly so in a longitudinal study, essentially stemming from the tension between maintaining participant compliance and retaining investigator objectivity. Such a tension may be exacerbated where measures of dependent variables such as self-esteem are used alongside the collection of physical data which is essential to the study, as in research into the psychological effects of short stature on children and young people. In this paper one particular project, the Wessex Growth Study, is used to illustrate the common threats to validity, both internal and external, of such research, and to consider future improvements in design. The Wessex Growth Study, set up in 1986, was designed to overcome some of the methodological problems found in earlier research with short stature children. It is following the growth and psychological development through their school years of a cohort of short children (below third centile for height when first identified) and case-matched controls (10th-90th centiles) recruited at school entry (ages 5/6). Findings have generally found only small differences between short and average height children. Though these results so far have mainly been presented cross-sectionally, the young people involved are followed up at 6-monthly intervals for height and other data to be collected, and thus to some extent the study also has the advantages and problems of a longitudinal research design. Using Campbell and Stanley's criteria this article makes clear the strain on both internal and external validity in the study, but argues that these problems are to some extent inherent in all longitudinal psychological research, and are outweighed in the present research by the collection of data on short stature which would not otherwise be available. Future data collection within the study will introduce further improvements in design.

Adolescent↗

E-state modeling of fish toxicity independent of 3D structure information.

Topological structure methods are used to model fish toxicity against three classes of organic chemicals. The models were obtained independent of 3D structure information. Further, no mechanism of partitioning was assumed, thus avoiding the problems associated with selection of partitioning system for computation of log P. QSAR models were developed for a set of 92 compounds, including phenols, anilines and substituted aromatic hydrocarbons, yielding excellent statistics: r2 = 0.87, s = 0.25 and q2 = 0.85 leave-one-out (LOO), that are better than those reported in the literature. The model is based on molecular connectivity valence chi-1 index [1chiv], the atom type E-State indices for chlorine [ST(-Cl)] and for ether oxygen [ST(-O-)], and the maximum hydrogen E-State atom value in a molecule [Hmax]. Each of the subgroups was also separately well modeled. The model for the full set is validated through use of external validation test sets and ten-fold cross-validation (repeated three times). The quality of the validation statistics supports the claim that the model may be used for estimation of pLC50 values for similar molecules. Detailed structure interpretation is given for the descriptors in the model. These four structure descriptors encode influence of molecular context of groups as well as counts of those groups, in addition to molecular skeletal structure.

Aniline Compounds↗

Validation of a QSAR model for acute toxicity.

In the present study, a quantitative structure--activity relationship (QSAR) model has been developed for predicting acute toxicity to the fathead minnow (Pimephales promelas), the aim being to demonstrate how statistical validation and domain definition are both required to establish model validity and to provide reliable predictions. A dataset of 408 heterogeneous chemicals was modelled by a diverse set of theoretical molecular descriptors by using multivariate linear regression (MLR) and Genetic Algorithm-Variable Subset Selection (GA-VSS). This QSAR model was developed to generate reliable predictions of toxicity for organic chemicals not yet tested, so particular emphasis was given to statistical validity and applicability domain. External validation was performed by using OECD Screening Information Data Set (SIDS) data for 177 High Production Volume (HPV) chemicals, and a good predictivity was obtained (=72.1). The model was evaluated according to the OECD principles for QSAR validation, and compliance with all five principles was established. The model could therefore be useful for the regulatory assessment of chemicals. For example, it could be used to fill data gaps within its chemical domain and contribute to the prioritization of chemicals for aquatic toxicity testing.

Animals↗

Evading detection on the MMPI-2: does caution produce more realistic patterns of responding?

Studies on MMPI and MMPI-2 malingering indexes often sacrifice generalizability in an attempt to control internal validity. This study improves external validity while still maintaining internal validity by providing graduate student participants with a realistic context for malingering on the MMPI-2 (n=94) and MMPI (n=30). Contextual parameters include a realistic life predicament, psychological knowledge, an incentive, the presence versus absence of a specific diagnosis, and a caution to be realistic. This study found that cautioning participants not to overexaggerate their responses significantly improves their ability to evade detection on the MMPI-2 and MMPI. Standard malingering indexes (Infrequency, F; Back Side, F, Fb; F-Correction, F-K; and Infrequency-Psychopathology, F(p)) were insufficiently sensitive in identifying simulators using common cutoff scores for these cautious simulators.

Adult↗