Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Developmental sequel from early nutritional deficiencies: conclusive and probability judgements.

Results from quasi-experimental longitudinal studies of children and from experimental research with animal models have led several investigators to state that early iron deficiency anemia leaves a permanent cognitive deficit. However, neither source of information provides a basis for such a claim. Some key confounders were not controlled by the quasi-experimental studies, and the external validity of the animal data is questionable. Further, three decades of research on the functional consequences of protein-energy malnutrition have shown that the social environment moderates the effects of an early nutritional insult; it can keep such effect unchanged, or increase or decrease its severity. The prediction of later life on the basis of a particular nutritional event carries a large error factor, which suggests that the search would be more fruitful if we tracked probability statements.

Anemia, Iron-Deficiency↗

Preoperative nomogram predicting the 10-year probability of prostate cancer recurrence after radical prostatectomy.

An existing preoperative nomogram predicts the probability of prostate cancer recurrence, defined by prostate-specific antigen (PSA), at 5 years after radical prostatectomy based on clinical stage, serum PSA, and biopsy Gleason grade. In an updated and enhanced nomogram, we have extended the predictions to 10 years, added the prognostic information of systematic biopsy results, and enabled the predictions to be adjusted for the year of surgery. Cox regression analysis was used to model the clinical information for 1978 patients treated by two high-volume surgeons from our institution. The nomogram was externally validated on an independent cohort of 1545 patients with a concordance index of 0.79 and was well calibrated with respect to observed outcome. The inclusion of the number of positive and negative biopsy cores enhanced the predictive accuracy of the model. Thus, a new preoperative nomogram provides robust predictions of prostate cancer recurrence up to 10 years after radical prostatectomy.

Aged↗

Empirically supported treatments in pediatric psychology: where is the diversity?

OBJECTIVE: To examine the extent to which studies used to support empirically supported treatments for asthma, cancer, diabetes, and obesity address issues of cultural diversity. METHOD: We chose original articles (71) of treatments used to support empirically supported treatments (ESTs) published as part of a special series on ESTs in the Journal of Pediatric Psychology. Trained coders reviewed each study to determine if the following were reported: race/ethnicity and socioeconomic status (SES) of the sample, moderating cultural variables, cultural assumptions or biases of the treatment, larger cultural issues, and measurement or procedure bias. RESULTS: Results revealed that few studies addressed cultural variables in any way. Only 27% of the studies reported the race or ethnicity and 18% reported the SES of research participants. Additionally, 6% discussed potential moderating cultural variables. The remaining variables were addressed in 7% or less of the studies. CONCLUSIONS: These data support the criticism that ESTs fail to address important issues of culture and call into question the external validity of ESTs to diverse populations. Future research should explicitly address cultural issues according to the nine recommendations described here.

Asthma↗

Three new consensus QSAR models for the prediction of Ames genotoxicity.

Three QSAR methods, artificial neural net (ANN), k-nearest neighbors (kNN), and Decision Forest (DF), were applied to 3363 diverse compounds tested for their Ames genotoxicity. The ratio of mutagens to non-mutagens was 60/40 for this dataset. This group of compounds includes >300 therapeutic drugs. All models were developed using the same initial set of 148 topological indices: molecular connectivity chi indices and electrotopological state indices (atom-type, bond-type and group-type E-state), as well as binary indicators. While previous studies have found logP to be a determining factor in genotoxicity, it was not found to be important by any modeling method employed in this study. The three models yielded an average training/test concordance value of 88%, with a low percentage of false positives and false negatives. External validation testing on 400 compounds not used for QSAR model development gave an average concordance of 82%. This value increased to 92% upon removal of less reliable outcomes, as determined by a reliability criterion used within each model. The ANN model showed the best performance in predicting drug compounds, yielding 97% concordance (34/35 drugs) after the removal of less reliable predictions. The appreciable commonality found among the top 10 ranked descriptors from each model is of particular interest because of the diversity in the learning algorithms and descriptor selection techniques employed in this study. Forty percent of the most important descriptors in any one model are found in one or two other models. Fourteen of the most important descriptors relate directly to known toxicophores involved in potent genotoxic responses in Salmonella typhimurium. A comparison of the validation results with those of MULTICASE and DEREK indicated that the new models presented in this work perform substantially better than the former models in predicting genotoxicity of therapeutic drugs. Substantially higher specificity was achieved with these new models as compared with MULTICASE or DEREK with comparable sensitivities among all models.

Algorithms↗

Critical review of the quality and development of randomized clinical trials (RCTs) and their influence on the treatment of advanced epithelial ovarian cancer.

Trials on chemotherapy of advanced ovarian cancer published between 1975-88 were systematically reviewed for quality (according to the method of Chalmers) and consistency of tested hypotheses with a view to a meta-analysis of all published studies in the field. Median overall, internal and external validity scores were 47%, 43% and 53%, respectively. No association was found between scores and key features of trials, such as percentage studies with significant results in response or survival or percentage studies with high or low follow-up retention (withdrawal rates less than or greater than or equal to 15%). Only 21% of trials reported a fully blind randomization procedure and only in 13% were drop-outs accounted for by the intent-to-treat method. Only 4 trials entered more than 150 patients per arm, a sample size consistent with detection of an absolute difference of 11% in mortality. The majority of trials (58%) investigated the role of combination regimens versus a single-agent control arm. The remaining trials tested different polychemotherapies. However, within these two general issues, treatment options were quite heterogeneous: seven subgroups were identified by whether cisplatin was present in either the treatment or the control arm. We conclude that the internal coherence and development of randomized clinical trials in advanced ovarian cancer and their methodologic soundness are quite poor. In this situation meta-analysis cannot go beyond a systematic attempt to answer a very general "treatment effectiveness" question.

Antineoplastic Agents↗

Health economic models: a question of balance--summary of an open discussion on the pharmacoeconomic evaluation of non-steroidal anti-inflammatory drugs.

Pharmacoeconomics and pharmacoeconomic models are increasingly being used to guide health care decisions. In designing and using these models, an appropriate balance must be struck between scientific rigour and model transparency. It is therefore important to consider carefully how the various model components, such as model perspective, internal and external validity, and choice of comparators and outcomes, should be integrated into the model. These factors are discussed in relation to the pharmacoeconomic evaluation of non-steroidal anti-inflammatory drugs.

Anti-Inflammatory Agents, Non-Steroidal↗

Epidemiology of schizophrenia: a European perspective.

Since its inception, the concept of dementia praecox and, later, of schizophrenia has been one of the most disputed entities in modern medicine. Schizophrenia was, and still is, defined by its clinical symptoms and their characteristic evolution over time. No external validating criteria for the diagnosis have been established, in spite of a host of suggestive biological findings, among which the genetic data carry most weight. This absence of clear-cut substrate markers and indicators underscores the importance of the epidemiological perspective in the study of the disorder. The European contributions to the epidemiological description and understanding of schizophrenic morbidity are numerous. They range from community surveys and studies of pedigrees to case-control designs for assessment of risk factors and long-term followup investigations of course and outcome. This review focuses on epidemiological approaches to schizophrenia that have attempted to highlight the essential attributes of a disease: its incidence and prevalence, ecology, and associated features. It is difficult to generalize about European epidemiological research in schizophrenia, because of the coexistence of a variety of "schools," traditions, and approaches. It is nevertheless possible to discern several clear trends in European epidemiological investigations of schizophrenia that can, to some extent, be contrasted to North American developments.

Adolescent↗

The multidimensionality of schizotypy.

We present an overview of self-report scales for measuring schizotypy and a review of factor-analytical studies of these scales. These studies show that schizotypy is a multidimensional construct consisting of three or four factors. Positive Schizotypy, Negative Schizotypy, Nonconformity, and possibly Social Anxiety/Cognitive Disorganization. Clinical and external validation studies provide support for the construct validity of the Positive Schizotypy and Negative Schizotypy factors, but as yet fail to support the Nonconformity and Social Anxiety/Cognitive Disorganization factors. In accordance with this multidimensional structure, the scales for measuring schizotypy can be classified as factor-specific scales. We consider the striking similarities between the multidimensionality of schizotypal traits and the multidimensionality of schizophrenic symptoms. We also look at the similarities and differences between schizotypy and normal personality traits. Some practical and theoretical implications of these relationships are discussed.

Diagnosis, Differential↗

Scaling community attitudes toward the mentally ill.

The measurement of public attitudes toward the mentally ill has taken on new significance since the introduction of community-based mental health care. Previous attitude scales have been constructed and applied primarily in a professional context. This article discusses the development and application of a new set of four scales explicitly designed to measure community attitudes toward the mentally ill. The scales represent dimensions included in previous instruments, specifically, authoritarianism, benevolence, social restrictiveness, and community mental health ideology, but are expressed in terms of an almost completely new set of items that emphasize community contact with the mentally ill and mental health facilities. Data from a study of community attitudes about neighborhood mental health facilities in Toronto are used to test the internal and external validity of the scales. Results of the analysis provide strong support for the validity of the scales and demonstrate their usefulness as explanatory and predictive variables for studying community response to mental health facilities.

Age Factors↗

Cancer patient accessions into clinical trials: a pilot investigation into some patient and physician determinants of entry.

This study investigated the external validity (or generalizability of results) of randomized clinical trials in cancer. Tao ECOG lung cancer chemotherapy protocols active in the early 1970s were studied using a case-control design. All lung cancer patients of the four specified cell types resident in Monroe County during the ECOG study period were identified from the Rochester Regional Tumor Registry. All of the patients entered into either protocol ("ECOG cases") and a random sample of the nonprotocol cases were examined by medical records review. Thirty-seven percent of the nonprotocol cases were determined to have been eligible for either of the two ECOG protocols, but not entered ("eligible controls"). A comparison of the ECOG cases (n = 65) and the eligible controls (n = 109) revealed that (1) ECOG cases were more likely than eligible controls to have been diagnosed at a hospital which participated in the University of Rochester Cancer Center's medical oncology program; (2) ECOG cases were of higher occupational status than eligible controls; (3) duration from diagnosis to protocol entry for ECOG cases was longer than duration from diagnosis to earliest date of eligibility for eligible controls. The implications of these findings for the conduct of cancer clinical trials are discussed.

Aged↗

Laparoscopic cholecystectomy versus mini-laparotomy cholecystectomy: a prospective, randomized, single-blind study.

OBJECTIVE: To analyze outcomes after open small-incision surgery (minilaparotomy) and laparoscopic surgery for gallstone disease in general surgical practice. METHODS: This study was a randomized, single-blind, multicenter trial comparing laparoscopic cholecystectomy (LC) to minilaparotomy cholecystectomy (MC). Both elective and acute patients were eligible for inclusion. All surgeons normally performing cholecystectomy, both trainees under supervision and consultants, operated on randomized patients. LC was a routine procedure at participating hospitals, whereas MC was introduced after a short training period. All nonrandomized cholecystectomies at participating units during the study period were also recorded to analyze the external validity of trial results. The randomization period was from March 1, 1997, to April 30, 1999. RESULTS: Of 1,705 cholecystectomies performed at participating units during the randomization period, 724 entered the trial and 362 patients were randomized to each of the procedures. The groups were well matched for age and sex, but there were fewer acute operations in the LC group than the MC group. In the LC group 264 and in the MC group 150 operations were performed by surgeons who had done more than 25 operations of that type. Median operating times were 100 and 85 minutes for LC and MC, respectively. Median hospital stay was 2 days in each group, but in a nonparametric test it was significantly shorter after LC. Median sick leave and time for return to normal recreational activities were shorter after LC than MC. Intraoperative complications were less frequent in the MC group, but there was no difference in the postoperative complication rate between the groups. There was one serious bile duct injury in each group, but no deaths. CONCLUSIONS: Operating time was longer and convalescence was smoother for LC compared with MC. Further analyses of LC versus MC are necessary regarding surgical training, surgical outcome, and health economy.

Bile Ducts↗

Comparing peer and faculty evaluations in an internal medicine residency.

PURPOSE: To compare in-training evaluations of residents by their peers with evaluations by faculty preceptors in an internal medicine residency. METHOD: The study group consisted of 22 residents enrolled in the core (three-year) internal medicine program at the University of Calgary in 1989-90 and 1990-91. At the end of each rotation, ratings of the residents were requested from faculty preceptors and from peers for several categories of clinical competence. The peer ratings were paired with faculty ratings, for a total of 74 pairs of ratings. The Wilcoxon matched-pair signed-rank procedure was used to compare the paired ratings. One-way analysis of variance was used to compare the peer and faculty ratings with the residents' scores on three other kinds of evaluation used by the residency. RESULTS: While there was no significant difference between peer and faculty ratings for overall competence or for several components of competence, there were significant differences for some components, with faculty tending to rate higher than peers. The latter components were physical examination, team relationships, industriousness and enthusiasm, teaching, physician-patient relationships, and case presentations. External validation of the ratings by comparing them with other kinds of evaluation yielded little meaningful information. CONCLUSION: That the faculty ratings were significantly higher than the peer ratings for some components of clinical competence suggests that there were differences in the quality of evaluation between the peers and faculty, or differences in the standards or expectations of the two groups.

Clinical Competence↗

Single-case experimental designs in medical education: an innovative research method.

This paper presents an argument for more extensive use of single-case experimental research designs in medical education research. Single-case experimental designs consist of a group of experimental techniques that are widely used in the social sciences but are just beginning to be utilized by medical researchers. The method emphasizes reliable observations of behavior, repeated measurements of outcome, and individualized tailoring of objectives for each subject; all of these occur within a system that allows an experimental analysis to be conducted. Single-case designs are particularly useful when only small numbers of participants are available for a relatively long period of time. Trends in medical education toward individualized instruction, adult-centered learning, and fine-grained analyses of medical skills and knowledge make this field especially amenable to single-case experimental designs. Issues of internal and external validity, generality, practicality, and ethics are discussed, and several typical designs are illustrated. While the emergence of qualitative research methods in medical education may prove useful, single-case designs can maintain experimental science's emphasis on methodologic rigor, while allowing the flexibility often needed to conduct research in applied settings.

Education, Medical↗

What predicts USMLE Step 3 performance?

BACKGROUND: Academic and other student-specific variables associated with United States Medical Licensing Examination (USMLE) Step 3 performance have not been fully defined. METHOD: We analyzed Step 3 scores in association with medical school academic-performance measures, gender, residency specialty, and first postgraduate year (PGY-l) of training program-director performance evaluations. RESULTS: There were significant first-order associations between Step 3 scores and each of USMLE Step 1 and Step 2 scores, third-year clerkships' grade point average (GPA), Alpha Omega Alpha election, Medical Scientist Training Program graduation, broad-based specialty residency training, and PGY-l performance evaluation score. In a multiple linear regression model accounting for over 50% of the total variance in Step 3 scores, Step 2 scores, broad-based-specialty residency training, and GPA independently predicted Step 3 scores. CONCLUSIONS: Individualized Step 3 scores provide medical schools with additional means to externally validate their educational programs and to enhance the scope of outcomes assessments for their graduates.

Clinical Clerkship↗

School based HIV prevention in Zimbabwe: feasibility and acceptability of evaluation trials using biological outcomes.

OBJECTIVE: To determine the feasibility and acceptability of conducting a community randomized trial (CRT) of an adolescent reproductive health intervention (ARHI) using biological measures of effectiveness. SETTING: Four secondary schools and surrounding communities in rural Zimbabwe. METHODS: Discussions were held with pupils, parents, teachers and community leaders to determine acceptability. A questionnaire and urine sampling survey was undertaken among Form 1 and 2 pupils. Studies were undertaken to inform likely participation and follow up in a future CRT. A community survey of 16-19-year-olds was conducted to determine levels of secondary school attendance and likely HIV prevalence at final follow up in the event of a trial. RESULTS: Form 1 and 2 pupils aged 12-18 years (n = 723; median age, 15 years) participated in the research. Prevalences of HIV, Chlamydia and gonorrhoea were 3.6% [95% confidence interval (CI), 2.3-5.3%], 0.4% (95% CI, 0.1-1.3%) and 1.9% (95% CI, 1.0-3.3%) respectively. There was poor correlation between biological evidence of sexual experience and questionnaire responses, due to concerns about confidentiality. Only 13% (95% CI, 4-27%) of those infected with HIV and/or a sexually transmitted disease admitted to having had sex. In the community survey of 573 adolescents aged 16-19 years, 6.6% (95% CI, 3.9-10.3%) of females and 5.1% (95% CI, 2.9-8.2%) of males were HIV positive. High participation and retention rates are achievable within a trial in this setting. CONCLUSIONS: It is acceptable and feasible to conduct randomized trials to establish the effectiveness of ARHIs. However, self-reported behavioural outcomes will probably be biased, emphasizing the importance of using externally validated biological outcome measures to determine effectiveness.

Adolescent↗

Single-subject research for nursing practice.

Presenting the characteristics of single-subject research, the authors advocate greater use of the method by clinical nurse specialists and explore issues implicit in the research design (external validity, rigor, statistical versus clinical significance, ethics, and service versus research orientation). Replicated single-subject designs seem well suited to link process-based nursing practice and research. Such designs may make better investigative use of daily observations and interventions that are part of nursing practice. They propose that, despite concerns raised by others, single-subject research as a viable adjunct to existing research methods be considered as a means for increasing the contribution of clinical nurse specialists in the development of clinical nursing knowledge.

Clinical Nursing Research↗

Are attentional-hyperactivity deficits unidimensional or multidimensional syndromes? Empirical findings from a community survey.

Factor analysis on teacher ratings of symptoms in a probability community sample of children aged 6 to 16 years (N = 614) yielded two factors: Inattention and Hyperactivity-Impulsivity. Subsequent cluster analyses on the scores of factorially derived scales for a subsample of 170 children with a diagnosis of attention deficit disorder with (ADDH) and without hyperactivity (ADDWO), or normals, resulted in five clusters that accounted for 88% of the variance. The existence of these clusters was confirmed using external validating criteria. The data support a bidimensional conceptualization of attention deficit disorder with hyperactivity, one dimension consisting of symptoms of inattention and another of symptoms of hyperactivity-impulsivity. The data also suggests that a condition very similar to the DSM-III-R description of undifferentiated attention-deficit disorder also exists as a distinct entity.

Adolescent↗

Attention deficit disorder with and without hyperactivity: a review and comparison of matched groups.

This paper compares attention deficit disorder (ADD) with hyperactivity (ADDH) and without hyperactivity (ADDWO). The literature is outlined, revealing the areas of possible differences to be not only the core symptoms, but also associated conduct and emotional symptoms, social relations functioning, learning, medical disorders, family history, and course and outcome of the disorder. Empirical data are presented comparing age and sex matched groups of children from a speech/language clinic sample with ADDH (N = 40) and ADDWO (N = 40). Although the methods of the present study are different from those of previous studies, they nonetheless support a number of previous findings, and, further, give support to the external validity of the ADDWO diagnostic category.

Adolescent↗