Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “External validity”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Problems in validating endogenous depression in the Arab culture by contemporary diagnostic criteria.

This study highlights the difficulties that may be encountered in attempting to apply the clinical construct of endogenous depression derived from western studies to depressed Arab patients. The agreement between 4 operational systems on the diagnosis of endogenous (melancholic) depression is explored in 100 patients with primary depressive disorder in Al-Ain, United Arab Emirates. The symptom characteristics of the 61 patients in whom all diagnostic systems agreed are then described quantitatively and qualitatively. Subjects were evaluated by the Newcastle scale, Hamilton's 21 item depression scale, global assessment of functioning scale, and the operational criteria of the diagnostic systems used. Diagnosis of endogenicity was derived by computer according to the respective criteria. The agreement between DSM-IV, ICD-10, and RDC criteria is moderately high (0.72). When the Newcastle Index is included, it is only moderate (0.58). Disagreements are related to differences in diagnostic criteria. Small differences affect concordance appreciably. DSM-IV agreed with a majority of external validators, differentiating a more homogeneous groups of patients. In the present study, endogenous depression identified by western criteria, was less likely to manifest by guilt feelings, a distinct quality of mood, and loss of libido. The descriptions of patients reveal that the mood component of depression is expressed differently, somatic metaphors are used frequently to express distress, religious elements influence the expression of symptoms, and depression may manifest in behaviours not directly indicative of the disorder. Endogenous depression may be identified in the Arab culture, but considerable variation in its component symptom frequencies and mode of expression needs to be taken in consideration for defining it in terms appropriate to the culture.

Adult↗

A hostility scale for the California Psychological Inventory: MMPI, observer Q-sort, and big-five correlates.

Using two samples, we developed and validated a hostility scale that can be scored from the California Psychological Inventory (CPI) and serves as an alternate for the Cook-Medley Hostility Scale (Ho; Cook & Medley, 1954). The CPI Hostility (H) scale consists of 33 items that are either duplicates or close equivalents of specific Ho items, and the two scales correlate at least .90 in samples differing in sex. The H and Ho scales show a similar pattern of correlations with conceptually relevant MMPI scales and with observer-rated personality attributes tapping Barefoot, Peterson, et al.'s (1991) five hostility categories of Hostile Affect, Cynicism, Aggressive Responding, Social Avoidance, and Hostile Attributions. These findings provide evidence for the equivalence of the two hostility scales, as well as external validation for those personality characteristics that are purported to underlie the construct of hostility as tapped by both the original Ho scale and the new CPI H scale.

Adolescent↗

Validation of Horne and Ostberg morningness-eveningness questionnaire in a middle-aged population of French workers.

As suggested by the authors, the Horne and Ostberg morning/evening questionnaire (MEQ) has never been adapted to evaluate a nonstudent population. The purpose of this study was to validate this MEQ in a sample of middle-aged workers by modifying only the cutoffs. It was administered in 566 non-shift-workers aged 51.2 to 3.2 years who presented no sleep disorders. According to the Home and Ostberg classification, the sample consisted of 62.1% morning type, 36.6% neither type, and 2.2% evening type. Multiple correspondence analysis, which determines the principal components, was performed on all MEQ items. Then an ascending hierarchical classification was applied to determine 3 clusters from these principal components. On the basis of these 3 clusters, new cutoffs were determined: evening types were considered as scoring under 53 and morning types above 64, thus giving 28.1% morning type, 51.7% neither type, and 20.2% evening type. As an external validation, eveningness was associated with later bedtime and waking-up time (more pronounced at the weekend), greater need for sleep, larger daily sleep debt, greater morning sleepiness, and ease of returning to sleep in the early morning. A positive correlation between age and morningness was again found. This study confirms that "owls" are not rare in a middle-aged sample. We conclude that this adapted MEQ could be useful when investigating age-related changes in sleep.

Biological Clocks↗

Treatment attrition during group therapy for social phobia.

Psychological group treatments, such as behavioral or cognitive-behavioral therapy, are generally effective interventions for social phobia. However, a substantial number of individuals discontinue these treatments prematurely. Participant attrition can threaten the validity of treatment outcome studies if attrition during therapy does not occur randomly. In order to examine this issue, we studied 133 individuals with a principal diagnosis of social phobia who initiated a 12-week behavioral or cognitive-behavioral group treatment for social phobia. Thirty-four participants discontinued therapy prematurely. These dropouts were compared to treatment completers in demographic characteristics, Axis I and II psychopathology, and their attitude toward treatment. The results only showed a small difference between treatment completers and dropouts in their attitude toward treatment: dropouts rated the treatment rationale as less logical than completers at the beginning of treatment. No other differences between dropouts and completers were observed. Therefore, dropouts are unlikely to present a serious threat to the external validity of treatment outcome studies for social phobia.

Adolescent↗

Development of a dedicated risk-adjustment scoring system for colorectal surgery (colorectal POSSUM).

BACKGROUND: The aim of the study was to develop a dedicated colorectal Physiological and Operative Severity Score for the enUmeration of Mortality and morbidity (CR-POSSUM) equation for predicting operative mortality, and to compare its performance with the Portsmouth (P)-POSSUM model. METHODS: Data were collected prospectively from 6883 patients undergoing colorectal surgery in 15 UK hospitals between 1993 and 2001. After excluding missing data and 93 patients who did not satisfy the inclusion criteria, 4632 patients (68.2 per cent) underwent elective surgery and 2107 had an emergency operation (31.0 per cent); 2437 operations (35.9 per cent) for malignant and 4267 (62.8 per cent) for non-malignant diseases were scored. Stepwise logistic regression analysis was used to develop an age-adjusted POSSUM model and a dedicated CR-POSSUM model. A 60:40 per cent split-sample validation technique was adopted for model development and testing. Observed and expected mortality rates were compared. RESULTS: The operative mortality rate for the series was 5.7 per cent (387 of 6790 patients) (elective operations 2.8 per cent; emergency surgery 12.0 per cent). The CR-POSSUM, age-adjusted POSSUM and P-POSSUM models had similar areas under the receiver-operator characteristic curves. Model calibration was similar for CR-POSSUM and age-adjusted POSSUM models, and superior to that for the P-POSSUM model. The CR-POSSUM model offered the best overall accuracy, with an observed : expected ratio of 1.000, 0.998 and 0.911 respectively (test population). CONCLUSION: The CR-POSSUM model provided an accurate predictor of operative mortality. External validation is required in hospitals different from those in which the model was developed.

Adult↗

Proteomics as a theranostic compass in BCR::ABL1-negative myeloproliferative neoplasms: Integrating biomarker discovery with therapeutic stratification.

Classic BCR::ABL1-negative myeloproliferative neoplasms (MPNs)-polycythaemia vera, essential thrombocythaemia, and primary myelofibrosis-are clonal haematopoietic stem cell disorders with marked heterogeneity in clinical phenotype, disease trajectory, and therapeutic response. Genomic stratification by driver and cooperating mutations only partially accounts for this variability, leaving gaps in predicting thrombotic risk, fibrotic progression, leukaemic transformation, and treatment benefit. Proteomics bridges this gap by providing function-proximal readouts of protein abundance, post-translational modifications, pathway activity, and intercellular signalling that genomics and transcriptomics cannot capture, positioning it as a theranostic platform in which the same molecular readouts simultaneously inform diagnostic stratification and therapeutic decision-making. We propose a five-stage translational framework spanning from discovery-scale mass spectrometry and affinity-based plasma profiling to targeted validation, multicentre standardisation, and machine learning-integrated clinical panels. Proteomic evidence is synthesised across the following four disease axes: clonal fitness in haematopoietic stem and progenitor cells; bone marrow microenvironmental remodelling and fibrosis; chronic inflammation and thrombosis; and leukaemic transformation. We further describe how phosphoproteomics reveals resistance mechanisms to JAK inhibitors, including AXL-MAPK bypass and PP2A-autophagy-mediated tolerance, and how protein-level biomarkers (BCL2-BCL-XL, RAS-ERK, CAMK2G, and ROCK1/2) can guide individualised therapeutic selection. Affinity-based platforms (Olink PEA and SomaScan) and spatially resolved technologies (CODEX and single-cell proteomics) complement discovery proteomics. At present, however, this evidence base is constrained by small and heterogeneous cohorts, limited cross-platform reproducibility, and a scarcity of independent external validation for candidate protein panels. Realising this vision will require multicentre standardisation, analytically validated panel assays, and prospective clinical studies that translate molecular findings into decision-grade tools for patients with MPNs.

Humans↗

Differential exoprotease activities confer tumor-specific serum peptidome patterns.

Recent studies have established distinctive serum polypeptide patterns through mass spectrometry (MS) that reportedly correlate with clinically relevant outcomes. Wider acceptance of these signatures as valid biomarkers for disease may follow sequence characterization of the components and elucidation of the mechanisms by which they are generated. Using a highly optimized peptide extraction and matrix-assisted laser desorption/ionization-time-of-flight (MALDI-TOF) MS-based approach, we now show that a limited subset of serum peptides (a signature) provides accurate class discrimination between patients with 3 types of solid tumors and controls without cancer. Targeted sequence identification of 61 signature peptides revealed that they fall into several tight clusters and that most are generated by exopeptidase activities that confer cancer type-specific differences superimposed on the proteolytic events of the ex vivo coagulation and complement degradation pathways. This small but robust set of marker peptides then enabled highly accurate class prediction for an external validation set of prostate cancer samples. In sum, this study provides a direct link between peptide marker profiles of disease and differential protease activity, and the patterns we describe may have clinical utility as surrogate markers for detection and classification of cancer. Our findings also have important implications for future peptide biomarker discovery efforts.

Amino Acid Sequence↗

Topologically distinct intratumoral heterogeneity scores for predicting high-risk pathological grades in invasive lung adenocarcinoma: A multicenter study across four institutions.

High-risk subtypes of invasive lung adenocarcinoma (IAC), particularly micropapillary- or solid-predominant patterns, are closely associated with poor prognosis. This multicenter retrospective study developed and validated a predictive model for the preoperative identification of these high-risk subtypes using topologically distinct intratumoral heterogeneity (ITH) scores derived from CT images. The study included 1,051 patients with IAC. Two complementary ITH scores were developed: a two-dimensional ITH score, which integrated local radiomics features with global pixel distribution patterns on the largest cross-sectional CT slice, and a three-dimensional ITH score, which extended this quantification across the entire tumor volume. Clinicoradiological features and ITH scores were incorporated as model inputs to construct six base machine learning classifiers and a final stacking ensemble classifier. Model interpretability and robustness were evaluated using SHapley Additive exPlanations (SHAP)-based ablation analyses. An independent dataset from The Cancer Imaging Archive (TCIA) was used for external validation to investigate associations between ITH scores and pathological characteristics, genomic features, recurrence-free survival, and overall survival. The stacking ensemble classifier achieved the best predictive performance, with an area under the receiver operating characteristic curve of 0.875, outperforming models based solely on radiomics features (0.834) or clinicoradiological features (0.792). SHAP analysis identified the 3D ITH score as the most influential contributor to model output, and TCIA validation showed that higher 3D ITH scores were associated with more aggressive tumor biology and poorer survival outcomes. The topologically distinct 3D ITH score may provide a clinically meaningful imaging biomarker for preoperative risk stratification in IAC.

Journal Article↗

Radiomics-based gradient boosting model on contrast-enhanced MRI for non-invasive prediction of epidermal growth factor receptor expression and therapeutic response to EGFR-targeted antibody-drug conjugates in high-grade glioma organoid models.

BACKGROUND: Epidermal growth factor (EGF) and its receptor EGF(EGFR) play crucial roles in glioblastoma (GBM) prognosis. However, non-invasive assessment of their expression remains challenging. This study aimed to determine whether radiomics features extracted from contrast-enhanced MRI could predict EGFR expression in high-grade gliomas (HGG) and to explore their associations with immune infiltration and therapeutic response of EGFR-Targeted antibody drug conjugates(EGFR-ADCs). METHODS: We extracted radiomic features from contrast-enhanced MRI of 298 GBM patients from The Cancer Imaging Archive (TCIA) and matched them with RNA-seq data from The Cancer Genome Atlas (TCGA). Feature selection was performed using minimum redundancy maximum relevance (mRMR) and recursive feature elimination (RFE). Machine learning models were built to predict EGF/EGFR expression. Radiogenomic associations were validated by immune infiltration analysis. Patient-Derived Tumor-Like Cell Clusters (PTC) were used to compare the antitumor efficacy of EGFR- ADCs and temozolomide. RESULTS: Elevated EGF/EGFR expression correlated with poor prognosis and increased infiltration of M2 macrophages, regulatory T cells, and CD4⁺ memory T cells. Pathway analysis demonstrated significant enrichment of the mechanistic target of rapamycin (mTOR) and Mitogen-Activated Protein Kinase (MAPK) signaling cascades. Radiomics-based prediction models achieved robust performance (AUC > 0.85) in stratifying EGFR expression status. In EGFR-positive tumor tissues, EGFR-ADCs exerted antitumor efficacy similar to that of temozolomide. CONCLUSIONS: EGF/EGFR expression is associated with immunosuppressive microenvironments and adverse outcomes in HGG. Radiomics may provide a non-invasive approach for estimating EGFR expression, although model performance requires external validation and EGFR-ADCs showed partial inhibitory activity within the tested range, though potency remains to be defined.These findings suggest a framework into radiogenomic stratification and targeted therapy in GBM.

Radiomics↗

Blood-based DNA methylation and exposure risk scores predict PTSD with high accuracy in military and civilian cohorts.

BACKGROUND: Incorporating genomic data into risk prediction has become an increasingly popular approach for rapid identification of individuals most at risk for complex disorders such as PTSD. Our goal was to develop and validate Methylation Risk Scores (MRS) using machine learning to distinguish individuals who have PTSD from those who do not. METHODS: Elastic Net was used to develop three risk score models using a discovery dataset (n&#x2009;=&#x2009;1226; 314 cases, 912 controls) comprised of 5 diverse cohorts with available blood-derived DNA methylation (DNAm) measured on the Illumina Epic BeadChip. The first risk score, exposure and methylation risk score (eMRS) used cumulative and childhood trauma exposure and DNAm variables; the second, methylation-only risk score (MoRS) was based solely on DNAm data; the third, methylation-only risk scores with adjusted exposure variables (MoRSAE) utilized DNAm data adjusted for the two exposure variables. The potential of these risk scores to predict future PTSD based on pre-deployment data was also assessed. External validation of risk scores was conducted in four independent cohorts. RESULTS: The eMRS model showed the highest accuracy (92%), precision (91%), recall (87%), and f1-score (89%) in classifying PTSD using 3730 features. While still highly accurate, the MoRS (accuracy&#x2009;=&#x2009;89%) using 3728 features and MoRSAE (accuracy&#x2009;=&#x2009;84%) using 4150 features showed a decline in classification power. eMRS significantly predicted PTSD in one of the four independent cohorts, the BEAR cohort (beta&#x2009;=&#x2009;0.6839, p=0.006), but not in the remaining three cohorts. Pre-deployment risk scores from all models (eMRS, beta&#x2009;=&#x2009;1.92; MoRS, beta&#x2009;=&#x2009;1.99 and MoRSAE, beta&#x2009;=&#x2009;1.77) displayed a significant (p&#x2009;<&#x2009;0.001) predictive power for post-deployment PTSD. CONCLUSION: The inclusion of exposure variables adds to the predictive power of MRS. Classification-based MRS may be useful in predicting risk of future PTSD in populations with anticipated trauma exposure. As more data become available, including additional molecular, environmental, and psychosocial factors in these scores may enhance their accuracy in predicting PTSD and, relatedly, improve their performance in independent cohorts.

Humans↗

Blood-based DNA methylation and exposure risk scores predict PTSD with high accuracy in military and civilian cohorts.

BACKGROUND: Incorporating genomic data into risk prediction has become an increasingly useful approach for rapid identification of individuals most at risk for complex disorders such as PTSD. Our goal was to develop and validate Methylation Risk Scores (MRS) using machine learning to distinguish individuals who have PTSD from those who do not. METHODS: Elastic Net was used to develop three risk score models using a discovery dataset (n = 1226; 314 cases, 912 controls) comprised of 5 diverse cohorts with available blood-derived DNA methylation (DNAm) measured on the Illumina Epic BeadChip. The first risk score, exposure and methylation risk score (eMRS) used cumulative and childhood trauma exposure and DNAm variables; the second, methylation-only risk score (MoRS) was based solely on DNAm data; the third, methylation-only risk scores with adjusted exposure variables (MoRSAE) utilized DNAm data adjusted for the two exposure variables. The potential of these risk scores to predict future PTSD based on pre-deployment data was also assessed. External validation of risk scores was conducted in four independent cohorts. RESULTS: The eMRS model showed the highest accuracy (92%), precision (91%), recall (87%), and f1-score (89%) in classifying PTSD using 3730 features. While still highly accurate, the MoRS (accuracy = 89%) using 3728 features and MoRSAE (accuracy = 84%) using 4150 features showed a decline in classification power. eMRS significantly predicted PTSD in one of the four independent cohorts, the BEAR cohort (beta = 0.6839, p-0.003), but not in the remaining three cohorts. Pre-deployment risk scores from all models (eMRS, beta = 1.92; MoRS, beta = 1.99 and MoRSAE, beta = 1.77) displayed a significant (p < 0.001) predictive power for post-deployment PTSD. CONCLUSION: Results, especially those from the eMRS, reinforce earlier findings that methylation and trauma are interconnected and can be leveraged to increase the correct classification of those with vs. without PTSD. Moreover, our models can potentially be a valuable tool in predicting the future risk of developing PTSD. As more data become available, including additional molecular, environmental, and psychosocial factors in these scores may enhance their accuracy in predicting the condition and, relatedly, improve their performance in independent cohorts.

DNA methylation↗

Validity of the inventory of interpersonal problems for predicting treatment outcome: an investigation with the Pennsylvania Practice Research Network.

In this study, we examined the relationship between treatment outcome and variables from the Inventory of Interpersonal Problems Circumplex scales (IIP-C; Horowitz, Alden, Wiggins, & Pincus, 2000) in the Pennsylvania Psychological Association's Practice Research Network (PRN; Borkovec, Echemendia, Ragusea, & Ruiz, 2001). The PRN was a naturalistic observation treatment outcome study conducted with clinicians who were providing outpatient therapy. Assessment instruments, including the Compass Assessment System (Howard, Brill, Lueger, O'Mahoney, & Grissom, 1993; Sperry, Brill, Howard, & Grissom, 1996) and the IIP-C, were used to assess outcome at the 7th session (N=73) and at termination (N=42). Significant associations were identified between seventh-session outcome and most of the IIP variables. Only IIP elevation and amplitude were related to termination outcome. Elevation, amplitude, and hostile submissive problems were related to treatment length. Ad hoc analyses indicated that the IIP elevation fully mediated the relationships between interpersonal problems and seventh-session outcome but not the relationship between amplitude and outcome. We discuss the results in relation to the external validity of the IIP.

Adult↗

Testing the EORTC Quality of Life Questionnaire on cancer patients with heterogeneous diagnoses.

This study aimed to contribute to the validation of the 30-item Quality of Life Questionnaire developed by the European Organization for Research and Treatment of Cancer Study Group (EORTC QLQ-C30). The sample consisted of 177 cancer patients with heterogeneous diagnoses. A series of scales representing various dimensions of quality of life were tested, including those proposed by the EORTC Study Group. Mokken's non-parametric latent trait model for unidimensional scaling was used as the basic scaling procedure. This model gives coefficients of scalability in addition to reliability coefficients. In terms of scalability measured by Loevinger's H, all EORTC Study Group scales, except the cognitive functioning scale were found to be quite satisfactory. The cognitive functioning scale and the role functioning scale were below the satisfactory level in terms of reliability (internal consistency). In total, our study strengthens the external validity of the EORTC QLQ-C30 and confirms that it may be used on cancer patients with various diagnoses.

Adult↗

Internal validity of an anxiety disorder screening instrument across five ethnic groups.

We tested the factor structure of the National Anxiety Disorder Screening Day instrument (n=14860) within five ethnic groups (White, Black, Hispanic, Asian, Native American). Conducted yearly across the US, the screening is meant to detect five common anxiety syndromes. Factor analyses often fail to confirm the validity of assessment tools' structures, and this is especially likely for minority ethnic groups. If symptoms cluster differently across ethnic groups, criteria for conventional DSM-IV disorders are less likely to be met, leaving significant distress unlabeled and under-detected in minority groups. Exploratory and confirmatory factor analyses established that the items clustered into the six expected factors (one for each disorder plus agoraphobia). This six-factor model fit the data very well for Whites and not significantly worse for each other group. However, small areas of the model did not appear to fit as well for some groups. After taking these areas into account, the data still clearly suggest more prevalent PTSD symptoms in the Black, Hispanic and Native American groups in our sample. Additional studies are warranted to examine the model's external validity, generalizability to more culturally distinct groups, and overlap with other culture-specific syndromes.

Adult↗

Development, evaluation and validation of an intelligent system for the management of labour.

Over the past 4 years our group has developed a prototype intelligent system which applies captured expert knowledge to support clinical decision-making during labour. This chapter presents a review of the system and the progress made to date. The system classifies the same features from the CTG as experienced clinicians using numerical algorithms and a small neural network. This hybrid approach has been shown to obtain a comparable performance with experts. The CTG information, together with the patient information and labour events, are collectively passed to an expert system for processing. The expert system interprets this combined data using a database of over 400 rules which are used to recommend action. Importantly, as the knowledge is rule-based, it allows the system to explain the reasoning which led it to recommend a certain action. In this way, the clinician is not expected to blindly follow the system's recommendations but can reach an informed judgement in the same way they might by discussing the case with an experienced informed colleague. After two internal evaluations had found the system obtained a performance comparable with local experts, an extensive external validation was undertaken. This study involved 17 experts from 16 leading centres within the UK. Each expert and the system reviewed 50 cases twice, at least one month apart which contained those CTGs considered most difficult to interpret selected from a database of 2400 high-risk labours. This study found that the majority of experts agreed well and were consistent in their management of the cases. The system obtained a performance that was indistinguishable from the experts, except it was more consistent, even when used by an engineer with little knowledge of labour management. This study demonstrates the potential for intelligent systems to transform the cardiotocograph from a difficult-to-use, ineffective recorder of fetal heart rate, to an interactive and effective decision support tool capable of raising the skills of staff.

Algorithms↗

Techniques and considerations for determining isoinertial upper-body power.

Power is an integral aspect of many sports. Although power output of the lower body is often measured during jumping and cycling movements, much less is known about power as pertains to the upper body musculature. Recently, isoinertial methods--with constant gravitational load--of power testing have become common, but little is known of the reliability and criterion validity of these tests as they pertain to sport performance. In addition, the varied methodology makes a lucid model more evasive. The aims of this review are to examine the various methods of assessing upper body power, to establish its role in predicting athletic performance, and to assess the body of literature that has assessed power output of the upper extremities by isoinertial methods. To our knowledge, only two studies on isoinertial upper-body power have shown a direct correlation to sporting ability (Baker, 2001; Baker et al., 2001); therefore, many unanswered questions exist as to the efficacy of these tests as predictors of athletic ability or as a method to track athletes' training over time. From this review we hope to allow the sport coach to assess the overall utility of these tests in terms of availability, safety and external validity.

Arm↗

A cross-national validation of the client satisfaction questionnaire: the Dutch experience.

A Dutch translation of the eight-term version of the Client Satisfaction Questionnaire (CSQ-8) was administered to community mental health outpatients in the Netherlands (n = 110). Data analyses indicate that the Dutch CSQ-8 has highly similar operating characteristics and psychometric properties compared with the English language version. Results also indicate that one general satisfaction factor was found in the Dutch CSQ-8 data. All eight of the scale items loaded heavily on this general factor, strong inter-item correlations were found, and the scale demonstrated high internal consistency. On these grounds, we can conclude that the Dutch CSQ-8 has the same properties as the original questionnaire and can be used as such in Holland. It was also found that those clients who decided to stop therapy on their own were less satisfied than other clients. Clients who made a common decision with their therapist or let him/her make that decision were more satisfied than clients who stopped by themselves. Additional research is planned to investigate whether clients who complete treatment goals are indeed more satisfied and whether this line of research could be a means of studying the external validity of the scale.

Analysis of Variance↗

Quality of guidelines for the laboratory management of diabetes mellitus.

BACKGROUND: There is increasing concern about the quality and reliability of practice guidelines, especially in the field of laboratory medicine, as most recommendations are developed by clinical specialty societies, often without involving laboratory professionals. Little information is available on the methodological quality of guidelines for the use of laboratory investigations in the care of specific diseases. We describe a pilot assessment of the most well-known guidelines for the diagnosis and monitoring of diabetes mellitus (DM). METHODS: Practice guidelines on DM published in English between 1999 and 2005 April were identified by systematic searching in Medline and international guideline databases. Fifty four DM guidelines were retrieved, of which 29 met our inclusion criteria. The four most widely used international guidelines (WHO, ADA, NACB, NICE) were selected for a critical appraisal of their methodological quality. This was carried out by seven independent assessors using a validated checklist, the AGREE Instrument. Twenty three guideline attributes arranged in six independent domains were investigated and the mean scores of assessors for each attribute and the aggregated scores for each domain were calculated. Cronbach's alpha and interclass correlations were calculated to measure internal consistency and reliability within each domain. The four guidelines were compared using one-way ANOVA and ANOVA using repeated measurements. RESULTS: The selected four guidelines on DM have significant shortcomings in demonstrating and/or reporting multidisciplinary stakeholder involvement in the guideline development process, evidence-based methodology for formulating recommendations, applicability of statements, and disclosing any conflicts of interest or reporting editorial independence. CONCLUSIONS: Poor quality and lack of explicitness of recommendations in laboratory medicine call for methodological standards of guideline development and reporting, and for an international collaboration of guideline development activities, to increase the internal and external validity of recommendations in laboratory practice.

Diabetes Mellitus↗