Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Assessment and integration of publicly available SAGE, cDNA microarray, and oligonucleotide microarray expression data for global coexpression analyses.

Large amounts of gene expression data from several different technologies are becoming available to the scientific community. A common practice is to use these data to calculate global gene coexpression for validation or integration of other "omic" data. To assess the utility of publicly available datasets for this purpose we have analyzed Homo sapiens data from 1202 cDNA microarray experiments, 242 SAGE libraries, and 667 Affymetrix oligonucleotide microarray experiments. The three datasets compared demonstrate significant but low levels of global concordance (rc<0.11). Assessment against Gene Ontology (GO) revealed that all three platforms identify more coexpressed gene pairs with common biological processes than expected by chance. As the Pearson correlation for a gene pair increased it was more likely to be confirmed by GO. The Affymetrix dataset performed best individually with gene pairs of correlation 0.9-1.0 confirmed by GO in 74% of cases. However, in all cases, gene pairs confirmed by multiple platforms were more likely to be confirmed by GO. We show that combining results from different expression platforms increases reliability of coexpression. A comparison with other recently published coexpression studies found similar results in terms of performance against GO but with each method producing distinctly different gene pair lists.

Gene Expression Profiling↗

Statistical validation of the EORTC prognostic model for malignant pleural mesothelioma based on three consecutive phase II trials.

PURPOSE: Malignant pleural mesothelioma (MPM) carries a poor prognosis due to chemoresistance. The European Organisation for Research and Treatment of Cancer (EORTC) prognostic model was reported to predict survival in MPM. Our retrospective analysis set out to test the validity of the model as a prognostic tool in patients treated in three phase II trials at St Bartholomew's Hospital (London, United Kingdom) between 1999 and 2003. PATIENTS AND METHODS: A total of 145 patients were treated in three phase II trials; vinorelbine (VIN; 70 patients), vinorelbine/oxaliplatin (VO; 26 patients), and irinotecan/cisplatin/mitomycin C (IPM; 49 patients). Two subgroups, high-risk and low-risk, were defined by EORTC prognostic score (EPS). EPS was determined by a five-parameter model incorporating age, sex, histology, probability of diagnosis, and leukocyte count. An EPS cutoff of less than 1.27 (low risk) or more than 1.27 (high risk) was used to stratify Kaplan-Meier survival curves. Each of the EPS variables exhibited either trends or significant stratification of overall survival (OS). RESULTS: Multivariate analysis confirmed leukocyte count, Eastern Cooperative Oncology Group performance status, and sarcomatous histology as independent prognostic variables. EPS stratified OS in both individual and pooled trial datasets. No association between objective tumor response and EPS classification was identified by multinomial logistic regression. EPS stratified progression-free survival for the VO and IPM cohorts, but not for VIN. CONCLUSION: This study validates the EPS system as a robust tool for stratifying small trials into low- and high-risk subgroups. EPS should facilitate patient selection and analysis in randomized clinical trials.

Adult↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗

History of the Cancer Information Service.

The Cancer Information Service (CIS) was established on July 1, 1975, following the mandate of the National Cancer Act of 1971 giving the National Cancer Institute (NCI) new responsibilities for educating the public, patients, and health professionals. Funded under a contract mechanism, the CIS has become one of the longest-running community programs in NCI. The CIS has been able to set up and maintain high-quality service, giving accurate, up-to-date medical information to cancer patients and their families and friends, to health professionals, and to the general public. The CIS network, which has taken more than 5 million calls since its inception, has weathered many changes, both at the national and the local level. Its current call volume, in excess of 500,000 calls per year, makes it one of the most heavily utilized health-related telephone helplines in the country. Using a standardized Call Record Form, data on calls have been recorded consistently since 1983; the dataset now contains information on more than 4.2 million calls. An outreach component that acts as NCI's field arm has been part of the CIS since its inception. The CIS has matured into a stable system that has been reconfigured into 19 regional offices, covering the entire country. These offices run the telephone service and serve as NCI's outreach arm, working with intermediaries to carry out NCI information and education programs in local communities.

History, 20th Century↗

Factors affecting the performance of maternal health care providers in Armenia.

BACKGROUND: Over the last five years, international development organizations began to modify and adapt the conventional Performance Improvement Model for use in low-resource settings. This model outlines the five key factors believed to influence performance outcomes: job expectations, performance feedback, environment and tools, motivation and incentives, and knowledge and skills. Each of these factors should be supplied by the organization in which the provider works, and thus, organizational support is considered as an overarching element for analysis. Little research, domestically or internationally, has been conducted on the actual effects of each of the factors on performance outcomes and most PI practitioners assume that all the factors are needed in order for performance to improve. This study presents a unique exploration of how the factors, individually as well as in combination, affect the performance of primary reproductive health providers (nurse-midwives) in two regions of Armenia. METHODS: Two hundred and eighty-five nurses and midwives were observed conducting real or simulated antenatal and postpartum/neonatal care services and interviewed about the presence or absence of the performance factors within their work environment. Results were analyzed to compare average performance with the existence or absence of the factors; then, multiple regression analysis was conducted with the merged datasets to obtain the best models of "predictors" of performance within each clinical service. RESULTS: Baseline results revealed that performance was sub-standard in several areas and several performance factors were deficient or nonexistent. The multivariate analysis showed that (a) training in the use of the clinic tools; and (b) receiving recognition from the employer or the client/community, are factors strongly associated with performance, followed by (c) receiving performance feedback in postpartum care. Other - extraneous - variables such as the facility type (antenatal care) and whether observation was on simulated vs. real patients (postpartum care) also had a role in observed performance. CONCLUSION: This study concludes that the antenatal and postpartum care performance of health providers in Armenia is strongly associated with having the practical knowledge and skills to use everyday tools of the trade and with receiving recognition for their work, as well as having performance feedback. The paper recognized several limitations and expects further studies will illuminate this important topic further.

Journal Article↗

[Virtual multislice computed tomography cystoscopy for evaluation of urinary bladder lesions].

The introduction of multislice computed tomography (MDCT) with the possibility of acquiring isotropic datasets has been an ideal prerequisite for development of virtual MDCT cystoscopy. Remarkable technical progress regarding post-processing of high-resolution 3D datasets as well as a considerable reduction of the time required for post-processing made it possible to introduce virtual MDCT cystoscopy into the clinical routine. 3D post-processing that often required 7-8 h when virtual endoscopy techniques were first developed can now be performed in less than 5 min after transfer of data to the 3D workstation. With the limitations and contraindications of conventional cystoscopy in mind, virtual MDCT cystoscopy may be seen as a valuable alternative to conventional cystoscopy for evaluation of hematuria.

Cystoscopy↗

Application of simple mathematical expressions to relate the half-lives of xenobiotics in rats to values in humans.

INTRODUCTION: Previous publications from GlaxoSmithKline and University of Toledo laboratories convey our independent attempts to predict the half-lives of xenobiotics in humans using data obtained from rats. The present investigation was conducted to compare the performance of our published models against a common dataset obtained by merging the two sets of rat versus human half-life (hHL) data previously used by each laboratory. METHODS: After combining data, mathematical analyses were undertaken by deploying both of our previous models, namely the use of an empirical algorithm based on a best-fit model and the use of rat-to-human liver blood flow ratios as a half-life correction factor. Both qualitative and quantitative analyses were performed, as well as evaluation of the impact of molecular properties on predictability. RESULTS: The merged dataset was remarkably diverse with respect to physiochemical and pharmacokinetic (PK) properties. Application of both models revealed similar predictability, depending upon the measure of stipulated accuracy. Certain molecular features, particularly rotatable bond count and pK(a), appeared to influence the accuracy of prediction. DISCUSSION: This collaborative effort has resulted in an improved understanding and appreciation of the value of rats to serve as a surrogate for the prediction of xenobiotic half-lives in humans when clinical pharmacokinetic studies are not possible or practicable.

Algorithms↗

Statistical analysis of real-time PCR data.

BACKGROUND: Even though real-time PCR has been broadly applied in biomedical sciences, data processing procedures for the analysis of quantitative real-time PCR are still lacking; specifically in the realm of appropriate statistical treatment. Confidence interval and statistical significance considerations are not explicit in many of the current data analysis approaches. Based on the standard curve method and other useful data analysis methods, we present and compare four statistical approaches and models for the analysis of real-time PCR data. RESULTS: In the first approach, a multiple regression analysis model was developed to derive DeltaDeltaCt from estimation of interaction of gene and treatment effects. In the second approach, an ANCOVA (analysis of covariance) model was proposed, and the DeltaDeltaCt can be derived from analysis of effects of variables. The other two models involve calculation DeltaCt followed by a two group t-test and non-parametric analogous Wilcoxon test. SAS programs were developed for all four models and data output for analysis of a sample set are presented. In addition, a data quality control model was developed and implemented using SAS. CONCLUSION: Practical statistical solutions with SAS programs were developed for real-time PCR data and a sample dataset was analyzed with the SAS programs. The analysis using the various models and programs yielded similar results. Data quality control and analysis procedures presented here provide statistical elements for the estimation of the relative expression of genes using real-time PCR.

Analysis of Variance↗

Intraclass correlation coefficient and outcome prevalence are associated in clustered binary data.

BACKGROUND AND OBJECTIVE: To describe the association between values for a proportion and the intraclass correlation coefficient (ICC). METHODS: Analysis of data obtained from the General Practice Research Database (GPRD) for variation between United Kingdom general practices and results from a Health Technology Assessment (HTA) review for a range of outcomes in community and health services settings. RESULTS: There were 188 ICCs from the GPRD, the median prevalence was 13.1% (interquartile range IQR 3.5 to 28.4%) and median ICC 0.051 (IQR 0.011 to 0.094). There were 136 ICCs from the HTA review, with median prevalence 6.5% (IQR 0.4 to 20.7%) and median ICC 0.006 (IQR 0.0003 to 0.036). There was a linear association of log ICC with log prevalence in both datasets (GPRD, regression coefficient 0.61, 95% confidence interval 0.53 to 0.69, P < 0.001; HTA, 0.91, 0.81 to 1.01, P < 0.001). When the prevalence was 1% the predicted ICC was 0.008 from the GPRD or 0.002 from the HTA, but when the prevalence was 40% the predicted ICC was 0.075 (GPRD) or 0.046 (HTA). CONCLUSION: The prevalence of an outcome may be used to make an informed assumption about the magnitude of the intraclass correlation coefficient.

Data Interpretation, Statistical↗

A tractable probabilistic model for Affymetrix probe-level analysis across multiple chips.

MOTIVATION: Affymetrix GeneChip arrays are currently the most widely used microarray technology. Many summarization methods have been developed to provide gene expression levels from Affymetrix probe-level data. Most of the currently popular methods do not provide a measure of uncertainty for the expression level of each gene. The use of probabilistic models can overcome this limitation. A full hierarchical Bayesian approach requires the use of computationally intensive MCMC methods that are impractical for large datasets. An alternative computationally efficient probabilistic model, mgMOS, uses Gamma distributions to model specific and non-specific binding with a latent variable to capture variations in probe affinity. Although promising, the main limitations of this model are that it does not use information from multiple chips and does not account for specific binding to the mismatch (MM) probes. RESULTS: We extend mgMOS to model the binding affinity of probe-pairs across multiple chips and to capture the effect of specific binding to MM probes. The new model, multi-mgMOS, provides improved accuracy, as demonstrated on some bench-mark datasets and a real time-course dataset, and is much more computationally efficient than a competing hierarchical Bayesian approach that requires MCMC sampling. We demonstrate how the probabilistic model can be used to estimate credibility intervals for expression levels and their log-ratios between conditions. AVAILABILITY: Both mgMOS and the new model multi-mgMOS have been implemented in an R package, which is available at http://www.bioinf.man.ac.uk/resources/puma.

Algorithms↗

Applied analysis of recurrent events: a practical overview.

STUDY OBJECTIVE: The purpose of this paper is to give an overview and comparison of different easily applicable statistical techniques to analyse recurrent event data. SETTING: These techniques include naive techniques and longitudinal techniques such as Cox regression for recurrent events, generalised estimating equations (GEE), and random coefficient analysis. The different techniques are illustrated with a dataset from a randomised controlled trial regarding the treatment of lateral epicondylitis. MAIN RESULTS: The use of different statistical techniques leads to different results and different conclusions regarding the effectiveness of the different intervention strategies. CONCLUSIONS: If you are interested in a particular short term or long term result, simple naive techniques are appropriate. However, if the development of a particular outcome is of interest, statistical techniques that consider the recurrent events and additionally corrects for the dependency of the observations are necessary.

Adrenal Cortex Hormones↗

BayGO: Bayesian analysis of ontology term enrichment in microarray data.

BACKGROUND: The search for enriched (aka over-represented or enhanced) ontology terms in a list of genes obtained from microarray experiments is becoming a standard procedure for a system-level analysis. This procedure tries to summarize the information focussing on classification designs such as Gene Ontology, KEGG pathways, and so on, instead of focussing on individual genes. Although it is well known in statistics that association and significance are distinct concepts, only the former approach has been used to deal with the ontology term enrichment problem. RESULTS: BayGO implements a Bayesian approach to search for enriched terms from microarray data. The R source-code is freely available at http://blasto.iq.usp.br/~tkoide/BayGO in three versions: Linux, which can be easily incorporated into pre-existent pipelines; Windows, to be controlled interactively; and as a web-tool. The software was validated using a bacterial heat shock response dataset, since this stress triggers known system-level responses. CONCLUSION: The Bayesian model accounts for the fact that, eventually, not all the genes from a given category are observable in microarray data due to low intensity signal, quality filters, genes that were not spotted and so on. Moreover, BayGO allows one to measure the statistical association between generic ontology terms and differential expression, instead of working only with the common significance analysis.

Bacteria↗

Anatomically based geometric modelling of the musculo-skeletal system and other organs.

Anatomically based finite element geometries are becoming increasingly popular in physiological modelling, owing to the demand for modelling that links organ function to spatially distributed properties at the protein, cell and tissue level. We present a collection of anatomically based finite element geometries of the musculo-skeletal system and other organs suitable for use in continuum analysis. These meshes are derived from the widely used Visible Human (VH) dataset and constitute a contribution to the world wide International Union of Physiological Sciences (IUPS) Physiome Project (www.physiome.org.nz). The method of mesh generation and fitting of tricubic Hermite volume meshes to a given dataset is illustrated using a least-squares algorithm that is modified with smoothing (Sobolev) constraints via the penalty method to account for sparse and scattered data. A technique ("host mesh" fitting) based on "free-form" deformation (FFD) is used to customise the fitted (generic) geometry. Lung lobes, the rectus femoris muscle and the lower limb bones are used as examples to illustrate these methods. Geometries of the lower limb, knee joint, forearm and neck are also presented. Finally, the issues and limitations of the methods are discussed.

Algorithms↗

Fire-related child deaths: what do newspaper case reports tell us?

OBJECTIVE: To describe the accuracy and public health relevance of newspaper accounts of child deaths from fire-related incidents. METHODS: Domestic fire-related deaths of children aged under 15 years in Auckland, New Zealand, over a 10-year period were retrospectively identified from fire service records and the national minimum mortality dataset. Forensic pathology and fire service records were reviewed and this information was compared with reports published within 3 days of the index event in the region's sole daily newspaper. RESULTS: All 14 fatal fire-related events (19 deaths) identified using fire service records and the national minimum dataset during the study period were reported in the newspaper with a high degree of detail and accuracy. Only four news items informed readers of specific measures that could prevent such events. CONCLUSIONS: Daily newspapers can provide reliable, useful and timely surveillance data on the incidence of fire-related childhood deaths. However, these reports often represented missed opportunities to disseminate public health messages that raised awareness of sources of risk and means of preventing fire-related deaths.

Accident Prevention↗

Identification of a claims data "signature" and economic consequences for treatment-resistant depression.

BACKGROUND: Major depressive disorder (MDD) is a debilitating condition with significant economic consequences. Conservative estimates indicate that between 10% and 20% of all individuals with MDD are treatment resistant. The objectives for this study were (1) to use current treatment strategies identified in the literature to evaluate the validity of studying treatment-resistant depression (TRD) using claims data and (2) to estimate cost differences between TRD-likely and TRD-unlikely patients identified by use of treatment patterns. METHOD: The data source consisted of medical, pharmaceutical, and disability claims from a Fortune 100 manufacturer for 1996 through 1998 (N = 125,242 continuously enrolled beneficiaries between the ages of 18 and 64 years). The sample included individuals with medical or disability claims for MDD (NMDD = 4186). A treatment pattern algorithm was applied to classify adult MDD patients into TRD-likely (NTRD = 487) and TRD-unlikely groups. Resource utilization and costs were compared among TRD-likely and TRD-unlikely patients and a random sample of average beneficiaries (i.e., 10% of all beneficiaries) for 1998. RESULTS: Consistent with the epidemiologic literature, the algorithm classified 12% of the MDD sample as TRD-likely. Mean annual costs were $10,954 for TRD-likely patients, $5025 for TRD-unlikely patients, and $3006 for average beneficiaries. TRD-likely patients used almost twice as many medical services as did TRD-unlikely patients and incurred significantly greater indirect costs (p < .0001). CONCLUSION: It is feasible to use an administrative dataset to develop a claim-based treatment algorithm to identify TRD-likely patients. Resource utilization by TRD-likely patients was substantial, not only for direct treatment of depression but also for treatment of comorbid medical conditions. Additionally, TRD imposed on employers substantial indirect costs resulting from high rates of depression-associated disability.

Adolescent↗

Association studies of cholesterol metabolism genes (CH25H, ABCA1 and CH24H) in Alzheimer's disease.

Recent studies have demonstrated that cholesterol metabolism has an important role in Alzheimer's disease (AD) pathogenesis, suggesting that cholesterol-related genes may be significant genetic risk factors for AD. Based on the results of genome-wide screens, along with biological studies, we selected three genes as candidates for AD risk factors: ATP-binding cassette transporter A1 (ABCA1), cholesterol 25-hydroxylase (CH25H) and cholesterol 24-hydroxylase (CH24H). Case-control of North American Caucasians and AD families of Caribbean Hispanic origin were examined. Although excellent biological candidates, the case-control dataset did not support the hypothesis that these three genes were associated with susceptibility to AD. Similarly, no association was found in the Caribbean Hispanic families for CH25H. However, we did observe a possible interaction between ABCA1 and APOE in the Hispanics.

ATP Binding Cassette Transporter 1↗

Incubation time for AIDS from French transfusion-associated cases.

Although incubation time is a key parameter of the epidemiology of AIDS, statistical estimates based on transfusion-associated AIDS cases have, up to now, used only the single dataset provided by the AIDS program of the Centers for Disease Control (CDC) in Atlanta. Using a new dataset provided by the Direction Générale de la Santé (DGS), of the French Ministry of Health1, we estimate the mean incubation time for AIDS (median in brackets) to be 5.3 years (5.3 years) with a 90% confidence interval ranging from 4.4 to 8.9 years (4.4 to 8.8 years), when a Weibull distribution is postulated for incubation time. The previously encountered problem of very large confidence intervals (range larger than 100 years), is not observed, indicating that an accurate estimate for mean incubation time will be obtainable in the near future.

Acquired Immunodeficiency Syndrome↗

Contribution of morphometry in the differential diagnosis of fine-needle thyroid aspirates.

BACKGROUND: Cytologic discrimination of cellular nodules, follicular adenoma, and follicular carcinoma in the thyroid is problematic. Methods are needed to achieve a reliable diagnosis. Some sophisticated tools, such as microarrays, offer great potential but lack accompanying morphologic information. METHODS: One hundred twelve samples obtained from patients with lesions histopathologically diagnosed as nodular goiter, follicular adenoma, follicular carcinoma, and papillary carcinoma were used. Eight geometric features, such as nuclear area and circular form factor, were measured. The dataset was divided into six overlapping groups to represent the frequently encountered situations in routine practice. Multivariate analysis of variance, Tukey's honestly significant differences test, and discriminant analysis were performed. Statistical analysis was carried out with two conceptually different approaches. In the first, data from all measured nuclei were used. In the second, a subset of data representing the most extreme values of variables was extracted from the entire dataset to simulate the "selection procedure" performed during conventional morphologic examination. RESULTS: When the selected dataset instead of data from all measured nuclei was used, the correct classification rates in discriminant analysis improved considerably. CONCLUSIONS: Morphologic examination is based primarily on selection. Using data obtained from all of the cells in morphometry may cause a dilution effect in diagnostically important features. Morphometric studies may also be planned with a proper selection "bias." This may be particularly helpful when isolated abnormal cells carry most of the diagnostic information.

Adenocarcinoma, Follicular↗