Search PubMedSearch

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Clinical trials in psychiatry: should protocol deviation censor patient data?

Clinical trials methodologists recommend counting all events regardless of adherence to protocol and comparing originally randomized groups. This strict "intent-to-treat" policy implies availability and use of outcome measures taken regardless of adherence to treatment protocol. However, outcome measurement in psychiatry requires the cooperation of the patient, and usually occurs in the context of treatment management. Consequently, the patient's or clinician's decision not to adhere to the treatment protocol may be design or default cause censorship of patient data by early truncation. This disables the analysis "by intent to treat" in the strict sense. Current methods applied to such nonrandomly truncated datasets are unsatisfactory ("last value" analysis, survival analysis) or worse (imputation by last value carried forward). I review the context of clinical experimentation in psychiatry, contrast the state of design and analysis with expert recommendations on general methods, review the current statistical strategies and propose that investigators should try to obtain complete follow-up data on all patients without regard to their adherence to treatment protocol.

Clinical Trials as Topic

Analysis of repeated categorical data using generalized estimating equations.

Moment methods for analysing repeated binary responses have been proposed by Liang and Zeger, and extended by Prentice and Zhao and Prentice. In these estimating equations, models are proposed for the correlation between the repeated binary responses. We extend Liang and Zeger's method to models for the correlation between repeated nominal or ordinal categorical responses; in particular, when the repeated responses are binary, our methods reduce to Liang and Zeger's method. Our method is illustrated with two datasets. One dataset contains repeated observations of self-assessment of arthritis, an ordered variable with three categories, collected during a randomized comparative study of alternative treatments of patients with rheumatoid arthritis. The second dataset is a longitudinal study of the health effects of air pollution, in which the repeated ordered multinomial response is the wheezing status (no wheeze, wheeze with cold, wheeze apart from cold) of a child at ages 9, 10, 11 and 12 years.

Adult

The hidden effect of time.

It is customary to regard datasets as homogeneous with respect to the order of collection of the measurements. Examples are given in which this assumption is breached. Hidden time trends have implications for the design of studies, their analysis and interpretation. It is suggested that, if the order of observations is known, a plot by time should be performed, perhaps using a cusum.

Aged

Calculation of benchmark doses from teratology data.

The benchmark dose approach has several potential advantages over the no observed adverse effect level (NOAEL) as a basis for risk assessment of toxic chemicals, based upon animal toxicity data. The practical use of the benchmark dose has been evaluated by applying dose-response models to an extensive historical database of teratology bioassays. Doses corresponding to 1 and 5% increases in incidence of lesions are calculated and compared to NOAELs. The statistical accuracy of these estimates was determined by calculating confidence intervals. The lower confidence limit on the 5% benchmark dose (LED05) is found to be comparable to the NOAEL for most datasets, and slightly higher on average. Benchmark doses at the 1% level could not be estimated accurately (i.e., they had wide confidence intervals) for a significant fraction of the datasets. LED01 values were lower on average than the NOAEL. Based on these results, it is concluded that benchmark doses for a 5% increases in incidence can be calculated for most datasets, and could be used as a satisfactory basis for risk assessment, e.g., to set reference doses or acceptable daily intakes. An exception occurs when the benchmark dose exceeds the highest dose of the study. This is only likely to occur when the chemical causes a small, but significant, increase in a finding that is uncommon in untreated animals.

Abnormalities, Drug-Induced

The Evolutionary Significance of Leaf Nodulation: Evidence from Ardisia and Its Relatives (Primulaceae: Myrsinoideae).

Interactions between plants and microorganisms have long been a central topic in biological research. Bacterial symbiosis on leaf surfaces represents a distinctive and mutually beneficial system within the phyllosphere microbiome. Leaf nodules are the visible manifestation of the symbiosis and confer ecological advantages to host plants by enhancing host resistance against pathogens and herbivores. It has been hypothesized that these advantages promote higher diversification rates in host lineages, but this remains uncertain. Ardisia subg. Crispardisia and its close relatives (Amblyanthopsis and Amblyanthus) within Primulaceae are typical plant groups with leaf nodule symbiosis, making them an ideal system for testing this hypothesis. In this study, we conducted extensive sampling of "Ardisioids" (Ardisia and its allies) and reconstructed their phylogenetic relationships and evolutionary history using plastid genomes and nuclear datasets (i.e., nuclear ribosomal DNA (nrDNA) and genome-wide single nucleotide polymorphisms (SNPs)). We clarified the phylogenetic positions of several "Ardisioids" genera (e.g., Sadiria, Tapeinosperma, Amblyanthus, and Amblyanthopsis) and multiple subgenera within Ardisia. We further detected a rapid radiation during the middle Miocene in Ardisia and its allies. Notably, we found that the leaf-nodulated clade appears to have originated during this period, approximately 11-8 Ma. BAMM (Bayesian Analysis of Macroevolutionary Mixtures) analyses revealed elevated diversification rates in leaf-nodulated lineages, while HiSSE (Hidden State Speciation and Extinction) analyses indicated that leaf nodule symbiosis might have increased speciation rates without significantly affecting extinction rates. These results provide strong evidence that leaf nodule symbiosis, together with other abiotic and biotic factors, represents a key evolutionary innovation that has promoted diversification in Ardisia and its close relatives.

diversification rate

An approach to the use of the DO IT Study Group guidelines for supporting the optimal implementation of information systems in diabetes care.

The usefulness of the 1992 DO IT Study Group guidelines for diabetes data information systems was assessed using two established diabetes databases designed for different purposes. The recommendations detailed in the guidelines, written in four separate but overlapping modules, were applied individually to each database in turn. Percentage compliance with the recommendation to collect the DIABCARE dataset was high, after discounting specialist areas. While on the whole the information systems complied with the guidelines within the purposes for which they were designed, areas highlighted as demanding further action in at least one of the two systems included password protection, data validation checks, screen design, and communication with those whose records were held on the systems. Application of the guidelines is already stimulating attention to some of these areas. Some of the guidelines proved rather vague in construction to be applied in any formal sense, while others (for example in relation to international accreditation of datasets) were not applicable to individual systems. The results suggest the importance of a structured approach to the design, development, and ongoing assessment of information systems in diabetes, but require the present guidelines to develop a more formal structure to be fully effective. The widespread adoption and further testing and refinement of these guidelines (both within and outside Europe) should promote the ultimate goal of improved diabetes care.

Communication

The NCIC-Manitoba Breast Tumor Bank: a resource for applied cancer research.

The NCIC-Manitoba Breast Tumor Bank is one of several tumour banks that have been established through the Molecular Epidemiology Program of the National Cancer Institute of Canada (NCIC). The NCIC-Manitoba Breast Tumor Bank is an example of one model developed to facilitate research designed to translate the findings of basic science into information useful in the clinical arena. The tumour bank's mandate is to provide a national resource that consists of a preassembled dataset of matched samples of paraffin-embedded and frozen tumour tissue with corresponding pathological and clinical data. In the first 3 years the tumour bank has accrued data and samples from over 1800 cases of breast cancer and has provided support for 20 research projects across Canada and the United States.

Academies and Institutes

Artificial intelligence for dental caries detection: An umbrella review.

Artificial intelligence (AI) has been proposed as a tool to improve dental caries detection across imaging modalities; however, its clinical value remains uncertain. This umbrella review aimed to synthesize and critically appraise systematic reviews evaluating AI for caries detection and diagnosis. An umbrella review was conducted following PRIOR guidance (PROSPERO CRD420261340728). Searches were performed in MEDLINE, Embase, Scopus, Web of Science, and Google Scholar up to 15 March 2026. Methodological quality was assessed using AMSTAR 2, and overlap of primary studies was quantified using the corrected covered area (CCA). Seventeen systematic reviews were included, of which five reported diagnostic test accuracy meta-analyses using bivariate or HSROC models. Across these meta-analyses, pooled sensitivity ranged from 0.76 to 0.94 and specificity from 0.85 to 0.91. Most systems were based on deep learning models applied to bitewing radiographs and intraoral photographs. However, substantial heterogeneity was observed in imaging modalities, lesion thresholds, analytical tasks, and evaluation metrics. In addition, a high degree of overlap across reviews and recurrent methodological limitations, including reliance on retrospective datasets, limited external validation, and inconsistent reporting, substantially weaken the reliability of the evidence. Although AI models demonstrate high diagnostic performance under experimental conditions, current evidence does not support their use as stand-alone diagnostic tools. Their clinical applicability remains limited, and implementation should be restricted to decision-support contexts until robust prospective validation demonstrates meaningful impact on clinical decision-making and patient outcomes.

Dental Caries

The hospital information system as a source for the planning and feed-back of specialized health care.

1. INTRODUCTION. In university hospitals, choices are made to which extend specialized health care will be supported. It is characteristic, for this type of care, that it takes place in a process of the continual advance of medical technology and the growing awareness by consumers and payors. Specialized healthcare contributes to the hospital qualifiers having a political and strategic impact. The hospital board needs information for planning and budgeting these new tasks. Much of the information will be based on data stored in the Hospital Information System (HIS). Due to load limitations, instant retrieval is not preferred. A separate executive information system, uploaded with HIS data, features statistics, on a corporate level, with the power to drill-down to detailed levels. However, the ability to supply information on new types of healthcare is limited since most of these topics require a flexible system for new dedicated cross-sections, like medical treatment from several specialisms and functional levels. 2. DATA RETRIEVAL AND DISTRIBUTION. During the information analysis, details were gathered on the necessary working procedures and the administrative organization, including the data registration in the HIS. In the next phase, all relevant data was organized in a relational datamodel. For each topic of care, dedicated views were developed at both low and high aggregation levels. It revealed that a matching change of the administrative organization was required, with an emphasis on financial registration aspects. For the selection of relevant data, a bottom-up approach was applied, which was based on the registrations starting from the patient administrative subsystem, through several transactional systems, ending at the general ledger in the HIS. Data on all levels was gathered, resulting in medical details presented in quantities, up to financial figures expressed in amounts of money. This procedure distinguishes from the predefined top-down techniques generally used for management and executive information systems. Data was regularly collected from the HIS, then converted and reorganized into relational datasets using XBase protocols. After having performed central quality controls and privacy protection measures, the datasets were distributed electronically to local PCs. Standard low-cost software packages enable analyses by user-friendly selection and presentation facilities. 3. EVALUATION. The method developed for data retrieval is flexible and easy to implement. If all basic data is registered in the HIS, the procedure can be applied for all strategic hospital functions that require planning and controlling during a certain time. Critical success factors and pitfalls will be presented in the poster. Using one consistent dataset, the information required about production and budget is presented at several functional levels and is quantified in units familiar to that level. The motivation for fast and accurate registration in the HIS was improved from the moment the medical and administrative staff recognized their own data in the feed-back on specialized health care.

Decision Making, Organizational

On the use of a hospital information system in evaluating clinical care: a case report.

In this paper we describe, as an example, how we obtained the information needed to evaluate a newly introduced protocol for ordering X-rays for ankle trauma patients. Extensive use was made of available data and facilities of the hospital information system (HIS). Procedures for collecting the required additional data, which were not recorded in the HIS but were needed to evaluate the protocol, were embedded in the current medical and administrative routine of the emergency room. These additional data were also stored in the HIS. Periodically all data were downloaded to a personal computer to analyse the impact of using the protocol on quality of care and costs. In total 1241 patients entered the study, and for 1149 patients a complete dataset was obtained. The sensitivity and specificity of the protocol at the threshold value which was used during the initial study period was 0.77 and 0.80. The reduction in the number of ankle X-rays due to the protocol was significant when compared with a strategy of ordering an X-ray for every ankle trauma patient visiting the emergency room.

Ankle Injuries

Optimum numerical integration methods for estimation of area-under-the-curve (AUC) and area-under-the-moment-curve (AUMC).

Eleven numerical methods for estimation of AUC (including 4 new methods) and 22 methods for AUMC (including 8 new methods) were tested on large simulated noisy datasets representing bolus, oral and infusion concentration-time profiles. Some methods were unacceptable because their mean error was large; these included a commonly recommended form of the linear trapezoidal rule for AUMC. Others, notably Lagrange and cubic spline methods, were unacceptable because the variance of their estimates was large. These methods should be abandoned. A simple and easily programmed new method, parabolas-through-the-origin then log-trapezoidal rule, performed especially well.

Numerical Analysis, Computer-Assisted

Some conceptual and statistical issues in analysis of longitudinal psychiatric data. Application to the NIMH treatment of Depression Collaborative Research Program dataset.

Longitudinal studies have a prominent role in psychiatric research; however, statistical methods for analyzing these data are rarely commensurate with the effort involved in their acquisition. Frequently the majority of data are discarded and a simple end-point analysis is performed. In other cases, so called repeated-measures analysis of variance procedures are used with little regard to their restrictive and often unrealistic assumptions and the effect of missing data on the statistical properties of their estimates. We explored the unique features of longitudinal psychiatric data from both statistical and conceptual perspectives. We used a family of statistical models termed random regression models that provide a more realistic approach to analysis of longitudinal psychiatric data. Random regression models provide solutions to commonly observed problems of missing data, serial correlation, time-varying covariates, and irregular measurement occasions, and they accommodate systematic person-specific deviations from the average time trend. Properties of these models were compared with traditional approaches at a conceptual level. The approach was then illustrated in a new analysis of the National Institute of Mental Health Treatment of Depression Collaborative Research Program dataset, which investigated two forms of psychotherapy, pharmacotherapy with clinical management, and a placebo with clinical management control. Results indicated that both person-specific effects and serial correlation play major roles in the longitudinal psychiatric response process. Ignoring either of these effects produces misleading estimates of uncertainty that form the basis of statistical tests of hypotheses.

Analysis of Variance

Epigenetic Profiling for Early Detection and Treatment Response Monitoring in Non-Small Cell Lung Cancer: Protocol for a Prospective Translational Biomarker Study.

BACKGROUND: Non-small cell lung cancer (NSCLC) is the leading cause of cancer-related mortality worldwide and continues to have poor survival outcomes, with most patients diagnosed at advanced stages of disease. In New Zealand, NSCLC contributes substantially to cancer inequities, with Māori communities experiencing disproportionately high incidence and mortality rates. Although low-dose computed tomography screening can improve early detection, major limitations remain, including false-positive findings, overdiagnosis, high infrastructure costs, and limited accessibility for rural and underserved populations. Liquid biopsy approaches using circulating tumor DNA (ctDNA), particularly DNA methylation profiling, have emerged as promising, minimally invasive strategies for improving cancer detection, treatment monitoring, and precision oncology. OBJECTIVE: This study aims to establish integrated genomic and epigenomic predictive and prognostic biomarkers using ctDNA, tumor tissue, and transcriptomic profiling to improve early detection, risk stratification, treatment selection and response prediction, and longitudinal monitoring, with particular emphasis on identifying molecular mechanisms associated with treatment resistance and disease progression. METHODS: This prospective observational translational biomarker study is being conducted through the University of Otago and associated respiratory and oncology services in New Zealand. The study will recruit participants with NSCLC (including squamous and nonsquamous subtypes), individuals referred to fast-track lung nodule assessment clinics, and nonmalignant respiratory controls. Serial peripheral blood sampling will be performed in selected participants at predefined clinical follow-up time points to evaluate treatment response and disease progression. The availability of formalin-fixed paraffin-embedded archival tissues will be recorded, but will not be mandatory for enrollment. Genome-scale DNA methylation profiling will be performed using cell-free reduced representation bisulfite sequencing (cfRRBS), while targeted genomic profiling and transcriptomic analyses will be conducted using targeted sequencing panels and RNA sequencing. Integrative bioinformatic analyses will be used to identify molecular biomarkers associated with early-stage disease, advanced disease, treatment response, and therapeutic resistance. RESULTS: Ethics approval for the study has been obtained from the New Zealand Health and Disability Ethics Committee (2022 EXP 12566). This study commenced in 2022, and recruitment and biospecimen collection are ongoing. The study aims to recruit approximately 450 participants, including patients with NSCLC, individuals referred through respiratory diagnostic pathways, and nonmalignant controls. As of July 31, 2026, 205 participants have been recruited, with recruitment continuing until the target sample size is reached. Molecular and data analyses are ongoing, with additional publications expected as the cohort matures. CONCLUSIONS: This study will generate one of the first integrated genomic, epigenomic, and transcriptomic liquid biopsy datasets for NSCLC in New Zealand. The findings are expected to support the development of sensitive, accessible, and equitable blood-based biomarkers for NSCLC detection and treatment monitoring while also contributing to improved precision oncology approaches and reducing NSCLC inequities among Māori populations.

Humans

Development of predictive models of laboratory animal growth using artificial neural networks.

Traditional regression analysis of body weight growth curves encounters problems when the data are extremely variable. While transformations are often employed to meet the criteria of the analysis, some transformations are inadequate for normalizing the data. Regression analysis also requires presuppositions regarding the model to be fit and the techniques to be used in the analysis. An alternative approach using artificial neural networks is presented which may be suitable for developing predictive models of growth. Neural networks are simulators of the processes that occur in the biological brain during the learning process. They are trained on the data, developing the necessary algorithms within their internal architecture, and produce a predictive model based on the learned facts. A dataset of Sprague-Dawley rat (Rattus norvegicus) weights is analyzed by both traditional regression analysis and neural network training. Predictions of body weight are made from both models. While both methods produce models that adequately predict the body weights, the neural network model is superior in that it combines accuracy and precision, being less influenced by longitudinal variability in the data. Thus, the neural network provides another tool for researchers to analyze growth curve data.

Algorithms

The influence of radiotherapy treatment time on the control of laryngeal cancer: a direct analysis of data from two British Institute of Radiology trials to calculate the lag period and the time factor.

This study analyses node-negative laryngeal tumour control data from two clinical trials conducted by the British Institute of Radiology in order to determine the time factors and the presence or absence of a lag period before the time factor takes effect. A direct maximum likelihood approach is used to fit a double-logarithmic model including a repopulation term which commences after an initial lag period, Tk. The analysis yields a time factor of 0.8 Gy per day (95% confidence interval 0.5-1.1 Gy per day) as the extra dose required to counteract the reduction in tumour control probability (TCP) with extension of the treatment time. The latter reduction amounted to between 5 and 12% TCP per week, depending on the stage and time period. With this dataset, where few patients were treated for short times, no statistically significant lag phase can be demonstrated. However, the best estimate of Tk is 21 days (95% confidence interval 0-27 days), which is consistent with estimates from other studies on other datasets. If a lag phase exists, this study would indicate that the duration is less than 27 days. Other studies have used retrospective data and are subject to a number of potential biases. The present study, using data from multicentre prospective randomized clinical trials, is free from some of these sources of bias. The fact that very similar estimates of the radiobiological parameters are obtained lends credence to these other studies and suggests that the potential biases may be small in practice.

Clinical Trials as Topic

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models

History of the Cancer Information Service.

The Cancer Information Service (CIS) was established on July 1, 1975, following the mandate of the National Cancer Act of 1971 giving the National Cancer Institute (NCI) new responsibilities for educating the public, patients, and health professionals. Funded under a contract mechanism, the CIS has become one of the longest-running community programs in NCI. The CIS has been able to set up and maintain high-quality service, giving accurate, up-to-date medical information to cancer patients and their families and friends, to health professionals, and to the general public. The CIS network, which has taken more than 5 million calls since its inception, has weathered many changes, both at the national and the local level. Its current call volume, in excess of 500,000 calls per year, makes it one of the most heavily utilized health-related telephone helplines in the country. Using a standardized Call Record Form, data on calls have been recorded consistently since 1983; the dataset now contains information on more than 4.2 million calls. An outreach component that acts as NCI's field arm has been part of the CIS since its inception. The CIS has matured into a stable system that has been reconfigured into 19 regional offices, covering the entire country. These offices run the telephone service and serve as NCI's outreach arm, working with intermediaries to carry out NCI information and education programs in local communities.

History, 20th Century

Incubation time for AIDS from French transfusion-associated cases.

Although incubation time is a key parameter of the epidemiology of AIDS, statistical estimates based on transfusion-associated AIDS cases have, up to now, used only the single dataset provided by the AIDS program of the Centers for Disease Control (CDC) in Atlanta. Using a new dataset provided by the Direction Générale de la Santé (DGS), of the French Ministry of Health1, we estimate the mean incubation time for AIDS (median in brackets) to be 5.3 years (5.3 years) with a 90% confidence interval ranging from 4.4 to 8.9 years (4.4 to 8.8 years), when a Weibull distribution is postulated for incubation time. The previously encountered problem of very large confidence intervals (range larger than 100 years), is not observed, indicating that an accurate estimate for mean incubation time will be obtainable in the near future.

Acquired Immunodeficiency Syndrome