Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Analysing spatially referenced public health data: a comparison of three methodological approaches.

In the analysis of spatially referenced public health data, members of different disciplinary groups (geographers, epidemiologists and statisticians) tend to select different methodological approaches, usually those with which they are already familiar. This paper compares three such approaches in terms of their relative value and results. A single public health dataset, derived from a community survey, is analysed by using 'traditional' epidemiological methods, GIS and point pattern analysis. Since they adopt different 'models' for addressing the same research question, the three approaches produce some variation in the results for specific health-related variables. Taken overall, however, the results complement, rather than contradict or duplicate each other.

Adult↗

Preliminary assessment of three-dimensional magnetic resonance imaging for various colonic disorders.

BACKGROUND: Improvements in magnetic resonance imaging (MRI) technology have enabled the acquisition of three-dimensional MRI datasets in a single breath hold. We adopted this technique to make a three dimensional intraluminal and extraluminal assessment of the colon in three patients with various colonic disorders. METHODS: One patient was studied after having a double-contrast barium enema. Two patients had MRI scans after colonoscopy, which showed three colonic tumours in one and multiple polyps in the ascending colon of the other. The process of rectal filling with 1.5-2.0 L water mixed with 15-20 mL 0.5 mol/L gadolinium-diethylenetriaminepentaacetic acid (Gd-DTPA) was monitored with MR fluoroscopic sequence. Three-dimensional datasets of the contrast-filled colon were taken with patients in prone (before and after intravenous administration of 0.1 mmol/kg bodyweight Gd-DTPA) and supine positions. 64 sections with a voxel-resolution of 2.0 x 2.0 x 1.25 mm3-were taken during a 28 s breath hold. Three-dimensional maximum intensity projection, multiplanar reconstruction, and virtual colonoscopic images of the colon were created from these. FINDINGS: Analysis of the coronal source images in conjunction with multiplanar reconstructions revealed all relevant abnormalities, including diverticula, carcinomas, and polyps. Three dimensional maximum-intensity projections gave a morphological overview of the whole colon. Targeted projections, made up of a limited number of coronal source images, showed diverticula and smaller polyps more clearly. After patients were given intravenous contrast all colonic mass lesions were enhanced. Datasets obtained in prone patients gave the best intraluminal views of the colon. Virtual magnetic resonance colonoscopy showed colonic haustra as well as the ileocaecal valve, but did not show clearly the diverticula. All intraluminal mass lesions, on the other hand, were easy to see. INTERPRETATION: The potential of three-dimensional colonic MRI to provide accurate, minimally invasive, cost-effective polyp screening, as well as comprehensive colonic tumour staging, warrants further investigation.

Aged↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗

History of the Cancer Information Service.

The Cancer Information Service (CIS) was established on July 1, 1975, following the mandate of the National Cancer Act of 1971 giving the National Cancer Institute (NCI) new responsibilities for educating the public, patients, and health professionals. Funded under a contract mechanism, the CIS has become one of the longest-running community programs in NCI. The CIS has been able to set up and maintain high-quality service, giving accurate, up-to-date medical information to cancer patients and their families and friends, to health professionals, and to the general public. The CIS network, which has taken more than 5 million calls since its inception, has weathered many changes, both at the national and the local level. Its current call volume, in excess of 500,000 calls per year, makes it one of the most heavily utilized health-related telephone helplines in the country. Using a standardized Call Record Form, data on calls have been recorded consistently since 1983; the dataset now contains information on more than 4.2 million calls. An outreach component that acts as NCI's field arm has been part of the CIS since its inception. The CIS has matured into a stable system that has been reconfigured into 19 regional offices, covering the entire country. These offices run the telephone service and serve as NCI's outreach arm, working with intermediaries to carry out NCI information and education programs in local communities.

History, 20th Century↗

Incubation time for AIDS from French transfusion-associated cases.

Although incubation time is a key parameter of the epidemiology of AIDS, statistical estimates based on transfusion-associated AIDS cases have, up to now, used only the single dataset provided by the AIDS program of the Centers for Disease Control (CDC) in Atlanta. Using a new dataset provided by the Direction Générale de la Santé (DGS), of the French Ministry of Health1, we estimate the mean incubation time for AIDS (median in brackets) to be 5.3 years (5.3 years) with a 90% confidence interval ranging from 4.4 to 8.9 years (4.4 to 8.8 years), when a Weibull distribution is postulated for incubation time. The previously encountered problem of very large confidence intervals (range larger than 100 years), is not observed, indicating that an accurate estimate for mean incubation time will be obtainable in the near future.

Acquired Immunodeficiency Syndrome↗

A body image scale for use with cancer patients.

Body image is an important endpoint in quality of life evaluation since cancer treatment may result in major changes to patients' appearance from disfiguring surgery, late effects of radiotherapy or adverse effects of systemic treatment. A need was identified to develop a short body image scale (BIS) for use in clinical trials. A 10-item scale was constructed in collaboration with the European Organization for Research and Treatment of Cancer (EORTC) Quality of Life Study Group and tested in a heterogeneous sample of 276 British cancer patients. Following revisions, the scale underwent psychometric testing in 682 patients with breast cancer, using datasets from seven UK treatment trials/clinical studies. The scale showed high reliability (Cronbach's alpha 0.93) and good clinical validity based on response prevalence, discriminant validity (P<0.0001, Mann-Whitney test), sensitivity to change (P<0.001, Wilcoxon signed ranks test) and consistency of scores from different breast cancer treatment centres. Factor analysis resulted in a single factor solution in three out of four analyses, accounting for >50% variance. These results support the clinical validity of the BIS as a brief questionnaire for assessing body image changes in patients with cancer, suitable for use in clinical trials.

Age Factors↗

Comparison of intron-containing and intron-lacking human genes elucidates putative exonic splicing enhancers.

Of the rules used by the splicing machinery to precisely determine intron-exon boundaries only a fraction is known. Recent evidence suggests that specific short sequences within exons help in defining these boundaries. Such sequences are known as exonic splicing enhancers (ESE). A possible bioinformatical approach to studying ESE sequences is to compare genes that harbor introns with genes that do not. For this purpose two non-redundant samples of 719 intron-containing and 63 intron-lacking human genes were created. We performed a statistical analysis on these datasets of intron-containing and intron-lacking human coding sequences and found a statistically significant difference (P = 0.01) between these samples in terms of 5-6mer oligonucleotide distributions. The difference is not created by a few strong signals present in the majority of exons, but rather by the accumulation of multiple weak signals through small variations in codon frequencies, codon biases and context-dependent codon biases between the samples. A list of putative novel human splicing regulation sequences has been elucidated by our analysis.

Alternative Splicing↗

Safety, permanency, and in-home services: applying administrative data.

This article describes the construction and use of safety and permanency indicators, two aspects of a full set of indicators that also includes child well-being and family functioning. The indicators were constructed from Philadelphia's Family and Child Tracking System and were used to examine the city's Services to Children in their Own Home (SCOH) program. Cohort datasets were constructed through the use of extract files, and two independent data file construction algorithms were employed to calibrate the accuracy of the data construction process. The primary unit of analysis was the "family" spell in SCOH services. Contextual variables included family structure, race, and service intensity. The indicators associated with SCOH spells included reports of maltreatment after service, founded maltreatment after service, and out-of-home placement after service. Event history techniques were used to conduct the data analysis. Baseline indicator data for Philadelphia are presented, and future uses for such data are discussed.

Child↗

NMR spectral quantitation by principal-component analysis. II. Determination of frequency and phase shifts.

This paper extends the use of principal-component analysis in spectral quantification to the estimation of frequency and phase shifts in a single resonant peak across a series of spectra. The estimated parameters can be used to correct the spectra accordingly, resulting in more accurate peak-area estimation. Further, the removal of the variations in phase and frequency cause by instrumental and experimental fluctuations makes it possible to determine more accurately the remaining variations, which bear biological significance. The procedure is demonstrated on simulated data, a 3D chemical-shift-imaging dataset acquired from a cylinder of inorganic phosphate (Pi), and a set of 736 31P NMR in vivo spectra taken from a kinetic study of rate muscle energetics. In all cases, the procedure rapidly and automatically identifies the frequency and phase shifts present in the individual spectra. In the kinetic study, the procedure is used twice, first to adjust the phase and frequency of a reference peak (phosphocreatine) and then to determine the individual frequencies of the Pi peak in each of the spectra which further can be used for estimation of pH changes during the experiment.

Computer Simulation↗

B-SPID: an object-relational database architecture to store, retrieve, and manipulate neuroimaging data.

We propose a hardware and software architecture to respond to crucial problems in the neuroimaging field: storage, retrieval, and processing of large datasets. The B-SPID project, here discussed, concerns the processing of neuroimages and attached components stored in an object-relational multimedia database management system (DBMS). Advanced bioinformation concepts are exploited in this project such as large scale data storage, high level graphical user interfaces and 3D graphical processing and display of data. Our database implementation is based on standard programming components, runs on several UNIX platforms and is written to be evolutive. Queries on this database are designed to obtain and display from neuroimaging data several types of results (pictures, text, or 3D graphical shapes) on heterogeneous systems.

Brain Mapping↗

Molecular phylogeny and evolutionary history of the tit-tyrants (Aves: Tyrannidae).

Tit-tyrants of the genus Anairetes presently consist of six species; five inhabit various regions along the Andean cordillera of South America and one is endemic to the Juan Fernandez Islands off the coast of Chile. Data from mtDNA ND2 and Cyt b sequences were used to construct a phylogeny for all Anairetes species as well as Uromyias agilis, a closely related genus, and Stigmatura as an outgroup, to determine their relationships and history of radiation in South America. Results strongly supported the following paired relationships: A. nigrocristatus-A. reguloides, A. flavirostris-A. alpinus, and A. parulus-A. fernandezianus. This dataset, however, could not resolve basal nodes; therefore relationships among these pairs remains obscure. Moreover the genus Uromyias, controversially separated on morphological criteria from Anairetes, fell within the Anairetes clade, although its exact position could not be ascertained with confidence. The molecular data indicate that this group probably radiated within the past 2 million years, concomitant with highly accentuated cycles of global climatic change. Certain high altitude areas within the Andes may have been stable during global climatic changes and may have served as refugia during the Plio-Pleistocene.

Animals↗

Comparison of chemotherapy and bone marrow transplants using two independent clinical databases.

Comparing the outcome of chemotherapy and bone marrow transplants in the absence of a randomized trial is difficult but necessary for diseases where small numbers of patients make such trials difficult if not impossible. To address this issue for adults with acute lymphoblastic leukemia in first remission, we created an empirical database using two separate datasets, one from the International Bone Marrow Transplant Registry and the other from two multicenter chemotherapy studies. Prior to combining the datasets, a study protocol was developed to define inclusion criteria, outcomes to be compared and statistical methods. The main problems of a non-randomized comparison are biases potentially introduced by differences in baseline composition of the two cohorts and differences in time-to-treatment. The source of the latter bias is different distributions of waiting times between achieving complete remission and receiving post-remission therapy. Several techniques to control these biases were evaluated; each gave qualitatively similar results. These methods can easily be applied to other clinical situations where randomized trials are not available.

Adolescent↗

The classification of subjects with joint complaints on incomplete biochemical and haematological datasets.

We performed a retrospective study on 163 subjects suffering from rheumatic fever (16), rheumatoid arthritis (36), lupus erythematosus (17), gout (21), arthrosis (50) and osteomyelitis (23). The number of variables evaluated was 39. These were all of a general biochemical and haematological nature. A feature reduction resulted in sixteen variables that matched well with those known from the literature. Linear discriminant analysis yielded poor results in classifying the six disease categories (with 18 variables 61.8%). A reduction to three disease categories improved the classification results remarkably. This, and the excellent discriminating power between patients and the reference group, shows that the selected variables are illustrative only for general clinical pictures, such as infection, and not for the desired differential diagnosis.

Arthritis, Rheumatoid↗

A new multiparameter flow cytometer: optical and electrical cell analysis in combination with video microscopy in flow.

BACKGROUND: Flow cytometers, which are commercially available, do not necessarily meet all demands of actual biomedical research. This is the case for the investigation of mechanisms involved in cell volume regulation, which requires electrical volume measurement and ratiometric multichannel fluorescence analysis for the simultaneous assessment of different physiologic parameters (intracellular pH and the intracellular concentration of calcium ions, etc). METHODS AND RESULTS: We describe the construction of a new nonsorting flow cytometer designed for the simultaneous acquisition of seven parameters including fluorescence signals, forward and perpendicular light scatter, cell volume according to the electrical Coulter principle, and flow cytometric imaging. The instrument is equipped with three different light sources. A tunable argon-ion laser generates efficient excitation of the most standard fluorescent probes in the visible spectral range, and an arc lamp provides the means for ultraviolet excitation at low cost. Because of the spatial filtering by the excitation and detection optics, two independent sets of dual fluorescence measurements can be performed, a prerequisite for flexible ratiometric fluorescence analysis. A flow video microscope integrated into the optical system optionally generates either brightfield or phase images of selected flowing particles. Only particles whose individual datasets meet predefined gating conditions are imaged in real time. To avoid smear effects, the motion of the object to be imaged (speed approximately 8 m/s) is frozen on the target of a CCD camera by flash illumination. For this purpose, a high radiance gas discharge lamp with 25-mJ electric pulse energy provides an illumination time of 18 ns (full width half maximum). Test results obtained from latex spheres and cells are shown. CONCLUSIONS: Test results indicate that our instrument can perform Coulter measurements in combination with flexible optical analysis. Moreover, integration of an adapted video microscope into a flow cytometer is an approach to overcome the gap between flow and image cytometry.

Animals↗

A population needs assessment profile for dementia.

The Tayside Profile for Dementia Planning is an instrument designed to obtain data for population needs assessment and planning. It provides a brief tool to collect a minimum dataset by non-specialists. Third-party informants-informal carers or involved professionals-are used as data sources. The key concept is the use of a descriptive profile rather than a summative score or categorization. The profile consists of a set of needs indicators, information on current service response and demographic and background data. Key levels of dependency are measured by time interval dependency. Validity, reliability, acceptability and usability are satisfactory, with the crucial exception that informal carers and professionals appear to perceive needs differently. Further research is needed to assess which type of informant provides the more useful data.

Caregivers↗

Quantification of 99Tcm-HMPAO brain SPET in two series of healthy volunteers using different triple-headed SPET configurations: normal databases and methodological considerations.

We evaluated the methodological issues underlying the assessment of normal confidence intervals, as used in clinically based region-of-interest (ROI) semi-quantification of 99Tcm-HMPAO brain SPET. At two different centres equipped with high-resolution, triple-headed gamma cameras, HMPAO SPET scans were performed on two groups of 24 and 15 healthy volunteers respectively. Together with an operator-defined analysis (ODA), a semi-automated analysis (SAA) was conducted on the normal datasets in one centre. Tests of intra- and inter-observer variability were performed. Repeat scans were performed within 72 h after the first to analyse short-term regional inter-study variations. The overall regional uptake showed significant differences in most regions between both normal datasets. Intra-observer and inter-observer reproducibility were on average within 4% for the ODA, while for the SAA it was less than 1%. Inter-study variations were excellent for both centres, ranging from -4% to +3% for most regions studied. The variability in clinical brain perfusion studies largely depends on the reproducibility of the data analysis technique. A semi-automated approach shows clear advantages over an entirely operator-defined approach. Intra-subject repeat studies show enough stability for use as reliable baseline measurements in the construction of a normal database or to allow activation studies with high sensitivity.

Adult↗

Evaluation of a novel infrared range vibration-based descriptor (EVA) for QSAR studies. 1. General application.

A novel molecular descriptor (EVA) based upon calculated infrared range vibrational frequencies is evaluated for use in QSAR studies. The descriptor is invariant to both translation and rotation of the structures concerned. The method was applied to 11 QSAR datasets exhibiting both a range of biological endpoints and various degrees of structural diversity. This study demonstrates that robust QSAR models can be obtained using the EVA descriptor and examines the effect of EVA parameter changes on these models; recommendations are made as to the appropriate choice of parameters. The performance of EVA was found to be comparable in statistical terms to that of CoMFA, despite the fact that EVA does not require the generation of a structural alignment. Models derived using semiempirical (MOPAC AM1 and PM3) and AMBER mechanics calculated normal mode frequencies are compared, with the overall conclusion that the semiempirical methods perform equally well and both outperform the AMBER-based models.

Computer Simulation↗

Comparison of alternative methods for assessing injury severity based on anatomic descriptors.

BACKGROUND: There is mounting confusion as to which anatomic scoring systems can be used to adequately control for trauma case mix when predicting patient survival. METHODS: Several Abbreviated Injury Scale (AIS) and International Classification of Disease Clinical (ICD-9CM)-based methods of scoring severity were compared by using data from the Pennsylvania Trauma Outcome Study. By using a design dataset, the probability of survival was modeled as a function of each score or profile. Resulting coefficients were used to derive expected probabilities in a test dataset; expected and observed probabilities were then compared by using standard measures of discrimination and calibration. RESULTS: The modified Anatomic Profile, Anatomic Profile, and New Injury Severity Score outperformed the International Classification of Disease-based Injury Severity Score. This finding remains true when AIS values are obtained by means of a conversion from International Classification of Disease to AIS. CONCLUSION: Results support the integrity of the AIS and argue for its continued use in research and evaluation. The modified Anatomic Profile, Anatomic Profile, and New Injury Severity Score, however, should be used in preference to the Injury Severity Score as an overall measure of severity.

Humans↗