Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Changes in physicians' sources of pharmaceutical information: a review and analysis.

Since 1952, 20 datasets have been generated through 17 studies in an attempt to describe the sources and importance and/or use of information about pharmaceuticals by physicians. The authors review the findings of the studies and subject them to three sequentially relevant, but different, meta-analytic procedures. The results of these analyses indicate significant changes in the sources and importance of various commercial/noncommercial and personal/nonpersonal information as they relate to physicians' prescribing behavior. Those changes over time have specific implications for marketers of pharmaceuticals.

Advertising↗

Evaluation of gestational age and admission date assumptions used to determine prenatal drug exposure from administrative data.

OBJECTIVE: Our aim was to evaluate the 270-day gestational age and delivery date assumptions used in an administrative dataset study assessing prenatal drug exposure compared to information contained in a birth registry. STUDY DESIGN AND SETTING: Kaiser Permanente Colorado (KPCO), a member of the Health Maintenance Organization (HMO) Research Network Center for Education and Research in Therapeutics (CERTs), previously participated in a CERTs study that used claims data to assess prenatal drug exposure. In the current study, gestational age and deliveries information from the CERTs study dataset, the Prescribing Safely during Pregnancy Dataset (PSDPD), was compared to information in the KPCO Birth Registry. Sensitivity and positive predictive value (PPV) of the claims data for deliveries were assessed. The effect of gestational age and delivery date assumptions on classification of prenatal drug exposure was evaluated. RESULTS: The mean gestational age in the Birth Registry was 273 (median = 275) days. Sensitivity of claims data at identifying deliveries was 97.6%, PPV was 98.2%. Of deliveries identified in only one dataset, 45% were related to the gestational age assumption and 36% were due to claims data issues. The effect on estimates of prevalence of prescribing during pregnancy was an absolute change of 1% or less for all drug exposure categories. For Category X, drug exposures during the first trimester, the relative change in prescribing prevalence was 13.7% (p = 0.014). CONCLUSION: Administrative databases can be useful for assessing prenatal drug exposure, but gestational age assumptions can result in a small proportion of misclassification.

Databases, Factual↗

GrainGenes, the genome database for small-grain crops.

GrainGenes, http://www.graingenes.org, is the international database for the wheat, barley, rye and oat genomes. For these species it is the primary repository for information about genetic maps, mapping probes and primers, genes, alleles and QTLs. Documentation includes such data as primer sequences, polymorphism descriptions, genotype and trait scoring data, experimental protocols used, and photographs of marker polymorphisms, disease symptoms and mutant phenotypes. These data, curated with the help of many members of the research community, are integrated with sequence and bibliographic records selected from external databases and results of BLAST searches of the ESTs. Records are linked to corresponding records in other important databases, e.g. Gramene's EST homologies to rice BAC/PACs, TIGR's Gene Indices and GenBank. In addition to this information within the GrainGenes database itself, the GrainGenes homepage at http://wheat.pw.usda.gov provides many other community resources including publications (the annual newsletters for wheat, barley and oat, monographs and articles), individual datasets (mapping and QTL studies, polymorphism surveys, variety performance evaluations), specialized databases (Triticeae repeat sequences, EST unigene sets) and pages to facilitate coordination of cooperative research efforts in specific areas such as SNP development, EST-SSRs and taxonomy. The goal is to serve as a central point for obtaining and contributing information about the genetics and biology of these cereal crops.

Alleles↗

Magnetic Resonance Spectroscopy of cancer-practicalities of multi-centre trials and early results in non-Hodgkin's lymphoma.

This review describes problems and solutions encountered in large scale multicentre trials of Magnetic Resonance Methods for monitoring cancer. It is illustrated with reference to the Multi-Institutional Group on Magnetic Resonance Spectroscopy (MRS) Applications to Cancer which was set up to perform a trial of 31P MRS for monitoring non-invasively chemotherapy of solid tumours. 31P MR spectra of non-Hodgkin's lymphoma (NHL) pre- and posttreatment, across nine Institutions, were acquired on either General Electric (GE) or Siemens 1.5T Clinical MR instruments. Development of the trial protocol, design of the Radio Frequency (RF) coils and Quality Control procedures necessary to ensure that the datasets acquired at each centre were comparable, are described. The data revealed that phosphomonoesters (PME)/nucleotide triphosphates (NTP) ratio decreased significantly after treatment in the Complete (P<0.001) and Partial (P<0.05) Responders but not in the Non-Responders (P>0.1). In addition, the PME/NTP ratio in the pre-treatment spectra correlated with the subsequent outcome of treatment indicating that PME/NTP levels are significant predictors of long-term clinical response and time-to-treatment failure in NHL.

Clinical Protocols↗

NIDDK data repository: a central collection of clinical trial data.

BACKGROUND: The National Institute of Diabetes and Digestive and Kidney Diseases have established central repositories for the collection of DNA, biological samples, and clinical data to be catalogued at a single site. Here we present an overview of the site which stores the clinical data and links to biospecimens. DESCRIPTION: The NIDDK Data repository is a web-enabled resource cataloguing clinical trial data and supporting information from NIDDK supported studies. The Data Repository allows for the co-location of multiple electronic datasets that were created as part of clinical investigations. The Data Repository does not serve the role of a Data Coordinating Center, but rather as a warehouse for the clinical findings once the trials have been completed. Because both biological and genetic samples are collected from many of the studies, a data management system for the cataloguing and retrieval of samples was developed. CONCLUSION: The Data Repository provides a unique resource for researchers in the clinical areas supported by NIDDK. In addition to providing a warehouse of data, Data Repository staff work with the users to educate them on the datasets as well as assist them in the acquisition of multiple data sets for cross-study analysis. Unlike the majority of biological databases, the Data Repository acts both as a catalogue for data, biosamples, and genetic materials and as a central processing point for the requests for all biospecimens. Due to regulations on the use of clinical data, the ultimate release of that data is governed under NIDDK data release policies. The Data Repository serves as the conduit for such requests.

Access to Information↗

Major structural determinants of transmembrane proteins identified by principal component analysis.

We identify amino acid characteristics important in determining the secondary structures of transmembrane proteins, and compare them with characteristics important for cytoplasmic proteins. Using information derived from multiple sequence alignments, we perform a principal component analysis (PCA) to identify the directions in the 20-dimensional amino acid frequency space that comprise the most variance within each protein secondary structure. These vectors represent the important position-specific properties of the amino acids for coils, turns, beta sheets, and alpha helices. As expected, the most important axis for most of the datasets was hydrophobicity. Additional axes, distinct from hydrophobicity, are surprising, especially in the case of transmembrane alpha helices, where the effects of aromaticity and beta-branching are the next two most significant characteristics. The axis representing beta-branching also has equal importance in cytoplasmic and transmembrane helices, a finding that contrasts with some experimental results in membrane-like environments. In a further analysis, we examine trends for some of the PCA axes over averaged transmembrane alpha helices, and find interesting results for aromaticity.

Amino Acids↗

Combining independent component analysis and correlation analysis to probe interregional connectivity in fMRI task activation datasets.

A new approach in studying interregional functional connectivity using functional magnetic resonance imaging (fMRI) is presented. Functional connectivity may be detected by means of cross correlating time course data from functionally related brain regions. These data exhibit high temporal coherence of low frequency fluctuations due to synchronized blood flow changes. In the past, this fMRI technique for studying functional connectivity has been applied to subjects that performed no prescribed task ("resting" state). This paper presents the results of applying the same method to task-related activation datasets. Functional connectivity analysis is first performed in areas not involved with the task. Then a method is devised to remove the effects of activation from the data using independent component analysis (ICA) and functional connectivity analysis is repeated. Functional connectivity, which is demonstrated in the "resting brain," is not affected by tasks which activate unrelated brain regions. In addition, ICA effectively removes activation from the data and may allow us to study functional connectivity even in the activated regions.

Adult↗

Analysing spatially referenced public health data: a comparison of three methodological approaches.

In the analysis of spatially referenced public health data, members of different disciplinary groups (geographers, epidemiologists and statisticians) tend to select different methodological approaches, usually those with which they are already familiar. This paper compares three such approaches in terms of their relative value and results. A single public health dataset, derived from a community survey, is analysed by using 'traditional' epidemiological methods, GIS and point pattern analysis. Since they adopt different 'models' for addressing the same research question, the three approaches produce some variation in the results for specific health-related variables. Taken overall, however, the results complement, rather than contradict or duplicate each other.

Adult↗

Risk assessment in oncology clinical practice. From risk factors to risk models.

Myelosuppression and neutropenia represent the major dose-limiting toxicity of cancer chemotherapy. Chemotherapy-induced neutropenia may be accompanied by fever, presumably due to life-threatening infection, which generally requires hospitalization for evaluation and treatment with empiric broad-spectrum antibiotics. The resulting febrile neutropenia is a major cause of the morbidity, mortality, and costs associated with the treatment of patients with cancer. Furthermore, the threat of febrile neutropenia often results in chemotherapy dose reductions and delays, which can compromise long-term clinical outcomes. Prophylactic colony-stimulating factor (CSF) has been shown to reduce the incidence, severity, and duration of neutropenia and its complications. Guidelines from the American Society of Clinical Oncology recommend the use of CSF on the basis of the myelosuppressive potential of the chemotherapy regimen. The challenge in ensuring the appropriate and cost-effective use of prophylactic CSF is to determine which patients would be most likely to benefit from it. A number of patient-, disease-, and treatment-related factors are associated with an increased risk of neutropenia and its complications. A number of clinical predictive models have been developed from retrospective datasets to identify patients at greater risk for neutropenia and its complications. Early studies have demonstrated the potential of such models to guide the targeted use of CSF to those patients who are most likely to benefit from the early use of these supportive agents. Additional prospective research is needed to develop more accurate and valid risk models and to evaluate the efficacy and cost-effectiveness of model-targeted use of CSF in high-risk patients.

Antineoplastic Agents↗

Pharmacoepidemiology--an Irish perspective.

The Irish healthcare system is a mixture of free, state-supported and private medicine. The state-supported General Medical Services (GMS) scheme maintains a large prescription database, which has been used to conduct pharmacoepidemiological studies in Ireland. The dataset is anonymized thus maintaining patient and prescriber confidentiality. Three recent studies using this data are described, two of which outline the effect of regulatory advice and the media on prescribing patterns and one which describes the development of an index of prescribing quality which may be applied to prescription data. The GMS prescription database is presently being complemented by a database for some 0.7 million people who seek reimbursement for prescriptions from individuals or families in excess of 42 Pounds per month which together will have an important role for the continued development of pharmacoepidemiology in Ireland.

Clinical Trials as Topic↗

Preliminary assessment of three-dimensional magnetic resonance imaging for various colonic disorders.

BACKGROUND: Improvements in magnetic resonance imaging (MRI) technology have enabled the acquisition of three-dimensional MRI datasets in a single breath hold. We adopted this technique to make a three dimensional intraluminal and extraluminal assessment of the colon in three patients with various colonic disorders. METHODS: One patient was studied after having a double-contrast barium enema. Two patients had MRI scans after colonoscopy, which showed three colonic tumours in one and multiple polyps in the ascending colon of the other. The process of rectal filling with 1.5-2.0 L water mixed with 15-20 mL 0.5 mol/L gadolinium-diethylenetriaminepentaacetic acid (Gd-DTPA) was monitored with MR fluoroscopic sequence. Three-dimensional datasets of the contrast-filled colon were taken with patients in prone (before and after intravenous administration of 0.1 mmol/kg bodyweight Gd-DTPA) and supine positions. 64 sections with a voxel-resolution of 2.0 x 2.0 x 1.25 mm3-were taken during a 28 s breath hold. Three-dimensional maximum intensity projection, multiplanar reconstruction, and virtual colonoscopic images of the colon were created from these. FINDINGS: Analysis of the coronal source images in conjunction with multiplanar reconstructions revealed all relevant abnormalities, including diverticula, carcinomas, and polyps. Three dimensional maximum-intensity projections gave a morphological overview of the whole colon. Targeted projections, made up of a limited number of coronal source images, showed diverticula and smaller polyps more clearly. After patients were given intravenous contrast all colonic mass lesions were enhanced. Datasets obtained in prone patients gave the best intraluminal views of the colon. Virtual magnetic resonance colonoscopy showed colonic haustra as well as the ileocaecal valve, but did not show clearly the diverticula. All intraluminal mass lesions, on the other hand, were easy to see. INTERPRETATION: The potential of three-dimensional colonic MRI to provide accurate, minimally invasive, cost-effective polyp screening, as well as comprehensive colonic tumour staging, warrants further investigation.

Aged↗

Correlating gene promoters and expression in gene disruption experiments.

MOTIVATION: Finding putative transcription factor binding sites in the upstream sequences of similarly expressed genes has recently become a subject of intensive studies. In this paper we investigate how much gene expression regulation can be attributed to the presence of various binding sites in the gene promoters by correlating the binding sites and the changes in gene expression resulting from gene disruptions (e.g. knockouts). RESULTS: We have developed a data analysis method for comparing mRNA measurements of gene disruption experiments with information about gene promoters. The method was applied to a well-known dataset to uncover correlations between known transcription factor binding site motifs in the upstream regions of all S. cerevisiae genes and the gene expression changes in various gene disruption experiments. The possible explanations of the correlations were categorized and analyzed using e.g. expression cascades. Several correlations turned out to be consistent with existing biological knowledge while some new ones suggest themselves for further study. AVAILABILITY: The resulting tables are available at http://www.cs.helsinki.fi/u/kpalin/CorrDisrupt/.

Algorithms↗

Development and preselection of criteria for short term improvement after anti-TNF alpha treatment in ankylosing spondylitis.

OBJECTIVE: To develop and compare candidate improvement criteria for anti-TNFalpha treatment in ankylosing spondylitis with optimal discriminating capacity between treatment and placebo. METHODS: Data from two randomised controlled trials which included 99 patients treated with infliximab or etanercept were used to evaluate 50 candidate improvement criteria. These were developed on the basis of pain, patient's global assessment, function, morning stiffness, spinal mobility, and C reactive protein. Different levels of improvement in each domain (20-60%) were used to define Boolean type criteria. These criteria were compared with different percentages of improvement on the BASDAI and with modified ASAS improvement criteria. Bootstrap methods were applied to calculate 95% confidence intervals (CI) of the chi(2) test values to select the best candidate improvement criteria. RESULTS: The best performing improvement criteria were "20% improvement in five of six domains" (chi(2) = 31.9 (95% CI, 18.0 to 46.9)) with a low placebo response of 2.9% and a high response to infliximab of 67.7%; and "ASAS 40% improvement" (chi(2) = 26.5 (13.3 to 41.1)), with response to placebo of 5.7% and response to infliximab of 64.7%. The good discriminating capacity of the two improvement criteria was confirmed by the combined dataset of the infliximab and etanercept trial. CONCLUSIONS: The "five of six" improvement criterion has the advantage of including the objective domains spinal mobility and acute phase reactants, but requires only 20% improvement. The ASAS 40% improvement criterion has the advantage of setting a high threshold, but only in patient reported outcomes. The choice between these improvement criteria needs to be based on further validation from upcoming trials.

Adult↗

Clinically validated genotype analysis: guiding principles and statistical concerns.

Whereas previously the output of HIV resistance tests has been based on therapeutically arbitrary criteria, there is now an ongoing move towards correlating test interpretation with virological outcomes on treatment. This approach is undeniably superior, in principle, for tests intended to guide drug choices. However the predictive accuracy of a given stratagem that links genotype or phenotype to drug response is strongly influenced by the study design, data capture and analytical methodology used to derive it. For genotyping, the most widely used resistance tool in clinical practice, these considerations are further complicated by the range of mutational patterns present in the treated population. There is no definitively superior methodology for generating a genotype-response association for use in interpreting a resistance test, and the various approaches used to date all have their strengths and weaknesses. This review discusses the processes involved in constructing such tools, with particular emphasis on establishing validated mutation score rules, and examines the key issues and confounding factors that influence predictive accuracy outside the originating dataset. Since the size of the sample is a key influence on the statistical power to determine an effect, it is hoped that a greater understanding of the influence of study design and methodology will assist the development of standardized outcome measures and reporting formats that allow data pooling at the international level.

Anti-HIV Agents↗

Genome-wide linkage and association mapping of disease genes with the GAW14 simulated datasets.

We combined the results of whole-genome linkage and association analyses to determine which markers were most strongly associated with Kofendrerd Personality Disorder. Using replicate 1 from the Genetic Analysis Workshop 14 Aipotu, Karangar, Danacaa, and New York City simulated populations, we determined that several markers showed significant linkage and association with disease status. We used both SNP and microsatellite markers to determine patterns and chromosomal regions of markers. Three consistently associated markers were C01R0050, C03R0280, and C10R0882. Using generalized linear mixed models, we modelled the effect of the three predefined phenotypic categories on disease status and concluded that the phenotypes defining the "anxiety-related" category best predicted the outcome.

Chromosome Mapping↗

Assessment and integration of publicly available SAGE, cDNA microarray, and oligonucleotide microarray expression data for global coexpression analyses.

Large amounts of gene expression data from several different technologies are becoming available to the scientific community. A common practice is to use these data to calculate global gene coexpression for validation or integration of other "omic" data. To assess the utility of publicly available datasets for this purpose we have analyzed Homo sapiens data from 1202 cDNA microarray experiments, 242 SAGE libraries, and 667 Affymetrix oligonucleotide microarray experiments. The three datasets compared demonstrate significant but low levels of global concordance (rc<0.11). Assessment against Gene Ontology (GO) revealed that all three platforms identify more coexpressed gene pairs with common biological processes than expected by chance. As the Pearson correlation for a gene pair increased it was more likely to be confirmed by GO. The Affymetrix dataset performed best individually with gene pairs of correlation 0.9-1.0 confirmed by GO in 74% of cases. However, in all cases, gene pairs confirmed by multiple platforms were more likely to be confirmed by GO. We show that combining results from different expression platforms increases reliability of coexpression. A comparison with other recently published coexpression studies found similar results in terms of performance against GO but with each method producing distinctly different gene pair lists.

Gene Expression Profiling↗

Statistical validation of the EORTC prognostic model for malignant pleural mesothelioma based on three consecutive phase II trials.

PURPOSE: Malignant pleural mesothelioma (MPM) carries a poor prognosis due to chemoresistance. The European Organisation for Research and Treatment of Cancer (EORTC) prognostic model was reported to predict survival in MPM. Our retrospective analysis set out to test the validity of the model as a prognostic tool in patients treated in three phase II trials at St Bartholomew's Hospital (London, United Kingdom) between 1999 and 2003. PATIENTS AND METHODS: A total of 145 patients were treated in three phase II trials; vinorelbine (VIN; 70 patients), vinorelbine/oxaliplatin (VO; 26 patients), and irinotecan/cisplatin/mitomycin C (IPM; 49 patients). Two subgroups, high-risk and low-risk, were defined by EORTC prognostic score (EPS). EPS was determined by a five-parameter model incorporating age, sex, histology, probability of diagnosis, and leukocyte count. An EPS cutoff of less than 1.27 (low risk) or more than 1.27 (high risk) was used to stratify Kaplan-Meier survival curves. Each of the EPS variables exhibited either trends or significant stratification of overall survival (OS). RESULTS: Multivariate analysis confirmed leukocyte count, Eastern Cooperative Oncology Group performance status, and sarcomatous histology as independent prognostic variables. EPS stratified OS in both individual and pooled trial datasets. No association between objective tumor response and EPS classification was identified by multinomial logistic regression. EPS stratified progression-free survival for the VO and IPM cohorts, but not for VIN. CONCLUSION: This study validates the EPS system as a robust tool for stratifying small trials into low- and high-risk subgroups. EPS should facilitate patient selection and analysis in randomized clinical trials.

Adult↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗