Search PubMedSearch

SEARCH · Search PubMed

Results for “Datasets as Topic”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Genetic analysis of IDDM: the GAW5 multiplex family dataset.

In a collaborative effort by 12 centers from Europe and North America, data were assembled from 94 multiplex families with insulin-dependent diabetes mellitus (IDDM) for analysis of genetic and other factors of possible etiological importance. The dataset contains information on the following genetic markers: HLA-DR beta and -DQ beta restriction fragment length polymorphisms (RFLPs), three RFLPs detected with two probes that map 5' to the insulin gene, the serologically defined HLA loci, and the immunoglobulin allotypes. Data also were included for auto-antibodies to insulin and pancreatic islet cells as possible indicators of pathogenesis and for antibodies to certain viruses that have been implicated as "triggering" agents in IDDM. Medical history of family members was obtained by means of a uniform questionnaire. Identical copies of the dataset were distributed to anyone wishing to participate in the analysis for the IDDM component of GAW5. The multiplex IDDM family dataset is now available on request for further analysis.

Adolescent

A Guide for Exploring Pleiotropic Associations in Genome-Wide Association Studies Using Summary Statistics.

Genome-wide association studies (GWAS) have shown that pleiotropy, whereby a single genetic variant or gene influences multiple traits, is common in complex human diseases. Detecting cross-phenotype associations from GWAS summary statistics remains challenging because of small effect sizes, extensive multiple testing, heterogeneous effects, and possible differences in effect direction across traits. Methods that jointly analyze multiple traits can improve the ability to detect pleiotropic signals while retaining the practical advantages of summary statistic-based analyses. Although a range of statistical approaches has been developed for this purpose, practical guidance on their application, assumptions, and interpretation remains limited. This tutorial reviews several widely used methods for pleiotropy detection from GWAS summary statistics, including ASSET, PLACO, GPA, CPBayes, and GCPBayes, and demonstrates their application using breast and thyroid cancer datasets. We also highlight the importance of accounting for effect heterogeneity, correlation, and biological group structure at the gene and pathway levels in the detection and interpretation of pleiotropic association signals.

Genome-Wide Association Study

A note on a generalized single step theory for any number of hierarchical genomic matrices.

BACKGROUND: The Single Step algorithm allows combining information from genotyped and un-genotyped individuals, provided they are connected by a pedigree. However, current single step theory is limited to a single list of markers. RESULTS: We present a generalized single step (GSS) method that can accommodate any number of hierarchical molecular datasets (e.g. sequence, high and low density arrays) and pedigree, avoiding imputation. We prove that a similar efficient inversion algorithm exists. The method is recursive, starting with the highest marker density scenario. We illustrate the method with simulation and show that GSS can increase predictive accuracy compared to standard single step. R code is provided so that custom scenarios can be easily compared, either with simulated or real data. CONCLUSION: The method developed generalizes extant single step theory to any number of hierarchical molecular relationship matrices, broadening the scenarios where single step can be applied. A topic of particular interest can be ecology field data or human populations where pedigree is not available, but where samples sequenced and genotyped at different densities can exist. GSS can also be a useful tool to optimize allocation of genotyping and / or sequencing resources.

Algorithms

Clinical trials in psychiatry: should protocol deviation censor patient data?

Clinical trials methodologists recommend counting all events regardless of adherence to protocol and comparing originally randomized groups. This strict "intent-to-treat" policy implies availability and use of outcome measures taken regardless of adherence to treatment protocol. However, outcome measurement in psychiatry requires the cooperation of the patient, and usually occurs in the context of treatment management. Consequently, the patient's or clinician's decision not to adhere to the treatment protocol may be design or default cause censorship of patient data by early truncation. This disables the analysis "by intent to treat" in the strict sense. Current methods applied to such nonrandomly truncated datasets are unsatisfactory ("last value" analysis, survival analysis) or worse (imputation by last value carried forward). I review the context of clinical experimentation in psychiatry, contrast the state of design and analysis with expert recommendations on general methods, review the current statistical strategies and propose that investigators should try to obtain complete follow-up data on all patients without regard to their adherence to treatment protocol.

Clinical Trials as Topic

The hidden effect of time.

It is customary to regard datasets as homogeneous with respect to the order of collection of the measurements. Examples are given in which this assumption is breached. Hidden time trends have implications for the design of studies, their analysis and interpretation. It is suggested that, if the order of observations is known, a plot by time should be performed, perhaps using a cusum.

Aged

The Evolutionary Significance of Leaf Nodulation: Evidence from Ardisia and Its Relatives (Primulaceae: Myrsinoideae).

Interactions between plants and microorganisms have long been a central topic in biological research. Bacterial symbiosis on leaf surfaces represents a distinctive and mutually beneficial system within the phyllosphere microbiome. Leaf nodules are the visible manifestation of the symbiosis and confer ecological advantages to host plants by enhancing host resistance against pathogens and herbivores. It has been hypothesized that these advantages promote higher diversification rates in host lineages, but this remains uncertain. Ardisia subg. Crispardisia and its close relatives (Amblyanthopsis and Amblyanthus) within Primulaceae are typical plant groups with leaf nodule symbiosis, making them an ideal system for testing this hypothesis. In this study, we conducted extensive sampling of "Ardisioids" (Ardisia and its allies) and reconstructed their phylogenetic relationships and evolutionary history using plastid genomes and nuclear datasets (i.e., nuclear ribosomal DNA (nrDNA) and genome-wide single nucleotide polymorphisms (SNPs)). We clarified the phylogenetic positions of several "Ardisioids" genera (e.g., Sadiria, Tapeinosperma, Amblyanthus, and Amblyanthopsis) and multiple subgenera within Ardisia. We further detected a rapid radiation during the middle Miocene in Ardisia and its allies. Notably, we found that the leaf-nodulated clade appears to have originated during this period, approximately 11-8 Ma. BAMM (Bayesian Analysis of Macroevolutionary Mixtures) analyses revealed elevated diversification rates in leaf-nodulated lineages, while HiSSE (Hidden State Speciation and Extinction) analyses indicated that leaf nodule symbiosis might have increased speciation rates without significantly affecting extinction rates. These results provide strong evidence that leaf nodule symbiosis, together with other abiotic and biotic factors, represents a key evolutionary innovation that has promoted diversification in Ardisia and its close relatives.

diversification rate

Artificial intelligence for dental caries detection: An umbrella review.

Artificial intelligence (AI) has been proposed as a tool to improve dental caries detection across imaging modalities; however, its clinical value remains uncertain. This umbrella review aimed to synthesize and critically appraise systematic reviews evaluating AI for caries detection and diagnosis. An umbrella review was conducted following PRIOR guidance (PROSPERO CRD420261340728). Searches were performed in MEDLINE, Embase, Scopus, Web of Science, and Google Scholar up to 15 March 2026. Methodological quality was assessed using AMSTAR 2, and overlap of primary studies was quantified using the corrected covered area (CCA). Seventeen systematic reviews were included, of which five reported diagnostic test accuracy meta-analyses using bivariate or HSROC models. Across these meta-analyses, pooled sensitivity ranged from 0.76 to 0.94 and specificity from 0.85 to 0.91. Most systems were based on deep learning models applied to bitewing radiographs and intraoral photographs. However, substantial heterogeneity was observed in imaging modalities, lesion thresholds, analytical tasks, and evaluation metrics. In addition, a high degree of overlap across reviews and recurrent methodological limitations, including reliance on retrospective datasets, limited external validation, and inconsistent reporting, substantially weaken the reliability of the evidence. Although AI models demonstrate high diagnostic performance under experimental conditions, current evidence does not support their use as stand-alone diagnostic tools. Their clinical applicability remains limited, and implementation should be restricted to decision-support contexts until robust prospective validation demonstrates meaningful impact on clinical decision-making and patient outcomes.

Dental Caries

Optimum numerical integration methods for estimation of area-under-the-curve (AUC) and area-under-the-moment-curve (AUMC).

Eleven numerical methods for estimation of AUC (including 4 new methods) and 22 methods for AUMC (including 8 new methods) were tested on large simulated noisy datasets representing bolus, oral and infusion concentration-time profiles. Some methods were unacceptable because their mean error was large; these included a commonly recommended form of the linear trapezoidal rule for AUMC. Others, notably Lagrange and cubic spline methods, were unacceptable because the variance of their estimates was large. These methods should be abandoned. A simple and easily programmed new method, parabolas-through-the-origin then log-trapezoidal rule, performed especially well.

Numerical Analysis, Computer-Assisted

Epigenetic Profiling for Early Detection and Treatment Response Monitoring in Non-Small Cell Lung Cancer: Protocol for a Prospective Translational Biomarker Study.

BACKGROUND: Non-small cell lung cancer (NSCLC) is the leading cause of cancer-related mortality worldwide and continues to have poor survival outcomes, with most patients diagnosed at advanced stages of disease. In New Zealand, NSCLC contributes substantially to cancer inequities, with Māori communities experiencing disproportionately high incidence and mortality rates. Although low-dose computed tomography screening can improve early detection, major limitations remain, including false-positive findings, overdiagnosis, high infrastructure costs, and limited accessibility for rural and underserved populations. Liquid biopsy approaches using circulating tumor DNA (ctDNA), particularly DNA methylation profiling, have emerged as promising, minimally invasive strategies for improving cancer detection, treatment monitoring, and precision oncology. OBJECTIVE: This study aims to establish integrated genomic and epigenomic predictive and prognostic biomarkers using ctDNA, tumor tissue, and transcriptomic profiling to improve early detection, risk stratification, treatment selection and response prediction, and longitudinal monitoring, with particular emphasis on identifying molecular mechanisms associated with treatment resistance and disease progression. METHODS: This prospective observational translational biomarker study is being conducted through the University of Otago and associated respiratory and oncology services in New Zealand. The study will recruit participants with NSCLC (including squamous and nonsquamous subtypes), individuals referred to fast-track lung nodule assessment clinics, and nonmalignant respiratory controls. Serial peripheral blood sampling will be performed in selected participants at predefined clinical follow-up time points to evaluate treatment response and disease progression. The availability of formalin-fixed paraffin-embedded archival tissues will be recorded, but will not be mandatory for enrollment. Genome-scale DNA methylation profiling will be performed using cell-free reduced representation bisulfite sequencing (cfRRBS), while targeted genomic profiling and transcriptomic analyses will be conducted using targeted sequencing panels and RNA sequencing. Integrative bioinformatic analyses will be used to identify molecular biomarkers associated with early-stage disease, advanced disease, treatment response, and therapeutic resistance. RESULTS: Ethics approval for the study has been obtained from the New Zealand Health and Disability Ethics Committee (2022 EXP 12566). This study commenced in 2022, and recruitment and biospecimen collection are ongoing. The study aims to recruit approximately 450 participants, including patients with NSCLC, individuals referred through respiratory diagnostic pathways, and nonmalignant controls. As of July 31, 2026, 205 participants have been recruited, with recruitment continuing until the target sample size is reached. Molecular and data analyses are ongoing, with additional publications expected as the cohort matures. CONCLUSIONS: This study will generate one of the first integrated genomic, epigenomic, and transcriptomic liquid biopsy datasets for NSCLC in New Zealand. The findings are expected to support the development of sensitive, accessible, and equitable blood-based biomarkers for NSCLC detection and treatment monitoring while also contributing to improved precision oncology approaches and reducing NSCLC inequities among Māori populations.

Humans

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models

Incubation time for AIDS from French transfusion-associated cases.

Although incubation time is a key parameter of the epidemiology of AIDS, statistical estimates based on transfusion-associated AIDS cases have, up to now, used only the single dataset provided by the AIDS program of the Centers for Disease Control (CDC) in Atlanta. Using a new dataset provided by the Direction Générale de la Santé (DGS), of the French Ministry of Health1, we estimate the mean incubation time for AIDS (median in brackets) to be 5.3 years (5.3 years) with a 90% confidence interval ranging from 4.4 to 8.9 years (4.4 to 8.8 years), when a Weibull distribution is postulated for incubation time. The previously encountered problem of very large confidence intervals (range larger than 100 years), is not observed, indicating that an accurate estimate for mean incubation time will be obtainable in the near future.

Acquired Immunodeficiency Syndrome

The classification of subjects with joint complaints on incomplete biochemical and haematological datasets.

We performed a retrospective study on 163 subjects suffering from rheumatic fever (16), rheumatoid arthritis (36), lupus erythematosus (17), gout (21), arthrosis (50) and osteomyelitis (23). The number of variables evaluated was 39. These were all of a general biochemical and haematological nature. A feature reduction resulted in sixteen variables that matched well with those known from the literature. Linear discriminant analysis yielded poor results in classifying the six disease categories (with 18 variables 61.8%). A reduction to three disease categories improved the classification results remarkably. This, and the excellent discriminating power between patients and the reference group, shows that the selected variables are illustrative only for general clinical pictures, such as infection, and not for the desired differential diagnosis.

Arthritis, Rheumatoid

Diverse prognosis in metastatic breast cancer: who should be offered alternative initial therapies?

In an attempt to clarify appropriate treatment options for women with stage IV breast cancer, we studied the survival experience of a large dataset of patients treated on Cancer and Leukemia Group B (CALGB) protocols. The study, restricted to women who had had no prior chemotherapy for metastatic disease, demonstrated a surprisingly poor prognosis, with an estimated median survival of 1.6 years and only 26% alive at 3 years. Analysis of prognostic factors permitted the identification of subsets with even shorter survival, such as women with estrogen receptor negative tumor in more than one metastatic site and prior adjuvant chemotherapy. We feel that an evaluation of intensive investigational treatment approaches, such as trials using autologous bone marrow transplantation, is justified for most stage IV breast cancer patients, in view of their poor prognosis.

Biomarkers, Tumor

Scientific and cost-effectiveness criteria in selecting batteries of short-term tests.

The scientific and cost-effectiveness criteria introduced in this paper can be applied to published datasets and current and proposed batteries of short-term tests. The reports in the current volume will provide a wealth of additional material for such evaluations, but more systematically obtained information will be necessary to assess both the internal and external validity of these tests. Individual tests and batteries of tests should be standardized, employ positive controls, generate results capable of quantitative analyses that may make dichotomous classification as "positive" and "negative" obsolete, be interpreted in light of mechanisms of action, and be cost-effective on a grand scale. For regulatory purposes our long-term goal should be to replace the whole animal lifetime bioassay with an appropriate and cost-effective set of short-term tests.

Animals

Distributed data analysis in a multicenter study: the CARDIA Study.

Unlike distributed data entry, which is used in many large epidemiologic studies and multicenter clinical trials, distributed data analysis is a relatively new concept. This paper reports on the usefulness of such a system in the Coronary Artery Risk Development in Young Adults (CARDIA) Study. CARDIA distributes the entire examination dataset to participating centers soon after completion of each round of data collection. The process was designed to encourage more numerous, diverse, and rapid publications, and to allow for more efficient use of the manpower and expertise in centers. Responsibilities of the coordinating center have changed from a conventional coordinating center but remain substantial due to the need for collating, monitoring, verifying, and documenting the distributed data analysis (DDA) system. DDA is successful from the standpoint of implementation and operation--21 manuscripts representing work analyzed at six participating centers had been submitted for publication within 3.5 years of the completion of the baseline examination.

Adolescent

A comparison of infection control software for use by hospital epidemiologists in meeting the new JCAHO standards.

To choose a microcomputer software package for our hospital epidemiology division, the two leading commercial software packages for infection control, AICE (ICPA, Inc., Austin, Texas) and NOS0-3 (Epi Systematics, Inc., Ft. Meyers, Florida), were compared for the types of epidemiologic analysis likely to be required to satisfy new Joint Commission on Accreditation of Healthcare Organizations (JCAHO) 1990 Infection Control Standards. The test dataset was a surgical database of 3,235 operations with 292 (9%) wound infections. Though NOSO-3 was more flexible in terms of the amount of data items one could record, it required seven times longer to learn, nine times more disk space to store and two times as long to enter cases than AICE. Six simple infection control reports (i.e., line listings, crosstabulations, stratified rates and graphs) required only seven computing steps and approximately 11 minutes to process with AICE, but 22 steps and over two hours with NOSO-3. All analytic results from AICE agreed with the results obtained with the Statistical Analysis System (SAS, SAS Institute, Inc., Cary, North Carolina), but analyses such as service-specific rates performed with NOSO-3 differed because of a design flaw in the NOSO-3 data structure.

Cross Infection

Evaluation of a 3D reconstruction algorithm for multi-slice PET scanners.

A fully 3D reconstruction algorithm based on filtered backprojection was evaluated for the reconstruction of data obtained with multi-slice positron emission tomography (PET) scanners which have had the septa removed. This algorithm uses forward-projection through the reconstructed images of a 2D subset of the data to complete the 3D dataset thus satisfying the condition of shift invariance. This is followed by 3D filtered backprojection. Axial sampling was doubled by combining adjacent polar angles, thus improving reconstructed axial resolution. The algorithm was tested using real and simulated datasets and gave high quality reconstructions without artifacts over a wide range of imaging conditions. Events are placed accurately throughout the imaging volume as determined by measurements with a MRI/PET registration phantom. The forward-projection step leads to degradation in image resolution due to insufficient axial and transaxial sampling. This effect is amplified if multiple iterations of the algorithm are used, with little decrease in image noise. Changing the filter employed in the initial 2D reconstruction can be used to alter the noise and resolution characteristics of the 3D images. This algorithm has proved very robust at reconstructing 3D PET data and is relatively fast. Those small problems which exist can be attributed to detector sampling problems, especially in the axial direction, which is a consequence of the geometry of these scanners, which are designed primarily for 2D data acquisition.

Algorithms

Evaluation of four in vitro genetic toxicity tests for predicting rodent carcinogenicity: confirmation of earlier results with 41 additional chemicals.

The effectiveness of four in vitro short-term tests (STT) for genetic toxicity, induction of mutations in Salmonella (SAL) and mouse lymphoma L5178Y cells (MLA), and induction of sister chromatid exchanges (SCE) and chromosome aberrations (ABS) in Chinese hamster ovary cells that are used for predicting rodent carcinogenicity were examined. The in vitro results were compared with the results from 41 rodent carcinogenicity studies performed by the National Toxicology Program. The predictive values of, and interrelationships among, the STT for these 41 chemicals were similar to those previously reported for 73 chemicals and confirm those earlier results [Tennant RW, Margolin BH, Shelby MD, Zeiger E, Haseman JK, Spalding J, Caspary W, Resnick M, Stasiewicz S, Anderson B, Minor R (1987): Science 236:933-941]. Because of this similarity among the two datasets, the chemicals were combined into a single dataset of 114. The results with 114 chemicals show that SAL had the lowest sensitivity (.48) and the highest specificity (.91), whereas MLA had the highest sensitivity (.72) and the lowest specificity (.40). The concordances of the test results with rodent carcinogenicity were .66, .61, .59, and .59, for SAL, ABS, SCE, and MLA, respectively. Salmonella was the most predictive for carcinogenicity; 89% of the chemicals mutagenic in SAL were carcinogenic in rodents, however a negative result in any or all of the STT was not indicative of noncarcinogenicity. The STT results reported here show good agreement with the potential electrophilicity of the chemicals, and the majority of carcinogens that are undetected by the STT do not have an electrophilic structure. There was no complementarity among the tests and no combination of the four tests was more effective than any single test for predicting carcinogenicity.

Animals