Search PubMedSearch

SEARCH · Search PubMed

Results for “Sample size estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Sample-size estimation: a sensitivity analysis in the context of a clinical trial for treatment of mild hypertension.

The effectiveness of treatment for mild hypertension (diastolic pressures of 85 to 105 mm Hg) has not been conclusively demonstrated. Both the costs of a carefully designed clinical trial and the likelihood that it will produce definitive answers will depend importantly on the sample size. This paper presents sample-size estimates under a variety of assumptions regarding the characteristics of the population to be studied, the degree of blood pressure control to be achieved, and the health benefits to be expected. Under a central set of assumptions, the estimated sample size per group is 22,700 with death as an endpoint and 14,000 with morbid events (CHD and stroke) as endpoints. As individual assumptions are varied one at a time, required sample sizes range from 10,900 to 101,100 and from 6,800 to 63,100 for the respective endpoints. Results are most sensitive to the degree of blood pressure control actually achieved to the expected health benefits from blood pressure control. They are also highly sensitive to the sex composition of the population and to expected dropout rates. The choice of sample size will depend on the decision maker's assessment of the likelihood that each assumption will be fulfilled and on the degree of willingness to risk an inconclusive study result. By making explicit the effect of variation in each assumption, decision making is rendered more susceptible to critical examination by outside reviewers.

Adult

An iterative approach to the analysis of EM autoradiographs. II. Estimates of sample sizes and confidence limits.

The errors inherent in EM autoradiography are discussed and certain of them deemed to be of particular practical significance in the quantitative assessment of preparations. A method is described for estimating the standard errors attributable to each of several sources of variation and thence for obtaining the overall standard error value to be attached to relative activity estimates obtained in the method of Downs & Williams (1978). In an appendix, a fully worked example is given illustrating clearly the strategy of the method and the magnitudes of error estimates that are to be attached to the final specific activity values.

Autoradiography

Effects of different training modalities on lower-limb explosive power, acceleration, 20-m sprint performance, and change-of-direction ability in youth soccer players: a systematic review and network meta-analysis.

BACKGROUND: Youth soccer players repeatedly perform explosive actions, short accelerations, linear sprints, decelerations, and multidirectional movements. However, the comparative effects of different structured physical-conditioning programmes remain uncertain. METHODS: Seven databases were searched from inception to 3 July 2026 using a final expanded search strategy encompassing plyometric, strength or resistance, sprint, acceleration, speed, change-of-direction, neuromuscular, multicomponent, and combined training. Randomised controlled trials involving healthy youth soccer players were eligible. Intervention arms were classified using operational, content-based node definitions. Construct-restricted primary networks and expanded sensitivity networks were analysed using frequentist random-effects network meta-analysis. Hedges' adjusted g was preferentially calculated from post-intervention or final-follow-up means, standard deviations, and sample sizes. Estimates were presented so that positive values indicated better performance. P-scores were treated as descriptive ranking summaries. Risk of bias was assessed using an adapted study-level application of the five-domain RoB 2 framework, and confidence in the evidence was assessed using CINeMA. A post hoc strict-age sensitivity analysis excluded two age-boundary studies. RESULTS: Eighty-nine studies were included in the expanded quantitative analysis, of which 74 contributed to at least one construct-restricted primary network. The primary lower-limb explosive-power, acceleration, 20-m sprint, and planned change-of-direction networks included 55, 20, 25, and 38 studies, respectively. Compared with usual soccer training, plyometric training combined with sprint and/or change-of-direction training showed favourable estimates for lower-limb explosive power (SMD 0.79, 95% CI 0.55 to 1.03), acceleration (1.19, 0.90 to 1.49), 20-m sprint performance (0.80, 0.33 to 1.28), and planned change-of-direction ability (1.46, 1.13 to 1.80). Corresponding I² values were 34.6%, 21.8%, 65.0%, and 41.0%. Between-design inconsistency was detected in the 20-m sprint (P = 0.0036) and change-of-direction (P = 0.0007) networks. CINeMA confidence for these four comparisons was low, low, very low, and low, respectively. Expanded sensitivity networks showed substantially greater heterogeneity. The highest-ranked intervention differed across outcome domains but remained consistent within each outcome across the three analysis sets. Excluding the two age-boundary studies did not materially alter the principal estimates. CONCLUSIONS: Plyometric training combined with sprint and/or planned change-of-direction training produced favourable comparative estimates across the four performance outcomes. However, evidence for several nodes and active-versus-active comparisons was sparse, heterogeneity in programmes and outcomes was present, inconsistency was detected in some networks, and confidence in the evidence was low or very low. These limitations do not support a conclusion that any training category is universally superior. The findings should be interpreted as provisional category-level signals rather than definitive training prescriptions. SYSTEMATIC REVIEW REGISTRATION: PROSPERO CRD420261347297, registered on 21 March 2026, https://www.crd.york.ac.uk/PROSPERO/view/CRD420261347297 .

Change-of-direction ability

Tobacco, nicotine, and cannabis use and exposure in an Australian Indigenous population during pregnancy: A protocol to measure parental and foetal exposure and outcomes.

BACKGROUND: The Australian National Perinatal Data Collection collates all live and stillbirths from States and Territories in Australia. In that database, maternal cigarette smoking is noted twice (smoking <20 weeks gestation; smoking >20 weeks gestation). Cannabis use and other forms of nicotine use, for example vaping and nicotine replacement therapy, are nor reported. The 2021 report shows the rate of smoking for Australian Indigenous mothers was 42% compared with 11% for Australian non-Indigenous mothers. Evidence shows that Indigenous babies exposed to maternal smoking have a higher rate of adverse outcomes compared to non-Indigenous babies exposed to maternal smoking (S1 File). OBJECTIVES: The reasons for the differences in health outcome between Indigenous and non-Indigenous pregnancies exposed to tobacco and nicotine is unknown but will be explored in this project through a number of activities. Firstly, the patterns of parental and household tobacco, nicotine and cannabis use and exposure will be mapped during pregnancy. Secondly, a range of biological samples will be collected to enable the first determination of Australian Indigenous people's nicotine and cannabis metabolism during pregnancy; this assessment will be informed by pharmacogenomic analysis. Thirdly, the pharmacokinetic and pharmacogenomic findings will be considered against maternal, placental, foetal and neonatal outcomes. Lastly, an assessment of population health literacy and risk perception related to tobacco, nicotine and cannabis products peri-pregnancy will be undertaken. METHODS: This is a community-driven, co-designed, prospective, mixed-method observational study with regional Queensland parents expecting an Australian Indigenous baby and their close house-hold contacts during the peri-gestational period. The research utilises a multi-pronged and multi-disciplinary approach to explore interlinked objectives. RESULTS: A sample of 80 mothers expecting an Australian Indigenous baby will be recruited. This sample size will allow estimation of at least 90% sensitivity and specificity for the screening tool which maps the patterns of tobacco and nicotine use and exposure versus urinary cotinine with 95% CI within &#xb1;7% of the point estimate. The sample size required for other aspects of the research is less (pharmacokinetic and genomic n = 50, and the placental aspects n = 40), however from all 80 mothers, all samples will be collected. CONCLUSIONS: Results will be reported using the STROBE guidelines for observational studies. FORWARD: We acknowledge the Traditional Custodians, the Butchulla people, of the lands and waters upon which this research is conducted. We acknowledge their continuing connections to country and pay our respects to Elders past, present and emerging. Notation: In this document, the terms Aboriginal and Torres Strait Islander and Indigenous are used interchangeably for Australia's First Nations People. No disrespect is intended, and we acknowledge the rich cultural diversity of the groups of peoples that are the Traditional Custodians of the land with which they identify and with whom they share a connection and ancestry.

Adult

An application of multivariate ratio methods for the analysis of a longitudinal clinical trial with missing data.

This paper presents an analysis of a longitudinal multi-center clinical trial with missing data. It illustrates the application, the appropriateness, and the limitations of a straightforward ratio estimation procedure for dealing with multivariate situations in which missing data occur at random and with small probability. The parameter estimates are computed via matrix operators such as those used for the generalized least squares analysis of catetorical data. Thus, the estimates may be conveniently analyzed by asymptotic regression methods within the same computer program which computes the estimates, provided that the sample size is sufficiently computer program which computes the estimates, provided that the sample size is sufficiently large.

Clinical Trials as Topic

Assessing data size requirements for training generalizable sequence-based TCR specificity models via pan-allelic MHC-I point-mutation ligandome evaluation.

Rapid identification of T cell receptors (TCRs) that specifically bind patient-unique neoepitopes is a critical challenge for personalized TCR-based therapies in oncology. Due to enormous diversity of both TCR and neoepitope repertoires, a machine learning predictor of TCR-pMHC specificity for personalized therapy must generalize to TCRs and epitopes not seen in the training data. We estimate the necessary size of such training data. We first confirm that published models fail to generalize beyond a single-residue dissimilarity to the epitope training set distribution. We then impute the point-mutation ligandome across the 34 most prevalent human MHC alleles and represent it as a graph based on our established dissimilarity cutoff. By finding the dominating set of this graph, we estimate that between one and 100 million epitopes are required to train a generalizable sequence-based TCR specificity prediction model-1000 times the size of current public data.

Humans

Harnessing Landscape Genomics to Evaluate Genomic Vulnerability and Future Climate Resilience in an East Asia Perennial.

In this era of rapid climate change, understanding the adaptive potential of organisms is imperative for buffering biodiversity loss. Genomic forecasting provides invaluable insights into population vulnerability and adaptive potential under diverse climatic conditions, thereby facilitating management interventions and bolstering shaping species-specific germplasm conservation strategies. We primarily employed landscape genomics approaches, leveraging single-nucleotide polymorphisms obtained through whole-genome resequencing of 201 individuals across 43 Rheum palmatum complex populations, to pinpoint adaptive variation and its significance in the context of future climates, delineate seed zones, and establish guidelines for ex situ germplasm conservation. The species complex exhibited strong signatures of local adaptation and differential genomic vulnerabilities across its distribution range, with eastern lineage populations facing significant maladaptation risks under future climate scenarios. Using diverse datasets of putatively adaptive loci and climate change scenarios, we delineated three distinct seed zones within the species' range, estimated varying sample sizes per zone to capture most adaptive diversity, and predicted shifts in seed zone centroids ranging from 48.3 to 359.3&#x2009;km from historical distributions to mitigate climate change impacts. Collectively, our findings underscore the importance of integrating genomic and environmental data to forecast the adaptive trajectory of an East Asian perennial under anticipated climate changes, guide seed zone delineation for germplasm conservation and enhance population resilience. These results provide a blueprint for designing targeted conservation strategies and restoration plans in other imperilled species.

Climate Change

Inavolisib for PIK3CA-mutated advanced endometrial cancer: a multicentric, phase II, MITO END-4 trial.

BACKGROUND: The phosphatase and tensin homolog-phosphoinositide 3-kinase (PI3K)-protein kinase B (AKT) pathway is frequently altered in gynecological tumors, notably in endometrial cancer where PIK3CA mutations are found in nearly half of patients. Despite this, evidence of clinical activity of PI3K inhibitors in endometrial cancer is poor and limited. Alpelisib, an oral PI3K alpha-selective inhibitor, showed encouraging preliminary activity in advanced gynecological tumors harboring PIK3CA alterations. Inavolisib is a highly potent and selective PI3K inhibitor. PRIMARY OBJECTIVE(S): The MITO END-4 trial aims to assess the efficacy and safety of inavolisib in patients with endometrial cancer who have received platinum-based chemotherapy and immunotherapy. The primary objective is to determine the anti-tumor activity (assessed by objective response rate) of inavolisib in patients with advanced endometrial cancer with PIK3CA mutated tumors. STUDY HYPOTHESIS: The study tests the hypothesis that inavolisib has superior anti-tumor activity compared to historically available standard therapies in previously treated patients with advanced endometrial cancer harboring a PIK3CA mutation. TRIAL DESIGN: This is a phase II, single-arm, multicenter trial in which advanced endometrial cancer patients whose tumors harbor a pathogenic PIK3CA mutation will receive inavolisib. MAJOR INCLUSION/EXCLUSION CRITERIA: Patients aged 18 years and older with documented evidence of PIK3CA mutated advanced endometrial cancer (endometrioid, serous, clear cell, carcinosarcoma or mixed histology) will be enrolled. Patients have previously received at least 1 platinum-based chemotherapy in any setting (adjuvant or advanced) with or without immune checkpoint inhibitor, alone or in combination. Not more than 4 lines of therapy are allowed. Key exclusion criteria include uterine sarcoma and prior treatment with any PI3K, AKT, or mechanistic target of rapamycin (mTOR) inhibitor. PRIMARY ENDPOINT(S): Objective response rate defined as a complete response or partial response by the Investigator using RECIST v1.1 criteria over the whole treatment period. SAMPLE SIZE: 48 patients. ESTIMATED DATES FOR COMPLETING ACCRUAL: May 2028. TRIAL REGISTRATION: MITO END-4; EU-CT NUMBER: 2025-522981-61-00; NCT07522697.

Endometrial cancer

APAV: An advanced pangenome analysis and visualization toolkit.

Traditional pangenome analysis focuses on gene presence/absence variations (gene PAVs). However, the current methods for gene PAV analysis are insensitive to detect small but valuable mutations within gene regions, and they overlook variations in intergenic regions. Additionally, the visual inspection of PAVs is an important but time-consuming step for pangenome analysis and result interpretation. To address these issues, we present APAV, an advanced toolkit designed for comprehensive PAV analysis and visualization. It integrates gene element-level PAV analysis and provides PAV analysis for arbitrary given regions in a genome. The resulted PAV profile can be visualized and investigated interactively with reports in HTML format, enabling researchers to conveniently verify sequencing read depth, target region coverage, and intervals of absence for each PAV. Furthermore, APAV offers various subsequent analysis and visualization functions based on the PAV profile table, including basic statistics, sample clustering, genome size estimation, and phenotype association analysis. We demonstrated the capability of APAV with pangenome analysis of tumor genomes and rice genomes. Performing PAV analysis at the element level not only provides more accurate information about the variations but also uncovers a larger number of variations for the phenotype-genotype association studies. In the rice genome analysis, we identified over twenty thousand distributed genes and more than fifty thousand distributed genetic elements. In the tumor genome analysis, element-level analysis revealed approximately three times as many phenotype-related genes as gene-level analysis. This indicates that altering the PAV unit from genes to smaller segments or elements can lead to more biological insights.

Software

A comparison of iron bioassay diets.

Two broiler chick experiments wer conducted to evaluate four basal diets for iron bioassay suitability. The test basal diets, identified according to their principal ingredients and iron content, were: (1) starch-skim milk - 15 p.p.m., (2) degerminated corn-skin milk - 18 p.p.m. (3) degermianted corn-fish meal-isolated soy - 45 p.p.m., and (4) degermianted corn-fish meal-dehulled soy - 59p.p.m. Significant differences between an iron source with known low availability (ferric oxide) and a highly available iron source (ferrous sulfate) were not detected with the degerminated corn-fish meal-isolated soy or the degerminated corn-fish meal-dehulled soy diets. Likewise, there were no significant differences found between supplemental iron levels, 10 and 20 p.p.m. The corn-skim milk and starch-skim milk diets were both found to be satisfactory for iron bioassays. However, the sample size needed to estimate the population mean was almost twice as great for the starch-skim milk fed groups, than was needed for the corn-skim milk fed groups which indicates the corn-skim milk diet obtained greater sensitivity in testing iron sources and levels. Mortality was excessively high in the starch-skim milk fed group. Ferrous sulfate was superior to ferric oxide as a source of iron.

Animals

Sample size needed for student ratings of instruction.

The number of evaluation forms students are asked to complete is multiplying. To reduce that number, the present study determines the minimum sample size needed for accurate student ratings of instruction. Typical questionnaire items using four- and seven-category ratings scales were studied. Data for four class sizes (40, 60, 100, 140) were sampled in graduated sizes, and a standard error of the mean was computed for each sample size. A permissible error index was computed to estimate the accuracy of ratings obtained from any sample size needed for the four different class sizes. Figures are presented from which minimum sample sizes necessary for accurate student evaluation of instruction can be computed. The figures show that sampling only one-third of classes of 100-140 students is sufficient to obtain accurate evaluations.

Evaluation Studies as Topic

dGAMLSS: an exact, distributed algorithm to fit Generalized Additive Models for Location, Scale, and Shape for privacy-preserving population reference charts.

MOTIVATION: There is growing interest in estimating population reference ranges across age and sex to better identify atypical clinically-relevant measurements throughout the lifespan. For this task, the World Health Organization recommends using Generalized Additive Models for Location, Scale, and Shape (GAMLSS), which can model non-linear growth trajectories under complex distributions that address the heterogeneity in human populations.Fitting GAMLSS models requires large, generalizable sample sizes, especially for accurate estimation of extreme quantiles, but obtaining such multi-site data can be challenging due to privacy concerns and practical considerations. In settings where patient data cannot be shared, privacy-preserving distributed algorithms for federated learning can be used, but no such algorithm exists for GAMLSS. RESULTS: We propose distributed GAMLSS (dGAMLSS), a distributed algorithm that can fit GAMLSS models across multiple sites without sharing patient-level data. This includes specific considerations for the fitting of smooth functions at varying levels of communication efficiency. We demonstrate the effectiveness of dGAMLSS in constructing population reference charts across clinical, genomics, and neuroimaging settings and show that dGAMLSS is able to reproduce pooled reference charts and inference down to numerical differences. AVAILABILITY AND IMPLEMENTATION: An R package providing examples of the dGAMLSS algorithm, as well as functions for sharing and aggregating site-specific parameters, is available at https://github.com/hufengling/dGAMLSS.

Algorithms

Determination of significant relative risks and optimal sampling procedures in prospective and retrospective comparative studies of various sizes.

Methods are given for determining the relative risks which it is possible to demonstrate as statistically significant with given probability in prospective and retrospective studies of a particular size. Also considered is the ratio of the sample sizes in the two groups being compared which will provide the most precise estimate of relative risk for a given total sample size.

Biometry

Reassessing Instrument Strength in Two-Sample Mendelian Randomization Analysis.

Mendelian randomization (MR) analysis is widely used to estimate causal relationships between risk factors and outcomes of interest. Two-sample MR approaches have gained increasing attention in genetic epidemiology due to the growing availability of Genome-Wide Association Study (GWAS) summary statistics from public databases. A critical step in two-sample MR is the selection of genetic variants as instrumental variables (IVs). Although genome-wide significant variants are typically preferred, the inclusion of variants with weaker association p-values is considered, as they may potentially improve power through an increased instrument number of instruments, while they may introduce weak instrument bias and attenuate effect estimates towards the null. Our simulation results show that even modest levels of pleiotropy substantially increase the variability of causal effect estimates, while the inclusion of weak IVs does not substantially affect the direction and variability of causal effect estimates in most cases. In real data analyses, we used two released versions of FinnGen GWAS summary statistics with different sample sizes as exposure GWASs to assess the influence of weak IVs. Here, the inclusion of IVs with higher exposure-association p-values resulted in weakened estimated effect sizes, particularly when the exposure GWAS sample size was small. These findings suggest that incorporating weak IVs is reasonable when the exposure GWAS sample size is large, but it poses a risk of falsely concluding null associations when the exposure GWAS sample size is small.

Journal Article

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results

[CK MB versus total CK in the estimation of myocardial necrosis. Comparative kinetic analysis and evaluation of an immunological method].

It has been suggested that the adoption of a relatively specific marker of the myocardial cell, such as creatine kinase MB isoenzyme, can yield improved accuracy in estimating infarct size by serial serum sampling and compartmental analysis. Nevertheless, current methods for the evaluation of isoenzyme activity are cumbersome and unsuitable for clinical use. We have therefore employed a new test for the rapid determination of CK MB activity, based on the immunological inhibition of M subunities. In 19 patients not submitted either to intramuscular injection or to repeated defibrillations, a good correlation was found between indexes of necrosis based on MB and total CK determination (r = 0.94), with the cumulative MB release amounting to 16 +/- 4% of total CK. Significant differences were observed in 3 patients submitted to external cardiac massage (MB = 9 +/- 1% of total CK) thus suggesting a considerable extracardiac source of total CK due to the trauma of the skeletal muscle. The comparative kinetic analysis shows substantial differences between the two isoenzymes, not only concerning the greater disappearance rate of CK MB but, more significantly, related to a faster release of this isoenzyme from the myocardium, which has not been previously reported. The good correlations found between maximal appearance rate and cumulative enzyme release (r = 0.86) suggest that the former may represent an index of the rate of degradation of cellular membranes. Practical implications of these data are discussed.

Creatine Kinase

The use of regression constants in estimating tooth size in a Negro population.

A study of a sample of 105 Negro children and adolescents, residents of Connecticut, was undertaken to determine the degree of correlation between mandibular tooth size and the size of the canines and premolars. The correlation between the total mesiodistal width of the mandibular permanent incisors and that of the maxillary or mandibular canine and first and second premolars was found to be 0.63 and 0.71, respectively. Further, regression constants were determined in an attempt to estimate the buccal segments from the mandibular incisors.

Adolescent