Search PubMedSearch

SEARCH · Search PubMed

Results for “statistical genetics”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Polygenic risk scores associate with asthma phenotypes and proteomic analyses implicate IL1R1 in two family-based studies.

Despite its high prevalence and the discovery of hundreds of genetic associations, the genetic determinants and heterogeneous manifestations of asthma remain incompletely understood. Incorporating polygenic risk scores (PRS) into asthma research offers a powerful approach to quantify inherited susceptibility, refine risk profiles, and advance mechanistic understanding of disease development. For this study, we leveraged whole-genome sequencing (WGS) data from two family-based cohorts of childhood asthma - the Genetics of Asthma in Costa Rica Study (GACRS) and the Childhood Asthma Management Program (CAMP) - to examine the transmission profiles of externally derived asthma PRS and their associations with clinical phenotypes in children with asthma. To further elucidate molecular mechanisms, we integrated large-scale external genome-wide association study (GWAS) summary statistics and genetic prediction models of protein abundance in a two-step proteome-wide association study (PWAS) of asthma. Our findings provide robust evidence supporting the validity of externally derived asthma PRS (asthma PRS association p-value p = 10-24 [GACRS and CAMP trios combined] for the Global Biobank Meta-analysis Initiative [GBMI]) and reveal consistent associations with spirometry measures and atopy markers across both studies, as 13 of 21 traits (62%) were significantly associated with the GBMI-PRS in the meta-analysis after multiple-testing correction. Moreover, the results of the integrative proteomic analysis implicate IL-1 signaling in the etiology of asthma, reinforcing the candidacy of IL1R1 antagonists for drug repurposing.

Journal Article

Extracting and calibrating evidence of variant pathogenicity from population biobank data.

Genomic medicine requires a robust evidence base of variant phenotypic impacts, which remains incomplete even in extensively studied genes with monogenic disease associations. Here, we evaluated the broad potential of using population cohort data to identify evidence that can be used in variant assessment. Across 41 genes related to 18 clinically actionable monogenic phenotypes, we calculated variant-level odds ratios of disease enrichment using data from 469,803 UK Biobank participants. We found significant differences in odds ratio values between ClinVar-labeled pathogenic and benign variants in 11 phenotypes, spanning both common and rare disorders. To facilitate clinical translation, we calibrated the strength of evidence provided by variant-level odds ratios to align with American College of Medical Genetics and Genomics and the Association for Molecular Pathology (ACMG/AMP) interpretation guidelines (PS4 criterion) and found that odds ratios may reach "moderate," "strong," or "very strong" evidence, varying by phenotype and gene. Overall, we found that 2.6% (N = 12,350) of participants harbor a rare variant of uncertain significance (VUS) with at least moderate evidence of pathogenicity-an indication of potentially unrecognized disease risk. Finally, by incorporating computational and functional data alongside population-based odds ratios, we identified variants that met the criteria for clinical reclassification. Notably, using this approach, we identified that 12.4% of rare VUSs in LDLR seen in participants meet diagnostic criteria to be classified as likely pathogenic, demonstrating its potential to scale the reclassification of VUSs.

Humans

Genetic evidence and cross-species functional characterization implicate CNN2 in age-related macular degeneration susceptibility.

Age-related macular degeneration (AMD) is a leading cause of irreversible visual impairment in the aging population globally. Although genome-wide association studies (GWAS) have identified many AMD susceptibility loci, the genes and mechanisms underlying many of these associations remain unresolved. Here, we integrated expression quantitative trait locus (eQTL) data with AMD GWAS to prioritize nine putative genes. Through in vivo screening in zebrafish, we demonstrated that the downregulation of cnn2 and sarm1 expression led to ocular structural abnormalities and visual functional impairment. Subsequent mouse model studies confirmed that Cnn2 deficiency affected photoreceptor structure and function, impaired contrast sensitivity, and caused abnormalities in cone cell immunostaining. Given that CNN2 is predominantly expressed in endothelial cells, we propose that endothelial dysfunction may cascade to impair photoreceptor function. Collectively, through in silico prioritization and cross-species functional characterization, we identify CNN2 as a candidate susceptibility gene in AMD pathogenesis, providing vital underlying mechanistic insights.

Animals

Statistical design of toxicity assays: role of genetic structure of test animal population.

This paper considers certain statistical aspects of the problem of among-strain differences in cancer susceptibilities and how these differences may affect the design of toxicity assays. First, in order to investigate the magnitude of within-study, between-strain differences in tumor induction, the data of Innes et al. (1969) were examined. It was found that although there was a very high overall association between mouse strains with respect to the induction of hepatomas, several compounds showed evidence of strain-to-strain variability. Next, a number of long-term carcinogenicity studies with DDT were considered, and among-strain differences in cancer susceptibility for this compound were noted. Finally, it was shown that if susceptible subgroups do exist and certain simplifying assumptions are made, then in many cases tumor increases can be detected more readily by using several inbred mouse strains for study rather than a single outbread stock.

Animals

TL-HDMR: a transfer learning framework for advancing equitable causal inference reveals metabolic signatures of stroke across multiple ancestries.

The limited genetic diversity in genome-wide association studies (GWAS) poses a significant challenge to the generalizability and equity of biomedical discoveries. Most causal inferences, particularly from high-dimensional phenomes (e.g. metabolomics), are primarily based on European populations, and their applicability to other ancestries remains uncertain. Traditional multivariable Mendelian randomization (MVMR) methods further struggle in high-dimensional and correlated settings due to collinearity and model instability. To bridge this gap, we present a two-step transfer learning framework for high-dimensional MR (TL-HDMR), designed to enhance causal exposure detection in understudied populations. Our approach leverages the Minimax Concave Penalty for asymptotically unbiased estimation amidst exposure correlations. Crucially, we introduce two novel pre-transfer procedures-HDMR.TSD for sourcing beneficial data and HDMR.PRESSO for filtering pleiotropic instruments-to ensure robust knowledge transfer. Extensive simulations demonstrated TL-HDMR's superior performance in ROC curves and mean absolute error over alternative methods. When applied to identify causal metabolites for stroke across multi-ancestry cohorts (European, East Asian, South Asian, and African), TL-HDMR successfully pinpointed both shared and ethnic-specific causal biomarkers, showcasing its unique capability for equitable causal inference. This work provides a powerful statistical tool that not only addresses critical methodological challenges but also promotes inclusivity and fairness in human health research.

Humans

Statistical design of toxicity assays: role of genetic structure of test animal population.

This paper concerns certain statistical aspects of the problem of among-strain differences in cancer susceptibility and how these differences may affect the design of toxicity assays. First, the data of Innes et al. (1969) were examined to investigate the magnitude of within-study, between-strain differences in tumor induction. Although there was a very high overall association between mouse strains with respect to the induction of hepatomas, evidence of strain-to-strain variability was found for several compounds. Next, a number of long-term carcinogenicity studies with DDT were considered, and among-strain differences in cancer susceptibility for this compound were noted. Finally, it was shown that if susceptible subgroups do exist, and certain simplifying assumptions are made, then in many cases tumor increases can be detected more readily by studying several inbred mouse strains rather than a single outbred stock.

Animals

An incorrect definition of fitness revisited.

I have attempted to show that a certain mistaken definition of fitness, which surfaces occasionally, may turn out to have some merit. No claim is made that it is an improvement on, or should replace, the conventional definition of fitness; but it is different and has its own validity. Its generality is intriguing, its application is not limited either to selection or one locus models, and it may be easier to measure experimentally.

Biological Evolution

Recombination values and their errors.

Two four-point testcrosses comprising 87,000 tomato plants were grown and the data collected from 28 subgroups. Each subgroup consisted of 2,000 or 5,000 plants and should give a valid estimate of the three recombination values. The 28 values for each interval give more outlyers (23% are outside the 95% limits set by the standard deviation calculated by the binomial formula square root of p q/n) than would be expected by chance. If each subgroup was regarded as the control and the other groups tested against this, then 42% of the time the two subgroups would be significantly different. It is suggested that there are many cases in the literature where this comparison has been made and the significant difference wrongly ascribed to treatment. While the causes of these changes in recombination value are unknown and therefore uncontrollable, they must be anticipated in all such studies. Control and treatment must be replicated enough that chance extreme values will not be attributed to treatment.

Crosses, Genetic

Dissection of a continuous distribution: red cell galactokinase activity in blacks.

A significant difference between blacks and whites in the distribution of red cell galactokinase (GALK) has been found by Tedesco et al. [2]. From the shapes of the distributions, it was inferred that whites are essentially all homozygous for one allele (GALKA), but blacks are polymorphic. A second allele (GALKP), for lower GALK activity, is presented at high frequency in blacks but rare or absent in whites. This paper presents a method which, assuming the genetic model presented, estimates the genotype composition of the black sample. We make some reasonable biochemical assumptions and fit a mixture of three normal distributions to the black data to obtain an estimate of p, the frequency of GALKA in blacks. The fit of the model to the data is excellent and the best estimate of p is .217 +/- .025. Since admixture of white genes in blacks from the United States is known to be about 20%, the value of p implies that virtually all GALKA alleles were introduced by admixture, and that the ancestral black population was monomorphic for GALKP. If whites are indeed monomorphic for GALKA, they differ from unmixed blacks by a full gene substitution at the locus for GALK.

Alleles

A comparison of two methods for making statistical inferences on Nei's measure of genetic distance.

The delta and jackknife methods can be used to estimate Nei's measure of genetic distance and calculate confidence intervals for this estimate. Computer stimulations were used to study the bias and variance of each estimator and the accuracy of the corresponding approximate 95% confidence intervals. The simulations were conducted using 3 sets of data and several sample sizes. The results showed: (1) the jackknife reduced bias; (2) in 8 out of 9 cases the variance and mean square error of the jackknife estimator were less; (3) a second order jackknife reduced the bias the most but suffered a corresponding increase in variance; (4) both the first order jackknife and delta methods yielded intervals whose confidence levels were approximately equal but less than 95%.

Alleles

From family trials to genomic mate allocation: statistical and genomic strategies to accelerate sugarcane genetic improvement.

Sugarcane (Saccharum spp.) underpins global sugar and bioenergy supply and is increasingly valued as a renewable biomass feedstock. Sustained improvement in commercial traits and resilience is constrained by long breeding cycles, clonal propagation, multi-stage testing, and a highly polyploid, heterozygous, and frequently aneuploid genome with substantial non-additive genetic variation. Genomic selection has demonstrated value for predicting elite-clone performance, yet its operational use remains limited at earlier decision points, including family selection, parent evaluation, and cross design. This review examines the biological, statistical, and genomic factors that shape these decisions, with emphasis on the Australian breeding context based on progeny assessment trials (PATs), clonal assessment trials (CATs), and final assessment trials (FATs). We evaluate challenges arising from family plot means, the use of different full-sib samples as nominal family replicates, spatial heterogeneity, competition, genotype-by-environment interaction, and the partitioning of additive and non-additive effects. We also assess the integration of pedigree and genomic relationship, genotype representation, allele-dosage estimation, aneuploidy, genomic prediction models, and training-population design. We then consider genomic prediction of cross performance and constrained mate allocation as approaches for improving expected family performance, accounting for cross-specific non-additive effects and managing relatedness. We propose a decision-centred framework that links family and clonal data across breeding stages, tracks the propagation of information and uncertainty, and supports parent recycling and cross allocation. We conclude with a practical research agenda for stage-integrated mixed-model and single-step analyses that connect early family evaluation with genomic prediction and cross-level decision support in sugarcane breeding.

Saccharum

Evaluation of the efficacy of optical genome mapping in prenatal diagnosis: a retrospective cohort study.

BACKGROUND: Optical genome mapping (OGM) is an emerging cytogenetic method for concurrently detecting structural variants (SVs) and copy number variants (CNVs). However, its clinical application in prenatal diagnosis remains underexplored. METHODS: This study retrospectively evaluated the clinical validity of OGM in prenatal diagnosis by comparing with two routine genetic testing methods: karyotyping and chromosomal microarray analysis (CMA). Both positive and negative cases detected by routine genetic methods were enrolled to evaluate the technical concordance of OGM and its capability to improve diagnostic rate in negative cases. The exclusion criteria were balanced centromeric translocations, mosaic cases with cellular fractions&#x2009;<&#x2009;20%, and loss of heterozygosity (LOH)&#x2009;<&#x2009;25&#xa0;Mb. All samples subjected to OGM testing were anonymized and analyzed blindly. The results from OGM were compared with those from routine genetic testing, and statistical analyses were performed to assess technical concordance and diagnostic rate. RESULTS: Of 217 samples (166 positive samples and 51 negative samples for routine genetic testing), all were successfully tested with OGM, including 2 umbilical cord blood samples, 4 chorionic villi samples, and 211 cultured amniotic fluid samples. Of the 207 reportable chromosomal aberrations from 166 positive samples, the blinded concordance between OGM and CMA, karyotyping, and combination of karyotyping plus CMA was 97.81%, 96.36%, and 97.10%, respectively. OGM missed six aberrations initially, including one LOH, two marker chromosomes, and three microdeletions. However, after reanalysis, its concordance improved to 100% with CMA and 99.03% with karyotyping plus CMA. OGM also diagnosed one additional case of a 3-kb deletion in 51 negative samples, improving the diagnostic rate by 1.96%. Moreover, OGM reclassified the pathogenicity of two microdeletions from pathogenic to uncertain significance in 2 positive cases. Furthermore, OGM clarified the diagnosis suspected by routine genetic testing and improved diagnostic accuracy in some cases. CONCLUSION: As far as we know, this is the largest retrospective study on OGM in prenatal diagnosis, and it includes a broad range of sample types. The results showed that OGM exhibits high concordance among the tested methods and increases the diagnostic rate. Thus, OGM has the potential to become a first-line technique for prenatal diagnosis in the future.

Humans

RAD-Seq-derived SNPs reveal no local population structure in the commercially important deep-sea queen snapper (Etelis oculatus) in Puerto Rico.

UNLABELLED: The queen snapper (Etelis oculatus Valenciennes in Cuvier & Valenciennes, 1828) is a deep-sea snapper whose commercial importance continues to increase in the US Caribbean. However, little is known about the biology and ecology of this species. In this study, the presence of a fine-scale population structure and genetic diversity of queen snapper from Puerto Rico was assessed through 16,188 SNPs derived from the Restriction site Associated DNA Sequencing (RAD-Seq) technique. Summary statistics estimated low genetic diversity (HO&#x2009;=&#x2009;0.333-0.264) and did not reveal population differentiation within our samples (F ST&#x2009;=&#x2009;-&#xa0;0.001-0.025). Principal component analysis and a model-based clustering method did not detect a fine-scale subpopulation structure among sampling sites, however, there was genetic variability within regions and sites. Our results have revealed comparable genetic and dispersal patterns to those observed in other shallow-water snapper species in Puerto Rico waters. It is crucial to further enhance our understanding of the ecological and biological aspect of the queen snapper to effectively manage and conserve this species as fishing pressure has been extended to deep water species in the US Caribbean. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s42995-025-00289-7.

Caribbean Fisheries

Is there a pattern of gene differentiation in the Indian populations.

Indian populations divided into a number of endogamous groups consisting of different castes, languages, religions, and tribes provide unique opportunities for examining the extent and nature of genetic differentiation at a microevolutionary stage. The genetic relationships between some of these Indian population groups have been examined using electrophoretic data from several biochemical loci in a gene diversity analysis. Does this type of analysis provide any insight into what causes such gene differentiation? What patterns of genetic variation emerge from these empirical findings? Answers are sought by relating the observed heterozygosity, genetic distance, and allied statistics to a mutation-drift hypothesis. The statistics used are: (1) interlocus mean and variance of heterozygosity, (2) mean and variance of genetic distance, and (3) correlation of heterozygosity and gene identity. The observed relationships between these sets of statistics agree well with the ones predicted by the hypothesis that different alleles at protein loci are selectively equivalent and gene frequency change occurs predominantly due to genetic drift.

Gene Frequency