Search PubMedSearch

SEARCH · Search PubMed

Results for “parameter estimation”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genome-wide association study of body weight and body size traits in Langya hens.

Langya chicken is a Chinese indigenous chicken breed with high genetic diversity. To systematically analyse the genetic basis of body size traits, eight traits (including BW, comb shape, and body size) of 2 952 Langya hens were measured at 130 days of age and at first egg of age. A total of 9 708 856 high-quality single-nucleotide polymorphisms (SNPs) were obtained through whole-genome resequencing and used for subsequent genetic parameter estimation and genome-wide association study (GWAS). The results of genetic parameter analysis revealed significant differences in the SNP heritability of different body size traits, with an overall range of 0.13-0.64. In particular, BW, comb length, comb height, and tibia length exhibited moderate-to-high heritability (0.34-0.64) during both developmental stages. GWAS revealed significantly associated SNP loci distributed across multiple chromosomal regions, indicating that body size traits have a complex multilocus genetic regulatory structure and that some chromosomal regions recur for different body size traits and during different developmental stages, showing potential pleiotropic effects or shared genomic regions. Notably, multiple stable body size trait-associated regions were identified on Gallus gallus autosome (GGA) 1, 4, and 27, including genomic regions on GGA1 (167.56-178.18 Mb), GGA4 (68.24-81.17 Mb), and GGA27 (5.22-6.73 Mb), in which significantly associated signals were repeatedly detected for multiple body size traits, such as BW and tibia length. The significant SNPs in the above regions were characterised by strong linkage disequilibrium and were associated with multiple body size traits, indicating that these SNPs may serve as important genetic hotspots for the regulation of chicken body shape and structure. Candidate genes annotated in these core regions include NCAPG, KPNA3, LDB2, PPARGC1A, FNDC3A, SOST, RB1, STON2, and TARP; the functions of these genes are involved mainly in the regulation of cell proliferation, energy metabolism, bone development, and tissue growth. NCAPG was consistently associated with multiple traits at both developmental stages. Functional enrichment analysis further revealed that these candidate genes were significantly enriched in the phosphatidylinositol, GnRH, energy metabolism, skeletal development and protein biosynthesis signalling pathways. The genetic characteristics of Langya chicken body size traits during the growth stage at the genome-wide level and the underlying molecular mechanisms were systematically revealed in this study. The findings provide important candidate gene resources and a theoretical basis for the screening of molecular markers for body size traits and the genomic breeding of regional chicken breeds.

Candidate genes

Quantifying uncertainty of predictions from cancer progression models.

MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).

Humans

Quantifying genetic and genotypic gain gaps in Eucalyptus: the hidden cost of ignoring inbreeding and dominance.

Understanding the mating system of Eucalyptus species is necessary to accurately estimate genetic parameters and improve breeding programs. Eucalyptus species often exhibit mixed mating systems, leading to complex relationships among progenies. Traditional tree breeding programs that assume a half-sibling relationship for open-pollinated (OP) trials may overestimate genetic gains by neglecting the effects of inbreeding and dominance. This study focuses on Eucalyptus pellita, a species with a mixed mating system, to quantify the impact of selfing and dominance on breeding strategies. We simulated OP trial growth data for 100 randomly selected families in a randomized complete block design, using published estimated parameters for diameter at breast height (DBH). Our analysis indicated that marker-based models, particularly those incorporating dominance effects, provide more accurate genetic parameter estimates and larger predicted genetic gains than pedigree-based models. These results reveal pronounced genetic and genotypic gain gaps when traditional models are employed, underscoring the imperative for integrating dominance-informed genomic selection strategies. Thus, our study provides essential guidance for optimizing breeding programs to sustainably enhance productivity and genetic quality in Eucalyptus plantations.

Eucalyptus

Modelling the effects of biological intervention in a dynamical gene network.

Cellular response to environmental and internal signals can be modeled by dynamical gene regulatory networks (GRN). In the literature, three main classes of gene network models can be distinguished: (1) non-quantitative (or data-based) models which do not describe the probability distribution of gene expressions; (2) quantitative models which fully describe the probability distribution of all genes co-expression; and (3) mechanistic models which allow for a causal interpretation of gene interactions. We propose two rigorous frameworks to model gene alteration in a dynamical GRN, depending on whether the network model is quantitative or mechanistic. We explain how these models can be used for design of experiment, or, if additional alteration data are available, for validation purposes or to improve the parameter estimation of the original model. We apply these methods to the Gaussian graphical model, which is quantitative but non-mechanistic, and to mechanistic models of Bayesian networks and penalized linear regression.

Gene Regulatory Networks

Bayesian reconstruction and differential testing of excised introns.

MOTIVATION: Characterizing the differential excision of introns is critical for understanding the functional complexity of a cell or tissue, from normal developmental processes to disease pathogenesis. Most transcript reconstruction methods infer full-length transcripts from high-throughput sequencing data. However, this is a challenging task due to incomplete annotations and the heterogeneous expression of transcripts across cell-types, tissues, and experimental conditions. Several recent methods circumvent these difficulties by considering local splicing events, but these methods lose transcript-level splicing information and may conflate similar, but distinct transcripts. RESULTS: In this work, we formalize a new transcript reconstruction problem that interpolates between the full-length and local splicing perspectives by considering sequences of exon-exon junctions (SEEJs) that co-occur in transcripts. We then present a hierarchical Bayesian admixture model and posterior inference algorithms for computing SEEJs (BSEEJ), and a generalized linear model for characterizing differential SEEJ usage based on model parameter estimates. We show that BSEEJ achieves high F1 score for reconstruction tasks and improved accuracy and sensitivity in differential splicing when compared with six transcript and local splicing methods on simulated data. Lastly, we evaluate BSEEJ on experimental data based on transcript reconstruction, novelty of transcripts produced, model sensitivity to hyperparameters, and a functional analysis of differentially expressed SEEJs. AVAILABILITY AND IMPLEMENTATION: BSEEJ is freely available at https://github.com/bayesomicslab/BSEEJ.

Bayes Theorem

A prospective pharmacokinetic interaction study between rifampicin and fusidic acid for the treatment of staphylococcal infections.

OBJECTIVE: Fusidic acid with rifampicin is used for the treatment of severe staphylococcal infection, particularly prosthetic joint infections. Previous studies using twice daily fusidic acid showed rifampicin increases fusidic acid clearance, potentially causing sub-therapeutic concentrations. It is uncertain whether this occurs with three-times daily dosing. This study sought to re-evaluate this potential drug-drug interaction. METHODS: In this prospective, open-label drug-drug interaction population pharmacokinetic (PK) study, participants were randomized to receive fusidic acid or rifampicin for 24 hours, followed by combination therapy for the duration of treatment. Drug concentration assays used liquid-chromatography mass spectroscopy on dried blood spots. Population PK models were built for fusidic acid, rifampicin and 25-desacetyl rifampicin. RESULTS: Ten participants were recruited. Inter-individual variability for both absorption (98.1%) and clearance (77.8%) were high for fusidic acid. A population pharmacokinetic model for fusidic acid revealed that early autoinhibition dominated over later rifampicin-mediated induction, resulting in a net decrease in fusidic acid clearance. The mean fusidic acid area under the curve during each dosing interval (AUCτ) at steady state was 1.76-fold [0.049, 34.918] higher relative to Day 1, despite high uncertainty.Large inter-individual (110%) variability in absorption was observed for rifampicin. There was no apparent effect on rifampicin metabolism by fusidic acid co-administration, however, fusidic acid decreased clearance of 25-desacetyl rifampicin. CONCLUSION: In patients treated with fusidic acid three times daily in combination with rifampicin, autoinhibition potentially counteracted rifampicin induction such that fusidic acid concentrations were not reduced. Rifampicin clearance was not affected by fusidic acid, but 25-desacetyl rifampicin clearance was decreased. There was large inter-individual variability in the observed concentrations and final parameter estimates.

Fusidic Acid

Incorporating Epidemiological Data into the Genomic Analysis of Partially Sampled Infectious Disease Outbreaks.

Pathogen genomic data are increasingly being used to investigate transmission dynamics in infectious disease outbreaks. Combining genomic data with epidemiological data should substantially increase our understanding of outbreaks, but this is highly challenging when the outbreak under study is only partially sampled, so that both genomic and epidemiological data are missing for intermediate links in the transmission chains. Here, we present a new dynamic programming algorithm to perform this task efficiently. We implement this methodology into the well-established TransPhylo framework to reconstruct partially sampled outbreaks using a combination of genomic and epidemiological data. We use simulated datasets to show that including epidemiological data can improve the accuracy of the inferred transmission links compared with inference based on genomic data only. This also allows us to estimate parameters specific to the epidemiological data (such as transmission rates between particular groups), which would otherwise not be possible. We then apply these methods to two real-world examples. First, we use genomic data from an outbreak of tuberculosis in Argentina, for which data was also available on the HIV status of sampled individuals, in order to investigate the role of HIV coinfection in the spread of this tuberculosis outbreak. Second, we use genomic and geographical data from the 2003 epidemic of avian influenza H7N7 in the Netherlands to reconstruct its spatial epidemiology. In both cases, we show that incorporating epidemiological data into the genomic analysis allows us to investigate the role of epidemiological properties in the spread of infectious diseases.

Humans

A Systematic Review of Spatial Epidemiological Modeling Approaches Applied During the COVID-19 Pandemic.

BACKGROUND: A wide range of epidemiological modeling approaches have been applied to the SARS-CoV-2 pandemic, which presents an opportunity to assess common approaches applied to specific research questions. Spatial models interrogate how heterogeneities and host movement dynamics influence local and regional patterns of disease, issues that were of great interest for understanding and controlling SARS-CoV-2. OBJECTIVE: Here we present a systematic review of spatial epidemiological modeling approaches of SARS-CoV-2. We describe common themes and highlight unique strategies, providing a foundation for researchers to devise spatial models most appropriate for future pathogens and epidemics. Our review also categorizes the research questions that were addressed with spatial models, highlights parameter estimation techniques, and describes the cyber infrastructure used for model development. METHODS: We conducted a systematic review using Web of Science and a standardized set of keywords, followed by thorough examination of abstracts and full texts to determine which studies met our inclusion criteria. To guide our description and comparisons of models, we developed a Geography, Population, Movement (GPM) framework that conceptualizes the interactions between three distinct subcomponents of any spatial model. The geographic model represents the physical arena in which the model is implemented, the intra-population model describes the transmission and disease processes that occur within distinct spatial units of the geography, and the movement model describes the algorithms that dictate how hosts move among spatial units within the geography. RESULTS: The search identified a total of 193 articles, of which 109 were included in our review. The most abundant intra-population modeling methods were agent-based (47.7%) and compartmental modeling (29.4%) approaches. Movement models ranged in complexity, with the most complex models implementing commuter movement among many points of interest in the geographic arena, which were sometimes parameterized by fine-scale mobility data. Geographic models ranged from describing microcosms, such as single classrooms, all the way up to multi-country models. Of the 63.3% of models studies that specified the programming language used, we detected ten different languages, with Matlab and Python being the most frequent, although only 30.6% of studies provided open-access code for their models. We also described eight specialized software systems that were used to construct agent-based or compartment models of COVID-19. CONCLUSIONS: Our review identified and characterized a variety of spatial modeling strategies and software that were usefully employed to address many relevant epidemiological questions for COVID-19. Future research is needed to quantitatively assess which modeling approaches are most appropriate in specific situations, to answer specific questions, or to apply to certain disease systems. Moreover, future cyberinfrastructure could help to modularize and standardize modeling approaches, which would increase transparency and reproducibility, and which would facilitate a detailed examination of which model attributes relate to model performance in a variety of contexts.

COVID-19

Modeling Early-Onset Cancer Kinetics Reveals Changes in Underlying Risk and the Impact of Population Screening.

UNLABELLED: Recent studies have reported increases in early-onset cancer cases (diagnosed less than 50 years of age) and raised questions about whether the increase is related to earlier diagnosis from nonspecific medical tests as reflected by decreasing tumor-size-at-diagnosis (apparent effects) or actual increases in underlying cancer risk (true effects), or both. The classic Multistage Clonal Expansion (MSCE) model assumes cancer detection at the first malignant cell's emergence, although later modifications have included lag-times or stochasticity in detection to represent the delay in tumor detection. In this study, we introduced an approach to explicitly incorporate tumor-size-at-diagnosis in the MSCE framework accounting for improvements in cancer detection over time to distinguish between apparent and true increases in early-onset cancer incidence. The model was structurally identifiable and provided better parameter estimation than the classic model. The model was applied to colorectal, breast, and thyroid cancers to examine changes in cancer risk while accounting for detection improvements over time in three representative birth cohorts (1950-1954, 1965-1969, and 1980-1984). The analyses suggested accelerated carcinogenic events and shorter mean sojourn times (the average time from the first malignant cell emergence to cancer detection) in more recent cohorts. Furthermore, using this model to examine the screening impact on the incidence of breast and colorectal cancers, for which both have established screening protocols, provided results that align with well-documented differences in screening effects between these cancers. These findings underscore the importance of incorporating tumor-size-at-diagnosis in cancer modeling and support true increases in early-onset cancer risk in recent years for breast, colorectal, and thyroid cancers. SIGNIFICANCE: A model of early-onset cancer trends that distinguishes true risk from detection effects accurately captures cancer kinetics, trends in cancer progression, and the impact of screening, which could inform cancer prevention strategies. This article is part of a special series: Driving Cancer Discoveries with Computational Research, Data Science, and Machine Learning/AI .

Humans

Trajectory inference from single-cell genomics data with a process time model.

Single-cell transcriptomics experiments provide gene expression snapshots of heterogeneous cell populations across cell states. These snapshots have been used to infer trajectories and dynamic information even without intensive, time-series data by ordering cells according to gene expression similarity. However, while single-cell snapshots sometimes offer valuable insights into dynamic processes, current methods for ordering cells are limited by descriptive notions of "pseudotime" that lack intrinsic physical meaning. Instead of pseudotime, we propose inference of "process time" via a principled modeling approach to formulating trajectories and inferring latent variables corresponding to timing of cells subject to a biophysical process. Our implementation of this approach, called Chronocell, provides a biophysical formulation of trajectories built on cell state transitions. The Chronocell model is identifiable, making parameter inference meaningful. Furthermore, Chronocell can interpolate between trajectory inference, when cell states lie on a continuum, and clustering, when cells cluster into discrete states. By using a variety of datasets ranging from cluster-like to continuous, we show that Chronocell enables us to assess the suitability of datasets and reveals distinct cellular distributions along process time that are consistent with biological process times. We also compare our parameter estimates of degradation rates to those derived from metabolic labeling datasets, thereby showcasing the biophysical utility of Chronocell. Nevertheless, based on performance characterization on simulations, we find that process time inference can be challenging, highlighting the importance of dataset quality and careful model assessment.

Single-Cell Analysis

Expanding and improving analyses of nucleotide recoding RNA-seq experiments with the EZbakR suite.

Nucleotide recoding RNA sequencing methods (NR-seq; TimeLapse-seq, SLAM-seq, TUC-seq, etc.) are powerful approaches for assaying transcript population dynamics. In addition, these methods have been extended to probe a host of regulated steps in the RNA life cycle. Current bioinformatic tools significantly constrain analyses of NR-seq data. To address this limitation, we developed EZbakR (https://github.com/isaacvock/EZbakR), an R package to facilitate a more comprehensive set of NR-seq analyses, and fastq2EZbakR (https://github.com/isaacvock/fastq2EZbakR), a Snakemake pipeline for flexible preprocessing of NR-seq datasets, collectively referred to as the EZbakR suite. Together, these tools generalize many aspects of the NR-seq analysis workflow. The fastq2EZbakR pipeline can assign reads to a diverse set of genomic features (e.g., genes, exons, splice junctions), and EZbakR can perform analyses on any combination of these features. EZbakR extends standard NR-seq mutational modeling to support multi-label analyses (e.g., s4U and s6G dual labeling), and implements an improved hierarchical model to better account for transcript-to-transcript variance in metabolic label incorporation. EZbakR also generalizes dynamical systems modeling of NR-seq data to support analyses of premature mRNA processing and flow between subcellular compartments. Finally, EZbakR implements flexible and well-powered comparative analyses of all estimated parameters via design matrix-specified generalized linear modeling. The EZbakR suite will thus allow researchers to make full, effective use of NR-seq data.

Software

Accounting for contact tracing in epidemiological birth-death models.

Phylodynamics bridges the gap between classical epidemiology and pathogen genome sequence data by estimating epidemiological parameters from time-scaled pathogen phylogenetic trees. The models used in phylodynamics typically assume that the sampling procedure is independent between infected individuals. However, this assumption does not hold for many epidemics, in particular for such sexually transmitted infections as HIV-1, for which contact tracing schemes are included in health policies of many countries. We extended phylodynamic multi-type birth-death (MTBD) models with contact tracing (CT), and developed a simulator to generate trees under MTBD and MTBD-CT models. We proposed a non-parametric test for detecting contact tracing in pathogen phylogenetic trees. Its application to simulated data showed that it is both highly specific and sensitive. For the simplest representative of the MTBD-CT family, the BD-CT(1) model, where only the last contact can be notified, we solved the differential equations and proposed a closed form solution for the likelihood function. We implemented a maximum-likelihood program, which estimates the BD-CT(1) model parameters and their confidence intervals from phylogenetic trees. It performed accurate parameter inference on BD and BD-CT(1) simulated data, and detected contact tracing in HIV-1 B epidemics in Zurich and the UK. Importantly, we showed that not accounting for contact tracing when it is present, leads to bias in parameter estimation with the BD model (overestimation of the becoming-non-infectious rate). This bias is also present, but greatly reduced, when the BD-CT(1) model is used on data where multiple contacts can be notified. Our CT test, MTBD-CT tree simulator and BD-CT(1) parameter estimator are freely available at GitHub (evolbioinfo/treesimulator and evolbioinfo/bdct).

Contact Tracing

Sparse Logistic Regression on Genomic Data for Prediction of Tumour Pathological Subtype.

The correct prediction of tumour subtype is critical for the treatment of cancer patients to maximise the chance of survival. The patients' genomic information, such as copy number alterations (CNA) profile, has increasingly become an important factor in the prediction to supplement the traditional pathological subtyping. The incorporation of the CNA information in a prediction model, such as logistic regression, faces two major statistical challenges: first, how to estimate the model parameters in the thousands and, second, how to deal with the correlation of CNA between genomic regions. To address them, we propose a sparse logistic regression model with random effects where some of its parameters are estimated to zero while the other parameters are non-zero. In effect, a variable selection is embedded in the modelling. To deal with the correlation of CNA across genomic regions, we extend further the model to incorporate an additional penalty in the corresponding likelihood function in the logistic regression. The results show that we can identify selected genomic regions that are informative to distinguish different tumour subtypes, while giving a good prediction ability. We illustrate the methodology using CNA dataset from a lung cancer cohort.

Journal Article

Informing agent-based models with spatial data using convolutional autoencoders.

MOTIVATION: Spatial computational models such as agent-based models (ABMs) offer powerful in silico tools to study tumor dynamics, yet imaging data are still rarely used to inform these models directly. RESULTS: We present an ABM optimization framework that leverages convolutional encoders to compare spatial patterns between experimental imaging data and ABM-generated outputs within a shared latent space. This quantitative comparison was used to estimate ABM parameters across three datasets, ranging from synthetic data to 3D tumoroid-T cell co-culture microscopy and histopathology images from The Cancer Genome Atlas skin cutaneous melanoma samples. Estimated parameters were evaluated using data-derived features and experimental knowledge, including experimental conditions and gene expressions. Simulations using optimized parameters reproduced key spatial features of the training images, such as tumor boundary complexity and tumor-tumor neighborhood structure. Together, these results demonstrate a flexible framework for ABM parameter optimization using spatial data across modalities, enabling systematic investigation of how spatial architecture influences tumor progression and immune interactions. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/SysBioOncology/ AutoencoderABM under the GPL-3.0 license, with corresponding data sets at https://zenodo.org/records/19022344.

Autoencoder

Genomic background of gestation length and calving-related traits in Holstein cattle.

The reproductive success of cows directly influences the profitability of dairy farms. Reproductive traits, particularly calving-related traits, generally have low heritability but sufficient additive genetic variance to enable genetic progress through genomic selection. Thus, the primary objectives of this study were to estimate genetic parameters and perform single-step genome-wide association studies (ssGWAS) for calf size, calving ease, gestation length, and stillbirth in Holstein cattle. Variance components were estimated based on animal models and Bayesian inference using a data set containing 226,717 animals with phenotypic records, 15,761 animals genotyped with 45,101 SNP markers, and 461,819 animals in the pedigree. SNP effects were estimated using the single-step GBLUP method. For direct and maternal genetic effects, heritability estimates (posterior standard deviation) ranged from 0.001 (0.002) for gestation length in heifers to 0.16 (0.001) for gestation length in cows. Genetic correlations ranged from -0.57 (0.01) between calving ease and stillbirth in heifers to 0.74 (0.01) between gestation length evaluated in heifers and cows. The ssGWAS results supported a highly polygenic architecture for calving-related traits, with most genomic signals not reaching genome-wide significance. A genome-wide significant association was detected for calving ease in cows on BTA23, highlighting FARS2 as a positional candidate gene. The strongest GWAS signals for each trait harbored additional biologically important candidate genes, including NPPA, NPPB, BCHE, EPHA4, DLD, and GTF2I. Given the generally low heritability estimates and the predominantly polygenic architecture observed for these traits, genomic selection may contribute to the genetic improvement of calving-related traits in Holstein cattle, with potential benefits for cow welfare, calf survival, and overall dairy production efficiency.

dairy cattle

SNP genotyping in Pseudotsuga menziesii and Pinus radiata using targeted genotyping-by-sequencing (GBS): improved Bayesian SNP calling using a beta-binomial distribution and other optimized input parameters.

BACKGROUND: Single-nucleotide polymorphism markers (SNPs) have important applications in gene conservation, breeding, and fundamental genetics research. Our long-term goal is to develop routine approaches for SNP genotyping in forest trees. Ideally, these approaches would be inexpensive, able to accommodate a wide range of samples and SNPs, available through commercial providers, and produce high-quality SNP data. RESULTS: Using targeted genotyping-by-sequencing (GBS), we developed SNP assays for two highly heterozygous tree species, Douglas-fir (Pseudotsuga menziesii) and radiata pine (Pinus radiata). Using Douglas-fir haploid and diploid data, we optimized Bayesian SNP calling by testing four input parameters: (1) allele and genotype prior probabilities, (2) Rho, the beta-binomial dispersion parameter, (3) estimated read error (BayesReadError), and (4) the logPO cutoff used to filter low confidence SNP calls. logPO is the Bayesian posterior odds ratio for a called SNP. Compared to assuming a binomial distribution of read counts (Rho = 0), the beta-binomial distribution (Rho = 0.33) substantially reduced call error and heterozygote undercalling. Compared to the other Bayesian parameters, genotype priors had little effect on genotyping success. For Douglas-fir, we tested 5,360 SNP assays, and then studied the performance of the best 4,000. For radiata pine, we tested 6,000 SNP assays, and then studied the performance of the best 4,570. In Douglas-fir and radiata pine, our Bayesian approach resulted in median call rates of 95% to 98% for the top-ranked SNPs, with an estimated call error of 1.60% for known homozygous genotypes and 2.27% for known heterozygotes. In radiata pine, median and mean call rates were above 91% for GBS and SNP genotyping using an Axiom fixed genotyping array. Additionally, the median correspondence between the GBS and Axiom genotypes was about 98% overall (mean 96%). CONCLUSIONS: By optimizing Bayesian SNP calling, selecting the best 4-5 K SNPs, and excluding samples with low DNA amounts, we substantially reduced call error and heterozygote undercalling, resulting in SNP genotypes that were nearly identical to genotypes obtained using the Axiom array. Furthermore, genotyping performance should increase even further if our SNP rankings were used to develop less complex probe pools that target fewer SNPs.

Pinus

Collective posterior inference from highly variable empirical replicates.

High-throughput experimental platforms now routinely generate data from dozens or hundreds of independent observations. Simulation-based inference (SBI) offers a powerful framework for estimating model parameters from such complex datasets, but standard methods struggle to scale to the noisy multiple-replicates regime without incurring prohibitive computational costs or careful hyperparameter tuning. Here, we introduce a new method for fast and robust collective posterior inference from multiple independent replicates using a robust product-of-experts aggregation scheme that automatically mitigates the influence of outliers. Evaluating it on synthetic and empirical evolutionary datasets, we find it achieves state-of-the-art estimation accuracy and computational efficiency, including inference from noisy observations. Our method is compatible with any SBI framework, providing a scalable, plug-and-play solution for inference from noisy multiple-replicate datasets.

Computational Biology

Genomic prediction and genome-wide association study for liver abscesses in crossbred beef cattle.

Liver abscesses are a concern in feedlot cattle, and little is known about the role of genetics in their development. This study aimed to estimate genetic parameters and to identify single-nucleotide polymorphisms (SNPs) associated with liver abscesses. Crossbred cattle representing 18 breeds in the U.S. Meat Animal Research Center Germplasm Evaluation Program were phenotyped for liver abscesses at slaughter (n&#x2005;=&#x2005;9,044). Seventeen percent of cattle had liver abscesses. These cattle had genotypes that were imputed to sequence variant genotypes. After filtering and quality control, 340,723 SNPs were used in the analysis. Liver abscess prevalence was modeled with a single-step genomic best linear unbiased prediction (ssGBLUP) threshold model using a Bayesian framework. The model included contemporary group (sex, treatment group, and slaughter date), additive genomic, and residual effects. Genomic heritability was 0.039 (95% highest posterior density&#x2005;=&#x2005;0.005, 0.081), which was very small. To assess prediction quality, a 5-fold random cross-validation structure was used. Method Linear Regression was used to assess accuracy, bias, and dispersion by comparing estimated breeding values (EBV) from full and reduced analyses. Cross-validation metrics showed EBV based on genotypes had 0.05 reliability (SD&#x2005;<&#x2005;0.01) with no bias relative to EBV based on genotypes and phenotypes. For the genome-wide association study, SNP effects were back calculated from the EBV solutions from ssGBLUP. No SNPs were associated with liver abscesses at a Benjamini-Hochberg adjusted 0.05 significance level. Although a large dataset was used, this result was because of the low genomic heritability and imprecise EBV used to calculate SNP effects. Based on these results, environmental factors contribute to most of the variation in liver abscesses. Genetic selection to reduce liver abscesses would be slow because of the low genomic heritability, measurement late in life, and inability to measure breeding animals. A faster approach would be finding additional environmental interventions that maintain animal performance.

Animals