Search PubMedSearch

SEARCH · Search PubMed

Results for “Bayesian model selection”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

On the origin of animals and placental mammals: a critique of literalist readings of the fossil record.

The fossil record is incomplete, as evidenced by the pervasive presence of ghost lineages throughout the Tree of Life. For example, across placental mammals, at least 720 Myr of basal lineages are ghost lineages, that is, lineages that have left no fossil evidence of their past history. In contrast, some studies have suggested that the fossil record is a faithful temporal archive of evolutionary history and thus the times of diversification of clades must be close to the ages of their oldest fossils. Such literalist interpretations have been contradicted by analysis of molecular datasets which, in many cases, indicate that groups including placental mammals and animals may have originated at times substantially older than their fossil records. Some of those studies have further argued that, in the case of animals and placental mammals, molecular clocks are uninformative, suffer from characteristic pathologies, and thus cannot distinguish between recent and ancient hypotheses of diversification. Here, we reexamine these two cases and show, using Bayesian model selection theory, that the explosive diversification models previously proposed for animals and placental mammals have a posterior probability of ∼0. We show the characteristic pathologies purportedly discovered do not exist, highlight errors in previous analyses, and provide advice on best practice for molecular-clock dating analysis.

Animals

Identifying multigenic modules under selection in the tumor genome.

MOTIVATION: Genomic alterations in cancer arise from selective pressures acting on hallmark molecular modules, layered over a background of random mutagenic events. Methods to detect selection at the level of modules, as opposed to genes or nucleotides, are relatively underdeveloped. RESULTS: Here we present CanSRMaPP (Cancer Selection Recovery by Maximum Posterior Probability), a Bayesian model of the cancer genome that infers mutational selection on single genes and multi-genic modules while simultaneously modeling background events. Applying CanSRMaPP to lung adenocarcinoma genomes, we identify positive selection on 63 modules, yielding a model that parsimoniously explains the observed pattern of genetic alterations observed in new cancer cohorts. We further show that CanSRMaPP is adaptable to more tumor types and to alternative module definitions. We show that these modules serve as an effective scaffold for translating the cancer genome to molecular states, with prediction of cancer biomarker status as demonstration. AVAILABILITY: CanSRMaPP is freely available on GitHub. SUPPLEMENTARY INFORMATION: Supplementary Figs. S1-5, Supplementary Tables S1-5, and Supplementary Notes 1 and 2 are available at Bioinformatics online.

Journal Article

Bayesian Modeling of Cancer Outcomes Using Genetic Variables Assisted by Pathological Imaging Data.

With the increasing maturity of genetic profiling, an essential and routine task in cancer research is to model disease outcomes/phenotypes using genetic variables. Many methods have been successfully developed. However, oftentimes, empirical performance is unsatisfactory because of a "lack of information." In cancer research and clinical practice, a source of information that is broadly available and highly cost-effective comes from pathological images, which are routinely collected for definitive diagnosis and staging. In this article, we consider a Bayesian approach for selecting relevant genetic variables and modeling their relationships with a cancer outcome/phenotype. We propose borrowing information from (manually curated, low-dimensional) pathological imaging features via reinforcing the same selection results for the cancer outcome and imaging features. We further develop a weighting strategy to accommodate the scenario where information borrowing may not be equally effective for all subjects. Computation is carefully examined. Simulations demonstrate competitive performance of the proposed approach. We analyze TCGA (The Cancer Genome Atlas) LUAD (lung adenocarcinoma) data, with overall survival and gene expressions being the outcome and genetic variables, respectively. Findings different from the alternatives and with sound properties are made.

Humans

Comparative effectiveness of game-based learning modalities in nursing and medical education: a systematic review and Bayesian network meta-analysis.

BACKGROUND: Game-based learning (GBL) is increasingly used in healthcare education, but educators must choose among diverse modalities (e.g., quiz platforms, apps, serious games and metaverse environments). Comparative evidence on which modalities perform best across learning domains (knowledge, attitudes, and practice) remains limited. AIM: To compare the effects of distinct GBL modalities on knowledge, attitudes, and practice outcomes in nursing and medical education and to explore whether comparative effects differ by learner group (pre-licensure students and in-service professionals). DESIGN: PRISMA-NMA-aligned systematic review and Bayesian network meta-analysis. METHODS: We searched eight databases and trial registries through September 2, 2024, for randomized controlled trials comparing GBL with traditional teaching (TT). Outcomes were transformed to a 0-100 scale and analysed as change from baseline in Bayesian consistency models; random-effects models were selected using deviance information criterion (DIC). Risk of bias was assessed using RoB 2. We report mean differences (MDs) with 95% credible intervals (CrIs) versus TT, ranking probabilities, and subgroup NMAs by learner group. RESULTS: Thirty-one RCTs (n = 3439) were included; 15 contributed complete data to the network. Risk of bias was low in 15 trials and raised some concerns in 16. The network was modest for knowledge (11 trials) and sparse for attitudes (3) and practice (4). Compared with TT, metaverse-based learning showed improved attitudes (MD 15; 95% CrI 12 to 18), based on a single trial. For knowledge and practice, Kahoot-based quizzes (MD 9.1; 95% CrI -8.9 to 27) and app-based learning (MD 4.6; 95% CrI -4.4 to 14) had the highest estimated mean improvements, but credible intervals were wide and included the null for most comparisons. Subgroup rankings differed by learner group, but several comparisons were imprecise and uncertainty was substantial, particularly in sparse networks. CONCLUSIONS: GBL modalities may improve learning outcomes compared with TT, but relative effects appear domain-specific and the certainty of rankings is limited by sparse evidence and imprecision. Future trials should prioritise head-to-head comparisons, robust outcome measurement, and longer-term retention and transfer outcomes in both student and in-service populations.

Humans

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569 bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A + T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU > 1) with a pronounced preference for A/C-ending codons. The ΔRSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean = 1.209 ± 0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae

Sugar kelp (Saccharina latissima) population genetics map onto geographic distance and oceanographic features across coastal Maine.

Sugar kelp (Saccharina latissima; order Laminariales) plays a vital role in kelp forest ecosystems, as well as an expanding kelp aquaculture industry, in the Gulf of Maine, United States. However, ocean warming is eroding the resilience of Maine's kelp forests and may be compromising their local genetic diversity, with impacts on population structure and gene flow. Here, we used genome-wide single nucleotide polymorphism (SNP) data to assess the genetic diversity, structure, and connectivity of S. latissima populations at 11 outer coastal sites spanning the historical range of kelp forests in Maine. Our analyses identified moderate genetic diversity and limited inbreeding within sites (average heterozygosity: 0.27). Further, they revealed that three clusters comprising four genetically distinct populations exist across the study region. Population structure was strongly associated with geographic distance and oceanographic features, as supported by principal coordinate analysis, FST calculations, Bayesian clustering, and spore dispersal modeling. Lastly, our outlier analysis identified genes potentially under selection. Thus, our findings highlight distinct, genetically unique kelp populations along Maine's coast and emphasize the need for regional management strategies that support both ecosystem resilience and sustainable aquaculture under climate change.

Gulf of Maine

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem

Ancient climate changes and relaxed selection shape cave colonization in North American cavefishes.

Extreme environments serve as natural laboratories for studying evolutionary processes, with caves offering replicated instances of independent colonizations. The timing, mode and genetic underpinnings underlying cave-obligate organismal evolution remain enigmatic. We integrate phylogenomics, fossils, palaeoclimatic modelling and newly sequenced genomes to elucidate the evolutionary history and adaptive processes of cave colonization in the study group, the North American Amblyopsidae fishes. Amblyopsid fishes present a unique system for investigating cave evolution, encompassing surface, facultative cave-dwelling and cave-obligate (troglomorphic) species. Using 1105 exon markers and total-evidence dating, we reconstructed a robust phylogeny that supports the nested position of eyed, facultative cave-dwelling species within blind cavefishes. We identified three independent cave colonizations, dated to the Early Miocene (18.5 Ma), Late Miocene (10.0 Ma) and Pliocene (3.0 Ma). Evolutionary model testing supported a climate-relict hypothesis, suggesting that global cooling trends since the Early-Middle Eocene may have influenced cave colonization. Comparative genomic analyses of 487 candidate genes revealed both relaxed and intensified selection on troglomorphy-related loci. We found more loci under relaxed selection, supporting neutral mutation as a significant mechanism in cave-obligate evolution. Our findings provide empirical support for climate-driven cave colonization and offer insights into the complex interplay of selective pressures in extreme environments.

Animals

Bayesian inference of fitness landscapes via tree-structured branching processes.

MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.

Bayes Theorem

Comparing ARG Inference Methods Under Transmission of Reproductive Success: Tree Imbalance Matters.

Inferring coalescent trees from genomic data has become a major subject in population genetics, particularly with the recent advances in tree sequence reconstruction methods. However, it remains unclear how well these methods perform for imbalanced genealogies. Such imbalances can arise from processes such as cultural transmission of reproductive success (CTRS) or positive selection. Using simulated genomic data, we benchmarked three major software packages, SINGER, Relate, and tsinfer, by comparing the imbalance of reconstructed trees by these methods with that of the true simulated trees, for three indices that quantify this imbalance. The three methods performed well under scenarios yielding balanced trees. However, their accuracy declined as imbalance increased. Performances also varied with mutation rate, recombination rate, and sample size. This study opens possibilities for applying these methods to infer CTRS or positive selection in large-scale genomic datasets, using simulation-based inference such as approximate Bayesian computation.

Models, Genetic

MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detection.

Molecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients.

Humans

Genomic background of gestation length and calving-related traits in Holstein cattle.

The reproductive success of cows directly influences the profitability of dairy farms. Reproductive traits, particularly calving-related traits, generally have low heritability but sufficient additive genetic variance to enable genetic progress through genomic selection. Thus, the primary objectives of this study were to estimate genetic parameters and perform single-step genome-wide association studies (ssGWAS) for calf size, calving ease, gestation length, and stillbirth in Holstein cattle. Variance components were estimated based on animal models and Bayesian inference using a data set containing 226,717 animals with phenotypic records, 15,761 animals genotyped with 45,101 SNP markers, and 461,819 animals in the pedigree. SNP effects were estimated using the single-step GBLUP method. For direct and maternal genetic effects, heritability estimates (posterior standard deviation) ranged from 0.001 (0.002) for gestation length in heifers to 0.16 (0.001) for gestation length in cows. Genetic correlations ranged from -0.57 (0.01) between calving ease and stillbirth in heifers to 0.74 (0.01) between gestation length evaluated in heifers and cows. The ssGWAS results supported a highly polygenic architecture for calving-related traits, with most genomic signals not reaching genome-wide significance. A genome-wide significant association was detected for calving ease in cows on BTA23, highlighting FARS2 as a positional candidate gene. The strongest GWAS signals for each trait harbored additional biologically important candidate genes, including NPPA, NPPB, BCHE, EPHA4, DLD, and GTF2I. Given the generally low heritability estimates and the predominantly polygenic architecture observed for these traits, genomic selection may contribute to the genetic improvement of calving-related traits in Holstein cattle, with potential benefits for cow welfare, calf survival, and overall dairy production efficiency.

dairy cattle

Oncotype DX-guided vs physician-directed chemotherapy and survival in HR+/HER2- breast cancer.

BACKGROUND: Oncotype DX testing guides adjuvant chemotherapy decisions in early-stage hormone receptor-positive/HER2-negative breast cancer, but testing is not universally performed, and outcomes associated with genomic-informed versus clinicopathologic-based chemotherapy decision pathways remain unclear. METHODS: Using the 2022 National Cancer Database Breast Participant User File, we identified women diagnosed from 2010 to 2022 with pathologic T1b-T2, node-negative, hormone receptor-positive/HER2-negative invasive breast cancer who received adjuvant chemotherapy and endocrine therapy. Patients were classified into an Oncotype-guided group, defined by Oncotype DX testing with a recurrence score of 26 or higher, and a physician-directed group, defined by receipt of chemotherapy without genomic testing. The primary outcome was overall survival. Analyses used multivariable Cox models, logistic-IPTW and MLP-IPTW, restricted mean survival time analysis, and a Bayesian latent confounding survival model. RESULTS: Among 56,625 women, 27,278 were in the Oncotype-guided group and 29,347 in the physician-directed group. Median ages were 59 and 56 years, respectively. The Oncotype-guided group had more favorable overall survival than the physician-directed group in multivariable Cox analysis (HR, 0.906; 95% CI, 0.856-0.959; P&#x202f;<&#x202f;0.001), with similar findings in IPTW analyses. The association was concentrated among patients aged 56 years or older (HR, 0.866; 95% CI, 0.809-0.927; P&#x202f;<&#x202f;0.001). The Bayesian model showed no strong residual confounding signal. CONCLUSIONS: Among chemotherapy-treated women, an Oncotype-guided pathway was associated with more favorable overall survival than a physician-directed pathway, particularly among older patients, which indicating prognostic heterogeneity selected using genomic versus conventional clinicopathologic information.

Humans

Non-destructive prediction of lead content in oilseed rape leaves by fluorescence hyperspectral technology based on neural network.

Based on fluorescence hyperspectral imaging (FHSI), this study targeted rapid, non-destructive quantification of lead (Pb) content in oilseed rape leaves treated with varying silicon (Si) concentrations, acquiring fluorescence spectra over the 484.43-1001.61&#xa0;nm wavelength range. To optimize spectral data quality, preprocessing methods (Savitzky-Golay smoothing, first derivative, detrending) were comprehensively compared. Characteristic wavelengths were then selected via interval variable iterative shrinkage, which effectively compressed data dimensionality and reduced computational load. A hybrid SE-CL1DA model, fusing a 1D convolutional neural network, a long short-term memory network and SE attention mechanism was constructed, with Bayesian optimization tuning hyperparameters to boost stability. The BO-SE-CL1DA outperformed both traditional machine learning and insufficiently optimized deep learning model (Rp2=0.9609, RMSE&#xa0;=&#xa0;0.0377&#xa0;mg/kg, RPD&#xa0;=&#xa0;5.1736), thus enabling accurate Pb estimation, supporting Si-regulated heavy metal stress management and facilitating agricultural contamination monitoring.

Plant Leaves

Spatially Contextualized Integrative Genomics Highlights Neuronal and Glial Regulatory Programs in Low Back Pain.

PURPOSE: Low back pain (LBP) is a heterogeneous pain condition with a measurable genetic contribution, but the genes, brain cell types, and spatial tissue contexts through which inherited risk is expressed remain unclear. We aimed to define cell-type-specific and spatially contextualized genetic mechanisms underlying LBP. METHODS: FinnGen R12 LBP GWAS summary statistics (42,521 cases and 353,224 controls) were integrated with brain single-nuclei eQTL data across eight major brain cell classes. We evaluated genome-wide polygenic signal using LDSC, prioritized genes using MAGMA and PoPS, and performed brain cell-type-specific eQTL-anchored Mendelian randomization, primarily based on single-instrument Wald ratio estimates, followed by Bayesian colocalization. Spatial genetic mapping was conducted using gsMap in an E16.5 murine embryonic atlas and two adult human lumbar spinal cord Visium sections. Selected candidates were assessed by RT-qPCR in neuronal-like and astroglial-like inflammatory cell models. RESULTS: LDSC supported interpretable polygenic signal for LBP. MAGMA and PoPS showed partial gene-level convergence, with TCF4 and TMEFF2 supported by both approaches. Across 1641 tested gene-cell type exposures, significant eQTL-anchored MR associations were concentrated in excitatory neurons, oligodendrocytes, inhibitory neurons, and astrocytes. Integrated eQTL-anchored MR, colocalization, and gene-prioritization evidence highlighted CLEC18A, QPRT, and GMPPB as higher-priority non-MHC candidates with moderate, but not strong, colocalization support. gsMap localized LBP-associated enrichment to neuroaxis-related embryonic regions, including brain, spinal cord, sympathetic nerve, and dorsal root ganglion, and to neuronal-like niches in adult lumbar spinal cord. RT-qPCR showed model-dependent expression changes, with QPRT and LGI4 preferentially responsive in neuronal-like SH-SY5Y cells and GMPPB and DPYSL5 responsive in astroglial-like U251 cells. CONCLUSION: These findings support neuronal and glial regulatory programs as plausible contributors to LBP genetic susceptibility and highlight CLEC18A, QPRT, and GMPPB as higher-priority non-MHC candidates with moderate colocalization support. The results provide a spatially contextualized framework for candidate prioritization in LBP, while emphasizing the need for larger cell-type-specific eQTL resources and functional validation before therapeutic or mechanistic conclusions can be drawn.

Mendelian randomization

Chromosome-scale genome assembly and genomic prediction of essential oil compounds in Atractylodes lancea for genomics-assisted breeding.

Atractylodes lancea rhizomes are used as crude drugs. Essential oil compounds, including atractylodin, hinesol, &#x3b2;-eudesmol, and atractylon, are key determinants of crude drug quality. Conventional breeding of A. lancea is difficult because of its perennial growth. In this study, a chromosome-scale reference genome of A. lancea (4.79 Gb) was generated, and genome-wide association studies (GWAS) and genomic predictions of essential oil compounds were conducted to explore the potential for genome-assisted breeding. Genotyping of 480 lines using double-digest restriction-site-associated DNA-sequencing yielded 29,136 high-quality SNPs. All the compounds showed high genomic heritability (h2 = 0.758-0.915), indicating strong genetic control. Despite the high genomic heritability, GWAS detected only one weak association with atractylon and no significant loci for the three compounds. However, genomic prediction achieved moderate to high accuracy across multiple models, particularly the ridge regression, genomic best linear unbiased prediction, and Bayesian approaches. The prediction accuracy, measured as the Pearson correlation coefficient between the observed and predicted values, exceeded 0.6 for all four essential oil compounds. These results demonstrate the efficacy of genomic selection for improving essential oil compound levels in A. lancea and provide a foundation for genome-assisted breeding of medicinal plants with long breeding cycles.

Atractylodes lancea

Polygenic Risk Scores for Incident Dementia in the Multi-Ethnic Study of Atherosclerosis.

Over 75 Alzheimer's disease (AD) and dementia-associated variants have been identified through genome-wide association studies, but the utility of polygenic risk scores (PRS) for predicting AD and dementia in diverse and admixed populations remains unclear. We compared how PRS approaches differing in p-value thresholds, variant weights, and source ancestry perform in predicting dementia in 6338 African American, Chinese, Hispanic, and White individuals from the Multi-Ethnic Study of Atherosclerosis. We tested clumping and thresholding (C+T) methods with varying parameters against Bayesian approaches (PRS-CS, PRS-CSx). We compared the ability of each method to predict incident dementia in all participants and in groups stratified by self-reported race/ethnicity. We additionally analyzed performance across groups stratified by estimated proportion of non-Finnish European (NFE)-like ancestry. Including more variants does not improve performance. We found comparable associations between dementia and PRS when comparing a C+T method with only 15 SNPs and PRS derived from Bayesian models that include >&#x2009;800,000 SNPs (HR5e-08 = 1.18, 95% CI: 1.08-1.28; HRCSx = 1.17, 95% CI: 1.07-1.27). The p&#x2009;<&#x2009;5e-08 C+T method was more strongly associated with incident dementia in populations genetically dissimilar from the source data (HRlowNFE_5e-08 = 1.27, 95% CI: 1.08-1.50; HRlowNFE_CSx = 1.12, 95% CI: 0.94-1.33). More selective PRS models using genome-wide significant SNPs may be preferable for dementia prediction in diverse populations.

Aged