Search PubMedSearch

SEARCH · Search PubMed

Results for “genomic prediction”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genomic prediction and genome-wide association studies of morphological traits and distraction index in Korean Sapsaree dogs.

The Korean Sapsaree dog is a native breed known for its distinctive appearance and historical significance in Korean culture. The accurate estimation of breeding values is essential for the genetic improvement and conservation of such indigenous breeds. This study aimed to evaluate the accuracy of breeding values for body height, body length, chest width, hair length, and distraction index (DI) traits in Korean Sapsaree dogs. Additionally, a genome-wide association study (GWAS) was conducted to identify the genomic regions and nearby candidate genes influencing these traits. Phenotypic data were collected from 378 Korean Sapsaree dogs, and of these, 234 individuals were genotyped using the 170k Illumina CanineHD BeadChip. The accuracy of genomic predictions was evaluated using the traditional BLUP method with phenotypes only on genotyped animals (PBLUP-G), another traditional BLUP method using a pedigree-based relationship matrix (PBLUP) for all individuals, a GBLUP method based on a genomic relationship matrix, and a single-step GBLUP (ssGBLUP) method. Heritability estimates for body height, body length, chest width, hair length, and DI were 0.45, 0.39, 0.32, 0.55, and 0.50, respectively. Accuracy values varied across methods, with ranges of 0.22 to 0.31 for PBLUP-G, 0.30 to 0.57 for PBLUP, 0.31 to 0.54 for GBLUP, and 0.39 to 0.67 for ssGBLUP. Through GWAS, 194 genome-wide significant SNPs associated with studied Sapsaree traits were identified. The selection of the most promising candidate genes was based on gene ontology (GO) terms and functions previously identified to influence traits. Notable genes included CCKAR and DCAF16 for body height, PDZRN3 and CNTN1 for body length, TRIM63, KDELR2, and SUPT3H for chest width, RSPO2, EIF3E, PKHD1L1, TRPS1, and EXT1 for hair length, and DDHD1, BMP4, SEMA3C, and FOXP1 for the DI. These findings suggest that significant QTL, combined with functional candidate genes, can be leveraged to improve the genetic quality of the Sapsaree population. This study provides a foundation for more effective breeding strategies aimed at preserving and enhancing the unique traits of this Korean dog breed.

Animals

Genomic prediction and genome-wide association study for liver abscesses in crossbred beef cattle.

Liver abscesses are a concern in feedlot cattle, and little is known about the role of genetics in their development. This study aimed to estimate genetic parameters and to identify single-nucleotide polymorphisms (SNPs) associated with liver abscesses. Crossbred cattle representing 18 breeds in the U.S. Meat Animal Research Center Germplasm Evaluation Program were phenotyped for liver abscesses at slaughter (n&#x2005;=&#x2005;9,044). Seventeen percent of cattle had liver abscesses. These cattle had genotypes that were imputed to sequence variant genotypes. After filtering and quality control, 340,723 SNPs were used in the analysis. Liver abscess prevalence was modeled with a single-step genomic best linear unbiased prediction (ssGBLUP) threshold model using a Bayesian framework. The model included contemporary group (sex, treatment group, and slaughter date), additive genomic, and residual effects. Genomic heritability was 0.039 (95% highest posterior density&#x2005;=&#x2005;0.005, 0.081), which was very small. To assess prediction quality, a 5-fold random cross-validation structure was used. Method Linear Regression was used to assess accuracy, bias, and dispersion by comparing estimated breeding values (EBV) from full and reduced analyses. Cross-validation metrics showed EBV based on genotypes had 0.05 reliability (SD&#x2005;<&#x2005;0.01) with no bias relative to EBV based on genotypes and phenotypes. For the genome-wide association study, SNP effects were back calculated from the EBV solutions from ssGBLUP. No SNPs were associated with liver abscesses at a Benjamini-Hochberg adjusted 0.05 significance level. Although a large dataset was used, this result was because of the low genomic heritability and imprecise EBV used to calculate SNP effects. Based on these results, environmental factors contribute to most of the variation in liver abscesses. Genetic selection to reduce liver abscesses would be slow because of the low genomic heritability, measurement late in life, and inability to measure breeding animals. A faster approach would be finding additional environmental interventions that maintain animal performance.

Animals

Mul-PheG2P: decoupled learning and prediction-space fusion enables robust and interpretable multi-phenotype genomic prediction.

Genomic prediction of multiple phenotypes is crucial in modern plant breeding; however, existing methods struggle with negative transfer and lack interpretability, particularly across high-dimensional small-sample data and diverse species. To address this, we propose Mul-PheG2P, a novel paradigm based on decoupled learning and predictive space fusion. It employs a two-stage design: first training phenotype-specific encoders using genetic data, then decoupling phenotype-specific learning from cross-phenotype aggregation via an interpretable prediction layer. Mul-PheG2P outperforms existing methods across diverse crop datasets, including maize (Zea mays), wheat (Triticum aestivum), and tomato (Solanum lycopersicum). It provides a multi-scale interpretability chain: at the macro level, it quantifies phenotypic contributions via attention-based weighting; at the micro level, Integrated Gradients reveal the genetic basis of predictions. Notably, the model successfully identified the CCT (CONSTANS, CO-like, and TOC) motif regulating photoperiodism and the SQUAMOSA (SQUAMOSA promoter binding protein) promoter for inflorescence development, confirming its ability to capture functional biological mechanisms. These results highlight the high performance and interpretability of Mul-PheG2P, showcasing its value for low-cost, large-scale screening to advance precision breeding.

Phenotype

Pedigree-assisted genotype imputation enables cost-effective genomic prediction in Penaeus vannamei.

Genomic selection in Penaeus vannamei has long been constrained by the high cost of dense genotyping. To address this limitation, we evaluated genotype imputation from a low-density 1&#xa0;K panel to a medium-density 55&#xa0;K panel of the "Yellow Sea Array No. 1" and examined its impact on genomic prediction for harvest body weight in P. vannamei. A four-generation pedigree including 30 great-grandparents, 39 grandparents, 100 parents, and 608 offspring was genotyped using the 55&#xa0;K panel. A two-step experimental design was implemented to (i) assess the performance of different imputation algorithms under reference population scenarios with varying proportions of siblings, and (ii) compare six alternative reference population structures incorporating parents, ancestors, and siblings. Genotype imputation using the pedigree-based method FImpute v3.0 consistently achieved higher accuracy than the population-based method Beagle v5.5. Using this pedigree-assisted approach, imputation accuracy increased from 0.73 when only parental genotypes were used to 0.84 with the inclusion of 10% siblings, and subsequently plateaued at 0.87-0.90 when sibling representation reached 20%. Across the six reference population structures, imputation accuracy was primarily driven by the availability of parental genotypes, ranging from 0.50 to 0.56 in the absence of parents to 0.88-0.89 when both parents and ancestral generations were included. Accuracy remained high when both parents were available (0.84-0.87 with siblings; 0.73 without siblings) but declined substantially when only one parent was genotyped (0.65-0.68). Imputation accuracy was positively associated with both minor allele frequency (MAF) and linkage disequilibrium (max r2LD), with LD exerting the stronger influence. Heritability estimates derived from imputed 55&#xa0;K genotypes were highly consistent with those obtained from the original 55&#xa0;K data (0.39&#x2009;&#xb1;&#x2009;0.14 vs. 0.41&#x2009;&#xb1;&#x2009;0.14), indicating that genotype imputation did not compromise variance component estimation. In predictive ability analyses, pedigree-based BLUP (PBLUP) achieved higher predictive ability than genomic BLUP (GBLUP) based on the 1&#xa0;K panel, with predictive abilities of 0.42-0.44 for PBLUP compared with 0.34-0.35 for GBLUP. Using imputed genotypes for genomic prediction further improved predictive ability relative to the true 1&#xa0;K panel, yielding values ranging from 0.35 to 0.47. Notably, when parental genotypes were included in the reference population, GBLUP based on imputed genotypes surpassed the predictive ability of PBLUP and approached that achieved with the original 55&#xa0;K genotypes (0.45-0.47). Collectively, these results provide the first empirical evidence that low- to medium-density genotype imputation, combined with pedigree information, can effectively support genomic prediction in P. vannamei. This study establishes a cost-efficient and scalable framework for implementing genomic selection in P. vannamei and provides a practical reference for the application of genomic selection in other aquaculture species with constrained breeding budgets.

Animals

Chromosome-scale genome assembly and genomic prediction of essential oil compounds in Atractylodes lancea for genomics-assisted breeding.

Atractylodes lancea rhizomes are used as crude drugs. Essential oil compounds, including atractylodin, hinesol, &#x3b2;-eudesmol, and atractylon, are key determinants of crude drug quality. Conventional breeding of A. lancea is difficult because of its perennial growth. In this study, a chromosome-scale reference genome of A. lancea (4.79 Gb) was generated, and genome-wide association studies (GWAS) and genomic predictions of essential oil compounds were conducted to explore the potential for genome-assisted breeding. Genotyping of 480 lines using double-digest restriction-site-associated DNA-sequencing yielded 29,136 high-quality SNPs. All the compounds showed high genomic heritability (h2 = 0.758-0.915), indicating strong genetic control. Despite the high genomic heritability, GWAS detected only one weak association with atractylon and no significant loci for the three compounds. However, genomic prediction achieved moderate to high accuracy across multiple models, particularly the ridge regression, genomic best linear unbiased prediction, and Bayesian approaches. The prediction accuracy, measured as the Pearson correlation coefficient between the observed and predicted values, exceeded 0.6 for all four essential oil compounds. These results demonstrate the efficacy of genomic selection for improving essential oil compound levels in A. lancea and provide a foundation for genome-assisted breeding of medicinal plants with long breeding cycles.

Atractylodes lancea

A leakage-aware genomic prediction pipeline for meropenem resistance in Klebsiella pneumoniae using transformer-based resistome representation learning.

MOTIVATION: Antimicrobial resistance (AMR) in Klebsiella pneumoniae, particularly to carbapenems such as meropenem, is a major global health problem. Machine learning is increasingly used to predict resistance from genomic markers; however, many models fail to capture high-level gene-gene interactions and may exhibit inflated performance due to lineage-biased prediction. Existing genomic prediction models largely rely on flat feature representations that fail to capture epistatic gene interactions, and commonly suffer from inflated performance estimates due to phylogenetic data leakage. To address these limitations simultaneously, a leakage-aware hybrid TabTransformer-CatBoost pipeline was developed, combining self-attention-based resistome representation learning with gradient boosting classification under clade-aware data partitioning. A self-attention encoder converts sparse gene presence-absence profiles into contextualized latent embeddings, which are subsequently classified using gradient boosting to capture lineage-aware AMR patterns. RESULTS: The proposed architecture outperformed classical baselines including Logistic Regression, Random Forest, XGBoost, and optimized CatBoost models. Internal accuracy reached 92.59% for the Chained Hybrid configuration (area under the receiver operating characteristic curve, AUROC = 0.8670, F1&#x2009;=&#x2009;0.8537). Performance gains primarily originated from the embedding stage, as confirmed by ablation analysis. External validation across independent multinational cohorts (n&#x2009;=&#x2009;305) demonstrated generalizability (AUROC = 0.8105; F1&#x2009;=&#x2009;0.7552). Permutation testing produced near-zero Matthews Correlation Coefficient (MCC)&#x2009;=&#x2009;0.0091, indicating predictions reflect genuine biological signal rather than noise. These results establish attention-based genomic embedding with gradient boosting as a scalable, interpretable, and leakage-aware framework for clinical AMR prediction. AVAILABILITY AND IMPLEMENTATION: The source code for the TabTransformer-CatBoost framework, including preprocessing pipelines and pre-trained embeddings, is available at https://github.com/SibelKervanci/kp-meropenem-tabtransformer.

Journal Article

Comparing artificial and convolutional neural networks with traditional models for Genomic prediction in wheat.

With the rapid development of sequencing technology, the application of genomic prediction has become more and more common in breeding schemes of livestocks and crops. Selecting an appropriate statistical model is of central importance to achieve high prediction accuracy. Recently, machine learning models have been expected to upgrade genomic prediction into a new era. However, the perspective still suffers from lack of evidence that machine learning models can generally outperform the traditional ones on empirical data sets. In this study, we compared two machine learning models based on artificial neural network (ANN) and convolutional neural network (CNN) with four traditional models, including genomic best linear unbiased prediction (GBLUP), Bayesian ridge regression (BRR), BayesA and BayesB, using three published data sets for grain yield in wheat. For each model, we considered two variants: modeling and ignoring the genotype-by-environment ([Formula: see text]) interaction. In the comparison, we considered two strategies of cross-validation: predicting genotypes that have not been evaluated in any environment (CV1) and predicting genotypes that have been tested in other environments (CV2). Our results showed that traditional Bayesian models (BayesA, BayesB, and BRR) outperformed GBLUP, ANN and CNN when considering [Formula: see text] interaction. The accuracies of ANN and CNN were higher than traditional models only in CV1 and when [Formula: see text] interaction was ignored. It was also found that the performance of the two machine learning models was significantly affected by the interaction between the CV strategy and the way of treating the [Formula: see text] interaction, while that of the four traditional models was only influenced by whether the [Formula: see text] interaction was considered or not. Thus, machine learning models can be a powerful complementary to the traditional ones and their superiority may depend on the prediction scenario. Among the two machine learning models, we observed that the accuracy of ANN was higher than CNN in most cases, indicating that it is still challenging to adapt complex machine learning models such as CNN to genomic prediction.

ANN

Genomic prediction of agronomic traits in perennial ryegrass (Lolium perenne L.) and genotype x environment interactions at the limit of the species distribution.

KEY MESSAGE: Perennial ryegrass shows extensive genotype x environment interactions at the limit of its ecological niche. Accounting for GxE may improve prediction even when environmental and genetic samples are highly diverse. BACKGROUND: In breeding the aim is to identify and accumulate beneficial variants. However, detection of these variants may be challenging in the presence of extensive genotype x environment interactions (GxE). METHODS: The study assesses the performance of 264 diploid perennial ryegrass accessions in a multi-environment field trial. We investigate the extent of GxE, for yield (total dry matter) and persistence traits under environmental conditions experienced in Nordic and Baltic regions at the limit of the species distribution. Two different approaches to modelling GxE were tested and validated under three different breeding scenarios. RESULTS: Our analysis documented the presence of significant GxE for all traits. Validation showed improvements in prediction accuracy when accounting for GxE: up to 4% for yield when predicting in unobserved environments, and up to 22% and 9% for spring cover and winter kill, respectively, when predicting unobserved germplasm. Genome-wide-association-studies (GWAS) were utilized to detect genetic variants with marginal effects (environment-independent effect) and conditional effects (environment-dependent effects). Results showed the presence of large-effect genetic variants with marginal effects, in addition to few Quantitative Trait Loci (QTL) whose effects were adaptive under specific environmental conditions while neutral or deleterious under different environmental conditions. CONCLUSION: This study demonstrates the usefulness and limitations of genomic prediction models for predicting GxE in highly diverse samples and describes the extent of GxE at the limit of species distribution for perennial ryegrass. Our study points towards adaptive variation which may enhance persistence of perennial ryegrass populations in Nordic and Baltic growing conditions.

Lolium

The potential of considering photosynthesis parameters in crop yield breeding by genomic prediction.

To meet the growing demand for agricultural products, optimizing photosynthesis is a promising strategy to improve crop yields. Phenotypic variance in photosynthesis has been observed within or between species. To explore the potential of integrating photosynthetic parameters into crop breeding programs, we explored the genetic variation in photosynthesis by assessing photosynthesis-related parameters across plant development in 631 barley recombinant inbred lines (RILs) from eight HvDRR subpopulations under field conditions. The genetic complexity of these parameters was resolved by analyses of bi-parental and multi-parental quantitative trait loci (QTLs). Finally, we examined the merit of integrating photosynthesis-related parameters in genomic prediction of yield and its components. Significant genotypic variations of the photosynthesis-related parameters were found among the RILs, with their heritability ranging from 0.38 to 0.54. The multiple QTLs and dynamic QTLs for photosynthesis observed across different developmental stages underlined the complexity of the genetics of photosynthesis in barley. The considerably higher percentage of phenotypic variance explained for genomic prediction than multi-parental QTL analysis illustrates that the photosynthesis-related parameters are inherited in a more complex way than classical agronomic traits. Notably, the prediction ability for yield was increased by integrating the photosynthesis-related parameters of some developmental stages into genomic prediction models. Thus, our results suggest a novel perspective on increasing the efficiency of crop breeding programs by integrating photosynthesis-related parameters into prediction models.

Photosynthesis

Dissecting genetic architecture and improving machine learning&#x2011;based genomic prediction of flowering time in Osmanthus fragrans by integrating structural variants.

Sweet osmanthus (Osmanthus fragrans), a traditional ornamental plant in China, exhibits substantial variation in autumn flowering time, which significantly affects landscape application and cultivation efficiency. Here, we performed a genome-wide association study on 127 resequenced accessions classified into early, intermediate, and late flowering types, using a set of 2,325,410 single-nucleotide polymorphisms (SNPs) and 246,824 structural variants (SVs). By integrating SNP/insertion and deletion (Indel) and SV data with weighted gene co-expression network analysis, machine learning, and genomic prediction, we dissected the genetic architecture of flowering time. We identified 24 associated SNP/Indels and six SVs, mapping to 30 candidate genes, including known flowering regulators FLK, LOS1, Y14, MIF2, and GID1B. These genes showed tissue-specific expression, with some responding to low temperature. The two hub genes, GUX1 and LYG027904, were located within modules of the co-expression network associated with low-temperature treatment. Haplotype analysis revealed a specific three-SNP haplotype associated with late flowering and linked to LOS1, and epistatic interactions among combined genotypes contributed to phenotypic variation. Notably, integrating SVs with SNP/Indels improved genomic prediction accuracy; the gradient boosting decision tree model outperformed other machine learning algorithms, achieving a mean accuracy of 0.859 and an AUC&#xa0;>&#xa0;0.8 (where AUC is area under receiver operating characteristic curve) for all flowering types. These findings provide insights into the genetic mechanisms underlying flowering time variation in O. fragrans, offer candidate genes and haplotypes for molecular breeding, and highlight the value of integrating SVs with machine learning for genomic prediction in woody ornamentals.

Machine Learning

Utilizing evolutionary conservation to detect deleterious mutations and improve genomic prediction in cassava.

INTRODUCTION: Cassava (Manihot esculenta) is an annual root crop which provides the major source of calories for over half a billion people around the world. Since its domestication ~10,000 years ago, cassava has been largely clonally propagated through stem cuttings. Minimal sexual recombination has led to an accumulation of deleterious mutations made evident by heavy inbreeding depression. METHODS: To locate and characterize these deleterious mutations, and to measure selection pressure across the cassava genome, we aligned 52 related Euphorbiaceae and other related species representing millions of years of evolution. With single base-pair resolution of genetic conservation, we used protein structure models, amino acid impact, and evolutionary conservation across the Euphorbiaceae to estimate evolutionary constraint. With known deleterious mutations, we aimed to improve genomic evaluations of plant performance through genomic prediction. We first tested this hypothesis through simulation utilizing multi-kernel GBLUP to predict simulated phenotypes across separate populations of cassava. RESULTS: Simulations showed a sizable increase of prediction accuracy when incorporating functional variants in the model when the trait was determined by<100 quantitative trait loci (QTL). Utilizing deleterious mutations and functional weights informed through evolutionary conservation, we saw improvements in genomic prediction accuracy that were dependent on trait and prediction. CONCLUSION: We showed the potential for using evolutionary information to track functional variation across the genome, in order to improve whole genome trait prediction. We anticipate that continued work to improve genotype accuracy and deleterious mutation assessment will lead to improved genomic assessments of cassava clones.

cassava (Manihot esculenta)

Assessment of genomic prediction and genetic gain in multi&#x2011;population half-sib families in the perennial grass crop intermediate wheatgrass.

The University of Minnesota has been domesticating the perennial forage intermediate wheatgrass (IWG) since 2011 using a combination of conventional methods and modern breeding tools such as genomic selection. Globally, most IWG selection nurseries are spaced-planted individuals of several hundred genotypes whereas commercial fields established for grain production are row-planted panmictic populations. This study evaluated genomic prediction models and estimated genetic gain in yield and agronomic performance of row-planted IWG half-sib families assessed over 3 years and 2 locations, Lamberton and St. Paul, MN, USA. The strongest trait correlation was negative (r&#x2009;=&#x2009;-0.49) between 2023 St. Paul height and 2022 St. Paul seed size. The three St. Paul environments were more similar for plant height and seed size and so were the Lamberton environments yet no specific trend was observed for grain yield. Evaluation of different univariate and multivariate genomic prediction models showed that multivariate models outperformed the best univariate models by 23 percentage points, yet no single multivariate model was the best predictor of all traits. Cross-environment predictions were the best among St. Paul environments and no single environment was the best predictor of the remaining environments. Genetic gain estimates indicated a 20 kg ha-1 increase in grain yield and 3 cm reduction in plant height per breeding cycle. While no single model predicted all traits with high accuracy, results obtained in this study suggest that evaluating IWG sibs in row plots followed by genomic trait predictions could lead to desired breeding progress for desired traits.

Poaceae

A comparative study highlights superiority of LSTM in crop genomic prediction.

We systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods, and found LSTM suitable for capturing additive and epistatic effects. Genomic prediction (GP) has been developed as an important method supporting crop breeding. By utilizing the phenotype values result from GP, breeders could make decisions in the seedling stage that consequently benefit for cost saving. In recent years, machine learning emerged as an efficient technology to solve modeling problems in many fields, including crop breeding. However, numerous modeling approaches have hindered the application of GP since breeders struggle to choose. Therefore, a comprehensively methodological research with guiding significance is extremely necessary. In the present study, we systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods. As for genomic feature processing, we found feature selection (SNP filtering approach) performed better than feature extraction (PCA method). Specifically, the feature relationship dependent methods (GBLUP, RNN, and LSTM) as well as DNN architecture showed superior performance with feature selection. Marker density analysis showed positive correlation with prediction accuracy in a limited threshold. Comparison on effect of population size demonstrated a positive correlation between trait genetic complexity and the optimal population size required. By testing fifteen modeling methods, we found LSTM network displayed superior performance, achieving the highest average STScore (0.967) across six datasets. Further research using all cell states or the latest cell states of LSTM inputs demonstrated its architecture particularly adept with capturing additive and epistatic QTL effects among SNPs. In conclusion, our findings provide basic principles for implementing GP in breeding project to maximize prediction accuracy while maintaining cost-effectiveness.

Plant Breeding

Assessment of genomic prediction capabilities of transcriptome data in a barley multi-parent RIL population.

Low-cost and high-throughput RNA sequencing data for barley RILs achieved GP performance comparable to or better than traditional SNP array datasets when combined with parental whole-genome sequencing SNP data. The field of genomic selection (GS) is advancing rapidly on many fronts including the utilization of multi-omics datasets with the goal of increasing prediction ability and becoming an integral part of an increasing number of breeding programs ensuring future food security. In this study, we used RNA sequencing (RNA-Seq) data to perform genomic prediction (GP) on three related barley RIL populations. We investigated the potential of increasing prediction ability by combining genomic and transcriptomic datasets, adding whole-genome sequencing (WGS) SNP data, functional annotation-based filtering, and empirical quality filtering. Our RNA-Seq data were generated cost-efficiently using small-footprint plant cultivation, high-throughput RNA extraction, and Library preparation miniaturization. We also examined sequencing depth reduction as an additional cost-saving measure. We used fivefold cross-validation to evaluate the prediction ability of the gene expression dataset, the RNA-Seq SNP dataset, and the consensus SNP dataset between the RNA-Seq and parental WGS data, resulting in prediction abilities between 0.73 and 0.78. The consensus SNP dataset performed best, with five out of eight traits performing significantly better compared to a 50K SNP array, which served as a benchmark. The advantage of the consensus SNP dataset was most prominent in the inter-population predictions, in which the training and validation sets originated from different RIL sub-populations. We were therefore able to not only show that RNA-Seq data alone are able to predict various complex traits in barley using RILs, but also that the performance can be further increased with WGS data for which the public availability will steadily increase.

Hordeum

Dissecting genetic variance structure and evaluating genomic prediction models for single-cross hybrids derived from Stiff Stalk and Non-Stiff Stalk maize heterotic groups.

The early 20th-century discovery of heterosis and the establishment of heterotic groups transformed maize (Zea mays L.) into a keystone of global agriculture. However, maize breeding faces two significant challenges: the gradual decline of general combining ability (GCA) variance within heterotic groups and the impracticality of testing all possible single crosses in the early stages of a breeding program. Here, we developed genomic best linear unbiased prediction (GBLUP)-based multikernel models, using additive and two alternative nonadditive genomic relationship matrices, to estimate the variance components associated with the general combining ability of Stiff Stalk (SS) and Non-Stiff Stalk (NSS) heterotic groups and the specific combining ability arising from their crosses. We further applied these models to predict the performance of untested single-cross combinations under varying levels of parental information. We showed that the SS and NSS groups retained significant GCA variance across traits in both early- and late-maturity groups. The SS group, in contrast, exhibited no detectable GCA variance in grain yield for the intermediate-flowering subset of hybrids, highlighting a limitation for future genetic improvement. Furthermore, our results showed that GBLUP-based multikernel models effectively identified superior hybrids when parental information was available. In the absence of this information, however, these models underperformed compared to covariance-based approaches. Both nonadditive matrices yielded similar results, indicating that they capture comparable genetic relationship patterns despite their distinct formulations. Overall, this study sheds light on the future use of US maize commercial germplasm and demonstrates how GBLUP-based multikernel models can improve the efficiency of hybrid breeding programs.

Zea mays

Whole-genome prediction of bacterial pathogenic capacity on novel bacteria using protein language models with PathogenFinder2.

MOTIVATION: Infectious diseases continue to be a leading cause of mortality and pose a significant global health threat. Thus, the development of tools for surveillance and early detection of emerging pathogens is needed. RESULTS: We introduce PathogenFinder2, a novel, alignment-free, taxonomy-agnostic model for predicting bacterial pathogenic capacity in humans using protein language models. It outperforms previous methods, particularly for novel taxa, and provides interpretable outputs by highlighting proteins most relevant to pathogenic potential. These insights aid the identification of virulence factors, vaccine targets, and infection-related metabolic pathways. Furthermore, we introduce the Bacterial Pathogenic Capacity Landscape, which reveals patterns linked to host condition, infection site, microbial antagonism, and environmental origin. AVAILABILITY: The model is freely available online at https://genepi.dk/pathogenfinder2, or as a standalone program (https://github.com/genomicepidemiology/PathogenFinder2).

Genome, Bacterial

A simple approach for multiple observations improves power to detect genetic effects and genomic prediction accuracy.

Many datasets, including widely used biobanks, have more than one observation of numerous phenotypes for at least a portion of their sample. The majority of GWAS utilize only a single observation per individual, even when more than one observation may be available, and apply a standard model in which the additive allelic effect being estimated is assumed to be constant across the age or time range in the sample. Here, we test a set of simple approaches to utilize multiple observations per individual, under this same assumption. We find that utilizing the mean or median of the available observations rather than a single observation improves power to detect associated loci and enriched gene sets and yields higher out-of-sample polygenic score prediction accuracy. Despite growing biobanks, many deeply phenotyped samples are relatively small but have multiple observations. While explicitly modeling age- or time-dependent genetic effects can estimate time- or age-specific genetic effects, most GWAS apply a standard, additive-only model; a simple approach of using the mean or median can improve power by reducing "noise" in the phenotype, utilize standard, optimized software, and be particularly impactful for smaller samples, including samples of diverse genetic ancestry currently existing in widely used biobanks.

Journal Article

Pretraining improves prediction of genomic datasets across species.

MOTIVATION: Recent studies suggest that deep neural network models trained on thousands of human genomic datasets can accurately predict genomic features, including gene expression and chromatin accessibility. However, training these models is computation- and time-intensive, and datasets of comparable size do not exist for most other organisms. RESULTS: Here, we identify modifications to an existing state-of-the-art model that improve model accuracy while reducing training time and computational cost. Using this streamlined model architecture, we investigate the ability of models pretrained on human genomic datasets to transfer performance to a variety of different tasks. Models pretrained on human data but fine-tuned on genomic datasets from diverse tissues and species achieved significantly higher prediction accuracy while significantly reducing training time compared to models trained from scratch, with Pearson correlation coefficients between experimental results and predictions as high as 0.8. Further, we found that including excessive training tasks decreased model performance and that this decrease could be partially but not completely rescued by fine-tuning. Thus, simplifying model architecture, applying pretrained models, and carefully considering the number of training tasks may be effective and economical techniques for building new models across data types, tissues, and species. AVAILABILITY AND IMPLEMENTATION: Code is available on GitHub and Figshare: https://github.com/optimizedlearning/genomicsML, https://doi.org/10.6084/m9.figshare.31796116.

Genomics